DNA Data Storage Encoding for Random Access Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

DNA-based data storage faces challenges such as high production costs, low data retrieval speed, errors in encoding, writing, storing, decoding, and reading, and difficulty in achieving random access to large datasets.

Innovation Solution

A novel 5-bit transcoding framework, combined with compression and error correction algorithms, is used to convert data into nucleotide sequences, allowing for efficient and reliable storage and retrieval, including random access to partial data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If DNA-based storage is used for large-scale archival storage, then storage density and long-term storage capability are improved, but data retrieval speed deteriorates due to sequencing requirements

Engineering Contradiction:
Improvestorage densityVSAvoiddata retrieval speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The patent divides the stored data into multiple independent DNA fragments, each with its own index sequence. This segmentation allows selective amplification and sequencing of only the required fragments rather than the entire dataset, thereby improving retrieval speed while maintaining high storage density.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent incorporates index sequences and barcodes into the DNA structure during the encoding phase. These preliminary markers enable rapid identification and targeted retrieval of specific data portions without requiring complete sequencing, thus resolving the speed-density tradeoff.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If error correction algorithms are implemented in the DNA storage process, then data reliability is improved, but device complexity increases due to additional processing steps

Engineering Contradiction:
Improvedata reliabilityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines error correction codes with the data encoding process by integrating redundancy information directly into the DNA sequence structure. This merging approach achieves reliable error correction while minimizing additional processing complexity through unified encoding operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent employs configurable error correction parameters that can be adjusted based on storage requirements. By changing the level of redundancy and correction strength, the system achieves reliable error handling while allowing flexibility to balance against processing complexity depending on specific application needs.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If random access to partial data is implemented in DNA storage, then data accessibility is improved, but production cost increases due to additional indexing and synthesis requirements

Engineering Contradiction:
Improvedata accessibilityVSAvoidproduction cost
Core Design Contradiction:
Ease of operationVSEase of manufacture

Solution Approach 1:

The patent introduces index sequences and barcode markers as intermediary elements between the data and the DNA storage medium. These intermediaries enable efficient random access by allowing targeted retrieval operations without requiring complete sequencing, thereby improving accessibility while managing synthesis costs through selective amplification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

By segmenting the data into independently addressable DNA fragments with unique indexes, the system enables random access to specific portions without synthesizing or sequencing the entire dataset, thus improving accessibility while controlling production costs through targeted operations.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12512185B2DNA-based data storage and retrieval
Publication Date: 2025.12.30 NANJING GENSCRIPT BIOTECH CO LTD
  • US12512185B2 patent drawing
  • US12512185B2 patent drawing
  • US12512185B2 patent drawing

AI summary

The present disclosure generally relates to DNA-based data storage. An exemplary method for storing input data on nucleic acid comprises: converting the input data into a set of nucleotide sequences and synthesizing a set of nucleic acids comprising the set of nucleotide sequences. The converting comprises a data processing step comprising converting the input data into a binary string, and a nucleotide encoding step comprising converting the binary string using a 5-bit transcoding framework to obtain the set of nucleotide sequences.