DNA Data Encoding With 5-Bit Transcoding for Random Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

DNA-based data storage faces challenges such as high production costs, low data retrieval speed, errors in encoding and decoding, and limited random access capabilities, making it unsuitable for frequent data access and large-scale applications.

Innovation Solution

A novel 5-bit transcoding framework combined with compression and error correction algorithms is used to convert data into nucleic acid sequences, allowing for efficient, reliable storage and retrieval, including random access to partial data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If DNA-based storage is used for large-scale archival storage, then storage density is improved, but data retrieval speed deteriorates

Engineering Contradiction:
Improvestorage densityVSAvoiddata retrieval speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The patent segments data into multiple DNA strands with unique index sequences, allowing parallel processing and selective retrieval of specific data portions without sequencing entire datasets, thus improving retrieval speed while maintaining high storage density

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If DNA synthesis is used for data encoding, then storage capacity is improved, but production cost deteriorates

Engineering Contradiction:
Improvestorage capacityVSAvoidproduction cost
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent performs preliminary data compression and error correction encoding before DNA synthesis, reducing the total amount of DNA required and optimizing the synthesis process, thereby lowering production costs while maintaining large storage capacity

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If traditional sequencing methods are used for data retrieval, then data accuracy is improved, but retrieval time deteriorates

Engineering Contradiction:
Improvedata accuracyVSAvoidretrieval time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and utilizes unique index sequences embedded in DNA strands to identify and retrieve specific data portions without requiring complete sequencing, significantly reducing retrieval time while maintaining data accuracy through targeted verification

Inventive Principle:
Principle #2Taking out (Extraction)

4Reliability

If error correction codes are added to DNA sequences, then data reliability is improved, but sequence complexity deteriorates

Engineering Contradiction:
Improvedata reliabilityVSAvoidsequence complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges error correction codes with data payload and index information into unified DNA sequences through systematic encoding, maintaining reliability while managing complexity through integrated design rather than separate components

Inventive Principle:
Principle #5Merging (Combining)

5Ease of operation

If random access is implemented in DNA storage, then data accessibility is improved, but process complexity deteriorates

Engineering Contradiction:
Improvedata accessibilityVSAvoidprocess complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments data into independently addressable DNA strands with unique indexes, enabling random access to specific segments through targeted retrieval processes, improving accessibility while managing complexity through modular organization

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3659147B1DNA-based data storage
Publication Date: 2025.11.05 NANJING GENSCRIPT BIOTECH CO LTD
  • EP3659147B1 patent drawingFigure 1
  • EP3659147B1 patent drawingFigure 2
  • EP3659147B1 patent drawingFigure 3A

AI summary

The present disclosure generally relates to DNA-based data storage. An exemplary method for storing input data on nucleic acid comprises: converting the input data into a set of nucleotide sequences and synthesizing a set of nucleic acids comprising the set of nucleotide sequences. The converting comprises a data processing step comprising converting the input data into a binary string, and a nucleotide encoding step comprising converting the binary string using a 5-bit transcoding framework to obtain the set of nucleotide sequences.