DNA Data Error Correction via Text Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for data error correction on DNA storage require redundant data storage, leading to high storage costs due to the need for repeated data storage to improve recovery success rates.

Innovation Solution

A data error correction method that decodes base sequences into text, performs word segmentation, detects errors in text units using natural language processing algorithms, and corrects errors without requiring redundant storage, thereby reducing storage costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If redundant data storage is increased to improve error correction capability, then data recovery success rate is improved, but storage cost increases

Engineering Contradiction:
Improvedata recovery success rateVSAvoidstorage cost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the DNA base sequence into multiple text units (words, phrases, or sentences) and performs error detection on each segment independently using natural language processing algorithms. This segmentation allows error correction without requiring redundant storage of entire data blocks, as each segment can be validated and corrected individually based on linguistic rules and patterns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces the traditional mechanical/redundant storage approach with an information-processing approach using natural language processing algorithms. Instead of physically storing multiple copies of data for error correction, the system uses computational methods (word segmentation, linguistic pattern recognition, and NLP-based error detection) to identify and correct errors in the DNA-encoded text, thereby reducing the need for redundant DNA storage.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If natural language processing algorithms are used for error detection, then error correction accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improveerror detection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs word segmentation and preliminary text processing on the decoded DNA sequence before error detection. By pre-processing the text into meaningful units (words, phrases, sentences) and establishing linguistic structures in advance, the system prepares the data for more efficient error detection using natural language processing algorithms, reducing the computational burden during the actual error correction phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240419897A1Data error correction method and apparatus, and electronic device
Publication Date: 2024.12.19 SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
  • US20240419897A1 patent drawing
  • US20240419897A1 patent drawing
  • US20240419897A1 patent drawing

AI summary

The present application is suitable for the technical field of data processing, and provides a data error correction method and apparatus and an electronic device. The method includes: decoding a base sequence to be subjected to error correction into a first text, the base sequence to be subjected to error correction being composed of a plurality of bases; performing word segmentation on the first text to obtain a plurality of text units; performing error detection on the plurality of text units to obtain a text unit having an error; and performing error correction on the base sequence to be subjected to error correction according to the text unit having the error. By means of the above method, error correction for data can be achieved, and the storage cost of DNA can also be reduced.