DNA Data Error Correction via Text Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for data error correction on DNA storage require redundant data storage, leading to high storage costs due to the need for repeated data storage to improve recovery success rates.
Innovation Solution
A data error correction method that decodes base sequences into text, performs word segmentation, detects errors in text units using natural language processing algorithms, and corrects errors without requiring redundant storage, thereby reducing storage costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundant data storage is increased to improve error correction capability, then data recovery success rate is improved, but storage cost increases
Solution Approach 1:
The patent segments the DNA base sequence into multiple text units (words, phrases, or sentences) and performs error detection on each segment independently using natural language processing algorithms. This segmentation allows error correction without requiring redundant storage of entire data blocks, as each segment can be validated and corrected individually based on linguistic rules and patterns.
Solution Approach 2:
The patent replaces the traditional mechanical/redundant storage approach with an information-processing approach using natural language processing algorithms. Instead of physically storing multiple copies of data for error correction, the system uses computational methods (word segmentation, linguistic pattern recognition, and NLP-based error detection) to identify and correct errors in the DNA-encoded text, thereby reducing the need for redundant DNA storage.
2Measurement precision
If natural language processing algorithms are used for error detection, then error correction accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent performs word segmentation and preliminary text processing on the decoded DNA sequence before error detection. By pre-processing the text into meaningful units (words, phrases, sentences) and establishing linguistic structures in advance, the system prepares the data for more efficient error detection using natural language processing algorithms, reducing the computational burden during the actual error correction phase.
Data Source
AI summary
The present application is suitable for the technical field of data processing, and provides a data error correction method and apparatus and an electronic device. The method includes: decoding a base sequence to be subjected to error correction into a first text, the base sequence to be subjected to error correction being composed of a plurality of bases; performing word segmentation on the first text to obtain a plurality of text units; performing error detection on the plurality of text units to obtain a text unit having an error; and performing error correction on the base sequence to be subjected to error correction according to the text unit having the error. By means of the above method, error correction for data can be achieved, and the storage cost of DNA can also be reduced.


