A neural network enhanced base drift type DNA storage error correction encoding method
Patent Information
- Application Number
- CN202511310867.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-09-15
AI Technical Summary
现有的RS、LDPC、BCH、汉明码等纠错编码虽能在一定程度上修正替换错误,但对碱基插入和删除错误的纠错能力有限,且冗余度与纠错能力近似呈线性关系:冗余过高会显著降低存储密度并增加合成成本,冗余不足则极易导致整链解码失败
[0035]1、本发明在不引入任何纠错码的条件下,即使插入、删除错误率高达1.1%仍然能够达到较高的解码恢复可能性,提高了DNA存储系统的纠错能力;
Smart Images

Figure CN121191583B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coding technology, and more specifically to a neural network-enhanced base-shifting DNA storage error correction coding method. Background Technology
[0002] With the exponential growth of global data, traditional storage media such as hard drives, magnetic tapes, and flash memory are approaching their physical limits in terms of capacity, energy consumption, and lifespan, urgently requiring new and transformative storage solutions. DNA, as a natural information carrier, possesses outstanding advantages such as a theoretical storage capacity of up to 215 PB per gram, stable preservation at room temperature for thousands of years, and near-zero energy consumption for maintenance, rapidly becoming a cutting-edge research area in the interdisciplinary fields of biology and information science. Since Church's team first encoded book content into DNA and fully read it in 2012, numerous research institutions worldwide, including Columbia University and Microsoft Labs, have successively achieved large-scale writing and recovery of data such as text, audio, video, and even operating systems. They have initially established a complete technology chain of "binary → error correction coding → base mapping → chemical synthesis → sequencing → decoding and restoration," and promoted the rapid iterative development of related reagents, instruments, and microfluidic synthesis platforms.
[0003] However, the practical application of DNA storage still faces constraints stemming from the interplay between biotechnological noise and coding redundancy. Synthesis, amplification, and sequencing processes randomly introduce various errors, including substitutions, insertions, and deletions, with highly uncertain locations and types. While existing error-correcting codes such as RS, LDPC, BCH, and Hamming codes can correct substitution errors to some extent, their ability to correct insertion and deletion errors is limited. Furthermore, redundancy and error-correction capability exhibit an approximately linear relationship: excessive redundancy significantly reduces storage density and increases synthesis costs, while insufficient redundancy easily leads to whole-strand decoding failure. Therefore, constructing a coding system with high robustness, low redundancy, and strong error correction capabilities for insertion / deletion errors has become a core technological bottleneck in the transition of DNA storage from laboratory research to petabyte-scale data center deployments. Summary of the Invention
[0004] To address the aforementioned deficiencies in existing technologies, this invention proposes a neural network-enhanced base-drift DNA storage error correction coding method. This method achieves stable mapping from binary data to DNA sequences while satisfying multidimensional biological constraints. It introduces a one-dimensional convolutional neural network (1D-CNN) to construct an insertion / deletion error classification model, and combines a dynamic window with a base-drift error correction strategy to achieve accurate repair of insertion / deletion errors.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a neural network-enhanced base-drift DNA storage error correction coding method, comprising:
[0006] Step 1: Convert the original text data into a continuous binary data stream and divide it into basic units of 8 bits each. For each type of basic unit after deduplication, generate DNA coding units that satisfy GC content constraints, homopolymer length constraints, orthogonality constraints, and repeating substring constraints respectively.
[0007] Step 2: Based on the above generation results, establish the mapping relationship between each binary basic unit and its corresponding DNA coding unit, thereby constructing the codebook.
[0008] Step 3: Use the codebook to convert the binary data of the file into DNA sequences that meet biological constraints, and set the length of each DNA sequence to 120 bases (nt).
[0009] Step 4: By simulating the processes of DNA synthesis, storage, and sequencing, base deletion and insertion errors are randomly introduced into each DNA sequence;
[0010] Furthermore, the method for determining whether a DNA sequence has base insertion and deletion errors based on a neural network includes the following construction process:
[0011] Step 5: Dataset Construction. Using the codebook, randomly select coding units and concatenate them to generate 10,000 DNA sequences, each 120 bases (nt) in length. ,according to P sub =4% P del =0.4%, P ins =0.4% introduces three types of errors and tags Introducing incorrect DNA sequences according to a function Numerical vector mapping, after zero-padding or truncation alignment, is divided into training / test sets in an 8:2 ratio;
[0012] Step 6: Extract local features of each DNA sequence using a one-dimensional convolutional (Conv) neural network with a stride of 8 and a kernel size of 8. The calculation method is shown in formula (1):
[0013] ;
[0014] in, X Given the input sequence, W For convolution kernel weights, b For bias terms, sThe activation function is denoted as . The position embedding vector is generated based on the sine / cosine function through the position encoding layer, as shown in equations (2) and (3):
[0015] ;
[0016] ;
[0017] in, pos For location index, i For dimensional indexing, d model For the model dimension, PE ( pos , 2 i ) and PE ( pos ,2 i +1) represents the position. pos Values are encoded at positions in even and odd dimensions.
[0018] Step 7: Construct the full-length dependency of the DNA sequence using a Transformer encoder containing 3 coding blocks, each with 4 self-attention mechanisms. The calculation method of the self-attention mechanism is shown in Equation (4):
[0019] ;
[0020] in, Q , K , V These are the query, key, and value matrices, obtained through linear transformation. d k The dimension of each attention head is defined. The encoder output is then input into the fully connected layer after global average pooling, and the error probability in the [0,1] interval is output through the sigmoid activation function, as shown in formula (5):
[0021] ;
[0022] in, h The output features of this network, W and b For the weights and biases of the fully connected layer, s It is the sigmoid activation function. P ∈[0,1] represents the error probability.
[0023] Step 8: Use the binary cross-entropy loss (BCELoss) to measure the prediction error. The calculation method is shown in formula (6):
[0024] ;
[0025] in, yFor true labels (0 or 1). p To predict probabilities, the optimizer uses Adam with an initial learning rate of 0.001 and prevents gradient explosion through gradient clipping.
[0026] Step 9: Input the DNA sequence with insertion / deletion errors introduced in Step 4 into the error detection network described in Step 7 for classification. If the prediction result is no insertion / deletion error, decode directly according to the Codebook; if the prediction result is insertion / deletion error, perform base drift error correction before decoding.
[0027] Step 10: For DNA sequences tagged with 1 S i Using window length w Divide the sequence using a sliding window of 8 to obtain sequence units. Each window unit is compared with the DNA coding units of the Codebook, and the matching unit with the smallest edit distance is selected. Within every three consecutive windows, the system records whether there are insertion or deletion errors based on the threshold change of the edit distance.
[0028] Step 11: If an insertion or deletion error is detected in three consecutive windows, then the sequence corresponding to that window is closed. DNA coding units matching the Codebook Perform a fuzzy alignment and calculate the cumulative offset based on the fuzzy alignment result. D If it's an insertion error, D ← D +1, otherwise delete the error. D ← D -1;
[0029] Step 12: Based on whether an insert or delete operation exists, prioritize using... structure ; the original sequence The segment before the error location, the sequence after the error location, and the cumulative offset. D Perform the splicing as shown in Formula 10:
[0030] ;
[0031] Step 13: Continue processing the newly generated sequence Proceed to step 10 to determine if any insertion or deletion errors still exist in the sequence. If the result indicates no errors, decode directly using the codebook to restore the original data.
[0032] Step 14: If insertion or deletion errors still exist in the sequence, then for deletion errors, sequentially check the sequence... S i Insert one or two bases at the incorrect position to generate candidate sequences. S i ′ To address insertion errors, one or two bases are removed sequentially to generate candidate sequences. S i ′ Calculate the matching score between all candidate sequences and the codebook, and select the sequence with the best score as the error correction result.
[0033] Step 15: Finally, according to the preset encoding rules Codebook, the DNA sequence is reverse-mapped unit by unit to convert it into raw binary data, thus completing the information recovery.
[0034] By adopting the above technical solution, the present invention can achieve the following technical effects:
[0035] 1. This invention achieves a high decoding and recovery probability even with an insertion and deletion error rate as high as 1.1% without introducing any error correction codes, thus improving the error correction capability of the DNA storage system.
[0036] 2. This method identifies sequences containing Indel errors through neural networks and initiates sliding window error correction only for sequences judged to be erroneous. At an error rate of 2%, the average decoding time is reduced by about 9%, significantly reducing computational overhead while maintaining accuracy. Attached Figure Description
[0037] Figure 1 This is a flowchart illustrating the implementation of a neural network-enhanced base-drift DNA storage error correction coding method. Detailed Implementation
[0038] The technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the protection scope of this invention.
[0039] The constraints involved in this invention include thermodynamic constraints, which are forcibly satisfied. (In this work, δ=10%) was set to control PCR amplification efficiency; (2) Synthesis constraints: limiting the length of homopolymers (2) Avoid fluorescence signal attenuation on the Illumina sequencing platform; (3) Orthogonal constraint: ensure edit distance between any two coding units (4) Repetitive substring constraint: prohibits the occurrence of tandem repeat sequences with a length greater than 2. Example
[0040] The embodiments of this invention are specific application examples based on the technical solution of this invention, intended to provide a detailed explanation of the technical implementation process of this invention, but not to limit the scope of protection of this invention. In this embodiment, a 60-bit binary data sequence is taken as an example. It is converted into a DNA sequence according to the Codebook, and insertion / deletion error interference is applied to the DNA sequence. Subsequently, the error detection and correction method proposed in this invention is used to correct the interference in the DNA sequence.
[0041] Step 1: According to the Codebook, encode the 64-bit binary data sequence into a 64-nt DNA sequence: TAGGCTAGGACACTTCTGATGGTCGATCCAGTGTACGACTGACTATGACATGACTGCTAGTGCT;
[0042] An insertion error of one A base was introduced at the 14th base position of this sequence, resulting in the interfered DNA sequence: TAGGCTAGGACACATTCTGATGGTCGATCCAGTGTACGACTGACTATGACATGACTGCTAGTGCT.
[0043] Step 2: Input the interfered DNA sequence obtained in Step 1 into a pre-trained neural network model to detect whether the sequence contains insertion or deletion errors. The model performs position-level prediction based on the base feature vector of the input sequence: an output of 0 indicates that there is no insertion / deletion error at that position, and an output of 1 indicates that there is an insertion / deletion error at that position. This step can preliminarily screen out sequences containing insertion or deletion errors.
[0044] Step 3: If the result of Step 2 indicates the presence of an insertion or deletion error, the DNA sequence affected by the error is localized using a sliding window. The window length is 8 nt, and the window slides sequentially in 8-nt increments. For each base fragment within the window, a matching coding unit is searched in a pre-defined DNA coding dictionary (Codebook), and the edit distance between the fragment and each coding unit is calculated to obtain the corresponding edit distance set. The set of edit distances for three consecutive windows , , , , , Threshold checks are performed separately. If the edit distances of the two subsequent positions at a given location in the three windows both exceed a preset threshold, then that location is marked as an error location. For example, in this embodiment, the result is that the position of the second window is incorrect.
[0045] Step 4: Perform a fuzzy alignment between the second window (containing 8 bases) identified as the error location in Step 3 and the coding units in the DNA coding dictionary (Codebook), obtaining a set of alignment results, for example: A1:GACACATT-, A2:GACAC-TTC. Count the first 8 characters of each alignment result, calculating the number of "-" symbols, and use this to determine the insertion or deletion error type and perform a preliminary count. For example, for alignment result A2, if one "-" symbol is found in the first 8 characters, it is determined to be an insertion error. Based on this determination, adjust the sequence offset Δ, where for insertion errors, Δ = Δ + 1.
[0046] Step 5: Based on the fuzzy alignment results and corresponding sequence offset Δ obtained in Step 4, adjust the position of the erroneous DNA sequence. For the first detected error position, the corresponding 8-base fragment remains unchanged, TAGGCTAG. For the DNA sequence after this position, the entire sequence is shifted according to the offset Δ. Taking an insertion error as an example, if Δ = +1, the original subsequent sequence: CTGATGGTCGATCCAGTGTACGACTCACTATGACATGACTGCTAGTGCT is shifted one position to the right, adjusted to: TGATGGTCGATCCAGTGTACGACTGACTATGACATGACTGCTAGTGCT. This adjustment can eliminate the misalignment effect caused by the base insertion error, thereby restoring the correct alignment of subsequent bases.
[0047] Step 6: After completing the offset correction in Step 5, the DNA sequence continues to be tested. Specifically, the state of the edit distance set corresponding to every three positions is checked sequentially: if most values in the edit distance set at that position are 0, the error at that position is considered to have been effectively mitigated; if a large number of values greater than 1 still exist in the edit distance set at that position, it indicates that the aforementioned fuzzy alignment correction effect is not ideal. For the latter, the system enters the brute-force base drift correction mode, that is, for the detected error position, various correction operations such as base insertion and base deletion are tried at different positions to generate multiple candidate corrected sequences. Subsequently, each candidate sequence is compared with the reference coding dictionary, its overall minimum edit distance set is calculated, the rationality of each correction scheme is verified, and finally the optimal scheme is selected as the corrected DNA sequence and output to the decoding module.
[0048] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A neural network-enhanced base-shift DNA storage error correction coding method, characterized in that, Includes the following steps: Step 1: Divide the raw binary data stream to be stored into basic units of 8 bits each; Step 2: For all the deduplicated basic units, generate DNA coding units that simultaneously satisfy GC content constraints, homopolymer length constraints, orthogonality constraints, and repeating substring constraints. Step 3: Establish a mapping relationship between each 8-bit binary basic unit and its unique DNA coding unit to form a coding dictionary; Step 4: Use the encoding dictionary to convert the file binary data into DNA sequences that meet biological constraints, with each DNA sequence having a fixed length of 120 nt; Step 5: By simulating the DNA synthesis, storage, and sequencing process, base deletion and insertion errors are randomly introduced into each DNA sequence; Step 6: Build and train a neural network-based error detection model to determine whether any input DNA sequence contains base insertion and / or deletion errors; Step 7: When the error detection model outputs no errors, directly decode using the encoding dictionary; when the output contains errors, process the sequence judged to have errors. S i Based on window length w = 8 is used for sliding segmentation to obtain a set of sequence units; Step 8: Compare the edit distance of each window unit with the DNA coding units in the coding dictionary, and select the unit with the smallest distance match; Step 9: Within three consecutive window ranges, determine whether there are insertion or deletion errors based on the changes in the edit distance threshold; Step 10 If S i If an error exists, the cumulative offset is calculated using fuzzy alignment. Δ Insertion error Δ ← Δ +1, delete error Δ ← Δ -1; Step 11 Based on offset Δ Reconstructed sequence: The DNA sequence fragment before the error position is spliced with the DNA coding unit in the coding dictionary that matches the error position, and then the DNA sequence after the error position is added, based on the stated offset. Δ By adjusting the base positions, the reconstructed sequence is obtained. S i ’ ’ Step 12: When it is determined that the error has been mitigated based on the reconstruction result of Step 11, the reconstructed sequence is decoded using an encoding dictionary to recover the original data; when the error has not been mitigated, one or two bases are inserted sequentially at the error position, or one or two bases are deleted sequentially if an insertion error exists, thereby generating multiple candidate correction sequences. S i ’’ ; step 13. Calculate the matching score between all candidate correction sequences and the encoding dictionary, and select the one with the best score as the error correction result. S ik ’’ ; Step 14: Finally, the original data is recovered by decoding using the encoded dictionary.
2. The neural network-enhanced base-shift DNA storage error correction coding method according to claim 1, characterized in that, In step 2, the GC content is constrained to 40%–60%, the homopolymer length is constrained to ≤4 consecutive identical bases, the orthogonality is constrained to Hamming distance ≥3 between any two coding units, and the repeating substring is constrained to have no completely repeating substrings of length ≥2.
3. The neural network-enhanced base-shift DNA storage error correction coding method according to claim 1, characterized in that, Step 9: Determine the change in distance threshold: If S i If the minimum edit distance of three consecutive windows is ≥2 at two or more positions within the window, then an insertion or deletion error is determined to exist in those three consecutive windows.
Citation Information
Patent Citations
Insertion, deletion and replacement error correction coding method applied to DNA storage
CN117271200A
DNA information storage method based on improved coding table and HMSA
CN117912563A