Effective Error Correction Method and System Based on Multiple Outputs of IDS Channels for DNA Data Storage

The proposed method addresses errors in DNA data storage by encoding information bits, mapping to DNA bases, and using parallel IDS channels with SPM and ENMS algorithms for iterative decoding, enhancing error correction and storage performance.

CN119649883BActive Publication Date: 2025-07-15GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411710097.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-07-15
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Existing DNA data storage methods are difficult to effectively correct errors in insertion, deletion and replacement forms during synthesis and sequencing, resulting in poor storage effects.

Method used

An effective error correction method based on the DNA data storage IDS channel is adopted. Through internal encoding, external encoding, mapping, segmentation and parallel channel transmission, combined with SPM algorithm, synchronous decoding and ENMS algorithm, the consensus sequence is calculated using the HMM model and Levenshtein distance to improve error correction performance.

Benefits of technology

Effectively avoid errors in DNA synthesis and sequencing, improve error correction performance, reduce bit error rate, and improve storage reliability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649883B_ABST
    Figure CN119649883B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information storage, and more particularly, to an effective error correction method and system for multiple outputs of an IDS channel based on DNA data storage. The method includes: performing internal encoding on an input information bit sequence and then performing external encoding to generate an external encoded bit sequence; performing mapping to obtain a DNA base sequence, and splitting it into M short sequences; obtaining M*N short output sequences through N parallel IDS channels; performing splicing through the SPM algorithm to obtain a consensus sequence; decoding the consensus sequence first through synchronous decoding and then through the ENMS algorithm; and performing iteration on the synchronous decoding and the ENMS algorithm through a joint iterative algorithm. The present invention obtains multiple outputs through parallel IDS channels, uses the segmented progressive alignment algorithm to combine the homologous information of multiple outputs to obtain a consensus sequence, which can effectively avoid errors in the DNA synthesis and sequencing processes, and then uses the synchronous decoding and ENMS algorithms for decoding and performs joint iteration, further improving the error correction performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information storage, and more specifically, to an effective error correction method and system for multiple outputs of an IDS channel based on DNA data storage. Background Art

[0002] With the rapid development of information technology, the global data volume has increased explosively. Currently, the main storage media in use include mechanical hard disks, solid-state drives, Blu-ray memories, and flash memories, etc., but they all have problems such as limited storage density, short storage life, and environmental pollution. However, with the progress of gene synthesis and assembly technology, research on using DNA molecules as storage media has attracted increasing attention in the academic community. As a molecule encoding biological information, DNA has the advantages of a long storage period, high storage density, and low energy consumption, and is considered a promising storage medium. DNA communication is different from binary communication. DNA data storage includes four major processes: encoding, synthesis, nanopore sequencing, and decoding. Among them, errors in the form of insertions, deletions, and substitutions may be introduced during the DNA synthesis and sequencing processes, resulting in a poor storage effect.

[0003] The prior art discloses an error correction method and system for DNA storage based on a generative adversarial network. Among them, the method includes: generating a DNA sequence data set based on a DNA template strand with uniform distributions of A, T, C, and G, and dividing it into a training set and a test set; constructing a generative adversarial network model GAN, training the generative adversarial network model GAN based on the training set, and obtaining a trained GAN model after testing through the test set; sequencing the stored DNA molecular sequence, performing clustering and screening according to the sequencing results to obtain DNA clusters; selecting appropriate DNA clusters according to preset rules, performing image conversion on the sequenced sequence information according to preset rules to generate corresponding sequence pictures; obtaining error-corrected picture information through the generator of the GAN model for the generated sequence pictures, and then restoring it to the encoded sequence according to the rules to complete error correction. However, this method still cannot well solve the errors in the form of insertions, deletions, and substitutions that may be introduced during the DNA synthesis and sequencing processes, resulting in a poor storage effect. Summary of the Invention

[0004] The purpose of the present invention is to disclose an effective error correction method and system for multiple outputs of an IDS channel based on DNA data storage with better storage effects.

[0005] To achieve the above purpose, the present invention provides an effective error correction method for multiple outputs of an IDS channel based on DNA data storage, including:

[0006] S1: Performing internal encoding on the input information bit sequence to obtain an internal encoded bit sequence, and then performing external encoding on the internal encoded bit sequence to generate an external encoded bit sequence;

[0007] S2: Map the external encoded bit sequence to obtain a DNA base sequence, and divide the DNA base sequence into M short sequences;

[0008] S3: Pass each of the M short sequences through N parallel IDS channels to obtain M * N short output sequences;

[0009] S4: Process the M * N short output sequences through the SPM algorithm to obtain M short consensus sequences; splice the M short consensus sequences to obtain a consensus sequence;

[0010] S5: Decode the consensus sequence first through synchronous decoding and then through the ENMS algorithm to obtain information; decoding first through synchronous decoding and then through the ENMS algorithm includes: iterating the synchronous decoding and the ENMS algorithm through a joint iterative algorithm.

[0011] Further, in step S1, internally encoding the input information bit sequence to obtain an internally encoded bit sequence includes: using the embedded tag code of (D E , D S ) to internally encode the input information bit sequence to obtain an internally encoded bit sequence c′=(c1, c2, …, c k ), D E represents the interval between two tag bit sequences, and D S represents the length of the tag code.

[0012] Further, in step S1, externally encoding the internally encoded bit sequence to generate an externally encoded bit sequence includes: using a QCLDPC code with a code type of [n, k] as the outer code for external encoding to generate an externally encoded bit sequence u=(u1, u2, …, u n ), where (n - k) is the length of the parity bits.

[0013] Further, in step S2, mapping the external encoded bit sequence to obtain a DNA base sequence includes: mapping the external encoded bit sequence to a DNA base sequence according to the mapping scheme of 00 - A, 01 - C, 10 - T, 11 - G, and the DNA base sequence contains four base symbols of ATCG.

[0014] Further, in step S3, passing one of the M short sequences through N parallel IDS channels to obtain N short output sequences includes:

[0015] One of the short sequences passes through one IDS channel to obtain an output sequence as follows:

[0016] Let w=(w1, w2, …, w nw is a short sequence to be transmitted through the channel, and the short sequence includes multiple base symbols w i , y = (y1, y2, …, y n′ ) is the output sequence; the base symbol w i When entering the channel, the base of the received sequence may have four situations:

[0017] The base symbol w i has an insertion error with a probability of p i , where a uniformly random symbol a ∈ {A, T, C, G} is inserted into the base y i′ of the output sequence, then the base symbol w i remains in the input channel and corresponds to the received symbol y i′+1 ;

[0018] The base symbol w i is deleted with a probability of p d , then no symbol is received, and subsequently the base symbol w i+1 is sent;

[0019] The probability of the base symbol w i being transmitted is p t = 1 - p i - p d , however, there is also a probability of p s of substitution error during the transmission. Therefore, the probability that the input base symbol is successfully transmitted and correctly received is p t (1 - p s ), and at this time the received base symbol y i′ = w i ;

[0020] The base symbol has a substitution error with a probability of p t p s , then a uniformly random symbol a is received, where a ≠ w i ;

[0021] When the channel receives the last symbol w n , the transmission of a short sequence ends, and an output sequence is obtained.

[0022] Furthermore, in step S4, it includes:

[0023] Quickly align any two of multiple output sequences, calculate the scoring matrix using the Levenshtein distance, where a match, a mismatch, and a gap are 1 point, -1 point, and 0 point respectively; thus obtain the total distance score of the alignment of any two sequences to obtain a distance matrix of size M×(M - 1) / 2, where M is the number of output sequences; establish a guide tree and calculate the weight of each branch. Each node in the guide tree corresponds to the order of comparison of multiple sequences. Finally, compare them pairwise one by one according to the sequence alignment order to obtain the base symbol weight of each sequence, deduce the base symbol based on the base weight, and thus obtain the consensus sequence.

[0024] Further, in step S5, the synchronous decoding algorithm includes:

[0025] Introduce an offset variable {O k} as the hidden state variable of the HMM model. When the k-th base symbol in the channel input sequence passes through the channel output, its position is offset. It is offset to the position of (k + d) in the output sequence, and the offset amount is d, denoted as O k = d. Assume the input sequence is w = (w1, w2,..., w t ), the output sequence is y = (y1, y2,..., y q ), t and q represent the lengths of the input sequence and the output sequence respectively. Only when the inserted base symbol matches the deleted base symbol, t and q are equal. Therefore, the offset amounts of the first base symbol w1 and the last base symbol w t are O1 = 0 and O t = q - t, and the output base symbol corresponding to the offset amount of the k-th base symbol is k + O k + O k-1 ;

[0026] During the transmission of the input base symbol w i , the probabilities of the four states of insertion, deletion, substitution, and transmission are represented by p i , p d , p s , and p t respectively; the maximum insertion length is I, and the maximum insertion length defines the offset limit between adjacent position steps. If the offset of the input symbol at position k + 1 (t k+1 ) is d and the offset of the input symbol at position k (t k ) is b, then the probability transition function P db = P(O k+1 = d|O k = b) is

[0027]

[0028] where λ I= 1 / 4 I (1 - (p i )) I ), P db represents two cases. When the base symbol of t i+1 is deleted, b - d + 1 base symbols are inserted between t i and t i+1 , while when the base symbol of t i+1 is not deleted, b - d base symbols are inserted between t i and t i+1 ;

[0029] In the HMM model, the output sequence y = (y1, y2,..., y q ) is the observation sequence in the HMM model. During the sequence transmission process, the probability matrix F(w k , y k+d ) for substitution error transfer between four base symbols. The probability matrix F(w k , y k+d ) is applicable when the base symbols w k-1 and w k have an offset of d, that is, O k-1 = O k = d, indicating no offset between the base symbols w k-1 and w k . The probability function of sending the base symbol from t k and receiving it at t k+d is defined as γ(y k+d ), expressed as:

[0030]

[0031] Among them, the APP of the unlabeled base p(w k ) is initialized as p(A) = 1 / 4, p(T) = 1 / 4, p(C) = 1 / 4, p(G) = 1 / 4. It is known that the labeled base is obtained through mapping the label bits. If w k = A, then p(A) = 1, p(T) = 0, p(C) = 0, p(G) = 0;

[0032] Based on the HMM model, the posterior probability of the observation sequence is derived by calculating the forward - backward algorithm: The forward function α k (d) represents the probability that O k is d and the first k - 1 + d base symbols are received. The initial condition for the forward recursion is α1(0) = 1, and the coefficient recursion relationship of the forward algorithm is calculated as:

[0033]

[0034] Similarly, the initial condition for the backward recursion is βt (q - t) = 1, and the coefficient recurrence relation of the backward algorithm is calculated as follows:

[0035]

[0036] Assume that the input base symbol is at t k The maximum allowed offset is d max , so the range of the offset is d ∈ {-d max , d max}, and by recursively calculating the coefficients of the forward - backward function, the posterior probability of each input base is calculated as follows:

[0037]

[0038] Convert the posterior probability of the base to LLR: The base sequence is mapped back to a binary sequence, and each base symbol corresponds to two binary bits. The first bit is defined as the low bit and the second bit is defined as the high bit. The LLR of the marked bit is set to the value with a larger absolute value. If the marked bit is 1, the corresponding LLR value is set to 100. If the marked bit is 0, the corresponding LLR value is set to - 100. For the unmarked bits, the LLRs of the low bit and the high bit are calculated as follows:

[0039]

[0040] The calculated output LLR sequence is L = (L1, L2, …, L n ).

[0041] Furthermore, in step S5, the ENMS algorithm includes: setting the information bit length as k, the code length as n, and the correction factor as σ; iter max represents the maximum number of iterations of LDPC decoding, D m represents the set of marked bits, L j represents the initial LLR value of the j - th bit, represents the total LLR, is the information passed from the check node i to the variable node j, is defined as the information passed from the variable node j to the i - th check node. C(j) represents the set of all check nodes connected to the j - th bit, V(i) represents the set of all variable nodes connected to the i - th check node, C(j) / i represents C(j) excluding the i - th check node, and V(i) / j represents V(i) excluding the j - th variable node;

[0042] S5.1 Initialization: Initialize the variable nodes according to the input LLR sequence:

[0043]

[0044] S5.2 Check Node Update: The check node does not pass information to the variable node corresponding to the marked bit. If the j-th bit belongs to the marked bits, then Otherwise, if the j-th bit does not belong to the marked bits, the check node is updated as follows:

[0045]

[0046] S5.3 Variable Node Update: The information passed from the variable node corresponding to the marked bit to the check node is not updated. If the j-th bit belongs to the marked bits, the variable node is not updated Otherwise, if the j-th bit does not belong to the marked bits, the check node is updated as follows:

[0047]

[0048] S5.4 Total LLR Calculation and Hard Decision: Calculate the total LLR. If the j-th bit belongs to the marked bits, the total LLR is not updated If the j-th bit does not belong to the marked bits, the total LLR is updated as follows:

[0049]

[0050] Perform a hard decision on the total LLR to obtain the decoded codeword. If Then If Then

[0051] S5.5 Judgment Iteration: Obtain the estimated codeword sequence through hard decision If Or the maximum number of iterations is reached, stop decoding; otherwise, go to step S5.2 to continue decoding.

[0052] Furthermore, in step S5, the joint iterative algorithm iterates the synchronous decoding and the ENMS algorithm, including:

[0053] When the ENMS algorithm does not completely correct the errors, the base symbols decoded by the ENMS algorithm will feedback and update the LLR of the synchronous decoding algorithm. In the iterative decoding, the iterative decoding ends until the errors are successfully corrected or the maximum number of iterations is reached: Let f denote the feedback LLR sequence, where f1 is the lower LLR and f2 is the higher LLR. When the base symbol is unmarked, the prior probabilities P(A), P(C), P(T), and P(G) of the base symbol are updated as follows:

[0054]

[0055] Then return to execute the synchronous decoding algorithm until the number of iterations is reached.

[0056] In addition, the present invention provides an effective error correction system based on multiple outputs of the IDS channel for DNA data storage, including:

[0057] Encoding module: internally encodes the input information bit sequence to obtain an internally encoded bit sequence, and then externally encodes the internally encoded bit sequence to generate an externally encoded bit sequence;

[0058] Mapping module: maps the externally encoded bit sequence to obtain a DNA base sequence, and divides the DNA base sequence into M short sequences;

[0059] Channel module: passes each of the M short sequences through N parallel IDS channels to obtain M*N short output sequences;

[0060] Consensus module: processes the M*N short output sequences through the SPM algorithm to obtain M short consensus sequences; splices the M short consensus sequences to obtain a consensus sequence;

[0061] Decoding module: decodes the consensus sequence first through synchronous decoding and then through the ENMS algorithm to obtain information; decoding first through synchronous decoding and then through the ENMS algorithm includes: iterating the synchronous decoding and the ENMS algorithm through a joint iterative algorithm.

[0062] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0063] The present invention obtains multiple outputs through parallel IDS channels, uses the segmented progressive alignment algorithm to combine the homologous information of multiple outputs to obtain a consensus sequence, which can effectively avoid errors in the DNA synthesis and sequencing processes, and then uses the synchronous decoding and ENMS algorithms for decoding and joint iteration, further improving the error correction performance. Description of the Drawings

[0064] Figure 1 It is a flowchart of an effective error correction method based on multiple outputs of the IDS channel for DNA data storage described in Embodiment 1;

[0065] Figure 2 It is a block diagram of an effective error correction system based on multiple outputs of the IDS channel for DNA data storage described in Embodiment 3;

[0066] Figure 3 It is a simulation result diagram of the bit error rate performance and the substitution error probability coefficient β when the sequencing read sequence outputs M = 3, 5, 7, 10 described in Embodiment 4;

[0067] Figure 4 It is a comparison diagram of the bit error rate performance between the traditional scheme and the proposed scheme under multiple sequence outputs of the channel described in Embodiment 4; Detailed Implementation Manner

[0068] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent;

[0069] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0070] Embodiment 1:

[0071] This embodiment provides an effective error correction method for multiple outputs of the IDS channel based on DNA data storage as shown in Figure 1 , including:

[0072] S1: Internally encode the input information bit sequence to obtain an internally encoded bit sequence, and then externally encode the internally encoded bit sequence to generate an externally encoded bit sequence;

[0073] S2: Map the externally encoded bit sequence to obtain a DNA base sequence, and divide the DNA base sequence into M short sequences;

[0074] S3: Pass each of the M short sequences through N parallel IDS channels to obtain M*N short output sequences;

[0075] S4: Process the M*N short output sequences through the SPM algorithm to obtain M short consensus sequences; splice the M short consensus sequences to obtain a consensus sequence;

[0076] S5: Decode the consensus sequence first through synchronous decoding and then through the ENMS algorithm to obtain information; decoding first through synchronous decoding and then through the ENMS algorithm includes: iterating the synchronous decoding and the ENMS algorithm through a joint iterative algorithm.

[0077] This embodiment obtains multiple outputs through parallel IDS channels, uses the segmented progressive alignment algorithm to combine the homologous information of multiple outputs to obtain a consensus sequence, which can effectively avoid errors in the DNA synthesis and sequencing processes, and then uses the synchronous decoding and ENMS algorithms for decoding and joint iteration, further improving the error correction performance.

[0078] Embodiment 2:

[0079] This embodiment makes a further disclosure on the basis of Embodiment 1:

[0080] Further, in step S1, internally encoding the input information bit sequence to obtain an internally encoded bit sequence includes: using the embedded tag code of (D E , D S ) to internally encode the input information bit sequence to obtain the internally encoded bit sequence c′ = (c1, c2,..., ck ),D E represents the interval between two marker bit sequences, D S represents the length of the marker code.

[0081] Furthermore, in step S1, performing outer coding on the inner coded bit sequence to generate an outer coded bit sequence includes: using a QCLDPC code with code type [n,k] as the outer code to perform outer coding to generate an outer coded bit sequence u = (u1, u2,..., u n ), where (n - k) is the length of the parity bits.

[0082] Furthermore, in step S2, mapping the outer coded bit sequence to obtain a DNA base sequence includes: mapping the outer coded bit sequence to a DNA base sequence according to the mapping scheme of 00 - A, 01 - C, 10 - T, 11 - G, and the DNA base sequence contains four base symbols A, T, C, and G.

[0083] Furthermore, in step S3, obtaining N short output sequences by passing one of the M short sequences through N parallel IDS channels includes:

[0084] One of the short sequences passes through one IDS channel to obtain an output sequence as follows:

[0085] Let w = (w1, w2,..., w n ) be a short sequence to be transmitted through the channel, and the short sequence includes multiple base symbols w i , y = (y1, y2,..., y n′ ) be the output sequence; the base symbol w i When entering the channel, the base of the received sequence may have four situations:

[0086] The base symbol w i has a probability of p i of occurring an insertion error, where a uniformly random symbol a ∈ {A, T, C, G} is inserted into the base y i′ of the output sequence, then the base symbol w i remains in the input channel and corresponds to the received symbol y i′+1 ;

[0087] The base symbol w i has a probability of p d of being sent and deleted, then no symbol is received, and subsequently the base symbol w i+1 is sent;

[0088] The probability of the base symbol w i being transmitted is p t = 1 - p i - pd , however, there is also a probability of substitution error during transmission. Therefore, the probability that the input base symbol is successfully transmitted and correctly received is p s (1 - p t ), and at this time the received base symbol y s = w i′ ; i ;

[0089] The base symbol has a probability of p t p s of substitution error, and then a uniform random symbol a is received, where a ≠ w i ;

[0090] When the channel receives the last symbol w n , the transmission of a short sequence ends, and an output sequence is obtained.

[0091] Furthermore, in step S4, it includes:

[0092] Quickly compare any two sequences among multiple output sequences, calculate the scoring matrix using the Levenshtein distance, where a match, a mismatch, and a gap are 1 point, -1 point, and 0 point respectively; thus obtain the total distance score of the comparison of any two sequences to obtain a distance matrix of size M×(M - 1) / 2, where M is the number of output sequences; establish a guide tree and calculate the weight of each branch. Each node in the guide tree corresponds to the order of comparison of multiple sequences. Finally, compare them pairwise one by one according to the sequence comparison order to obtain the base symbol weight of each sequence, deduce the base symbol based on the base weight, and then obtain the consensus sequence.

[0093] Furthermore, in step S5, the synchronous decoding algorithm includes:

[0094] Introduce an offset variable {O k} as the hidden state variable of the HMM model. When the k-th base symbol in the channel input sequence passes through the channel output, its position has an offset and offsets to the position of (k + d) in the output sequence. The offset amount is d, denoted as O k = d. Assume the input sequence is w = (w1, w2,..., w t ), the output sequence is y = (y1, y2,..., y q ), t and q respectively represent the lengths of the input sequence and the output sequence. Only when the inserted base symbol matches the deleted base symbol, t and q are equal. Therefore, the offset amounts of the first base symbol w1 and the last base symbol w t are O1 = 0 and O t = q - t, and the output base symbol corresponding to the offset amount of the k-th base symbol is k + O k + O k-1 ;

[0095] During the transmission of the input base symbol w i , the probabilities of the four states of insertion, deletion, substitution, and transmission are represented by p i , p d , p s , and p t respectively; the maximum insertion length is I, and the maximum insertion length defines the offset limit between adjacent position steps. If the input symbol offset at position k + 1 (t k+1 ) is d and the input symbol offset at position k (t k ) is b, then the probability transition function P db = P(O k+1 = d|O k = b) is

[0096]

[0097] where λ I = 1 / 4 I (1 - (p i )) I , and P db represents two cases. When the base symbol at t i+1 is deleted, b - d + 1 base symbols are inserted between t i and t i+1 , while when the base symbol at t i+1 is not deleted, b - d base symbols are inserted between t i and t i+1 ;

[0098] In the HMM model, the output sequence y = (y1, y2,..., y q ) is the observed sequence in the HMM model. During the sequence transmission, the probability matrix F(w k , y k+d ) of substitution error transfer occurring between four base symbols, and the probability matrix F(w k , y k+d ) is applicable to the case where the base symbols w k-1 and w k have an offset of d, that is, O k-1 = O k = d, indicating that there is no offset between the base symbols w k-1 and w k . The probability function of sending the base symbol from t k and receiving it at t k+d is defined as γ(y k+d ), which is expressed as:

[0099]

[0100] Among them, the non-marked base p(w k ) of APP is initialized as p(A) = 1 / 4, p(T) = 1 / 4, p(C) = 1 / 4, p(G) = 1 / 4. It is known that the marked base is obtained by mapping the marked bits. If w k = A, then p(A) = 1, p(T) = 0, p(C) = 0, p(G) = 0;

[0101] Based on the HMM model, the posterior probability of the observed sequence is deduced by calculating the forward-backward algorithm: the forward function α k (d) represents the probability that O k is d and the first k - 1 + d base symbols are received. The initial condition for the recursion is α1(0) = 1, and the coefficient recursion relationship of the forward algorithm is calculated as:

[0102]

[0103] Similarly, the initial condition for the backward recursion is β t (q - t) = 1, and the coefficient recursion relationship of the backward algorithm is calculated as:

[0104]

[0105] Assume that the input base symbol is at t k The maximum allowed offset is d max , so the range of the offset is d ∈ {-d max , d max}. Through the recursion of the forward-backward function coefficients, the posterior probability of each input base is calculated as:

[0106]

[0107] The posterior probability of the base is converted into LLR: the base sequence is mapped back to a binary sequence, and each base symbol corresponds to two binary bits. The first bit is defined as the low bit and the second bit is defined as the high bit. The LLR of the marked bit is set to the value with a larger absolute value. If the marked bit is 1, the corresponding LLR value is set to 100. If the marked bit is 0, the corresponding LLR value is set to -100. For the unmarked bits, the LLRs of the low bit and the high bit are calculated as follows:

[0108]

[0109] The calculable output LLR sequence is L = (L1, L2, …, L n ).

[0110] Further, in step S5, the ENMS algorithm includes: setting the information bit length to k, the code length to n, and the correction factor to σ; iter max represents the maximum number of iterations of LDPC decoding, D m represents the set of marked bits, L j represents the initial LLR value of the j-th bit, represents the total LLR, is the information passed from the check node i to the variable node j, is defined as the information passed from the variable node j to the i-th check node. C(j) represents the set of all check nodes connected to the j-th bit, V(i) represents the set of all variable nodes connected to the i-th check node, C(j) / i represents C(j) excluding the i-th check node, and V(i) / j represents V(i) excluding the j-th variable node;

[0111] S5.1 Initialization: Initialize the variable nodes according to the input LLR sequence:

[0112]

[0113] S5.2 Check node update: The check node does not pass information to the variable node corresponding to the marked bit. If the j-th bit belongs to the marked bits, then Otherwise, if the j-th bit does not belong to the marked bits, the check node is updated as:

[0114]

[0115] S5.3 Variable node update: The information passed from the variable node corresponding to the marked bit to the check node is not updated. If the j-th bit belongs to the marked bits, the variable node is not updated Otherwise, if the j-th bit does not belong to the marked bits, the check node is updated as:

[0116]

[0117] S5.4 Total LLR calculation and hard decision: Calculate the total LLR. If the j-th bit belongs to the marked bits, the total LLR is not updated If the j-th bit does not belong to the marked bits, the total LLR is updated as:

[0118]

[0119] Perform a hard decision on the total LLR to obtain the decoded codeword. If then If then

[0120] S5.5 Judgment Iteration: Obtain the estimated codeword sequence through hard decision If Or reach the maximum number of iterations, stop decoding; otherwise, go to step S5.2 to continue decoding.

[0121] Furthermore, in step S5, the joint iteration of the synchronous decoding and the ENMS algorithm by the joint iteration algorithm includes:

[0122] When the ENMS algorithm does not completely correct the error, the base symbols decoded by the ENMS algorithm will feedback and update the LLR of the synchronous decoding algorithm. In the iterative decoding, the iterative decoding ends until the error is successfully corrected or the maximum number of iterations is reached: Let f represent the feedback LLR sequence, where f1 is the lower LLR, f2 is the higher LLR. When the base symbol is unlabeled, the prior probabilities P(A), P(C), P(T), and P(G) of the base symbol are updated as follows:

[0123]

[0124] Then return to execute the synchronous decoding algorithm until the number of iterations is reached.

[0125] In this embodiment, multiple outputs are obtained through parallel IDS channels, and the progressive alignment algorithm is used to combine the homologous information of multiple outputs to obtain a consensus sequence, which can effectively avoid errors in the DNA synthesis and sequencing processes. Then, the synchronous decoding and the ENMS algorithm are used for decoding and joint iteration, further improving the error correction performance.

[0126] Embodiment Three:

[0127] This embodiment provides an effective error correction system based on multiple outputs of the IDS channel for DNA data storage as Figure 2 shown, including:

[0128] Encoding module: Internally encode the input information bit sequence to obtain an internally encoded bit sequence, and then externally encode the internally encoded bit sequence to generate an externally encoded bit sequence;

[0129] Mapping module: Map the externally encoded bit sequence to obtain a DNA base sequence, and divide the DNA base sequence into M short sequences;

[0130] Channel module: Pass each of the M short sequences through N parallel IDS channels to obtain M*N short output sequences;

[0131] Consensus module: Process the M*N short output sequences through the SPM algorithm to obtain M short consensus sequences; splice the M short consensus sequences to obtain a consensus sequence;

[0132] Decoding module: The consensus sequence is first decoded by synchronous decoding and then by the ENMS algorithm to obtain information; first decoding by synchronous decoding and then by the ENMS algorithm includes: iterating the synchronous decoding and the ENMS algorithm through a joint iterative algorithm.

[0133] In this embodiment, multiple outputs are obtained through parallel IDS channels, and the progressive alignment algorithm is used to combine the homologous information of the multiple outputs to obtain a consensus sequence, which can effectively avoid errors in the DNA synthesis and sequencing processes. Then, synchronous decoding and the ENMS algorithm are used for decoding and joint iteration, further improving the error correction performance.

[0134] Embodiment 4:

[0135] In this embodiment, Microsoft Visual Studio is used as the simulation platform, and the bit error rate is used as the performance evaluation criterion. First, during the sequencing process, the segmentation progressive alignment algorithm and the traditional sequence alignment algorithm proposed in this embodiment are used to compare the performance of different sequencing output sequences M. Secondly, to verify the reliability of the decoding algorithm, this embodiment simulates the influence of the decoding scheme proposed in this embodiment and the traditional scheme on the decoding effect under the same conditions.

[0136] The selected IDS channel model for simulation is based on nanopore sequencing. In the concatenated coding scheme with a storage efficiency R of 0.72, a QC LDPC code with a code rate of 0.9 (4096, 4544) is used as the outer code, and an embedded tag code with (D E = 4, D S = 18) is used as the inner code, and the tag symbols are set as AGTC... AGTC. The maximum number of iterations iter max of the ENMS decoding is set to 20, and the segmented sequence length is set to N = 110. The probability of insertion and deletion errors during the synthesis process is 0.0005 - 0.005. Therefore, the maximum offset of the IDS channel is selected to set the probabilities p i = 0.0005 and p d = 0.0065.

[0137] Such as Figure 3The simulation results of the bit error rate performance and the substitution error probability coefficient β are shown for the sequencing read sequence outputs M = 3, 5, 7, and 10. First, when the error transfer probability is β = 0.02, the substitution error probability of bases AG is 0.04, and the substitution error probability of bases CT is 0.12. It can be calculated that the bit error rate of M = 10 is reduced by 64.2% compared to the bit error rate of M = 7, and as the number of alignment times M increases, the error correction performance is better. In addition, as the output sequence M for comparison increases, the decoding complexity increases exponentially. Balancing the complexity and the decoding performance is crucial for practical applications.

[0138] Under the condition that the sequencing read sequence output M is 5 and 7 and the split progressive alignment algorithm is used, the encoding and decoding scheme proposed in this embodiment and the traditional scheme using the watermark code as the inner code use the bit error rate as the performance evaluation criterion, as Figure 4 shown. The insertion and deletion error probabilities remain unchanged, and the substitution error coefficient β is simulated at 0.015 - 0.05. Compared with the traditional scheme, the joint iterative decoding method proposed in this embodiment reduces the bit error rate by 21.72% - 99.75%. When β = 0.04, the joint iterative decoding algorithm at J m = 1 and J m = 3 reduces the bit error rate by approximately 54% and 98% respectively compared with the traditional scheme. It can be seen that the decoding performance of this embodiment for the IDS channel with a relatively large error probability is very good. Therefore, the joint iterative decoding scheme based on the embedded tag code as the inner code in this embodiment combines the soft information of the embedded tag code and iteratively updates to reduce the bit error rate and improve the decoding performance.

[0139] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An effective error correction method based on multiple outputs of an IDS channel for DNA data storage, characterized in that, Including: S1: Internally encode the input information bit sequence to obtain an internally encoded bit sequence, and then externally encode the internally encoded bit sequence to generate an externally encoded bit sequence; S2: Map the externally encoded bit sequence to obtain a DNA base sequence, and divide the DNA base sequence into M short sequences; S3: Pass each of the M short sequences through N parallel IDS channels to obtain M * N short output sequences; S4: Process the M * N short output sequences through the SPM algorithm to obtain M short consensus sequences; splice the M short consensus sequences to obtain a consensus sequence; S5: Decode the consensus sequence first through synchronous decoding and then through the ENMS algorithm to obtain information; Decoding first through synchronous decoding and then through the ENMS algorithm includes: Iterating the synchronous decoding and the ENMS algorithm through a joint iterative algorithm.

2. The effective error correction method for multiple outputs of the IDS channel based on DNA data storage according to claim 1, characterized in that In step S1, internal encoding of the input information bit sequence to obtain an internal encoded bit sequence includes: using the embedded tag code of (D E , D S ) to perform internal encoding on the input information bit sequence to obtain the internal encoded bit sequence c′ = (c1, c2,..., c k ), D E represents the interval between two tag bit sequences, and D S represents the length of the tag code.

3. The effective error correction method for multiple outputs of the IDS channel based on DNA data storage according to claim 1, characterized in that, In step S1, performing external encoding on the internal encoded bit sequence to generate an external encoded bit sequence includes: using a QCLDPC code with a code pattern of [n, k] as the outer code to perform external encoding to generate an external encoded bit sequence u = (u1, u2,..., u n ), where (n - k) is the length of the parity bits.

4. The effective error correction method for multiple outputs of the IDS channel based on DNA data storage according to claim 1, wherein In step S2, mapping the externally encoded bit sequence to obtain a DNA base sequence includes: Mapping the externally encoded bit sequence to a DNA base sequence according to the mapping scheme of 00 - A, 01 - C, 10 - T, 11 - G, and the DNA base sequence contains four base symbols A, T, C, G.

5. The effective error correction method for multiple outputs of an IDS channel based on DNA data storage according to claim 1, wherein In step S3: Passing one of the M short sequences through N parallel IDS channels to obtain N short output sequences includes: One of the short sequences passes through one IDS channel to obtain an output sequence as follows: Let \(w=(w_1, w_2, \ldots, w n )\) be a short sequence to be transmitted through a channel. The short sequence includes multiple base symbols \(w i \), and \(y=(y_1, y_2, \ldots, y n′ )\) be the output sequence; When the base symbol \(w i \) enters the channel, the base of the received sequence may have four situations: Base symbol w i Inserts an error with probability p i where a uniformly random symbol a ∈ {A, T, C, G} is inserted into base y of the output sequence i′ then base symbol w i remains in the input channel and corresponds to the received symbol y i′+1 ; Base symbol w i is sent with probability p d and is deleted, no symbol is received, and then base symbol w is sent i+1 ; Base symbol w i The transmission probability is p t = 1 - p i -p d , however, there is also a probability of p s of substitution error during transmission. Therefore, the probability that the input base symbol is successfully transmitted and correctly received is p t (1 - p s ), and at this time the received base symbol y i′ = w i ; The base symbol is replaced with p t p s with a probability of substitution error, and the unified random symbol a is received, where a ≠ w i ; When the channel receives the last symbol w n the transmission of a short sequence ends, resulting in an output sequence.

6. The effective error correction method based on multiple outputs of the IDS channel for DNA data storage according to claim 1, wherein, In step S4, it includes: Quickly compare any two of the multiple output sequences, calculate the scoring matrix using the Levenshtein distance, where a match, a mismatch, and a gap are 1 point, -1 point, and 0 point respectively; thus obtain the total distance score of the comparison of any two sequences to obtain a distance matrix of size M×(M - 1) / 2, where M is the number of output sequences; establish a guide tree and calculate the weight of each branch, each node in the guide tree corresponds to the order of comparison of multiple sequences, and finally compare them pairwise one by one according to the sequence comparison order to obtain the base symbol weight of each sequence, and deduce the base symbol based on the base weight, and then obtain the consensus sequence.

7. The effective error correction method for multiple outputs of an IDS channel based on DNA data storage according to claim 1, characterized in that, In step S5, the synchronous decoding algorithm includes: Introduce an offset variable {O k} as the hidden state variable of the HMM model. When the k-th base symbol in the channel input sequence passes through the channel output, its position is offset. In the output sequence, it is offset to the position of (k + d), and the offset amount is d, denoted as O k = d. Assume the input sequence is w = (w1, w2,..., w t ), and the output sequence is y = (y1, y2,..., y q ). Let t and q represent the lengths of the input sequence and the output sequence respectively. Only when the inserted base symbol matches the deleted base symbol, t and q are equal. Therefore, the offset amounts of the first base symbol w1 and the last base symbol w t are O1 = 0 and O t = q - t. The output base symbol corresponding to the offset amount of the k-th base symbol is k + O k + O k-1 ; During the transmission of the input base symbol w i , the probabilities of the four states of insertion, deletion, substitution, and transmission are represented by p i , p d , p s , and p t respectively; the maximum insertion length is I, and the maximum insertion length defines the offset limit between adjacent position steps. If the input symbol offset at position k + 1 (t k+1 ) is d and the input symbol offset at position k (t k ) is b, then the probability transition function P db = P(O k+1 = d|O k = b) is where λ I = 1 / 4 I (1 - (p i )) I ), P db represents two cases. When the base symbol of t i+1 is deleted, b - d + 1 base symbols are inserted between t i and t i+1 , while when the base symbol of t i+1 is not deleted, b - d base symbols are inserted between t i and t i+1 ; In the HMM model, the output sequence y = (y1, y2, …, y q ) is the observation sequence in the HMM model. During the sequence transmission process, the probability matrix F(w k , y k+d ) of substitution error transfer occurring between the four base symbols, and the probability matrix F(w k , y k+d ) is applicable to the base symbols w k-1 and w k with an offset of d, that is, O k-1 = O k = d, indicating that there is no offset between the base symbols w k-1 and w k . The probability function of sending the base symbol from t k and receiving it at t k+d is defined as γ(y k+d ), which is expressed as: where the non-tagged base p(w k ) of APP is initialized as p(A)=1 / 4, p(T)=1 / 4, p(C)=1 / 4, p(G)=1 / 4. It is known that the tagged base is obtained by mapping the tag bit. If w k =A, then p(A)=1, p(T)=0, p(C)=0, p(G)=0; Based on the HMM model, the posterior probability of the observation sequence is derived by calculating the forward-backward algorithm: the forward function α k (d) represents O k is the probability that d and the first k - 1 + d base symbols are received. The initial condition for the recursion is α1(0) = 1, and the coefficient recursion relation of the forward algorithm is calculated as: Similarly, the initial condition for backward recursion is β t (q - t)=1, and the coefficient recurrence relation of the backward algorithm is calculated as: Assume the input base symbol is at t k The maximum allowed offset is d max , so the range of the offset is d ∈ {-d max , d max} and the posterior probability of each input base is calculated by the recursion of the forward-backward function coefficients as follows: Convert the posterior probability of the base to LLR: The base sequence is mapped back to a binary sequence, and each base symbol corresponds to two binary bits. The first bit is defined as the low bit and the second bit is defined as the high bit. Set the LLR of the marked bit to the value with the larger absolute value. If the marked bit is 1, set the corresponding LLR value to 100. If the marked bit is 0, set the corresponding LLR value to -100. For the unmarked bits, calculate the LLR of the low bit and the high bit as follows: The computable LLR sequence of the output is \(L=(L_1,L_2,\ldots,L\) n ).

8. The effective error correction method for multiple outputs of an IDS channel based on DNA data storage according to claim 1, wherein In step S5, the ENMS algorithm includes: setting the information bit length to k, the code length to n, and the correction factor to σ; iter max represents the maximum number of iterations of LDPC decoding, D m represents the set of marked bits, L j represents the initial LLR value of the j-th bit, represents the total LLR, is the information passed from the check node i to the variable node j, is defined as the information passed from the variable node j to the i-th check node. C(j) represents the set of all check nodes connected to the j-th bit, V(i) represents the set of all variable nodes connected to the i-th check node, C(j) / i represents C(j) excluding the i-th check node, and V(i) / j represents V(i) excluding the j-th variable node; S5.1 Initialization: Initialize the variable nodes according to the input LLR sequence; S5.2 Check Node Update: The check node does not pass information to the variable node corresponding to the marked bit. If the j-th bit belongs to the marked bit, then Otherwise, if the j-th bit does not belong to the marked bit, the check node is updated as follows: S5.3 Variable Node Update: The information passed from the variable node corresponding to the marked bit to the check node is not updated. If the j-th bit belongs to the marked bits, the variable node is not updated. Otherwise, if the j-th bit does not belong to the marked bits, the check node is updated as follows: S5.4 Total LLR calculation and hard decision: Calculate the total LLR. If the j-th bit belongs to the marked bit, the total LLR is not updated. If the j-th bit does not belong to the marked bit, the total LLR is updated as follows: Make a hard decision on the total LLR to obtain the decoded codeword. If then If then S5.5 Decision Iteration: Obtain the estimated codeword sequence through hard decision If Or the maximum number of iterations is reached, stop decoding; otherwise, go to step S5.2 to continue decoding.

9. The effective error correction method for multiple outputs of an IDS channel based on DNA data storage according to claim 1, wherein In step S5, iterating the synchronous decoding and the ENMS algorithm through a joint iterative algorithm includes: When the ENMS algorithm does not completely correct the error, the base symbols decoded by the ENMS algorithm will feedback and update the LLR of the synchronous decoding algorithm. In iterative decoding, the iterative decoding ends until the error is successfully corrected or the maximum number of iterations is reached. Let f denote the feedback LLR sequence, where f1 is the lower LLR and f2 is the higher LLR. When the base symbol is unlabeled, the prior probabilities P(A), P(C), P(T), and P(G) of the base symbol are updated as follows: Then return to execute the synchronous decoding algorithm until the number of iterations is reached.

10. An effective error correction system based on multiple outputs of an IDS channel for DNA data storage, characterized in that, Including: Coding module: internally encodes the input information bit sequence to obtain an internally encoded bit sequence, and then externally encodes the internally encoded bit sequence to generate an externally encoded bit sequence; Mapping module: maps the externally encoded bit sequence to obtain a DNA base sequence, and divides the DNA base sequence into M short sequences; Channel module: passes each of the M short sequences through N parallel IDS channels to obtain M * N short output sequences; Consensus module: processes the M * N short output sequences through the SPM algorithm to obtain M short consensus sequences; Concatenates the M short consensus sequences to obtain a consensus sequence; Decoding module: decodes the consensus sequence first through synchronous decoding and then through the ENMS algorithm to obtain information; Decoding first through synchronous decoding and then through the ENMS algorithm includes: Iterating the synchronous decoding and the ENMS algorithm through the joint iterative algorithm.

Citation Information

Patent Citations

  • Coding and decoding method for integrity check and error correction of DNA sequence

    CN112802549A

  • DNA data storage decoding method and device

    CN116192160A