A DNA storage method for rapid readout of short fragments by nanopore sequencing
By adopting a probability-based multi-sequence merging method and coding strategy in nanopore sequencing, the high error rate problem when reading short fragment DNA sequences is solved, and fast and error-free data recovery is achieved.
Patent Information
- Application Number
- CN202411284902.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The prior art is difficult to efficiently and quickly read out short fragment DNA sequences obtained by nanopore sequencing, especially when dealing with high insertion abridge errors.
The probability-based multi-sequence merging method is adopted, and the data file is encoded into DNA sequences through the encoding strategy of grouping error correction code and watermark code. The read segment clustering is achieved using double-ended primer recognition and highly reliable label recognition, and the symbol probability is calculated through the soft judgment forward-backward algorithm to achieve correction of insertion and abridge errors.
It realizes fast and error-free recovery of DNA data, reduces the error rate of nanopore sequencing reads, and improves the practicality of data storage.
Smart Images

Figure CN119274664B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNA storage, and particularly to a DNA storage method for rapidly reading out short fragments by nanopore sequencing. Background Art
[0002] With the continuous deepening of the digital, networked, and intelligent trends in information technology, the data volume has shown an explosive growth. Traditional media face many challenges such as cost and energy consumption and cannot meet the rapidly growing data storage needs. DNA data storage, with its ultra-high physical storage density and long-term stability of the medium, has become a very promising future storage medium. With the rapid development of DNA high-throughput synthesis technology and sequencing technology, it has become technically feasible to store digital information using DNA sequences and recover the original digital information from sequencing reads. In 2021, the Semiconductor Industry Association of the United States released the "Semiconductor 10-Year Plan", and DNA storage became one of the four main media for future large-scale data storage. DNA is a biological macromolecule with good stability. Compared with the existing storage media based on magnetism, light, and electricity, DNA as a data storage medium has the characteristics of small volume, high density, and long available time of the medium. In terms of density, Columbia University published a paper in Science confirming that the storage density of DNA can reach 215 PB / g; Tianjin University used a lower-complexity product coding to synthesize more than 270,000 short-chain DNAs, achieving a DNA data storage density of 125 PB / g (Science China Life Sciences, 2020). In terms of storage stability, Grass et al. confirmed that if the DNA encoding information is encapsulated in materials such as silica, it can be preserved for thousands of years.
[0003] DNA data storage is to convert binary digital information into DNA sequences using modern information encoding technology, write data into physical DNA through high-throughput DNA synthesis, and achieve parallel reading of a large number of DNA molecules through high-throughput DNA sequencing technology. The storage form of its medium is an oligonucleotide pool, which can be stored in the form of dry powder, solution, or other enhanced encapsulation. When reading data, generally the second-generation high-throughput sequencing such as the Illumina and Ion Torrent platforms is used, both of which are based on the sequencing-by-synthesis technology, and the sequencing process is time-consuming. The third-generation nanopore sequencing can achieve real-time single-molecule sequencing, and the R10 chip can read at a rate of 400 bases per second, which can rapidly read out short fragment DNA sequences. In recent years, with the rapid development of solid-state nanopore technology, due to the continuous increase in the throughput of nanopore sequencing, it has become a very feasible solution to use nanopores to rapidly read out a large-scale DNA molecule pool.
[0004] However, nanopore sequencing has a relatively high initial error rate, including intractable base insertion and deletion errors. Yazdi et al. adopted an encoding scheme of constrained coding and homopolymer check code for long DNA sequences, achieving error-free data recovery at a sequencing coverage of 200×. Tianjin University proposed a watermark code constructed by superimposing LDPC codes on pseudo-random sequences. This method calculates the offset path using the forward-backward algorithm for a yeast artificial chromosome with a length of 254 kb, identifies and corrects base insertion and deletion errors, and can recover data at a sequencing coverage of 16.8×. This method is based on the methods of watermarking and sparse coding, can correct a large number of insertion and deletion methods, and is a systematic method. The researchers realized the rapid nanopore sequencing to read out the data stored in a yeast artificial chromosome with a length of 254 kb. Press et al. developed a concatenated code encoding scheme, where the inner code is a hash code (HEDGES) decoded by greedy search, which can also be used to correct insertion and deletion errors, but this method has not been tested in the reagent environment of third-generation nanopore sequencing.
[0005] Regarding nanopore sequencing, the researchers proposed a rolling circle concatemer consensus (R2C2) method, which can generate a highly accurate nanopore sequencing consensus sequence. The researchers at Microsoft used polymerase chain reaction and Gibson assembly techniques to assemble short DNA sequences into large DNA fragments. This method achieved data recovery at a sequencing coverage of 22×, but the complex sample preparation reduced the practicality of DNA data storage. However, there is no effective method for directly sequencing and reading short DNA sequences using nanopore sequencing. Summary of the Invention
[0006] The present invention provides a method for rapid nanopore sequencing to read out short DNA storage. The present invention performs probability-based multiple sequence merging on the sequencing read copies obtained by nanopore sequencing, achieving rapid and error-free data recovery, as described in detail below:
[0007] A method for rapid nanopore sequencing to read out short DNA storage, the method comprising the following steps:
[0008] (1) Encode the data to be stored using a block error-correcting code C1(n1,k1), where k1 is the length of the data to be stored and n1 is the length of the block error-correcting codeword. After interleaving the codewords, segment them. Sparsify the sequences of some segmented codewords and then superimpose pseudo-noise sequences. Combine other segmented codewords of equal length and map them to a data DNA sequence according to a preset mapping rule. The labeled part uses a short block error-correcting code C2(n2,k2) accompanied by a pseudo-random sequence, where k2 is the bit length before encoding the labeled part and n2 is the length of the short block error-correcting codeword. The labeling range is To distinguish data blocks, primer sequences are added to both ends of the encoded DNA sequence;
[0009] (2) The synthesized oligonucleotide pool is subjected to polymerase chain reaction to achieve sample amplification, and the DNA sample is subjected to library construction and third-generation sequencing;
[0010] (3) Identify the dual-end primers of the sequencing reads, intercept the valid data part of the reads, use the watermark sequence w2 in the label part to cyclically shift to identify and correct the insertion and deletion errors in the label sequence. The label identification is used for read clustering. Each read within the cluster is demapped into a double-layer bit sequence, where the upper-layer bit sequence set is R and the lower-layer bit sequence set is S. Use the data part watermark sequence w1 to deduce the coding codeword probability V of the upper-layer bit sequence within the cluster, and achieve multi-sequence merging based on the method of multi-probability merging, correct the insertion and deletion errors, obtain the consensus upper-layer sequence, use it to correct the lower-layer bit sequence within the cluster, obtain the consensus lower-layer sequence, and send the corrected codeword into the decoder to perform soft decision error correction decoding to restore the original data.
[0011] Among them, the specific steps of step (1) are as follows:
[0012] (1.1) Encode the data to be stored with length k1 using the block error-correcting code C1(n1,k1) to generate a codeword sequence with length n1. After interleaving the codeword sequence with length n1, it is divided into p segments with equal length;
[0013] (1.2) The segmented codewords are divided into upper and lower layer bit sequences. The length of the upper layer bit sequence is 4m, and the length of the lower layer bit sequence is 5m. The upper layer bit sequence is sparsely processed in the way of converting 4 bits to 5 bits, and then an equal-length watermark sequence is superimposed. n1 is an integer multiple of 9m;
[0014] (1.3) Map the upper layer bit sequence and the lower layer bit sequence to the encoded DNA sequence according to the mapping rule: 00→A, 01→T, 10→G, 11→C;
[0015] (1.4) Encode the label with a range of using the short block error-correcting code C2(n2,k2). k2 is the bit length of the label part before encoding, and n2 is the codeword length of the short block error-correcting code. The encoded codeword sequence and the watermark sequence are mapped according to the same mapping rule as in step (1.3) to obtain the label DNA sequence with strong error-correcting ability;
[0016] (1.5) Add the label DNA sequence to the left end of the data DNA sequence, and add primer sequences for amplification at both ends.
[0017] Among them, the specific steps of step (3) are as follows:
[0018] (3.1) Screen the sequencing reads according to the edit distance between the read segments and the dual primers, intercept the label part and the data part located between the two primers, demap the label part to obtain the part corresponding to the known pseudo-random sequence, use circular shift and dynamic programming to identify and correct the insertion and deletion errors in the label part, and then perform error correction decoding, and cluster the sequencing reads;
[0019] (3.2) Demap the c read copies in the cluster into a two-layer bit sequence, where the upper-layer bit sequence set R = { r 1 , r 2 , …, r c}, the lower-layer bit sequence set S = { s 1 , s 2 , …, s c}, sequentially take r i from the set R as the observation vector, combine with the watermark sequence w1, perform the forward-backward algorithm based on the hidden Markov model, estimate the forward metric and backward metric of each state, as well as the intermediate metric containing symbol soft information, and output the symbol probability results V = {v1, v2, …, v c} with consistent lengths of c read segments in the cluster. Further use the multi-sequence merging strategy based on probability merging to obtain the consensus symbol probability v c . According to the maximum possible probability at each symbol position, infer the upper-layer symbols to obtain the consensus upper-layer bit sequence;
[0020] (3.3) Compare the error-corrected consensus upper-layer bit sequence with the upper-layer bit sequence set obtained from the c read segments in the cluster, identify the error positions, sequentially correct the insertion and deletion errors in the lower-layer bit sequence set, perform majority voting on the corrected lower-layer bit sequences in the cluster to obtain the consensus lower-layer bit sequence, splice it with the error-corrected upper-layer bit sequence, generate the probability information of soft decision decoding, and perform error correction decoding to restore the original data.
[0021] Among them, the specific steps of the step (3.1) are:
[0022] (3.1.1) Compare the sequencing reads with the designed dual primers, screen the sequencing reads, and determine the boundary positions of the dual primers in the sequencing reads;
[0023] (3.1.2) Intercept the label part, use correlation detection to identify the watermark sequence window, and demap the label sequence corresponding to the window to obtain the damaged watermark sequence r 0 and the damaged label bit sequence s 0 ;
[0024] (3.1.3) Circularly shift the damaged watermark sequence r 0 and the error-free watermark sequence w2 in turn, and use dynamic programming to identify the insertion and deletion positions for corresponding correction of the damaged labeled bit sequence s 0 . The corrected result is subjected to error correction decoding, and the reads are clustered accordingly.
[0025] Among them, the specific steps of the step (3.2) are as follows:
[0026] (3.2.1) According to the error characteristics of nanopore sequencing, estimate the insertion error probability P i of the sequencing read, the deletion error probability P d , and the substitution error probability P s to construct an error transmission model. The upper-layer read sequence r i , i ∈ [1, c], is used as the observation vector, and the base offset is used as the hidden state. Perform the forward-backward algorithm based on the hidden Markov model to estimate the forward metric and backward metric of each state, as well as the intermediate metric containing symbol soft information;
[0027] (3.2.2) For a DNA sequence of length N, the number of symbols is N / 5. Use the soft decision forward-backward algorithm to calculate the symbol probability distribution v = {p1, p2,..., p N / 5} of each read within the cluster, where p j is the probability distribution of the jth symbol, p j = (p j,0 , p j,1 ..., p j,k ),
[0028] (3.2.3) According to the probability information V = {v1, v2,..., v c} calculated within each cluster, multiply the corresponding positions in turn and normalize to output the consensus symbol probability v c = {p1′, p2′,..., p N / 5 ′} of each merged cluster, where p j ′ is the probability distribution of the jth symbol, p j ′ = (p j,0 ′, p j,1 ′..., p j,k ′),
[0029] (3.2.4) Based on the maximum possible probability max(p j,0 ′, p j,1 ′,..., p j,15 ′) at each symbol position, infer the upper-layer symbol sequence to obtain the consensus upper-layer bit sequence.
[0030] Among them, the specific steps of step (3.3) are as follows:
[0031] (3.3.1) Compare the corrected consensus upper-layer bit sequence with the original error-carrying upper-layer bit sequence obtained by demapping the c sequencing reads within the cluster, and identify the positions where insertion and deletion errors occur in each read.
[0032] (3.3.2) According to the identified insertion and deletion error positions, assist in correcting the damaged codeword sequences of the corresponding lower layer, perform majority voting on all the corrected lower-layer bit sequences within the cluster, and calculate the consensus lower-layer bit sequence.
[0033] (3.3.3) Combine the consensus upper-layer bit sequence and the consensus lower-layer bit sequence, and calculate the probability soft information for soft decision decoding.
[0034] (3.3.4) Send the probability soft information into the decoder, execute the logarithmic domain belief propagation decoding algorithm to correct the remaining substitution errors, obtain the decoded codeword, and restore the original data.
[0035] The beneficial effects of the technical solution provided by the present invention are as follows:
[0036] 1. The present invention uses an encoding strategy with high encoding efficiency for the data part and strong error correction ability for the label part to encode the data file into a DNA sequence; synthesizes an oligonucleotide pool as a data storage medium, and uses nanopore sequencing to quickly obtain sequencing reads; uses dual-end primer recognition and highly reliable label recognition to achieve rapid clustering of sequencing reads.
[0037] 2. The present invention aims at the problem of high insertion and deletion errors, and uses the symbol probability calculated by the soft decision forward-backward algorithm to achieve probability-based multi-sequence merging, enabling the error correction algorithm to play a more accurate role.
[0038] 3. The present invention provides a solution for realizing rapid reading using nanopore sequencing technology for oligonucleotide pool storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a block diagram of a DNA storage method for rapid reading of short fragments by nanopore sequencing provided by the present invention;
[0040] Figure 2 It is a design flow chart of a data DNA sequence with high encoding efficiency of the present invention;
[0041] Figure 3 It is a design flow chart of a label DNA sequence with strong error correction ability of the present invention;
[0042] Figure 4 It is a data reading flow chart of the present invention;
[0043] Figure 5 This is the flow chart for the present invention to perform dual - primer recognition and highly reliable label recognition;
[0044] Figure 6 This is the flow chart for multi - sequence merging based on probability merging in the present invention;
[0045] Figure 7 This is the flow chart for assisting lower - layer bit error correction in the present invention;
[0046] Figure 8 This is the length distribution diagram of the effective data part after intercepting the sequencing reads provided by the present invention;
[0047] Figure 9 This is the distribution diagram of the number of intra - cluster copies of the reads after label recognition and read clustering provided by the present invention;
[0048] Figure 10 This is the initial error distribution diagram of the sequencing reads provided by the present invention;
[0049] Figure 11 This is the bit error rate distribution diagram after multi - sequence merging based on probability merging provided by the present invention. Detailed implementation manners
[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0051] In view of the above problems, the embodiments of the present invention propose an encoding strategy for a watermark code constructed by superimposing single - layer pseudo - random sequences of error - correcting codes. The encoded DNA sequence is synthesized into short oligonucleotides as a storage medium. After linking library construction and nanopore signal recognition, sequencing reads with many base insertion and deletion errors are obtained. When reading data, combining the multi - copy characteristics of the encoded DNA sequence, a strategy for deriving the probability of the encoded codeword using watermark is proposed to implement an error - correction strategy for multi - sequence merging in a multi - probability merging manner. Compared with directly using the MSA algorithm, it has the advantages of better matching the encoding algorithm and stronger error - correction ability, and supports the short - fragment DNA data storage mode read out quickly by nanopore sequencing.
[0052] The embodiments of the present invention propose an encoding and recovery method for storing information using short - fragment DNA. Using an encoding strategy with high encoding efficiency for the data part and strong error - correction ability for the label part, the data file is encoded into a data DNA sequence, and an oligonucleotide pool is synthesized as a data storage medium. Nanopore sequencing is used to quickly obtain sequencing reads; the reads are clustered through dual - primer recognition and highly reliable label recognition, and the symbol probability calculated by the soft - decision forward - backward algorithm is used to implement probability - based multi - sequence merging, correct the insertion and deletion errors of the nanopore sequencing reads, and at the same time the symbol likelihood information also improves the error - correction decoding performance to achieve error - free recovery.
[0053] The following will make a detailed description of the implementation manner of the embodiment of the present invention in conjunction with the accompanying drawings.
[0054] This embodiment will make a detailed description of a short - fragment DNA data storage method for rapid read - out of nanopore sequencing proposed for the embodiment of the present invention. Refer to Figure 1 , and specifically includes the following steps:
[0055] (1) The data to be stored is encoded using a block error - correcting code C1(n1,k1), where k1 is the length of the data to be stored, n1 is the length of the block error - correcting codeword. After the codewords are interleaved, they are divided into p segments of equal length. After the segmented codeword sequences are superimposed with a pseudo - noise sequence, they are mapped into a data DNA sequence according to a specific mapping rule. The labeling part uses a short block error - correcting code C2(n2,k2) accompanied by a pseudo - random sequence. k2 is the bit length before encoding of the labeling part, n2 is the length of the short block error - correcting codeword, and the labeling range is To distinguish data blocks, primer sequences are added to both ends of the encoded DNA sequence.
[0056] (2) The synthesized oligonucleotide pool undergoes multiple rounds of PCR thermal cycling reactions to amplify the target DNA fragments. Further, a library is constructed for the DNA fragments and third - generation sequencing is performed.
[0057] (3) Identify the dual - end primers of the sequencing reads, intercept the valid data part of the reads, use the watermark sequence w2 of the labeling part to cyclically shift to identify and correct the insertion and deletion errors in the labeling sequence, achieve highly reliable labeling recognition for read clustering. Each read within the cluster is demapped into a two - layer bit sequence, where the upper - layer bit sequence set is R and the lower - layer bit sequence set is S. Use the watermark sequence w1 of the data part to deduce the coding codeword probability V of the upper - layer bit sequence within the cluster, and achieve multi - sequence merging based on the method of multi - probability merging, correct the insertion and deletion errors, obtain a consistent upper - layer sequence, use it to correct the lower - layer bit sequence within the cluster, obtain a consistent lower - layer sequence, and send the corrected codewords to the decoder to perform soft - decision error - correcting decoding to restore the original data.
[0058] Among them, as Figure 2 and Figure 3 shown, the data to be stored is encoded using a block error - correcting code C1(n1,k1), where k1 is the length of the data to be stored, n1 is the length of the block error - correcting codeword. After the codewords are interleaved, they are divided into p segments of equal length. After the segmented codeword sequences are superimposed with a pseudo - noise sequence, they are mapped into a data DNA sequence according to a specific mapping rule. The labeling part uses a short block error - correcting code C2(n2,k2) accompanied by a pseudo - random sequence. k2 is the bit length before encoding of the labeling part, n2 is the length of the short block error - correcting codeword, and the labeling range is To distinguish data blocks, primer sequences are added to both ends of the encoded DNA sequence. The specific steps are as follows:
[0059] (1.1) Use the block error-correcting code C1(n1,k1) to encode the data to be stored with length k1, generating a codeword sequence with length n1. After interleaving the codeword sequence with length n1, it is divided into p segments with equal length;
[0060] (1.2) The segmented codewords are divided into upper and lower layer bit sequences according to the proportionality coefficient m. The length of the upper layer bit sequence is 4m, and the length of the lower layer bit sequence is 5m. The upper layer bit sequence is sparsified in the way of converting 4 bits to 5 bits, and then an equal-length watermark sequence is superimposed;
[0061] (1.3) According to the mapping rule: 00→A, 01→T, 10→G, 11→C, map the upper layer bit sequence and the lower layer bit sequence to the encoded DNA sequence;
[0062] (1.4) Use the short block error-correcting code C2(n2,k2) to encode the label in the range of , where k2 is the bit length of the label part before encoding, and n2 is the codeword length of the short block error-correcting code. The encoded codeword sequence and the watermark sequence are mapped according to the same mapping rule as in step (1.3) to obtain a label DNA sequence with strong error-correcting ability;
[0063] (1.5) Add the label DNA sequence to the left end of the data DNA sequence, and add primer sequences for amplification to both ends.
[0064] Among them, as Figure 4 shown, identify the double-end primers of the sequencing reads, intercept the valid data part of the reads, use the watermark sequence w2 of the label part to cyclically shift to identify and correct the insertion and deletion errors in the label sequence, realize highly reliable label recognition for read clustering. Each read in the cluster is demapped into a double-layer bit sequence, where the set of upper layer bit sequences is R, and the set of lower layer bit sequences is S. Use the data part watermark sequence w1 to deduce the coding codeword probability V of the upper layer bit sequence in the cluster, and realize multi-sequence merging based on the multi-probability merging method, correct the insertion and deletion errors, obtain a consistent upper layer sequence, use it to correct the lower layer bit sequence in the cluster, obtain a consistent lower layer sequence, and send the corrected codewords into the decoder to perform soft decision error-correcting decoding to restore the original data. The specific steps are as follows:
[0065] (3.1) According to the edit distance between the reads and the double-end primers, screen the sequencing reads, intercept the label part and the data part located between the two primers. Demap the label part to obtain the part corresponding to the known pseudo-random sequence. After identifying and correcting the insertion and deletion errors in the label part by cyclic shift and dynamic programming, perform error-correcting decoding to cluster the sequencing reads;
[0066] (3.2) Demap the c read segment copies within the cluster into a two-layer bit sequence, where the upper-layer bit sequence set R = {r 1 , r 2 , …, r c}, the lower-layer bit sequence set S = {s 1 , s 2 , …, s c}. Take r i from the set R in sequence as the observation vector, combine it with the watermark sequence w1, execute the forward-backward algorithm based on the hidden Markov model, estimate the forward metric and backward metric of each state, as well as the intermediate metric containing symbol soft information, and output the symbol probability results V = {v1, v2, …, v c} with consistent lengths for the c read segments within the cluster. Further, use the multi-sequence merging strategy based on probability merging to obtain the consensus symbol probability v c . According to the maximum possible probability at each symbol position, infer the upper-layer symbols to obtain the consensus upper-layer bit sequence;
[0067] (3.3) Compare the error-corrected consensus upper-layer bit sequence with the upper-layer bit sequence set obtained from the c read segments within the cluster, identify the error positions, sequentially correct the insertion and deletion errors of the lower-layer bit sequence set, perform majority voting on the corrected lower-layer bit sequences within the cluster to obtain the consensus lower-layer bit sequence, splice it with the error-corrected upper-layer bit sequence, generate the probability information of soft-decision decoding, and perform error-correction decoding to restore the original data.
[0068] Among them, as Figure 5 shown, according to the edit distance between the read segment and the dual primers, screen the sequencing read segments, intercept the label part and data part located between the two primers, demap the label part to obtain the part corresponding to the known pseudo-random sequence, use cyclic shift and dynamic programming to identify and correct the insertion and deletion errors of the label part, and then perform error-correction decoding. Cluster the sequencing read segments. The specific steps are as follows:
[0069] (3.1.1) Compare the sequencing read segments with the designed dual primers, screen the sequencing read segments, and determine the boundary positions of the dual primers in the sequencing read segments;
[0070] (3.1.2) Intercept the label part, use correlation detection to identify the watermark sequence window, and demap the label sequence corresponding to the window to obtain the damaged watermark sequence r 0 and the damaged label bit sequence s 0 ;
[0071] (3.1.3) Cyclically shift the damaged watermark sequence r 0 and the error-free watermark sequence w2 in sequence, use dynamic programming to identify the insertion and deletion positions, and use them to correspondingly correct the damaged label bit sequence s0 , the corrected result is subjected to error correction decoding, and based on this, the reads are clustered.
[0072] Among them, as Figure 6 shown, the c read segment copies within the cluster are demapped into a two-layer bit sequence, where the upper-layer bit sequence set R = {r 1 , r 2 , …, r c}, and the lower-layer bit sequence set S = {s 1 , s 2 , …, s c}. Take r i from the set R in sequence as the observation vector, combine with the watermark sequence w1, execute the forward-backward algorithm based on the hidden Markov model, estimate the forward metric and backward metric of each state, as well as the intermediate metric containing symbol soft information, and output the symbol probability results V = {v1, v2, …, v c} with consistent lengths for the c read segments within the cluster. Further use the multi-sequence merging strategy based on probability merging to obtain the consensus symbol probability v c . According to the maximum possible probability at each symbol position, infer the upper-layer symbols to obtain the consensus upper-layer bit sequence. The specific steps are as follows:
[0073] (3.2.1) According to the error characteristics of nanopore sequencing, estimate the insertion error probability P i , deletion error probability P d , and substitution error probability P s to construct an error transmission model. The upper-layer read sequence r i , i ∈ [1, c], is used as the observation vector, and the base offset is used as the hidden state. Execute the forward-backward algorithm based on the hidden Markov model to estimate the forward metric and backward metric of each state, as well as the intermediate metric containing symbol soft information;
[0074] (3.2.2) For a DNA sequence of length N, the number of symbols is N / 5. Use the soft-decision forward-backward algorithm to calculate the symbol probability distribution v = {p1, p2, …, p N / 5} of each read segment within the cluster, where p j is the probability distribution of the jth symbol, and p j = (p j,0 , p j,1 ..., p j,k ),
[0075] (3.2.3) According to the probability information V = {v1, v2, …, v c} calculated within each cluster, multiply the corresponding positions in sequence and normalize to output the consensus symbol probability v c={p1′, p2′, …, p N / 5 ′}, where p j ′ is the probability distribution of the j-th symbol, p j ′ = (p j,0 ′, p j,1 ′..., p j,k ′).
[0076] (3.2.4) Based on the maximum possible probability max(p j,0 ′, p j,1 ′,..., p j,15 ′) at each symbol position, infer the upper-layer symbol sequence to obtain the consensus upper-layer bit sequence.
[0077] Among them, as Figure 7 shown, compare the error-corrected consensus upper-layer bit sequence with the set of upper-layer bit sequences obtained from c reads within the cluster to identify the error positions, successively correct the insertion and deletion errors in the lower-layer bit sequence set, perform majority voting on the corrected lower-layer bit sequences within the cluster to obtain the consensus lower-layer bit sequence, splice it with the error-corrected upper-layer bit sequence to generate the probability information for soft decision decoding, perform error-correction decoding, and recover the original data. The specific steps are as follows:
[0078] (3.3.1) Compare the error-corrected consensus upper-layer bit sequence with the original error-carrying upper-layer bit sequences unmapped from c sequencing reads within the cluster to identify the positions of insertion and deletion errors in each read;
[0079] (3.3.2) According to the identified positions of insertion and deletion errors, assist in correcting the damaged codeword sequences of the corresponding lower layer, perform majority voting on all the corrected lower-layer bit sequences within the cluster, and calculate to obtain the consensus lower-layer bit sequence;
[0080] (3.3.3) Combine the consensus upper-layer bit sequence and the consensus lower-layer bit sequence, and calculate the probability soft information for soft decision decoding;
[0081] (3.3.4) Send the probability soft information into the decoder, perform the logarithmic domain belief propagation decoding algorithm to correct the remaining substitution errors, obtain the decoded codeword, and recover the original data. Specific embodiments
[0083] The following gives specific embodiments to illustrate the feasibility of a method for fast readout of short fragment DNA data storage in nanopore sequencing proposed by the present invention.
[0084] In the embodiment of the present invention, a coding method of concatenating LDPC codes and watermark codes is used to encode six segments of audio with sizes of 1,320,907 bytes, 1,320,646 bytes, 1,321,986 bytes, 1,320,237 bytes, 1,320,010 bytes, and 1,320,580 bytes into DNA sequences and synthesize them into oligonucleotide pools respectively. The oligonucleotide pools are named pool1, pool2, pool3, pool4, pool5, and pool6 respectively. The code length of the LDPC code is 64,800 bits, the information bit length is 54,000 bits, the code rate is 5 / 6, the code rate of the watermark code is 4 / 5, and the overall efficiency is 4 / 3 bits represented by each base.
[0085] The coding method of concatenating LDPC codes and watermark codes is specifically as follows:
[0086] (1.1) Encode the 54,000-bit information sequence with LDPC codes. After encoding, the length of the codeword sequence is 64,800 bits. Each segment of audio is encoded into 196 codewords. Each codeword is segmented into fixed-length 423-bit segments after interleaving, and each codeword is divided into 153 segments;
[0087] (1.2) Take the first 188 bits of each segment of the bit sequence as the upper-layer bit sequence and the last 235 bits as the lower-layer bit sequence. After the upper-layer bit sequence is sparsified from 4 bits to 5 bits, it is superimposed with the watermark sequence to obtain the upper-layer bit sequence of the superimposed watermark of 235 bits. Then, one bit is selected from the upper-layer bit sequence and the lower-layer bit sequence in turn to form a bit pair;
[0088] (1.3) Map the bit pairs to bases according to the mapping rule of 00→A, 01→T, 10→G, 11→C to obtain a data DNA sequence with a length of 235 bases. The process of superimposing the watermark on the first segment of 423 bits of the first codeword and mapping it to a base sequence is shown in Table 1;
[0089] (1.4) Adopt the coding strategy of BCH code (31, 21) to encode the binary label sequence in the range of 1-2 15 The encoded codeword is combined with the watermark sequence of the same length, and after mapping, a labeled DNA sequence is obtained and added to the left end of the data DNA sequence. Primer sequences with a length of 20 bases are added to the left end of the labeled DNA sequence and the right end of the data DNA sequence respectively, and finally an encoded DNA sequence with a length of 300 bases is obtained.
[0090] Table 1 The process of superimposing the watermark on the codeword sequence and mapping it to a base sequence
[0091]
[0092]
[0093] In the embodiment of the present invention, a DNA sequence is synthesized into an oligonucleotide pool as a storage medium. After library construction and third-generation nanopore sequencing, nanopore sequencing reads are obtained. The method for recovering data from the nanopore sequencing reads in the embodiment of the present invention is specifically as follows:
[0094] (2.1) Determine the start and end positions of the payload sequence by comparing the edit distance between the duplex and the sequencing reads, and intercept the sequencing reads according to the start and end positions. In the embodiment of the present invention, the length distribution after data interception is as Figure 8 shown. It can be seen that the read length is about 260 bases;
[0095] (2.2) Intercept the label part and demap it to obtain the damaged label bit sequence and the damaged watermark sequence. Use cyclic shift and dynamic programming to identify and correct the insertion and deletion errors in the label part, perform error correction decoding on the label part, and cluster the sequencing reads with the same label according to the recognition results. In the embodiment of the present invention, the label recognition is respectively performed on the sequencing reads corresponding to 6 oligonucleotide pools, and the recognition results are shown in Table 2. Among them, 5x sequencing data is used, and the coverage distribution of the reads within the cluster after label recognition is as Figure 9 shown. It can be seen that the average coverage of the reads within the cluster is 4.5x, and the highest coverage reaches 21x. Among them, two reads corresponding to the encoded DNA sequence in Table 1 after label recognition are shown in Table 3. The length of read 1 is 235 bases, and the length of read 2 is 236 bases;
[0096] Table 2 Success rate of label recognition
[0097] 3x 4x 5x 6x pool1 87.76% 87.76% 87.81% 87.82% pool2 89.13% 89.09% 89.34% 89.22% pool3 88.14% 88.22% 88.16% 88.14% pool4 89.17% 89.29% 89.31% 89.34% pool5 88.88% 88.94% 88.90% 88.90% pool6 90.16% 90.04% 90.2% 90.14%
[0098] Table 3 Two reads within a cluster after label recognition
[0099]
[0100]
[0101] (2.3) Within each cluster, the upper-layer bit sequence set R successively takes r i as the observation vector, combines with the watermark sequence w1, executes the forward-backward algorithm based on the hidden Markov model, estimates the forward metric and backward metric of each state, and the intermediate metric containing symbol soft information. Use the multi-sequence merging strategy based on probability merging to obtain the consensus symbol probability within the cluster, and infer the upper-layer symbol according to the maximum possible probability at each symbol position to obtain the consensus upper-layer bit sequence;
[0102] Figure 10shows the initial error probability distribution of the DNA sequences of 6 oligonucleotide pools. It can be seen that the total base error rates of each oligonucleotide pool are basically equivalent and all lower than 1.4%. Based on the error characteristics of the above sequencing reads, a hidden Markov model is established, and the insertion error probability P i is set to 0.004, the deletion error probability P d is set to 0.005, the substitution error probability P s is set to 0.006, the maximum offset X max is set to 10, and the maximum number of consecutive insertions I = 2. The soft decision forward-backward algorithm is used to estimate the symbol likelihood probability, and the upper-layer symbol probabilities of each cluster are combined to obtain the consensus and normalized symbol probabilities. Based on the maximum possible probability of each symbol position, the upper-layer symbols are inferred to obtain the consensus upper-layer bit sequence. Among them, the symbol likelihood probabilities are calculated for the two reads shown in Table 3 and are used for the multi-sequence merging process based on probability merging as shown in Table 4;
[0103] Table 4 Calculation and merging of symbol probabilities of reads within the same cluster
[0104]
[0105]
[0106]
[0107] (2.4) For the lower-layer bits of the reads, the upper-layer bit sequences before and after error correction are compared to identify the error positions of the reads, and the insertion and deletion errors of the lower-layer bits are corrected at the corresponding positions. The corrected lower-layer bit sequences within the cluster are subjected to majority voting to obtain the consistent lower-layer bit sequence. Among them, the lower-layer bit error correction process for the reads in Table 3 is shown in Table 5. The corrected lower-layer bit sequence is concatenated with the corrected upper-layer bit sequence to generate the probability information of soft decision decoding, and error correction decoding is performed to recover the original data.
[0108] Figure 11 shows the comparison of the upper and lower layer bit error rates of the sequencing reads of 6 oligonucleotide pools in the embodiment of the present invention after error correction. It can be seen that the upper and lower layer bit error rates of the 6 groups of sequencing reads after error correction are basically equivalent and both lower than the initial error rate of 1.5% preset in the embodiment of the present invention, which fully verifies that the probability-based multi-sequence merging method adopted can effectively cope with the problem of high error rates in nanopore sequencing.
[0109] Table 5 Lower-layer bit error correction process
[0110]
[0111]
[0112] In the embodiments of the present invention, the proposed error correction method that uses probability merging to correct upper-layer bit errors and assist in correcting lower-layer bit errors is compared with the multi-sequence merging method. The multi-sequence merging-based method first uses MAFFT for multiple sequence alignment, and then obtains the merged DNA sequence through majority voting decision. The merged DNA sequence further corrects insertion and deletion errors using the forward-backward algorithm and is input into the decoder for decoding.
[0113] The bit error rates after error correction of the two methods are shown in Table 6. It can be seen that for the data of six oligonucleotide pools, at a 10x coverage, the proposed error correction method achieves a lower error rate than the multi-sequence merging-based method. The running times of the two methods for error correction are shown in Table 7. The multi-sequence merging-based method is faster than the proposed error correction method. Each oligonucleotide pool in the embodiments of the present invention corresponds to 196 codewords. The codeword decoding recovery situations after error correction of the two methods are shown in Table 8, which shows the number of correctly decodable codewords in each oligonucleotide pool under different sequencing coverages. It can be seen that the multi-sequence merging-based method depends on high sequencing coverage, and the probability-merging-based error correction method proposed in the embodiments of the present invention has a better error correction effect under low sequencing coverage than the multi-sequence merging-based method.
[0114] The embodiments of the present invention correct the short DNA sequences obtained by sequencing in the literature of the DNA information storage scheme based on nanopore sequencing to verify the feasibility of the present invention.
[0115] Table 6 Comparison of error rates between the present method and the multi-sequence merging-based method at 10x coverage
[0116]
[0117] Table 7 Comparison of running times (unit: seconds) between the present method and the multi-sequence merging-based method at 10x coverage
[0118]
[0119] Table 8 Comparison of the number of successfully decoded codewords between the present method and the multi-sequence merging-based method
[0120]
[0121] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for storing DNA for rapid readout of short fragments by nanopore sequencing, characterized in that: The method comprises the following steps: (1) The data to be stored is encoded using a block error correction code C1(n1,k1), where k1 is the length of the data to be stored, n1 is the length of the block error correction code codeword, and the codewords are interleaved and segmented. The codewords of some segments are sparsely sequenced and superimposed with a pseudo-noise sequence, and other segments of equal length are combined and mapped into a data DNA sequence according to a preset mapping rule. The label part uses a short block error correction code C2(n2,k2) accompanied by a pseudo-random sequence, where k2 is the bit length before the label part is encoded, and n2 is the length of the short block error correction code codeword. The label range is To distinguish data blocks, primer sequences are added at both ends of the coding DNA sequence; (2) The synthesized oligonucleotide pool is subjected to polymerase chain reaction to achieve sample amplification, and the DNA sample is library constructed and sequenced; (3) Identify the double-ended primers of the sequencing read segment, intercept the valid data portion of the read segment, and use the watermark sequence of the labeled portion w 2. Cyclic shift recognition and correction of insertion and deletion errors in the label sequence. Label recognition is used to cluster read segments. Each read segment in the cluster is demapped into a double-layer bit sequence, where the upper layer bit sequence set is R and the lower layer bit sequence set is S. The data part watermark sequence is used w 1. The probability V of the encoded codeword of the upper bit sequence in the cluster is derived, and multiple sequences are merged based on the multi-probability merging method, and the insertion and deletion errors are corrected to obtain a consistent upper sequence, which is used to correct the lower bit sequence in the cluster to obtain a consistent lower sequence. The corrected codeword is sent to the decoder to perform soft decision error correction decoding to restore the original data; The specific steps of step (3) are as follows: (3.1) According to the edit distance between the read and the double-ended primer, the sequencing reads are screened, the label part and the data part between the two end primers are intercepted, the label part is demapped to obtain the part corresponding to the known pseudo-random sequence, and the insertion and deletion errors of the label part are identified and corrected by circular shift and dynamic programming, and then error correction decoding is performed to cluster the sequencing reads; (3.2) Demap the c read segment replicas in the cluster into two-layer bit sequences, where the upper layer bit sequence set R = { r 1 , r 2 ,…, r c }, the lower layer bit sequence set S = { s 1 , s 2 ,…, s c }, take r from the set R in turn i As the observation vector, combined with the watermark sequence w 1. Execute the forward-backward algorithm based on the hidden Markov model to estimate the forward metric and backward metric of each state, as well as the intermediate metric containing symbol soft information, and output the symbol probability result V = {v1, v2, ..., v c }, and further use the multi-sequence merging strategy based on probability merging to obtain the probability v of the consensus symbol within the cluster c , according to the maximum possible probability of each symbol position, the upper layer symbol is inferred to obtain the consensus upper layer bit sequence; (3.3) Compare the corrected consensus upper-layer bit sequence with the set of upper-layer bit sequences obtained from the c read segments in the cluster, identify the error location, correct the insertion and deletion errors of the lower-layer bit sequence set in turn, perform majority voting on the corrected lower-layer bit sequence in the cluster to obtain a consensus lower-layer bit sequence, concatenate it with the corrected upper-layer bit sequence, generate probability information for soft decision decoding, perform error correction decoding, and restore the original data; Wherein, the specific steps of step (3.2) are: (3.2.1) According to the error characteristics of nanopore sequencing, estimate the insertion error probability P of the sequencing read segment i , deletion error probability P d , and the substitution error probability P s Construct the error transmission model, the upper layer reads the sequence r i , i∈[1,c], as the observation vector and the base offset as the hidden state, and perform the forward-backward algorithm based on the hidden Markov model to estimate the forward metric and backward metric of each state, as well as the intermediate metric containing the soft information of the symbol; (3.2.2) For a DNA sequence of length N, the number of symbols is N / 5. Using the soft-decision forward-backward algorithm, the symbol probability distribution v = {p1, p2, …, p N / 5 }, where p j is the probability distribution of the jth symbol, p j =(p j,0 ,p j,1 ...,p j,k ), (3.2.3) According to the probability information V = {v1, v2, …, v c }, the corresponding positions are multiplied and normalized in sequence, and the consensus symbol probability v after each cluster is merged is output c ={p1′,p2′,…,p N / 5 ′}, where p j ′ is the probability distribution of the jth symbol, p j ′=(p j,0 ′,p j,1 ′...,p j,k ′), (3.2.4) According to the maximum possible probability max(p j,0 ′,p j,1 ′,…,p j,15 ′), infer the upper-layer symbol sequence and obtain the consensus upper-layer bit sequence.
2. A method for storing DNA for rapid readout of short fragments by nanopore sequencing according to claim 1, characterized in that: The specific steps of step (1) are as follows: (1.1) The block error correction code C1(n1,k1) is used to encode the data to be stored with a length of k1 to generate a codeword sequence with a length of n1. After the codeword sequence with a length of n1 is interleaved, it is divided into p segments of equal length; (1.2) The segmented codeword is divided into two layers of bit sequences, the upper layer bit sequence length is 4m, and the lower layer bit sequence length is 5m. The upper layer bit sequence is thinned by converting 4 bits to 5 bits, and then a watermark sequence of equal length is superimposed, and n1 is an integer multiple of 9m; (1.3) Map the upper layer bit sequence and the lower layer bit sequence into a coding DNA sequence according to the mapping rules: 00→A, 01→T, 10→G, 11→C; (1.4) Using the short block error correction code C2(n2,k2) for the range The label is encoded, k2 is the bit length of the label part before encoding, n2 is the codeword length of the short packet error correction code, and the encoded codeword sequence and the watermark sequence are mapped according to the same mapping rule as step (1.3) to obtain a label DNA sequence with strong error correction capability; (1.5) Add the labeled DNA sequence to the left end of the data DNA sequence, and add primer sequences for amplification at both ends.
3. A method for storing DNA for rapid readout of short fragments by nanopore sequencing according to claim 1, characterized in that: The specific steps of step (3.1) are: (3.1.1) Align the sequencing reads with the designed double-ended primers, screen the sequencing reads, and determine the boundary positions of the double-ended primers in the sequencing reads; (3.1.2) Cut off the label part, use correlation detection to identify the watermark sequence window, and demap the label sequence corresponding to the window to obtain the damaged watermark sequence r 0 and the damaged label bit sequence s 0 ; (3.1.3) The damaged watermark sequence r 0 and error-free watermark sequence w 2 Circular shift in sequence, using dynamic programming to identify the insertion and deletion positions, which are used to correct the damaged label bit sequence s 0 , the corrected results are subjected to error correction decoding, and the read segments are clustered accordingly.
4. A method for storing DNA for rapid readout of short fragments by nanopore sequencing according to claim 1, characterized in that: The specific steps of step (3.3) are: (3.3.1) Compare the consensus upper bit sequence after error correction with the original upper bit sequence carrying errors demapped from the c sequencing reads in the cluster to identify the location where the insertion and deletion errors occurred in each read; (3.3.2) According to the identified insertion and deletion error positions, assist in correcting the damaged codeword sequence of the corresponding lower layer, perform majority voting on all corrected lower layer bit sequences in the cluster, and calculate the consensus lower layer bit sequence; (3.3.3) combining the consensus upper layer bit sequence and the consensus lower layer bit sequence, and calculating the probabilistic soft information for soft decision decoding; (3.3.4) The probabilistic soft information is sent to the decoder, and the logarithmic domain belief propagation decoding algorithm is executed to correct the remaining substitution errors, obtain the decoded codeword, and restore the original data.