A Fast Reading Method for Watermark-Overlaid Encoded Large-Fragment DNA Data Storage

By using watermark overlay encoding and sliding-related algorithms in large fragment DNA data storage, the problems of high-complexity de novo assembly and high error rate indel error processing in the prior art are solved, and error-free data recovery and cost reduction under low sequencing coverage are achieved.

CN119479733BActive Publication Date: 2025-06-24TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411592389.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-06-24
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

In the rapid readout process of large fragment DNA data storage, the prior art faces high complexity of de novo assembly and high error rate indel error processing, which makes it difficult to achieve error-free data recovery under low sequencing coverage, increasing readout costs.

Method used

Large fragment DNA encoded by watermark superposition is used for second-generation high-throughput sequencing, the starting position of the sequencing read segment is determined through sliding correlation and sequence consensus algorithms, and the majority voting judgment algorithm is used to generate a consistent base sequence, and finally, efficient and reliable data recovery is achieved through error correction decoding.

Benefits of technology

The error-free recovery of large fragment DNA data is achieved with low sequencing coverage, reducing data readout costs and avoiding complex de novo assembly operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479733B_ABST
    Figure CN119479733B_ABST
Patent Text Reader

Abstract

The present invention discloses a rapid reading method for watermark-overlaid encoded large-fragment DNA data storage, belonging to the field of DNA data storage. First, for the large-fragment DNA with watermark-overlaid encoding, after fragmentation, a library is constructed for second-generation high-throughput sequencing to obtain corresponding sequencing data; then, the watermark sequence hidden in the noisy sequencing reads is used to perform sliding correlation calculation with the locally known watermark sequence to obtain the cross-correlation peak value; according to the corresponding position of the cross-correlation peak value, the position of the read in the large-fragment DNA sequence is determined, and then the majority voting decision algorithm is used to generate a consensus sequence, and further, the reliable recovery of the stored data is realized through an efficient error correction and deletion decoding algorithm. The advantages of the present invention are that it supports DNA lengths from Kb to Mb, uses simple sliding correlation and sequence consensus to achieve read positioning, effectively avoids the high-complexity de novo assembly, and can exclude the interference of non-coding DNA reads, and can achieve error-free recovery of the original data at a relatively low sequencing coverage.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the field of DNA data storage, and particularly to a fast reading method for watermark superposition encoding of large fragment DNA data storage. Background Art

[0002] With the rapid development of information technology towards digitalization, networking, and intelligence, the global data volume has been growing rapidly, and the value of long-term data storage has become increasingly prominent with the development of intelligent technologies such as large models. How to store massive data in the long term has become an important requirement for the development of information technology. With the explosive growth of the data storage scale, traditional data storage media and storage technologies are facing challenges. Existing media face problems such as limited service life, high daily maintenance costs, and slow growth of storage density restricted by physical limits. In particular, long-term data storage often requires data migration to ensure long-term reliability, resulting in a significant increase in maintenance costs. Traditional storage media cannot meet the storage needs of rapidly growing massive data, and exploring new storage media and storage modes has become one of the keys to the long-term healthy development of information technology.

[0003] Compared with traditional data storage media, synthetic DNA molecules have advantages such as high storage density, low maintenance costs, and a large amount of accumulated high-throughput reading and writing technologies, and are becoming a new form of storage media with great potential in future massive data archiving storage. In recent years, data storage systems using various forms of synthetic DNA media have been evaluated, mainly divided into: short fragment oligonucleotide pool storage methods and large fragment DNA data storage methods. Different from the storage scheme based on short fragment oligonucleotide pools, large fragment DNA data storage can achieve assembly and amplification through cell proliferation, with low data replication costs and high reliability, and has potential application value in scenarios of large-scale data distribution; the processing of large fragment DNA can better utilize the processing mechanisms in organisms for DNA assembly, production, etc., with higher reliability and better flexibility.

[0004] The reading of large fragment DNA is similar to the reading of cell DNA sequencing, and traditional cell genome sequencing technologies and data processing technologies can be used.

[0005] On the one hand, large DNA fragments can be sequenced using third-generation nanopore sequencing. However, the current error rate of third-generation sequencing is relatively high (about 10%), and the error types are relatively complex, including base insertions and deletions that are difficult to process. Handling such errors is usually challenging, and complex insertion and deletion processing algorithms are required to recover the stored data. For example, the research team at Tianjin University proposed a scheme that combines sparse LDPC codes and watermark sequence superposition coding, designed and synthesized a yeast artificial chromosome with a length of 254,886 bases, stored a 37.8 KB data file, and after third-generation nanopore sequencing, proposed a decoding strategy that combines genome assembly and an efficient indel (insertion and deletion) error correction and decoding algorithm, achieving error-free recovery of the data at a sequencing coverage of 16.8×. Researchers from Shanghai Jiao Tong University and Tianjin University used six large and uniform DNA sequences of 6,750 bases to store a 6.5 KB picture file and recovered the original data using 53.76× third-generation nanopore sequencing reads through the sequence alignment software MAFFT. The above methods all utilize the characteristics of third-generation nanopore sequencing such as portability and speed, but they also face the problem of dealing with complex indel errors.

[0006] On the other hand, very similar to genome sequencing, large DNA fragments can also use second-generation high-throughput sequencing technology based on shotgun sequencing to read out data, obtaining a large number of short sequencing reads. Similar to de novo genome sequencing, the read assembly operation is complex, and a "perfect" consensus sequence cannot be obtained at low sequencing coverage, posing a great challenge to low-cost error-free data recovery. The BGI Research Institute used a 54,520-base DNA fragment for intracellular data storage, and used the de novo genome assembly software SOAPdenovo to assemble the second-generation sequencing reads during the data readout process to recover the stored data file; the Tianjin University team performed second-generation high-throughput sequencing readout on the designed and synthesized yeast artificial chromosome, and through the second-generation sequencing read assembly software Velvet, using a DNA assembly method based on graph theory, further combined with the method of locating autonomous replication sequences (ARS), achieving error-free data recovery at a sequencing coverage of 23.5×; at the same time, the researchers designed and simulated a 2.5 Mb-long circular chromosome through a coding scheme that interleaves multiple very high-rate Reed-Solomon (RS) codewords and performed second-generation high-throughput sequencing readout, and used the assembly software Velvet and ABySS based on the de Brujin graph to achieve data recovery at a sequencing coverage of 20×. Similar to de novo genome sequencing, when using second-generation high-throughput sequencing to read out data from large DNA fragments, the data processing complexity of read assembly is high, and the required sequencing coverage is generally high, usually dozens of times or more.

[0007] Previously, researchers proposed a coding scheme based on the superposition of a watermark sequence and a sparse error-correcting code for long-fragment DNA data storage. The watermark sequence is hidden in the sequencing reads and can be used to identify complex insertions and deletions introduced by third-generation nanopore sequencing. Further combined with substitution errors, it is possible to achieve the readout of sequencing data based on third-generation nanopores. For such large-fragment DNA with watermark superposition coding, when read out using second-generation high-throughput sequencing, it is restricted by the de novo assembly of reads with high complexity and cannot achieve data recovery at low sequencing coverage, increasing the cost of reading data from long-fragment DNA. Summary of the Invention

[0008] The present invention provides a fast readout method for watermark superposition coding large-fragment DNA data storage. The present invention supports DNA lengths from Kb to Mb, uses simple sliding correlation and sequence consensus to achieve read positioning, effectively avoids high-complexity de novo assembly, and can exclude the interference of non-coding DNA reads, and can achieve error-free recovery of the original data at low sequencing coverage. See the following description for details:

[0009] A fast readout method for watermark superposition coding large-fragment DNA data storage, the method comprising the following steps:

[0010] (1) Demap the double-layer bit sequence of the sequencing reads obtained by fragmenting large-fragment DNA with watermark superposition coding and performing second-generation high-throughput sequencing according to the "base-bit" mapping rule, perform sliding correlation on the bit sequence and the local known double-layer watermark sequence, and calculate the cross-correlation peak value to determine the source of the sequencing reads and their starting positions on the large-fragment DNA sequence;

[0011] (2) After placing the sequencing reads identified by sliding correlation in the correct positions, use the majority voting decision algorithm to obtain a consensus base sequence;

[0012] (3) Transcode the consensus base sequence into a damaged bit sequence and then send it to a decoder for error correction to decode and recover the data.

[0013] Among them, step (1) is:

[0014] (1.1) The sequencing reads with a length of n base pairs obtained by fragmenting large-fragment DNA with watermark superposition coding and performing second-generation high-throughput sequencing r Demap the double-layer bit sequence with a length of n bits according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)} r 1 and r 2;

[0015] (1.2) The double-layer bit sequence r 1 andr 2 and a known double-layer watermark sequence of length L bits w 1 and w Convert the "0" in 2 to "-1", and keep the bit "1" unchanged;

[0016] (1.3) Calculate the double-layer bit sequence using a sliding window of length n r 1 and r The cross-correlation value between 2 and the known double-layer watermark sequence w 1 and w 2, where the cross-correlation value V at the i-th position i The calculation formula is expressed as

[0017]

[0018] Obtain L - n + 1 cross-correlation values; then compare the calculated cross-correlation values at each position to obtain the cross-correlation peak value V peak = max(V0, V1, … V L-n ) and the corresponding position;

[0019] (1.4) Use the cross-correlation peak threshold to screen the sequencing reads that meet the requirements and record the peak position, that is, the starting position of the sequencing reads on the large fragment DNA sequence;

[0020] (1.5) Retrieve the sequencing reads according to the cross-correlation peak position for the recovery of different data.

[0021] Among them, step (2) is:

[0022] (2.1) If the length of the data DNA sequence is l bases, construct a 5 * l counting buffer corresponding to the number of five bases {A, T, G, C, N} at each position;

[0023] (2.2) Slide the starting position of the sequencing reads obtained by correlation recognition. According to the cross-correlation peak position, place the recognized sequencing reads at the corresponding starting position, and perform an addition counting operation on the corresponding position of the counting buffer according to the base type of each position of the read, to obtain a counting buffer containing the number of five bases at each position;

[0024] (2.3) Perform a per-base majority voting decision on the counting buffer. Judge the base with the largest number in the buffer at the i-th position as the base at that position. If there is no sequencing data filled at that position, or a majority voting decision cannot be made, mark it as "X", that is, an erased error, to obtain a consensus sequence of length l bases s .

[0025] Among them, after transcoding the consensus base sequence into a damaged bit sequence, it is sent to a decoder for error correction, and the decoder restores the data. Step (3) is as follows:

[0026] (3.1) The consensus sequence with a length of l bases s is demapped into a double-layer bit sequence according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)}, and the base X is marked as an erasure;

[0027] (3.2) Perform a bit-by-bit exclusive OR operation on the blurred watermark sequence superimposed with the damaged error correction codeword sequence and the known watermark sequence to remove the superimposed watermark sequence, and obtain a damaged sparse error correction codeword sequence;

[0028] (3.3) Perform sparse decoding on the damaged sparse error correction codeword sequence to obtain a damaged error correction codeword sequence, and mark the erasure positions;

[0029] (3.4) Send the damaged error correction codeword sequence into the corresponding decoder, and perform error correction and erasure decoding to restore the original data.

[0030] The double-layer bit sequence is calculated using a sliding window of length n r 1 and r 2 and the known double-layer watermark sequence w 1 and w 2 cross-correlation values are obtained to get L - n + 1 cross-correlation values; then the cross-correlation values at each position calculated are compared to obtain the cross-correlation peak value V peak = max(V0, V1,... V L-n ) and the corresponding positions. There are two different sliding correlation calculation schemes for the cross-correlation peak value, which are respectively:

[0031] (1) Sequentially select sequencing reads from the sequencing data containing a large number of reads, demap them, and perform bit-by-bit sliding correlation with the known double-layer watermark sequence of length L bits to sequentially obtain the cross-correlation values corresponding to each position;

[0032] (2) Split the sequencing data containing a large number of sequencing reads with a length of n bases to obtain multiple sub-files of sequencing data; then segment the known double-layer watermark sequence of length L bits, and keep an overlapping segment of n bits between adjacent short segments; use a multi-threaded method to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, and achieve the acceleration of the sliding correlation calculation process.

[0033] When using the cross-correlation peak threshold to screen the sequencing reads that meet the requirements, it is necessary to determine and optimize the cross-correlation peak threshold according to whether the sequencing data contains DNA fragment sequencing data with non-preset watermark superposition encoding:

[0034] (1) For the application scenario of DNA fragment sequencing data that does not contain non - preset watermark superposition encoding, the sequencing read alignment is directly achieved according to the cross - correlation peak position.

[0035] (2) For the application scenario of DNA fragment sequencing data that contains non - preset watermark superposition encoding, including the second - generation high - throughput sequencing data of the host cell genome and the second - generation high - throughput sequencing data of DNA fragments with other encoding methods, first set different cross - correlation peak thresholds to identify and screen the sequencing reads; then use the identified sequencing reads for large - number merging to obtain the corresponding consensus sequence and calculate the base error rate; finally, select the threshold corresponding to the minimum base error rate as the optimal cross - correlation peak threshold to reduce the influence of interfering reads.

[0036] The beneficial effects of the technical solution provided by the present invention are as follows:

[0037] 1. The present invention proposes a fast read - out method for storing large - fragment DNA data with watermark superposition encoding. The proposed method first performs second - generation high - throughput sequencing on the large - fragment DNA sequence with watermark superposition encoding to obtain the corresponding sequencing data; then calculates the cross - correlation peak by performing sliding correlation between the fuzzy watermark sequence hidden in the noisy sequencing reads and the locally known watermark sequence; according to the cross - correlation peak and the corresponding position, uses the sequence consensus algorithm to obtain the consensus base sequence, and further realizes the efficient and reliable recovery of the stored data through an efficient error - correcting and deletion - correcting decoding algorithm.

[0038] 2. The present invention supports storing data with DNA lengths ranging from Kb to Mb. It uses simple sliding correlation and sequence consensus to achieve read alignment, effectively avoiding the high - complexity de novo assembly, and can exclude the interference of non - coding DNA reads. It can recover the data stored in the DNA sequence without error at a relatively low sequencing coverage, reducing the data read - out cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is the system block diagram for implementing the solution in the present invention;

[0040] Figure 2 It is the data recovery flow chart from NGS noisy reads in the present invention;

[0041] Figure 3 It is the read alignment flow chart of sliding correlation between the sequencing reads and the reference watermark sequence in the present invention;

[0042] Figure 4 It is the flow chart for generating the consensus sequence by large - number merging in the present invention;

[0043] Figure 5 It is the flow chart for error - correcting decoding by converting the consensus sequence into a damaged codeword sequence in the present invention;

[0044] Figure 6 It is a flowchart of the parallel sliding related solution in the present invention;

[0045] Figure 7 It is a distribution diagram of specific data recovery in Specific Embodiment 1 of the present invention;

[0046] Figure 8 It is a base error distribution diagram in the majority voting decision process in Specific Embodiment 1 of the present invention;

[0047] Figure 9 It is the acceleration performance of the parallel sliding related solution in Specific Embodiment 1 of the present invention;

[0048] Figure 10 It is a distribution diagram of hybrid sequencing data in Specific Embodiment 2 of the present invention;

[0049] Figure 11 It is a distribution diagram of specific data recovery in Specific Embodiment 2 of the present invention;

[0050] Figure 12 It is the acceleration performance of the parallel sliding related solution in Specific Embodiment 2 of the present invention;

[0051] Figure 13 It is a base error distribution diagram in the majority voting decision process in Specific Embodiment 2 of the present invention;

[0052] Figure 14 It is the base coverage distribution under the minimum sequencing coverage of error-free recovery in the present invention;

[0053] Figure 15 It is a comparison between the data recovery solution proposed in the present invention and the traditional read-based de novo assembly solution.

[0054] Table 1 shows the actual second-generation sequencing data parameters adopted in the present invention;

[0055] Table 2 shows the experimental results of parallel data recovery of different sequencing samples in Specific Embodiment 1 of the present invention;

[0056] Table 3 shows the experimental results of parallel data recovery of different sequencing samples in Specific Embodiment 2 of the present invention;

[0057] Table 4 shows a comparison between the present invention and other large fragment DNA data storage solutions. Specific Embodiments

[0058] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.

[0059] The present invention belongs to the field of DNA data storage, and proposes a fast reading method for watermark-overlaid encoded large-fragment DNA data storage. First, for the large fragments of DNA encoded by watermark overlay, after fragmentation and library construction, second-generation high-throughput sequencing is performed to obtain corresponding sequencing data. Then, the watermark sequences hidden in the noisy sequencing reads are used to perform sliding correlation calculation with the locally known watermark sequences to obtain the cross-correlation peak values. According to the corresponding positions of the cross-correlation peak values, the positions of the reads in the large-fragment DNA sequence are determined, and then the majority voting decision algorithm is used to generate a consensus sequence, and further, the reliable recovery of the stored data is realized through an efficient error correction and deletion decoding algorithm. The advantages of the present invention are that it supports DNA lengths from Kb to Mb, uses simple sliding correlation and sequence consensus to achieve read positioning, effectively avoids the high complexity of de novo assembly, and can exclude the interference of non-coding DNA reads, and can achieve error-free recovery of the original data at a lower sequencing coverage, greatly reducing the data reading cost.

[0060] Attached Figure 1 The system block diagram implemented by the solution of the present invention is shown. Attached Figure 2 The flowchart for recovering data from NGS noisy reads is shown. The solution of the present invention includes the following steps:

[0061] (1) Demap the double-layer bit sequence of the sequencing reads obtained by fragmenting the large-fragment DNA encoded by watermark overlay and performing second-generation high-throughput sequencing according to the "base-bit" mapping rule, perform sliding correlation on the bit sequence and the locally known double-layer watermark sequence, calculate the cross-correlation peak value, so as to determine the source of the sequencing reads and their starting positions on the large-fragment DNA sequence;

[0062] (2) After placing the sequencing reads identified by sliding correlation in the correct positions, use the majority voting decision algorithm to obtain a consensus base sequence;

[0063] (3) Transcode the consensus base sequence into a damaged bit sequence and send it to the decoder for error correction to decode and recover the data.

[0064] Among them, as shown in the attached Figure 3 The system block diagram implemented by the solution of the present invention is shown. In step (1) of the present invention, the sequencing reads obtained by fragmenting the large-fragment DNA encoded by watermark overlay and performing second-generation high-throughput sequencing are demapped into a double-layer bit sequence according to the "base-bit" mapping rule, perform sliding correlation on the bit sequence and the locally known double-layer watermark sequence, calculate the cross-correlation peak value, so as to determine the source of the sequencing reads and their starting positions on the large-fragment DNA sequence. The specific steps are as follows:

[0065] (1.1) The sequencing reads with a length of n base pairs obtained by fragmenting the large-fragment DNA encoded by watermark overlay and performing second-generation high-throughput sequencing rDemap the double-layer bit sequence of length n bits according to the mapping rule of "base-bit" {A→(00), T→(01), G→(10), C→(11)} r 1 and r 2;

[0066] (1.2) For the double-layer bit sequence r 1 and r 2 and the known double-layer watermark sequence of length L bits w 1 and w 2, convert the "0" to "-1", and keep the bit "1" unchanged;

[0067] (1.3) Use a sliding window of length n to calculate the cross-correlation value between the double-layer bit sequence r 1 and r 2 and the known double-layer watermark sequence, where the cross-correlation value V w 1 and w 2 at the i-th position is calculated as follows, i The calculation formula is expressed as,

[0068]

[0069] Obtain L - n + 1 cross-correlation values; then compare the calculated cross-correlation values at each position to obtain the cross-correlation peak value V peak = max(V0, V1,... V L-n ) and the corresponding position;

[0070] (1.4) Use the cross-correlation peak threshold to screen the sequencing reads that meet the requirements and record the peak position, that is, the starting position of the sequencing reads on the large fragment DNA sequence;

[0071] (1.5) Retrieve the sequencing reads according to the cross-correlation peak position for the recovery of different data.

[0072] Among them, as shown in the appendix Figure 4 After placing the sequencing reads obtained by sliding correlation recognition in the correct position in step (2) of the present invention, the majority voting decision algorithm is used to obtain the consensus base sequence. The specific steps are as follows:

[0073] (2.1) If the length of the data DNA sequence is l bases, construct a 5*l counting buffer to correspond to the number of five bases {A, T, G, C, N} at each position;

[0074] (2.2) The starting positions of the sequenced reads obtained by sliding correlation recognition are different. According to the cross-correlation peak positions, the recognized sequenced reads are placed at the corresponding starting positions, and an addition counting operation is performed on the corresponding position counting buffer according to the base types at each position of the reads, resulting in a counting buffer containing the quantities of the five bases at each position;

[0075] (2.3) Perform a per-base majority voting decision on the counting buffer. The base with the largest quantity in the buffer at the i-th position is judged as the base at that position. If there is no sequencing data filled at this position or a majority voting decision cannot be made, it is marked as "X", that is, an erasure error, to obtain a consensus sequence of length l bases s .

[0076] Among them, as shown in the appendix Figure 5 In the step (3) of the present invention, the consensus sequence is transcoded into a damaged codeword sequence and then sent to a decoder for error correction decoding to recover data. The specific steps are as follows:

[0077] (3.1) Demap the consensus sequence s of length l bases into a double-layer bit sequence according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)}, and the base X is marked as an erasure;

[0078] (3.2) Perform a bit-by-bit exclusive OR operation on the superimposed damaged error correction codeword sequence and the known watermark sequence to remove the superimposed watermark sequence, obtaining a damaged sparsified error correction codeword sequence;

[0079] (3.3) Perform sparse decoding on the damaged sparsified error correction codeword sequence to obtain a damaged error correction codeword sequence, and mark the erasure error positions;

[0080] (3.4) Send the damaged error correction codeword sequence into the corresponding decoder for error correction and erasure decoding to recover the original data.

[0081] Among them, in step (1.3) of the present invention, a sliding window of length n is used to calculate the cross-correlation values between the double-layer bit sequence r 1 and r 2 and the known double-layer watermark sequence w 1 and w 2, obtaining L - n + 1 cross-correlation values; then compare the calculated cross-correlation values at each position to obtain the cross-correlation peak V peak = max(V0, V1,... V L-n ) and the corresponding position. Specifically, there are two different sliding correlation calculation schemes for the cross-correlation peak, which are respectively:

[0082] (1) After selecting sequencing reads one by one from sequencing data containing a large number of reads and unmapping them, perform a bit-by-bit sliding correlation with a known double-layer watermark sequence of length L bits, and sequentially obtain the cross-correlation values corresponding to each position.

[0083] (2) Split the sequencing data containing a large number of sequencing reads of length n bases to obtain multiple sub-files of sequencing data; then segment the known double-layer watermark sequence of length L bits, with an overlap segment of n bits between adjacent short segments; use a multi-threaded approach to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, achieving acceleration of the sliding correlation calculation process.

[0084] Among them, as shown in the appendix Figure 6 The present invention splits the sequencing data containing a large number of sequencing reads of length n bases to obtain multiple sub-files of sequencing data; then segments the known double-layer watermark sequence of length L bits, with an overlap segment of n bits between adjacent short segments; uses a multi-threaded approach to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, achieving acceleration of the sliding correlation calculation process. The specific steps are as follows:

[0085] (4.1) Split the second-generation high-throughput sequencing data containing m reads of length n bases to obtain q sub-files, each containing m / q sequencing reads.

[0086] (4.2) Segment the known double-layer watermark sequence of length L bits w ={ w 1, w 2} into k short watermark segments { w (0) , w (1) ,… w (k-1)}, with an overlap segment of n bits between adjacent short segments.

[0087] (4.3) Start q*k threads, which are respectively used for the sliding correlation between the sequencing reads in each sub-file and each short watermark segment, and calculate q*k cross-correlation values in parallel at the same time.

[0088] Among them, in step (1.4) of the present invention, the cross-correlation peak threshold is used to screen the sequencing reads that meet the requirements, and the cross-correlation peak threshold needs to be preferably selected and determined according to whether the sequencing data contains non-coding DNA sequencing reads:

[0089] (1) For application scenarios that do not contain non-coding DNA sequencing reads, directly locate the sequencing reads according to the cross-correlation peak position.

[0090] (2) For the application scenarios involving non-coding DNA sequencing reads, including the second-generation high-throughput sequencing data of the host cell genome and the second-generation high-throughput sequencing data of DNA fragments with other coding methods, first set different cross-correlation peak thresholds to identify and screen the sequencing reads; then use the identified sequencing reads for large number merging to obtain the corresponding consensus sequence and calculate the base error rate; finally, select the threshold corresponding to the minimum base error rate as the optimal cross-correlation peak threshold to reduce the influence of interfering reads. Specific embodiment

[0092] A fast readout method for storing large fragment DNA data with watermark superposition coding designed by the present invention performs second-generation high-throughput sequencing on the large fragment DNA sequence with watermark superposition coding, and then uses the fuzzy watermark sequence hidden in the noisy sequencing reads to perform sliding correlation calculation with the locally known watermark sequence to achieve read segment alignment. The sequence consensus algorithm and efficient error correction and deletion decoding can be used to recover the stored data without error, effectively avoiding the relatively high complexity of de novo assembly, and can exclude the interference of non-coding DNA sequencing reads. The data stored in the DNA sequence can be recovered without error at a relatively low sequencing coverage, reducing the data readout cost. Two specific embodiments are given below to illustrate the feasibility of the present invention. Combining the given embodiments, the purpose, features and advantages of the invention can be further understood.

[0093] Embodiment 1

[0094] The first specific embodiment of the present invention is given below. This embodiment is for the application scenario that does not contain other non-coding DNA sequencing reads. Based on the 254,886-base yeast artificial chromosome with the sparse low-density parity-check (LDPC) code and watermark sequence superposition coding scheme designed in the previous work, second-generation high-throughput sequencing technology is used for data readout. Nucleic acid extraction and library construction are directly performed on yeast cells, and finally, second-generation high-throughput sequencing data with a length of 150 bases is obtained through fragmentation and second-generation high-throughput sequencing, including the second-generation high-throughput sequencing data of the artificial chromosome and the second-generation high-throughput sequencing data of the host cell genome.

[0095] Among them, two different coding methods are used to construct corresponding information blocks in the 254,886-base yeast artificial chromosome, specifically including:

[0096] (1) Use a non-binary LDPC(64,512,2,256) code with a code rate of 1 / 2 to encode the data, and after sparse operation and superposition transcoding with the watermark sequence, obtain 1 information sub-block with a length of 40,320 bases, which is used to store a 4KB picture 2 file;

[0097] (2) The data is encoded using a non-binary LDPC(64,800,54,000) code with a code rate of 5 / 6. After sparsification operation and superposition transcoding with the watermark sequence, 5 information sub-blocks with a length of 40,500 bases are obtained. Among them, 1 information sub-block is used to store the 6.5KB picture 1 file, and the remaining 4 information sub-blocks are used to store the 25.5KB video file;

[0098] (3) The data is encoded using a non-binary LDPC(64,800,16,200) with a code rate of 2 / 3. After sparsification operation and superposition transcoding with the watermark sequence, 1 information sub-block with a length of 40,500 bases is obtained, which is used to store the 5.4KB picture 3 file.

[0099] To meet the application scenario that does not contain other non-coding DNA sequencing reads, different sequencing data are separated by preprocessing the sequencing data. The specific steps are as follows: Use the sequence alignment software BWA to align the sequencing reads with the known host cell genome sequence, and screen out the second-generation high-throughput sequencing reads belonging to the host cell genome from the alignment results to obtain the second-generation high-throughput sequencing data that only contains the large fragment DNA with watermark superposition encoding.

[0100] The sequencing reads obtained by fragmenting the large fragment DNA with watermark superposition encoding and performing second-generation high-throughput sequencing are demapped to the double-layer bit sequence according to the "base-bit" mapping rule. The bit sequence is slid-correlated with the local known double-layer watermark sequence, and the peak value of the cross-correlation value is calculated to determine the source of the sequencing reads and its starting position on the large fragment DNA sequence. The specific steps are as follows:

[0101] (1.1) The 150-base sequencing reads obtained by fragmenting the large fragment DNA with watermark superposition encoding and performing second-generation high-throughput sequencing r are demapped to a 150-bit double-layer bit sequence according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)} r 1 and r 2;

[0102] (1.2) Convert the "0" in the double-layer bit sequence r 1 and r 2 and the known 254,886-bit double-layer watermark sequence w 1 and w 2 to "-1", and keep the bit "1" unchanged;

[0103] (1.3) Use a 150-base sliding window to calculate the double-layer bit sequence r 1 and r 2 and the known double-layer watermark sequence wThe cross-correlation values between 1 and w 2 are obtained, resulting in 254,737 cross-correlation values; then, the cross-correlation values at each position calculated are compared to obtain the cross-correlation peak value and the corresponding position;

[0104] (1.4) Retrieve the sequencing reads according to the cross-correlation peak position for the recovery of different data.

[0105] Among them, a sliding window with a length of 150 bases is used to calculate the double-layer bit sequence r The cross-correlation values between 1 and r 2 and the known double-layer watermark sequence w The cross-correlation values between 1 and w 2 are obtained, resulting in 254,737 cross-correlation values; then, the cross-correlation values at each position calculated are compared to obtain the cross-correlation peak value and the corresponding position. Specifically, there are two different sliding correlation calculation schemes for the cross-correlation peak value, which are respectively:

[0106] (1) After demapping the sequencing reads selected one by one from the sequencing data containing a large number of reads with a length of 150 bases, perform a bit-by-bit sliding correlation with the known double-layer watermark sequence with a length of 254,886 bits, and sequentially obtain the cross-correlation value corresponding to each position;

[0107] (2) Split the sequencing data containing a large number of reads with a length of 150 bases to obtain multiple sub-files of sequencing data; then, perform segmented processing on the known double-layer watermark sequence with a length of 254,886 bits, and keep an overlapping segment of 150 bits between adjacent short segments; use a multi-threaded method to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, realizing the acceleration of the sliding correlation calculation process.

[0108] Among them, the specific steps for splitting the sequencing data containing a large number of reads with a length of 150 bases to obtain multiple sub-files of sequencing data; then, performing segmented processing on the known double-layer watermark sequence with a length of 254,886 bits, and keeping an overlapping segment of 150 bits between adjacent short segments; using a multi-threaded method to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, realizing the acceleration of the sliding correlation calculation process are as follows:

[0109] (2.1) Split the second-generation high-throughput sequencing data containing m reads with a length of 150 bases to obtain q sub-files, each containing m / q sequencing reads;

[0110] (2.2) Segment the known double-layer watermark sequence with a length of 254,886 bits w ={ w 1, w 2} into segments to obtain k short watermark segments { w (0) , w(1) ,… w (k-1)}, with a 150-bit overlap segment maintained between adjacent short segments;

[0111] (2.3) Start q*k threads, which are respectively used for the sliding correlation between the sequencing reads in each sub-file and each short watermark segment, and calculate q*k cross-correlation values in parallel at the same time.

[0112] After placing the sequencing reads identified by the sliding correlation in the correct positions, the specific steps to obtain the consensus base sequence using the majority voting decision algorithm are as follows:

[0113] (3.1) If the length of the data DNA sequence is l bases (information sub-block with a code rate of 1 / 2, l = 40,320; information sub-block with a code rate of 5 / 6, l = 40,500), construct a 5*l counting buffer corresponding to the quantities of the five bases {A, T, G, C, N} at each position;

[0114] (3.2) Since the starting positions of the sequencing reads identified by the sliding correlation are different, place the identified sequencing reads at the corresponding starting positions according to the cross-correlation peak positions, and perform an addition counting operation on the corresponding positions of the counting buffer according to the base types at each position of the reads, to obtain a counting buffer containing the quantities of the five bases at each position;

[0115] (3.3) Perform a majority voting decision on each base of the counting buffer, and determine the base with the largest quantity in the buffer at the i-th position as the base at that position. If there is no sequencing data filled at this position, or a majority voting decision cannot be made, mark it as "X", that is, an erasure error, to obtain a consensus sequence with a length of l bases s .

[0116] The obtained consensus sequence with a length of l bases s After transcoding it into a damaged LDPC codeword sequence, the specific steps to send it to the corresponding decoder for error correction decoding to recover the data are as follows:

[0117] (4.1) Map the consensus sequence with a length of l bases s according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)} into a double-layer bit sequence, and mark the base X as an erasure;

[0118] (4.2) Perform a bit-by-bit exclusive OR operation on the blurred watermark sequence superimposed with the damaged error correction codeword sequence and the known watermark sequence to remove the superimposed watermark sequence, and obtain a damaged sparsified error correction codeword sequence;

[0119] (4.3)Sparsely decode the damaged sparse error-corrected codeword sequence to obtain the damaged error-corrected codeword sequence, and mark the erased error positions;

[0120] (4.4)Send the damaged error-corrected codeword sequence into the corresponding decoder, and perform error correction and erasure decoding to recover the original data.

[0121] First, based on the proposed data error correction and readout scheme, multiple independent repeated experiments are carried out under different sequencing coverages. The experimental data parameters are shown in Table 1, to explore the reliability of the proposed scheme and the minimum sequencing coverage required for error-free data recovery. As Figure 7 shown, error-free recovery of different data stored in artificial chromosomes can be achieved under the sequencing coverage of 1.1× - 5×, and the required sequencing coverage is less than the current large-fragment DNA data recovery scheme based on read segment assembly. Secondly, to explore the specific impact of sequencing coverage on the error distribution of the majority voting decision process, Figure 8 the corresponding base error rate distributions are statistically analyzed under different sequencing coverages. It can be found that the base error rate gradually decreases with the increase of sequencing coverage, and mainly includes base erasure errors and base substitution errors. In this embodiment, complex de novo assembly operations are effectively avoided through sliding correlation, and with the help of a simple sequence consensus algorithm and a reliable error correction coding algorithm, error-free data recovery can be achieved under low sequencing coverage in application scenarios without interference from other non-coding large-fragment DNA sequencing data. To prove the stability of the proposed scheme, multiple parallel data recovery experiments are carried out using the sequencing data samples in Table 1, and the results are all consistent, as shown in Table 2.

[0122] Then, this embodiment analyzes and verifies the two sliding correlation methods proposed in the scheme of the present invention. For the test of the parallel sliding correlation scheme, 60 (6×10) threads are selected, the sequencing data is divided into 6 sub-files, and the reference watermark sequence is equally divided into 10 segments, with a 150-bit overlapping segment between adjacent segments. Two schemes are respectively used to perform sliding correlation operations on different numbers of sequencing reads. Figure 9 shows the running time comparison of the two sliding correlation schemes and the speed improvement effect of the parallel scheme under different numbers of sequencing reads. From Figure 9 it can be seen that in the application scenario of DNA sequencing reads without other non-coding, the number of reads corresponding to the effective sequencing coverage is small, and the parallel acceleration scheme can improve the read alignment process by about 20 times under 60 threads. The effectiveness of the different sliding correlation schemes proposed in the scheme of the present invention is verified.

[0123] Embodiment 2

[0124] The second specific embodiment of the present invention is given below. This embodiment is for the application scenario of sequencing data of other large non-coding DNA fragments. Based on the sparse low-density parity-check (LDPC) code and the watermark sequence superposition coding scheme designed in the previous work for a 254,886-base yeast artificial chromosome, second-generation high-throughput sequencing technology is used for data reading. Nucleic acid extraction and library construction are directly carried out on yeast cells, and finally, second-generation high-throughput sequencing data with a length of 150 bases is obtained through fragmentation and second-generation high-throughput sequencing, including second-generation high-throughput sequencing data of the artificial chromosome and second-generation high-throughput sequencing data of the host cell genome.

[0125] Among them, two different coding methods are used to construct corresponding information blocks in the 254,886-base yeast artificial chromosome, specifically including:

[0126] (1) The data is encoded using a non-binary LDPC(64,512,32,256) code with a code rate of 1 / 2. After sparsification operation and superposition transcoding with the watermark sequence, 1 information sub-block with a length of 40,320 bases is obtained, which is used to store a 4KB picture 2 file;

[0127] (2) The data is encoded using a non-binary LDPC(64,800,54,000) code with a code rate of 5 / 6. After sparsification operation and superposition transcoding with the watermark sequence, 5 information sub-blocks with a length of 40,500 bases are obtained. Among them, 1 information sub-block is used to store a 6.5KB picture 1 file, and the remaining 4 information sub-blocks are used to store a 25.5KB video file;

[0128] The sequencing reads obtained by fragmenting and second-generation high-throughput sequencing of the large fragment DNA with watermark superposition coding are demapped into a double-layer bit sequence according to the "base-bit" mapping rule. The bit sequence is slid-correlated with the local known double-layer watermark sequence, and the peak value of the cross-correlation value is calculated to determine the source of the sequencing reads and their starting positions on the large fragment DNA sequence. The specific steps are as follows:

[0129] (1.1) The 150-base sequencing read r obtained by fragmenting and second-generation high-throughput sequencing of the large fragment DNA with watermark superposition coding is demapped into a 150-bit double-layer bit sequence according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)} r 1 and r 2;

[0130] (1.2) The double-layer bit sequence r 1 and r 2 and the known 254,886-bit double-layer watermark sequence w 1 andw The "0" in 2 is converted to "-1", and the bit "1" remains unchanged;

[0131] (1.3) Calculate the double-layer bit sequence using a sliding window of 150 bases in length r 1 and r 2 and the known double-layer watermark sequence w 1 and w 2 to obtain 254,737 cross-correlation values; then compare the cross-correlation values at each position calculated to obtain the cross-correlation peak and the corresponding position;

[0132] (1.4) Use the cross-correlation peak threshold to screen the sequencing reads that meet the requirements and record the peak position, that is, the starting position of the sequencing reads on the large fragment DNA sequence;

[0133] (1.5) Retrieve the sequencing reads according to the cross-correlation peak position for the recovery of different data.

[0134] Among them, calculate the double-layer bit sequence using a sliding window of 150 bases in length r 1 and r 2 and the known double-layer watermark sequence w 1 and w 2 to obtain 254,737 cross-correlation values; then compare the cross-correlation values at each position calculated to obtain the cross-correlation peak and the corresponding position. Specifically, there are two different sliding correlation calculation schemes for cross-correlation peaks, which are respectively:

[0135] (1) After selecting each sequencing read from the sequencing data containing a large number of reads of 150 bases in length and unmapping it, perform a bit-by-bit sliding correlation with the known double-layer watermark sequence of 254,886 bits in length, and sequentially obtain the cross-correlation value corresponding to each position;

[0136] (2) Split the sequencing data containing a large number of reads of 150 bases in length to obtain multiple sub-files of sequencing data; then perform segmented processing on the known double-layer watermark sequence of 254,886 bits in length, and keep an overlapping segment of 150 bits between adjacent short segments; use a multi-threaded method to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, realizing the acceleration of the sliding correlation calculation process.

[0137] Among them, the specific steps of splitting the sequencing data containing a large number of reads of 150 bases in length to obtain multiple sub-files of sequencing data; then performing segmented processing on the known double-layer watermark sequence of 254,886 bits in length, and keeping an overlapping segment of 150 bits between adjacent short segments; using a multi-threaded method to perform parallel sliding correlation to calculate multiple cross-correlation values at the same time, realizing the acceleration of the sliding correlation calculation process are as follows:

[0138] (2.1) Split the next-generation high-throughput sequencing data containing m reads with a length of 150 bases into q sub-files, each containing m / q sequencing reads;

[0139] (2.2) Segment the known double-layer watermark sequence with a length of 254,886 bits w ={ w 1, w 2} into k short watermark fragments { w (0) , w (1) ,… w (k-1)}, with a 150-bit overlapping segment between adjacent short fragments;

[0140] (2.3) Start q*k threads, which are respectively used for the sliding correlation between the sequencing reads in each sub-file and each short watermark fragment, and calculate q*k cross-correlation values in parallel at the same time.

[0141] The sequencing data obtained in this embodiment contains interference from other non-coding large fragment DNA sequencing data. Therefore, it is necessary to optimize and determine the cross-correlation peak threshold to screen the sequencing reads that meet the requirements. The specific steps for determining the optimal cross-correlation peak are as follows:

[0142] (3.1) Select the values in the set {66, 72, 78, …, 150} one by one as the cross-correlation peak threshold, and use different thresholds to identify and screen the sequencing reads to obtain the sequencing reads whose cross-correlation peaks meet the threshold requirements;

[0143] (3.2) Perform large number merging on the screened sequencing reads under different sequencing coverages to obtain the corresponding consensus sequences, and count the base error rates in the sequences;

[0144] (3.3) Select the threshold with the smallest base error rate as the optimal cross-correlation peak threshold.

[0145] Use the next-generation high-throughput sequencing reads to perform sliding correlation calculation with the known watermark sequence to obtain the cross-correlation peak and position. After further identification and screening by the cross-correlation peak threshold, reduce the interference of the host cell genome next-generation high-throughput sequencing data, and obtain relatively pure artificial chromosome sequencing data. Then, the specific steps for obtaining the consensus sequence by majority voting according to the peak position are as follows:

[0146] (4.1) If the length of the data DNA sequence is l bases (information sub-block with a coding rate of 1 / 2, l = 40,320 bp; information sub-block with a coding rate of 5 / 6, l = 40,500 bp), construct a 5*l counting buffer corresponding to the quantities of the five bases {A, T, G, C, N} at each position;

[0147] (4.2) Since the starting positions of the sequencing reads obtained by sliding correlation recognition are different, place the recognized sequencing reads at the corresponding starting positions according to the cross-correlation peak positions, and perform an addition counting operation on the corresponding positions of the counting buffer according to the base types at each position of the reads, to obtain a counting buffer containing the quantities of the five bases at each position;

[0148] (4.3) Perform a per-base majority voting decision on the counting buffer, and determine the base with the largest quantity in the buffer at the i-th position as the base at that position. If there is no sequencing data filling at this position, or a majority voting decision cannot be made, mark it as "X", that is, an erasure error, to obtain a consensus sequence s with a length of l bases.

[0149] The obtained consensus sequence with a length of l bases s The specific steps for transcoding it into a damaged LDPC codeword sequence and then sending it to the corresponding decoder for error correction decoding to recover the data are as follows:

[0150] (5.1) Demap the consensus sequence s with a length of l bases into a double-layer bit sequence according to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)}, and mark the base X as an erasure;

[0151] (5.2) Perform a bit-by-bit exclusive OR operation on the blurred watermark sequence superimposed on the damaged error correction codeword sequence and the known watermark sequence to remove the superimposed watermark sequence, to obtain a damaged sparsified error correction codeword sequence;

[0152] (5.3) Perform sparse decoding on the damaged sparsified error correction codeword sequence to obtain a damaged error correction codeword sequence, and mark the erasure error positions;

[0153] (5.4) Send the damaged error correction codeword sequence to the corresponding decoder, and perform error correction and erasure decoding to recover the original data.

[0154] The sequencing data obtained in this embodiment contains second-generation high-throughput sequencing data of the host cell genome, such as Figure 10As shown, the proportion of artificial chromosome sequencing data is only 3%. In the absence of the host genome, all the noise reads come from the artificial chromosomes carrying the data. To determine the optimal cross-correlation peak recognition threshold, values from the set {66, 72, 78, …, 150} are selected as the cross-correlation peak thresholds, which are respectively used for the recognition and screening of the reads containing host genome sequencing data. And at different sequencing coverages, the screened reads are used for large-scale merging, and the average base error rate in the corresponding consensus sequences is statistically analyzed. According to the error rate distribution under different thresholds, six information sub-blocks are analyzed, and finally the optimal cross-correlation peak thresholds are determined to be 108, 108, 102, 100, 102, and 102 respectively.

[0155] First, based on the proposed data retrieval and readout scheme, multiple independent repeated experiments are carried out at different sequencing coverages to explore the reliability of the proposed scheme in different application scenarios and the minimum sequencing coverage required for error-free data recovery. As Figure 11 shown, error-free recovery of different data stored in artificial chromosomes can be achieved at a sequencing coverage of 1.2× - 8.4×, and the required sequencing coverage is less than the current large-fragment DNA data recovery scheme based on read assembly. This further verifies that the scheme of the present invention, through sliding correlation, sequence consensus algorithm and reliable error correction and deletion correction decoding, etc., avoids complex de novo assembly and realizes accurate data recovery at low coverage, thus reducing the complexity of biochemical operations and the workload of data processing. To prove the stability of the proposed scheme, multiple parallel data recovery experiments are carried out using different sequencing data samples, and the results are all consistent, as shown in Table 3.

[0156] Second, this embodiment analyzes and tests the parallel sliding correlation method proposed by the scheme of the present invention. In the test of the parallel sliding correlation scheme, the number of threads selected is 60 (6×10). The sequencing data is divided into 6 sub-files, and the reference watermark sequence is equally divided into 10 segments, with an overlapping segment of 150 bits between adjacent segments. Two schemes are respectively used to perform sliding correlation processing on different numbers of sequencing reads. Figure 12 shows the running time comparison of the serial-parallel and parallel sliding correlation schemes in the case of different numbers of sequencing reads and the speed improvement effect of the parallel scheme. It can be seen from Figure 12 that in the application scenario of sequencing data containing other non-coding DNA fragments, for a large number of second-generation high-throughput sequencing reads corresponding to the same effective sequencing coverage, the read alignment operation scale based on sliding correlation is large, and the parallel acceleration scheme can increase the speed of large-scale read alignment based on sliding correlation by about 45 times under 60 threads.

[0157] Figure 13Figure 0 shows the base error distribution map of the majority voting decision process in Example 2. Since the interference data cannot be completely excluded, the base error rate increases at the same sequencing coverage. Figure 14 The base coverage distribution under the minimum sequencing coverage for error-free data recovery in Example 2 was explored. It can be seen that when the sequencing coverage is at least 1.1×, the original data can be successfully recovered without the host genome. Compared with the current large-fragment DNA data recovery scheme based on de novo assembly of reads, the proposed data recovery method has the advantage that the known watermark sequence at the read end is similar to the reference genome sequence in genome resequencing, which supports the direct and accurate positioning and identification of sequencing reads, effectively avoiding the de novo assembly operation of reads. Further, during the sequence reconstruction process, this method can efficiently connect short sequencing reads at different starting positions through a simple majority voting decision algorithm, avoiding complex operations such as graph construction and optimization, and does not require obtaining long continuous fragments. With a reliable error correction coding scheme, error-free data recovery can be achieved at low sequencing coverage, as Figure 15 shown.

[0158] Finally, Table 4 compares with other schemes for data storage using large DNA fragments. It can be seen that compared with other schemes for large DNA fragment readout, the recovery method of the present invention requires a lower coverage rate for error-free data recovery, only 1.1× or 5×, thus effectively reducing the readout cost.

[0159] Table 1

[0160]

[0161] Table 2

[0162]

[0163] Table 3

[0164]

[0165]

[0166] Table 4

[0167]

[0168] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment. The above serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0169] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for rapid reading of large-fragment DNA data stored by watermark superposition coding, characterized in that: The method comprises the following steps: (1) The sequencing reads obtained by interrupting and second-generation high-throughput sequencing of the large DNA fragments encoded with watermarks are demapped into double-layer bit sequences according to the "base-bit" mapping rule, and the bit sequences are slidingly correlated with the locally known double-layer watermark sequences to calculate the cross-correlation peaks, thereby determining the source of the sequencing reads and their starting positions on the large DNA fragment sequence; (2) After placing the sequencing reads obtained by sliding correlation recognition at the correct position, a majority voting decision algorithm is used to obtain the consensus base sequence; (3) The consistent base sequence is converted into a damaged bit sequence and then sent to a decoder for error correction and decoding to recover the data; Wherein, the step (1) is: (1.1) The large DNA fragments with watermark superposition encoding are fragmented and sequenced by second-generation high-throughput sequencing to obtain sequencing reads with a length of n base pairs. r According to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)}, demap the double-layer bit sequence of length n bits r 1 and r 2; (1.2) Double-layer bit sequence r 1 and r 2 and a known double-layer watermark sequence with a total length of L bits w 1 and w The bit "0" in 2 is converted to "-1", and the bit "1" remains unchanged; (1.3) Use a sliding window of length n to calculate the double-layer bit sequence r 1 and r 2 and the known double-layer watermark sequence w 1 and w 2, where the cross-correlation value V at the i-th position i The calculation formula is expressed as: Get L-n+1 cross-correlation values; then compare the calculated cross-correlation values ​​of each position to obtain the cross-correlation peak value V peak =max(V0,V1,…V L-n ) and the corresponding positions; (1.4) Using the cross-correlation peak threshold to select the sequencing reads that meet the requirements, and record the peak position, that is, the starting position of the sequencing read on the large DNA sequence; (1.5) The sequencing reads are retrieved according to the cross-correlation peak positions for recovery of different data.

2. According to claim 1, a method for fast reading of watermark superimposed coded large-fragment DNA data storage, characterized in that: The step (2) is: (2.1) If the length of the data DNA sequence is l bases, construct a 5*l counting buffer corresponding to the number of the five bases {A, T, G, C, N} at each position; (2.2) The starting positions of the sequencing reads obtained by sliding correlation recognition are different. The identified sequencing reads are placed at the corresponding starting positions according to the cross-correlation peak positions, and the counting buffers at the corresponding positions are added and counted according to the base types at each position of the reads to obtain a counting buffer containing the number of five bases at each position; (2.3) Perform a base-by-base majority vote on the counting buffer, and determine the base with the largest number in the buffer at the i-th position as the base at that position. If there is no sequencing data at that position, or a majority vote cannot be performed, it is marked as "X", which means an erasure error, and a consensus sequence of length l bases is obtained. s .

3. A method for rapid reading of watermark superimposed coded large-fragment DNA data storage according to claim 2, characterized in that: The step (3) is: (3.1) The consensus sequence of length l bases s According to the "base-bit" mapping rule {A→(00), T→(01), G→(10), C→(11)}, it is demapped into a double-layer bit sequence, and base X is marked as erasure; (3.2) Perform bit-by-bit XOR operation on the fuzzy watermark sequence superimposed with the damaged error correction codeword sequence and the known watermark sequence, remove the superimposed watermark sequence, and obtain the damaged sparse error correction codeword sequence; (3.3) Sparsely decode the damaged sparse error correction codeword sequence to obtain a damaged error correction codeword sequence, and erase the error position for marking; (3.4) The damaged error correction codeword sequence is sent to the corresponding decoder, and the original data is restored after error correction and deletion decoding.

4. According to claim 1, a method for rapid reading of watermark superposition coded large-fragment DNA data storage is characterized in that: The double-layer bit sequence is calculated using a sliding window of length n. r 1 and r 2 and the known double-layer watermark sequence w 1 and w 2, and obtain L-n+1 cross-correlation values; then compare the calculated cross-correlation values ​​of each position to obtain the cross-correlation peak value V peak =max(V0,V1,…V L-n ) and the corresponding position, there are two different sliding correlation calculation cross-correlation peak schemes, namely: (1) After demapping, the sequencing reads are selected one by one from the sequencing data containing a large number of reads, and then bit-by-bit sliding correlation is performed with a known double-layer watermark sequence of length L bits to sequentially obtain the cross-correlation value corresponding to each position; (2) Sequencing data containing a large number of sequencing read segments with a length of n bases are split to obtain multiple sequencing data sub-files; then a known double-layer watermark sequence with a length of L bits is segmented, and n bits of overlapping segments are maintained between adjacent short segments; multiple cross-correlation values ​​are calculated in parallel using a multi-threaded method at the same time, thereby accelerating the sliding correlation calculation process.

5. According to claim 1, a method for rapid reading of watermark superposition coded large-fragment DNA data storage, characterized in that: The method of using the cross-correlation peak threshold to screen the sequencing reads that meet the requirements includes determining the cross-correlation peak threshold as follows according to whether the sequencing data contains non-coding DNA sequencing reads: (1) For application scenarios that do not contain non-coding DNA sequencing reads, sequencing read positioning is achieved directly based on the cross-correlation peak position; (2) For application scenarios containing non-coding DNA sequencing reads, including the second-generation high-throughput sequencing data of the host cell genome and the second-generation high-throughput sequencing data of other coding DNA fragments, different cross-correlation peak thresholds are first set to identify and screen the sequencing reads; then, the identified sequencing reads are merged in large numbers to obtain the corresponding consensus sequence and calculate the base error rate; finally, the threshold corresponding to the minimum base error rate is selected as the optimal cross-correlation peak threshold to reduce interfering reads.

Citation Information

Patent Citations

  • Frame preamble structure design method for power line communication and synchronous detection method and device

    CN103684699A

  • Insertion, deletion and segmentation identification method for long DNA sequence storage

    CN113300720A