A DNA storage method with composite code adjoint addressing encoding

By using the method of composite code accompanying addressing encoding in DNA data storage, the difficulty of efficient addressing and fast readout of DNA data storage is solved, and the rapid and robust recognition of large-scale short-segment DNA data is achieved and fragment loss is tolerated, which improves readout efficiency and reduces computational complexity.

CN119476422BActive Publication Date: 2025-06-13TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411594023.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-06-13
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

Existing DNA data storage technologies have difficulties in efficient addressing and fast readouts, especially in long-fragment DNA storage modes, where high sequencing coverage is required to avoid gaps and improve accurate positioning of the de novo assembly process, resulting in increased computational complexity.

Method used

Using the method of compound code accompanying addressing encoding, multiple pseudo-random sequence logic are logically combined to form a composite code as an address sequence, and a DNA sequence is constructed in combination with the data part. When reading, the sliding correlation detection and Chinese residual calculation are used to quickly identify the unique position mark of the read segment, and correct errors through majority voting judgments and decoders to restore the original user data.

Benefits of technology

It realizes fast and robust recognition of large-scale short-fragment DNA data, tolerate partial loss of fragments, reduces computational complexity, is suitable for short-fragment and long-fragment storage modes, and improves the readout efficiency of DNA data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476422B_ABST
    Figure CN119476422B_ABST
Patent Text Reader

Abstract

The present invention discloses a DNA storage method with composite code adjoint addressing coding. A composite code constructed from multiple pseudo-random sub-codes is used as an address sequence, which is adjointly coded and mapped with an error-corrected coded data bit sequence to combine and obtain long DNA fragments or short oligonucleotide molecules. During reading, the long fragment DNA is randomly fragmented or the oligonucleotide pool is read through next-generation high-throughput sequencing technology and demapped into two-layer bit sequences. Then, multiple pseudo-random sub-codes are used for sliding correlation capture of the sub-code period with the damaged composite code, and the unique address is determined by solving according to the Chinese Remainder Theorem. Finally, consensus and decoding are performed on the damaged data sequence. The method proposed by the present invention mainly solves the problems of large addressing difficulty, high assembly complexity, and DNA breakage when long fragment DNA in next-generation sequencing is randomly fragmented into short fragments, supports fast addressing and recognition of arbitrarily intercepted short fragments in shotgun sequencing of long fragment DNA, is also suitable for robust addressing of short fragments, and tolerates partial deletion of fragments, especially deletion at the beginning and end of the sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of DNA storage, and particularly relates to a DNA storage method with composite code accompanying addressing coding. Background Art

[0002] The rapid development of information technology has accelerated the generation of data. According to the prediction of International Data Corporation (IDC), the total global data volume will reach 175 ZB in 2025. Entering the digital economy era, data, as a new type of production factor, plays an increasingly important role in social production activities. According to the "Semiconductor Decadal Plan" released by the Semiconductor Industry Association (SIA) and Semiconductor Research Corporation (SRC) of the United States in 2021, more than 60% of all stored data will become cold data that does not need to be frequently accessed. According to the report released by Horizons, the value of data has the characteristic that its value becomes important again as time goes by. With the development of artificial intelligence technology based on big data, the value of long-term data storage has become increasingly prominent. Long-term data storage has brought the demand for long-term storage of massive data. However, the service life of existing data storage methods based on magnetism, optics, and electricity generally does not exceed several decades, the maintenance cost is relatively high, it is difficult to improve the density, and it is difficult to meet the demand for data storage with explosive growth.

[0003] DNA data storage uses deoxyribonucleic acid (DNA) as a new type of data storage medium, which has core advantages such as high storage density, long available time, and low maintenance cost. It is a very promising new model for long-term stable storage of massive data, especially with broad application prospects in the long-term storage of archived data. According to the "Semiconductor Synthetic Biology Roadmap" released by the National Institute of Standards and Technology (NIST) and SRC of the United States, compared with the high-capacity magnetic tape with the highest current storage density, the DNA data storage method has an improvement of about 7 orders of magnitude; according to the prediction of the Intelligence Advanced Research Projects Agency (IARPA) of the United States, it is expected that the DNA storage in an exabyte-level data center can reduce the power consumption from 200 MW to 200 KW. At the same time, DNA data storage can achieve offline storage, has high security, is easy to backup and store, and is expected to become a potential solution for long-term stable storage of future massive archived data. DNA data storage is a very cutting-edge emerging technology. In 2023, the UK Research and Innovation Agency listed DNA data storage as one of the 50 emerging technologies with great potential in the future, and it is one of the 8 technologies in the "Artificial Intelligence, Digital and Computing Technologies" direction. In 2023, the "International Device and System Roadmap" released by the Institute of Electrical and Electronics Engineers (IEEE) and others all regard DNA storage as one of the main storage media for future massive data.

[0004] In recent years, a series of breakthroughs have been made in proof-of-concept research on DNA data storage. Currently, the mainstream DNA data storage modes mainly include short-fragment DNA storage mode and long-fragment DNA storage mode. The short-fragment DNA storage mode stores data using an oligonucleotide pool composed of a large number of short-chain DNA molecules (usually 100-300 bases), and the reading-out is usually achieved by means of next-generation high-throughput sequencing. The reading-out of data in the long-fragment DNA storage mode can be realized by means of genome sequencing technologies similar to those used in genome sequencing, including next-generation high-throughput sequencing and third-generation nanopore sequencing.

[0005] Regardless of which mode, DNA data storage adopts the method of dispersing information into a large number of short DNA fragments for storage, and restoring the correct order of all fragments is the key to reliable data recovery. Therefore, how to achieve efficient addressing and fast reading-out of DNA-stored data remains a key technical problem that urgently needs to be solved in DNA data storage.

[0006] For the long-fragment DNA data storage mode, if next-generation high-throughput sequencing is used for reading-out, similar to genome sequencing, the enriched DNA is usually randomly fragmented into a large number of short fragments by "shotgun" sequencing, and the original data sequence is reconstructed by means of de novo assembly strategies. However, the fragmentation positions of shotgun sequencing and the lengths of the generated short DNA fragments are random, resulting in difficult addressing and the need to adopt complex graph methods based on overlapping relationships. To improve the accurate positioning of contigs during de novo assembly and avoid gaps as much as possible, the sequencing coverage required for data recovery is usually relatively high. Currently, the software for assembling short fragments by next-generation sequencing is mainly based on the de Bruijn graph algorithm, and a high sequencing coverage will lead to an increase in the number of k-mer nodes during graph construction, further increasing the computational complexity. Researchers have respectively used the de Bruijn graph-based assembly software Velvet and AbySS to achieve de novo assembly of next-generation sequencing reads, and with the assistance of a guiding sequence, the payload data is restored by means of sequence alignment algorithms. The sequencing coverage required to achieve error-free data recovery using these two assembly software exceeds 20×. Researchers use the minimum sequencing coverage required for assembly (usually at least not less than 20×) to perform de novo assembly and recovery of long-fragment DNA data. Researchers use the graph-based short-sequence assembly software SOAPdenovo2 for de novo assembly, and combine contig alignment and consensus to achieve data sequence reconstruction, but the processing complexity is very high.

[0007] For the short - fragment DNA data storage mode, that is, the storage mode of short - fragment DNA oligonucleotides with regular segmentation, a large number of oligonucleotide molecules have disorder characteristics, and the order during reading is random. Therefore, generally, an additional small segment of base sequence (index) needs to be attached to each oligonucleotide molecule to identify the position of the data carried by the molecule in the whole data. At the same time, during the actual DNA storage and reading process, there is a risk that DNA molecules degrade, resulting in a decrease in quality, and it is extremely easy to hydrolyze in an aqueous solution, leading to frequent DNA strand breaks, thus threatening the reliability of DNA data storage. For the problem of DNA strand breaks, researchers have proposed a sequence reconstruction and error - correction algorithm based on the de Bruijn graph theory, and verified the high robustness of data recovery through accelerated aging experiments of sample breakage and degradation. However, the sequence reconstruction algorithms based on the de Bruijn graph generally rely on high - sequencing coverage, have high computational complexity, and are difficult to meet the requirements of rapid data reading. Summary of the Invention

[0008] To overcome the deficiencies of the prior art, the present invention provides a DNA storage method with composite - code adjoint addressing encoding. The present invention realizes the rapid and robust recognition of large - scale short - fragment data DNA at arbitrary positions with very low complexity, tolerates partial deletions within the fragments, and this method is applicable to both short - fragment and long - fragment storage modes, as described in detail below:

[0009] A DNA storage method with composite - code adjoint addressing encoding, the method comprising the following steps:

[0010] (1) At the data - writing end, select n pseudo - random sequences as sub - codes, where n is a positive integer, and construct a composite code as the address sequence through combinatorial logic. Encode the user data using error - correction encoding to obtain an encoded data bit sequence, then perform adjoint encoding by bit - by - bit combination of the composite code and the data bit sequence to form a double - layer bit sequence, and transcode it into a data DNA sequence according to the mapping rule between bit pairs and bases. Finally, complete data writing by artificially synthesizing DNA;

[0011] (2) Data readout terminal: Obtain sequencing reads of short DNA fragments through the second-generation high-throughput sequencing technology, preprocess the sequencing reads, perform demapping according to the same mapping rules as the write terminal, extract the damaged composite code and the damaged data sequence. Secondly, use the known n sub-codes to perform sliding correlation detection with the damaged composite code respectively to obtain the correlation peak phase estimation of each sub-code. Then, perform calculation through the Chinese Remainder Theorem to determine the unique position label of the read segment. Next, calculate the in-phase autocorrelation value of the composite code and perform threshold verification. Keep the read segments that pass the verification unchanged, while for the read segments that fail the verification, re-identify them in enhanced mode, and perform insertion / deletion error location and correction on the damaged data sequence. Finally, according to the identified position label, use multiple copies of the sequences at the same position to perform bit-by-bit majority voting decision to generate a consistent data sequence, and send it to the decoder for decoding to restore the original user data.

[0012] The specific steps of step (1) are as follows:

[0013] (1.1) Select n pseudo-random sequences as sub-codes to form a sub-code set {C 1 , C 2 , …, C n}. Then, perform periodic extension on each sub-code respectively to obtain sequences with a length of P = p 1 p 2 …p n . And calculate bit by bit according to the preset Boolean combination logic to construct a composite code a with a length of P = p 1 p 2 …p n , and represent it in binary form. Among them, C n represents sub-code n, and its code length is p n , n ≥ 1 and is an integer;

[0014] (1.2) Use m linear block codes (N, K) to perform block error correction coding on user data to obtain an encoded data bit sequence b. Among them, m is the number of codewords, m ≥ 1 and is an integer, N is the length of the encoded codeword, and K is the length of the information bit. Since the composite code period can be designed to be very long, a single composite code can carry multiple block code codewords, satisfying mN ≤ p 1 p 2 …p n ;

[0015] (1.3) Perform adjoint coding on the composite code a and the data bit sequence b by bit by bit combination to construct a double-layer bit sequence, and transcode it into a DNA sequence according to the mapping rule between bit pairs and bases {00 → A, 01 → T, 10 → G, 11 → C}, and obtain long DNA fragments or short oligonucleotide molecules with the help of artificial synthesis technology;

[0016] Among them, the preset Boolean combination logic specifically includes: logical multiplication, modulo-two sum, and majority decision;

[0017] Among them, n pseudo-random sequences with short periods are selected as sub-codes, satisfying: the periods of the selected pseudo-random sequences are strictly co-prime to each other, and each has a sharp periodic autocorrelation characteristic;

[0018] Among them, the step (1.3) is specifically:

[0019] For long-fragment DNA data storage, first intercept the composite code to obtain a shortened composite code with the same length as the data bit sequence, then align the shortened composite code with the data bit sequence head to tail to form a double-layer bit sequence, and directly perform mapping to obtain a long DNA fragment; for short-fragment oligonucleotide data storage, first divide the data bit sequence into mN / l groups evenly according to the length l, and at the same time perform sliding interception on the composite code according to the step length l to obtain mN / l shortened composite codes, then pair the shortened composite codes with a length of l with the data bit sequence in sequence to construct mN / l groups of double-layer bit sequences, and after mapping, finally generate mN / l short-fragment oligonucleotides; where l is a positive integer.

[0020] The step (2) is specifically:

[0021] (2.1) For the short-fragment oligonucleotide molecule pool, directly construct a library and then read it through the second-generation high-throughput sequencing technology to obtain the sequencing reads of the data DNA; for the long-fragment data DNA, perform shotgun sequencing by high-throughput sequencing. Specifically, randomly fragment the enriched DNA molecules and read them through the second-generation high-throughput sequencing to obtain the sequencing reads;

[0022] (2.2) Preprocess the sequencing reads, perform demapping according to the same mapping rule as the writing end, and extract the damaged composite code and the damaged data bit sequence Specifically, for the sequencing reads of short-fragment oligonucleotide molecules, trim the primer parts at both ends by aligning with the amplification primers, and perform demapping on the remaining payload part; for the reads obtained by shotgun sequencing of long-fragment DNA, directly perform demapping on the sequencing reads;

[0023] (2.3) Send the damaged composite code into the cross-correlation detectors of each sub-code, use the known sub-code set {C 1 , C 2 , …, C n} for sliding correlation capture, search for the positions of the correlation peaks of each sub-code (i.e., the corresponding phases of in-phase correlation), obtain the estimated phases of each sub-code, and then solve through the Chinese Remainder Theorem to determine the unique position label of the read;

[0024] (2.4) Estimate the phase using each of the sub - codes obtained above, regenerate an ideal composite code a' that has the same length and is in - phase as the damaged composite code, calculate the autocorrelation value of the composite code. If the autocorrelation value meets the preset conditions, retain the read segment; for the read segments that fail the verification, in the enhanced mode, use the method based on edit - distance comparison for re - identification, determine the position label, and retain the read segment. At the same time, identify and correct the insertion and deletion error patterns according to the alignment path.

[0025] (2.5) Considering that there may be multiple copies of the data sequence at the same position, perform majority - vote decision on each bit of the data sequence to obtain a consistent data sequence. If the vote decision is not unique, use the same erasure marker as above to fill the bit at that position to obtain a codeword, and then send it to the decoder to correct the residual errors and recover the original user data.

[0026] The specific steps of step (2.3) are as follows:

[0027] (3.1) Perform periodic extension on each sub - code in the known sub - code set. Slide - correlate each sub - code with the damaged composite code. The length of the correlation window is the same as the length of the damaged composite code, and the sliding step is the sub - code period. The cross - correlation function between sub - code i and the damaged composite code can be expressed as:

[0028]

[0029] Among them, is the periodic extension sequence of sub - code i, is the damaged composite code, τ is the sub - code phase offset, L is the length of the correlation window, 1 ≤ i ≤ n, and the periodic extension sequence and the composite code are represented by the symbols {+1, - 1}, and the mapping relationship {+1→0, - 1→1} is satisfied between them and the binary - form sequence;

[0030] (3.2) According to the sliding - correlation values between each sub - code and the damaged composite code, search for the position where the correlation peak is located to obtain the phase estimation of each sub - code. The phase τ corresponding to the correlation peak of sub - code i i can be expressed as:

[0031]

[0032] Among them, argmax represents obtaining the position corresponding to the correlation peak;

[0033] (3.3) Use the estimated phases of each sub - code obtained by the search to solve the unique position label of the corresponding read segment through the Chinese Remainder Theorem. The estimated phases of each sub - code satisfy

[0034]

[0035] Among them, is the unique position label of the decoded read segment, τ 1 , τ 2 , …, τ n respectively correspond to the estimated phases of n sub - codes.

[0036] The specific steps of step (2.4) are as follows:

[0037] (4.1) According to the estimated phases of the above - mentioned sub - codes, regenerate an ideal composite code a′ with the same phase and the same length as the damaged composite code , and then calculate the in - phase correlation value between the damaged composite code and the ideal composite code a′. The obtained correlation value can be expressed as:

[0038]

[0039] Among them, r and r′ are the damaged composite code and the ideal composite code respectively, both are represented by the symbols {+1, - 1}, and satisfy the mapping relationship {+1→0, - 1→1} with the binary - form sequence;

[0040] (4.2) If the in - phase autocorrelation value of the composite code satisfies R r,r′ ≥ωL, the check passes and the current read segment is retained. Where ω represents a preset correlation threshold, 0≤ω≤1.

[0041] (4.3) For the read segments that fail the above autocorrelation check, in the enhanced mode, use the method based on edit - distance comparison to compare the damaged composite code with the known complete composite code a, search for the alignment position where the minimum edit distance is located as the position label of the read segment. At the same time, use the alignment path to locate the positions where insertion and deletion errors occur in the damaged composite code, and according to the characteristic that the positions of insertion / deletion errors in the double - layer sequence structure are exactly the same, correct the insertion and deletion error patterns in the damaged data sequence; discard the bits at the positions where insertion errors occur, mark the bits at the positions where deletion errors occur as erased, and fill them with determined marks; if the enhanced mode is not adopted, skip the current step.

[0042] Compared with the prior art, the beneficial effects of the technical solution provided by the present invention are:

[0043] (1) The present invention makes full use of the single - cycle ambiguity - free characteristic of the composite code and the good cross - correlation characteristic between the sub - code and the composite code, uses the composite code generated by multiple short - cycle sub - codes as the address sequence, and performs adjoint coding with the original data sequence to form a DNA sequence; when reading data, perform sliding correlation detection with a finite number of sub - code cycle lengths, and then use the Chinese Remainder Theorem to calculate the actual position of the read segment, thereby quickly identifying the unique position label of the sequencing read segment, and significantly improving the addressing efficiency of large - scale data storage read - out and recovery.

[0044] (2) The present invention is applicable to the storage mode of long DNA fragments and short oligonucleotide pools, supports robust identification of DNA fragments at any position, and avoids the complex assembly operations required for sequence reconstruction in long - fragment shotgun sequencing; for oligonucleotide pools with a fixed starting position, partial deletions within the fragment are allowed, thereby improving the read - out efficiency of the DNA oligonucleotide pool. Brief Description of the Drawings

[0045] Figure 1 It is a system implementation block diagram of a DNA storage method with composite - code adjoint addressing coding provided by the present invention;

[0046] Figure 2 It is an implementation flow chart of composite - code adjoint addressing coding for constructing a long DNA fragment or short oligonucleotide pool storage mode in the present invention;

[0047] Figure 3 It is a flow block diagram for read - segment position label recognition and data recovery for second - generation sequencing read - out provided by the present invention;

[0048] Figure 4 It is a composite - code construction flow chart provided by the present invention;

[0049] Figure 5 It is an implementation flow chart for quickly identifying read segments based on sliding correlation detection of short - cycle sub - codes in the present invention;

[0050] Figure 6 It is a performance comparison graph of read - segment recognition success rate and available read - segment ratio under different normalization thresholds in Specific Embodiment 1 of the present invention;

[0051] Figure 7 It is a performance graph of fragment recognition success rate when Specific Embodiment 1 of the present invention adopts the enhanced mode at different autocorrelation normalization thresholds;

[0052] Figure 8 It is an error - rate performance graph after multi - copy merging under different sequencing coverages in Specific Embodiment 1 of the present invention;

[0053] Figure 9 It is a time - complexity comparison graph of Specific Embodiment 1 of the present invention;

[0054] Figure 10 It is a performance comparison chart of the read segment recognition success rate and the proportion of available read segments of Specific Embodiment 2 of the present invention under different normalization thresholds;

[0055] Figure 11 It is an error rate performance chart after multi-copy merging of Specific Embodiment 2 of the present invention under different sequencing coverages;

[0056] Figure 12 It is the error-free recovery situation of Specific Embodiment 3 of the present invention under different sequencing coverages with 50 groups of repeated experiments respectively. Detailed implementation manners

[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.

[0058] Aiming at the problem of low sequence addressing efficiency in the application scenarios of short and long DNA fragments, the present invention provides a new encoding method that is generally applicable to short DNA fragment data storage and long DNA fragment data storage, a DNA storage method of composite code accompanied addressing encoding. This method constructs a DNA sequence by combining the segments of a composite code formed by logically combining multiple pseudo-random sequences with the data part. When reading, it can quickly and efficiently address both DNA sequencing read segments starting from any position (shotgun long fragment sequencing mode) and starting from a fixed position (oligonucleotide pool sequencing readout mode), realizing the rapid recovery of data, and can cope with partial deletions of fixed segmented short DNA molecular sequences, thereby improving the reading efficiency.

[0059] The following further describes the embodiments of the present invention in detail with reference to the accompanying drawings:

[0060] This embodiment details a DNA storage method of composite code accompanied addressing encoding proposed by the present invention. For the overall system implementation block diagram, see Figure 1 which specifically includes the following steps:

[0061] (1) At the data writing end, see Figure 2 , select n pseudo-random sequences as sub-codes, where n is a positive integer, and construct a composite code as the address sequence through combinatorial logic. Encode the user data using error correction coding to obtain the encoded data bit sequence. Then perform an accompanying encoding by combining the composite code and the data bit sequence bit by bit to form a double-layer bit sequence. According to the mapping rule between bit pairs and bases, transcode to obtain the data DNA sequence, and finally complete the data writing by artificially synthesizing DNA;

[0062] (2) At the data reading end, see Figure 3, obtaining sequencing reads of short DNA fragments by means of second-generation high-throughput sequencing technology, preprocessing the sequencing reads, demapping according to the same mapping rules as the writing end, extracting damaged composite codes and damaged data sequences. Secondly, using n known subcodes to perform sliding correlation detection with the damaged composite codes respectively to obtain the correlation peak phase estimation of each subcode. Then, performing calculation through the Chinese Remainder Theorem to determine the unique position label of the read. Next, calculating the in-phase autocorrelation value of the composite code and performing threshold verification. For the reads that pass the verification, they remain unchanged, while for the reads that fail the verification, they are re-identified in enhanced mode, and the damaged data sequences are subjected to insertion / deletion error location and correction. Finally, according to the identified position label, multiple copies of the sequences at the same position are used for bit-by-bit majority voting decision to generate a consistent data sequence, which is sent to the decoder for decoding to restore the original user data.

[0063] Among them, step (1) is specifically as follows:

[0064] (1.1) As Figure 4 shown, select n pseudo-random sequences as subcodes to form a subcode set {C 1 , C 2 , …, C n}, and then perform periodic extension on each subcode respectively to obtain sequences with a length of P = p 1 p 2 …p n , and calculate bit by bit according to the preset Boolean combination logic to construct a composite code a with a length of P = p 1 p 2 …p n , and represent it in binary form; among them, C n represents subcode n, whose code length is p n , n ≥ 1 and is an integer;

[0065] (1.2) Use m linear block codes (N, K) to perform block error correction coding on user data to obtain an encoded data bit sequence b; where m is the number of codewords, m ≥ 1 and is an integer, N is the length of the encoded codeword, and K is the length of the information bit; since the composite code period can be designed to be very long, a single composite code can carry multiple block code codewords, satisfying mN ≤ p 1 p 2 …p n ;

[0066] (1.3) Perform adjoint coding on the composite code a and the data bit sequence b by bit by bit combination to construct a double-layer bit sequence, and transcode it into a DNA sequence according to the mapping rule of bit pairs and bases {00→A, 01→T, 10→G, 11→C}, and obtain a long DNA fragment or short oligonucleotide molecule by means of artificial synthesis technology;

[0067] Among them, the preset Boolean combinatorial logic specifically includes: logical multiplication, modulo-two sum, majority decision, etc.;

[0068] Among them, n pseudo-random sequences with short periods are selected as sub-codes, satisfying that the periods of the selected pseudo-random sequences are strictly relatively prime to each other, and each has a sharp periodic autocorrelation characteristic;

[0069] Among them, step (1.3) is specifically:

[0070] For long-fragment DNA data storage, first intercept the composite code to obtain a shortened composite code with the same length as the data bit sequence, then align the shortened composite code with the data bit sequence head-to-tail to form a double-layer bit sequence, and directly perform mapping to obtain a long DNA fragment; for short-fragment oligonucleotide data storage, first evenly divide the data bit sequence into mN / l groups according to the length l, and at the same time perform sliding interception on the composite code according to the step length l to obtain mN / l shortened composite codes, then pair the shortened composite codes with the data bit sequence of length l in sequence to construct mN / l groups of double-layer bit sequences, and after mapping, finally generate mN / l short-fragment oligonucleotides; where l is a positive integer.

[0071] Among them, step (2) is specifically:

[0072] (2.1) For the short-fragment oligonucleotide molecule pool, directly construct a library and then read it through the second-generation high-throughput sequencing technology to obtain the sequencing reads of the data DNA; for the long-fragment data DNA, use the high-throughput shotgun sequencing technology for sequencing readout. Specifically, randomly fragment the enriched DNA molecules and read them through the second-generation high-throughput sequencing technology to obtain the sequencing reads;

[0073] (2.2) Preprocess the sequencing reads, perform demapping according to the same mapping rule as the writing end, and respectively extract the damaged composite code and the damaged data bit sequence Specifically, for the sequencing reads of short-fragment oligonucleotide molecules, trim off the primer parts at both ends by aligning with the amplification primers, and perform demapping on the remaining payload part; for the reads obtained by shotgun sequencing of long-fragment DNA, directly perform demapping on the sequencing reads;

[0074] (2.3) As Figure 5 shown, send the damaged composite code to the cross-correlation detectors of each sub-code, and utilize the known sub-code set {C 1 , C 2 , …, C n}Perform sliding-related capture, search for the positions of the correlation peaks of each sub-code (i.e., the in-phase correlation corresponding phases), obtain the estimated phases of each sub-code, and then solve through the Chinese Remainder Theorem to determine the unique position label of the read segment;

[0075] (2.4) Utilize the estimated phases of each sub-code obtained above to regenerate an ideal composite code a' that has the same length and is in-phase with the damaged composite code, calculate the autocorrelation value of the composite code. If the autocorrelation value meets the preset conditions, retain the read segment; for the read segments that fail the verification, in the enhanced mode, use the method based on edit distance comparison for re-identification, determine the position label, and retain the read segment. At the same time, identify and correct the insertion and deletion error patterns according to the alignment path;

[0076] (2.5) Considering that there may be multiple copies of the data sequence at the same position, perform majority voting decision bit by bit on the data sequence to obtain a consistent data sequence. If the voting decision does not have uniqueness, use the same erasure marker as above to fill the bits at that position to obtain a codeword, and then send it to the decoder to correct the residual errors and restore the original user data.

[0077] Among them, step (2.3) is specifically as follows:

[0078] (3.1) Perform periodic extension on each sub-code in the known sub-code set. Each sub-code is respectively subjected to a sliding correlation operation with the damaged composite code. The length of the correlation window is the same as the length of the damaged composite code, and the sliding step is the sub-code period. The cross-correlation function between sub-code i and the damaged composite code can be expressed as:

[0079]

[0080] Among them, is the periodic extension sequence of sub-code i, is the damaged composite code, τ is the sub-code phase offset, L is the length of the correlation window, 1 ≤ i ≤ n, and the periodic extension sequence and the composite code are represented by the symbols {+1, -1}, and they satisfy the mapping relationship {+1 → 0, -1 → 1} with the binary form sequence;

[0081] (3.2) According to the sliding correlation values between each sub-code and the damaged composite code, search for the positions of the correlation peaks to obtain the phase estimates of each sub-code. The phase τ corresponding to the correlation peak of sub-code i i can be expressed as:

[0082]

[0083] Among them, argmax represents obtaining the position corresponding to the correlation peak;

[0084] (3.3) Estimate the phase using each sub - code obtained by searching, and solve through the Chinese Remainder Theorem to obtain the unique position label of the corresponding read segment. The estimated phases of each sub - code satisfy:

[0085]

[0086] Among them, is the unique position label of the read segment obtained by solving, and τ 1 、τ 2 、…、τ n correspond to the estimated phases of n sub - codes respectively.

[0087] Among them, step (2.4) is specifically:

[0088] (4.1) According to the estimated phases of the above - mentioned each sub - code, regenerate an ideal composite code a′ with the same phase and the same length as the damaged composite code , and then calculate the in - phase correlation value between the damaged composite code and the ideal composite code a′. The obtained correlation value can be expressed as:

[0089]

[0090] Among them, r′ are the damaged composite code and the ideal composite code respectively, both are represented by the symbols {+1, - 1}, and satisfy the mapping relationship {+1→0, - 1→1} with the binary - form sequence;

[0091] (4.2) If the in - phase autocorrelation value of the composite code satisfies , then the check passes, and the current read segment is retained. Where ω represents the preset correlation threshold, 0 ≤ ω ≤ 1.

[0092] (4.3) For the read segments that fail the above - mentioned autocorrelation check, in the enhanced mode, use the method based on edit - distance comparison to compare the damaged composite code with the known complete composite code a, search for the alignment position where the minimum edit distance is located as the position label of the read segment. At the same time, use the alignment path to locate the positions where insertion and deletion errors occur in the damaged composite code, and according to the characteristic that the positions of insertion / deletion errors in the double - layer sequence structure are exactly the same, correct the insertion and deletion error patterns in the damaged data sequence; specifically, discard the bits at the positions where insertion errors occur, mark the bits at the positions where deletion errors occur as erased, and fill them with determined marks; if the enhanced mode is not adopted, skip the current step.

[0093] A DNA storage method with composite - code - accompanied addressing coding proposed by the present invention utilizes the good correlation characteristics of the composite code and the distinguishable characteristics of any phase within a single period. The composite code constructed from multiple short - period pseudo - random sub - codes is used as the address sequence, and combined with the data sequence to form a double - layer sequence structure, which is mapped to obtain the DNA sequence. When reading the data, combined with the Chinese Remainder Theorem, only a finite number of sub - codes are needed for the sliding correlation detection of the sub - code period to achieve rapid read - segment recognition, supporting robust recognition of DNA fragments at any position, and tolerating partial deletions within the fragment, which is generally applicable to the storage modes of long DNA fragments and short - fragment oligonucleotide pools for next - generation sequencing readout.

[0094] The following gives two specific embodiments, and in combination with the attached drawings, illustrates the feasibility of the present invention, and further understands the purpose, features and advantages of the invention. In the three specific embodiments given, the Boolean combination logic for constructing the composite code selects majority decision, and its expression can be represented as where represents the symbol sequence corresponding to sub - code n, which is composed of {+1, - 1}, and satisfies the mapping relationship {-1→1, +1→0} with the bit sequence. The number of sub - codes n satisfies n>1 and is odd.

[0095] Embodiment 1

[0096] In this embodiment, the composite code generated by constructing with 5 sub - codes is used as the address sequence of the data DNA. By simulating the recognition performance of sequencing reads with a length of 130 bases, the low complexity and high robustness of the addressing method proposed by the present invention are verified. Select 5 pseudo - random sequences with periods of 31, 35, 43, 47, and 59 respectively to form a sub - code set. Among them, the sub - code with a period of 31 is an m - sequence, the sub - code with a period of 35 is a twin - prime sequence, and the sub - codes with periods of 43, 47, and 59 are Legendre sequences. Table 1 gives the bit sequences corresponding to each sub - code.

[0097] Table 1 Five pseudo - random sub - codes constituting the composite code

[0098]

[0099] The specific implementation of the DNA storage method with composite - code - accompanied addressing coding includes the following steps:

[0100] (1) Perform period extension on the selected 5 sub - codes respectively to obtain the period - extended sequences of each sub - code, with a length of 129374315. Then select majority decision as the Boolean combination logic for constructing the composite code, and perform bit - by - bit operations to obtain a composite code with a period P = 129374315, and represent it in binary form;

[0101] (2) Determine the starting phase of the composite code, then segment the data sequence and the composite code into groups with a length of 130 bits, obtaining 995,187 shortened composite codes and data sequences respectively. After pairwise pairing in sequence, a double-layer bit sequence is formed;

[0102] (3) According to the mapping rule {00→A, 01→T, 10→G, 11→C}, convert the bit pairs into bases, and finally obtain 995,187 DNA sequences with a length L = 130 bases;

[0103] (4) Add random errors to each sequence to simulate the generation of simulated sequencing reads, where the insertion (Ins.), deletion (Del.), and substitution (Sub.) error rates satisfy P Sub. = 2P Ins. = 2P Del. = 0.0050, and perform Monte Carlo repeated trials;

[0104] (5) Directly demap the sequencing reads to obtain the damaged composite code and data bit sequence;

[0105] (6) Use the known 5 sub-codes to perform sliding correlation detection with the damaged composite code respectively, and a total of 31 + 35 + 43 + 47 + 59 = 215 sub-code correlation operations are executed. Search for the estimated phase where the correlation peak of each sub-code is located, and then solve the unique position label of the current read segment through the Chinese Remainder Theorem;

[0106] (7) Use the known sub-codes and the estimated phases of each identified sub-code to regenerate an ideal composite code with a length of 130 bits, and perform in-phase autocorrelation operation with the damaged composite code. Compare the normalized autocorrelation value with the preset threshold, and the threshold is set to 0.8. If the autocorrelation value is greater than or equal to the preset threshold, the position label recognition is valid, otherwise discard the read segment; in the enhanced mode, retain the read segments with autocorrelation values less than the preset threshold, and use the method based on edit distance alignment to align the damaged composite code with the known complete composite code, search for the alignment position corresponding to the minimum edit distance as the position label, and identify and correct the insertion and deletion error patterns according to the alignment path;

[0107] (8) According to the identified data position information, perform consensus and decoding on the damaged data sequence to recover the original data.

[0108] Figure 6 The performance comparison of the read segment recognition accuracy and the proportion of available read segments under different thresholds is given. The simulation results show that when the number of sub-codes is 5 and the verification threshold is set to 0.8, the read segment recognition accuracy can reach more than 98%, and the proportion of available read segments is nearly 80%, verifying that the method proposed in the present invention can achieve high-robust recognition in the scenario including insertion / deletion errors.

[0109] Figure 7 It shows the proportion of correctly identified reads in the enhanced mode. The simulation results indicate that when the threshold is set to 0.8 or above, the proportion of correctly identified reads among all reads increases to 97.69%, demonstrating that the enhanced mode of the method proposed in the present invention can effectively increase the number of available valid reads while ensuring the accuracy of read identification.

[0110] Figure 8 It shows the bit error rate performance of the method proposed in the present invention under different sequencing coverages. The simulation results indicate that after setting threshold verification, the bit error rate of the data after multi-copy merging can be significantly improved. When the enhanced mode is adopted, as the sequencing coverage increases, the error rate performance is further improved, approaching the method based on reference sequence alignment.

[0111] Figure 9 It shows the comparison of time complexity performance. The running time is identified and counted for 1,399,667 reads with a length of 130 bases. The results indicate that the processing speed of the method proposed in the present invention is approximately 230 times faster than that of the method based on reference sequence alignment, and the read identification efficiency is significantly improved. After adopting the enhanced mode, the processing speed is also approximately 5 times faster than that of the method based on reference sequence alignment, verifying the low implementation complexity of the method proposed in the present invention.

[0112] Example 2

[0113] In this example, a composite code generated by constructing 3 sub-codes is used as the address sequence of the data DNA. By simulating the identification performance of 130-base-long sequencing reads, it is further verified that the addressing method proposed in the present invention still has low complexity and high robustness when the number of sub-codes is 3. Three pseudo-random sequences with periods of 499, 503, and 511 are selected to form a sub-code set. Among them, the sub-codes with periods of 499 and 503 are Legendre sequences, and the sub-code with a period of 511 is an m-sequence. Table 1 shows the bit sequences corresponding to each sub-code.

[0114] Table 2 Three pseudo-random sub-codes constituting the composite code

[0115]

[0116]

[0117] First, the three selected sub - codes are respectively subjected to periodic extension, and the combined logic of majority decision is adopted to construct a composite code with a period P = 128259467, which is represented in binary form. Secondly, the data sequence and the composite code are segmented into groups with a length of 130 bits, respectively obtaining 986611 shortened composite codes and data sequences. After pairing them in turn, a double - layer bit sequence is formed, and 986611 DNA fragments with a length L = 130 bases are obtained after mapping.

[0118] Further, the sequencing read - out process is simulated. Random errors are added to each sequence to generate simulated sequencing reads, where the insertion (Ins.), deletion (Del.), and substitution (Sub.) error rates satisfy P Sub. = 2P Ins. = 2P Del. = 0.0050.

[0119] The specific implementation of data read - out refers to steps (5) - (8) of Embodiment 1, with the difference that in this embodiment, 3 known sub - codes are used for sliding correlation detection.

[0120] Figure 10 The performance comparison of the read - segment recognition accuracy and the proportion of available read - segments under different thresholds is given. The simulation results show that when the number of sub - codes is 3 and the verification threshold is set to 0.8, the read - segment recognition accuracy can reach more than 98%, and the proportion of available read - segments is nearly 80%, verifying that the method proposed in the present invention can still achieve high - robustness recognition in the scenario of insertion / deletion errors when the number of sub - codes is 3.

[0121] Figure 11 The bit - error - rate performance of the method proposed in the present invention under different sequencing coverages is given. The simulation results show that after setting the threshold verification, when the sequencing coverage is about 5×, the error - rate performance of the non - enhanced mode and the enhanced mode is basically consistent with the method based on reference - sequence alignment, verifying that the proposed method can obtain performance similar to that of the method based on reference - sequence alignment when the number of sub - codes is 3.

[0122] Embodiment 3

[0123] This embodiment demonstrates the read - out recovery of long - fragment DNA for next - generation sequencing. A sub - code set composed of 5 sub - codes in Embodiment 1 is selected, and the block error - correcting code uses a binary LDPC(22680,7560) code with a code rate R = 1 / 3.

[0124] For the long - fragment DNA storage mode, the specific implementation of the composite - code - accompanied addressing coding is as follows:

[0125] The compound code construction process is the same as that in Embodiment 1, obtaining a compound code with a period P = 129374315. The user data is encoded using a binary LDPC(22680,7560) code to obtain a codeword with a length of 22680 bits, which forms a data sequence. The starting phase of the compound code is determined and the compound code is shortened to obtain a shortened compound code with a length of 22680 bits. Then, it forms a double-layer bit sequence with the data sequence and is transcoded according to the mapping rule between bits and bases {00→A, 01→T, 10→G, 11→C}, finally obtaining a single long DNA fragment with a length of 22680 bases.

[0126] The process of randomly fragmenting and sequencing the long DNA fragment is simulated using the second-generation genomic sequencing data simulation software ART. The single-end 150 sequencing mode is set, and the insertion, deletion, and substitution error rates satisfy P Sub. = 2P Ins. = 2P Del. = 0.0050, thus obtaining short read sequencing reads with a length of 150 bases. The specific operations for readout recovery are consistent with steps (5)-(8) in Embodiment 1.

[0127] Figure 12 The error-free recovery situations of 50 groups of repeated experiments are given under different sequencing coverages. The results show that the method proposed by the present invention can achieve error-free recovery at a sequencing coverage of 2.6× (a well-known term in the art). After adopting the enhanced mode, it can achieve error-free recovery at a sequencing coverage of 2.2×, and has the same performance as the method based on reference sequence alignment, proving that the method proposed by the present invention can achieve complete error-free recovery of the original data at a lower sequencing coverage.

[0128] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment. The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0129] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A DNA storage method with composite code and addressing encoding, characterized in that: The method comprises the following steps: (1) At the data writing end, n pseudo-random sequences are selected as subcodes, where n is a positive integer, and a composite code is obtained as an address sequence through combinatorial logic construction. The user data is encoded using error correction coding to obtain an encoded data bit sequence. The composite code and the data bit sequence are then concomitantly encoded bit by bit to form a double-layer bit sequence. The data DNA sequence is then transcoded according to the mapping rule between bit pairs and bases. Finally, the data writing is completed through artificially synthesized DNA. (2) At the data reading end, the sequencing reads of short DNA fragments are obtained by means of second-generation high-throughput sequencing. The sequencing reads are preprocessed and demapped according to the same mapping rules as the writing end to extract the damaged composite code and the damaged data sequence. Then, the known n subcodes are used to perform sliding correlation detection with the damaged composite code to obtain the correlation peak phase estimation of each subcode. Then, the Chinese remainder theorem is used to solve the problem and determine the unique position label of the read. Then, the in-phase autocorrelation value of the composite code is calculated and a threshold check is performed. The reads that pass the check are kept as they are, while the reads that fail the check are re-identified using the enhanced mode, and the insertion / deletion errors of the damaged data sequence are located and corrected. Finally, according to the identified position label, multiple copies of the same position sequence are used to perform a bit-by-bit majority voting decision to generate a consistent data sequence, which is sent to the decoder for decoding to restore the original user data.

2. The DNA storage method with composite code and addressing encoding according to claim 1, characterized in that: The step (1) is specifically: (1.1) Select n pseudo-random sequences as subcodes to form a subcode set {C1, C2, …, C n }, and then extend the period of each subcode to obtain a length of P = p1p2…p n The sequence is constructed by calculating the length of P = p1p2…p bit by bit according to the preset Boolean combination logic. n The composite code a is expressed in binary form; where C n represents subcode n, whose code length is p n , n ≥ 1 and is an integer; (1.2) Use m linear block codes (N, K) to perform group error correction coding on the user data to obtain the coded data bit sequence b; where m is the number of code words, m ≥ 1 and is an integer, N is the length of the coded code word, K is the information bit length, and satisfies mN ≤ p1p2…p n ; (1.3) The composite code a and the data bit sequence b are concomitantly encoded bit by bit to construct a double-layer bit sequence, and then converted into a DNA sequence according to the mapping rule between bit pairs and bases {00→A, 01→T, 10→G, 11→C}, and a long DNA fragment or a short oligonucleotide molecule is obtained by artificial synthesis; The preset Boolean combination logic specifically includes: logical multiplication, modulo-two sum, and majority decision; Among them, n pseudo-random sequences with short periods are selected as subcodes, which satisfy: the periods of the selected pseudo-random sequences are strictly mutually prime, and each of them has a sharp periodic autocorrelation characteristic; Wherein, the step (1.3) is specifically as follows: For long-fragment DNA data storage, the composite code is first truncated to obtain a shortened composite code of the same length as the data bit sequence, and then the shortened composite code is aligned with the data bit sequence to form a double-layer bit sequence, and directly mapped to obtain a long DNA fragment; for short-fragment oligonucleotide data storage, the data bit sequence is first divided into mN / l groups according to the length l, and the composite code is slidingly truncated according to the step length l to obtain mN / l shortened composite codes, and then the shortened composite code of length l is paired with the data bit sequence in sequence to construct mN / l groups of double-layer bit sequences, and after mapping, mN / l short-fragment oligonucleotides are finally generated; wherein l is a positive integer.

3. The DNA storage method with composite code and addressing encoding according to claim 1, characterized in that: The step (2) is specifically: (2.1) For short-fragment oligonucleotide pools, the library is directly constructed and then read by second-generation high-throughput sequencing to obtain sequencing reads of the data DNA; for long-fragment data DNA, high-throughput shotgun sequencing is used for sequencing reads; (2.2) Preprocess the sequencing reads, demap them according to the same mapping rules as the writing end, and extract the damaged composite codes from the double-layer structure. With damaged data bit sequence For sequencing reads of short-fragment oligonucleotide molecules, the primers at both ends are trimmed off by alignment with the amplification primers, and the retained payload is demapped; for reads obtained by long-fragment DNA shotgun sequencing, the sequencing reads are directly demapped; (2.3) The damaged composite code The cross-correlation detector of each subcode is sent to the known subcode set {C1, C2, ..., C n } Perform sliding correlation capture, search for the location of the correlation peak of each subcode, obtain the estimated phase of each subcode, and then determine the unique position label of the read segment through the Chinese residual theorem; (2.4) Using the phase estimates of each subcode obtained above, regenerate an ideal composite code a′ with the same length and the same phase as the damaged composite code, calculate the autocorrelation value of the composite code, and retain the read segment if the autocorrelation value meets the preset conditions; for the read segment that fails the verification, use the method based on edit distance alignment in the enhanced mode to re-identify it, determine the position label, and retain the read segment, and at the same time identify and correct the insertion and deletion error patterns according to the alignment path; (2.5) Considering that there are multiple copies of the data sequence at the same position, a majority vote is performed on the data sequence bit by bit to obtain a consistent data sequence. If the voting decision is not unique, the same erasure mark is used to fill the bit at that position to obtain a codeword, which is then sent to the decoder to correct the residual errors and restore the original user data.

4. The DNA storage method with composite code and addressing encoding according to claim 3 is characterized in that: The step (2.3) is specifically: (3.1) Each subcode in the known subcode set is periodically extended, and each subcode is subjected to sliding correlation operation with the damaged composite code. The correlation window length is the same as the length of the damaged composite code, and the sliding step length is the subcode period. The cross-correlation function between subcode i and the damaged composite code is expressed as: in, is the periodic extension sequence of subcode i, is the damaged composite code, τ is the subcode phase offset, L is the length of the correlation window, 1≤i≤n, and the periodic extension sequence With composite code It is represented by the symbol {+1, -1}, which satisfies the mapping relationship {+1→0, -1→1} with the binary sequence; (3.2) According to the sliding correlation value between each subcode and the damaged composite code, the position of the correlation peak is searched to obtain the phase estimation of each subcode. The correlation peak of subcode i corresponds to the phase τ i It is expressed as: Among them, argmax represents the position corresponding to the correlation peak; (3.3) Using the estimated phases of each subcode obtained by the search, the unique position label of the corresponding read segment is obtained by the Chinese remainder theorem. The estimated phases of each subcode satisfy in, is the unique position label of the read segment obtained by solution, τ1, τ2, ..., τ n They correspond to the estimated phases of n subcodes respectively.

5. The DNA storage method with composite code and addressing encoding according to claim 3 is characterized in that: The step (2.4) is specifically: (4.1) Based on the estimated phase of each subcode, regenerate a composite code with the same phase and the same length as the damaged one The same ideal composite code a′ is then used to calculate the damaged composite code The in-phase correlation value between the ideal composite code a′ and the obtained correlation value is expressed as: in, r′ is a damaged composite code and an ideal composite code, both of which are represented by symbols {+1, -1}, and satisfy the mapping relationship {+1→0, -1→1} with the binary sequence; (4.2) If the in-phase autocorrelation value of the composite code satisfies If the verification is successful, the current read segment is retained, where ω represents the preset correlation threshold, 0≤ω≤1; (4.3) For the reads that fail the above autocorrelation check, the damaged composite code is replaced by the edit distance comparison method in the enhanced mode. Compare with the known complete composite code a, search for the alignment position where the minimum edit distance is located, and use it as the position label of the read segment. At the same time, use the alignment path to locate the position where the insertion and deletion errors in the damaged composite code occur, and according to the characteristic that the positions of the insertion / deletion errors in the double-layer sequence structure are exactly the same, correct the insertion and deletion error patterns in the damaged data sequence; discard the position bits where the insertion errors occur, mark the position bits where the deletion errors occur as erased, and fill them with the determined markers; if the enhanced mode is not used, skip the current step.

Citation Information

Patent Citations

  • DNA storage method for pseudo noise sequence adjoint coding

    CN116707541A

  • Space layering DNA storage method of large-scale oligonucleotide pool

    CN118841094A