A multi-level directory DNA storage encoding and decoding method based on modulation technology

Through the multi-level directory DNA storage encoding and decoding method based on modulation technology, addressing sequences are designed and file management systems are built, which solves the problems of high capacity and high accessibility in DNA storage, and realizes efficient and flexible file management and highly reliable data reading and writing.

CN116364185BActive Publication Date: 2025-08-29YAMI TECH (GUANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310173414.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-27
Publication Date
2025-08-29
Estimated Expiration
2043-02-27

AI Technical Summary

Technical Problem

Existing DNA storage technologies are difficult to achieve high capacity and high access at the same time. Addressing sequence design is time-consuming and lacks unified standards. Base errors lead to low data reliability, affecting the real-time read-out throughput and speed of data.

Method used

A multi-level directory DNA storage codec method based on modulation technology is adopted to design addressing sequences to satisfy biological constraints, provide similar computer file management methods, and build a primer sequence and file absolute path set by generating a modulation code table for primer and file data, synthesize DNA sequences and store them, and decode the data by editing the principle of closest distance.

Benefits of technology

It realizes efficient and flexible DNA storage file management, supports high-throughput and high-reliability reading and writing of massive data, reduces the impact of base errors, and improves data reading speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116364185B_ABST
    Figure CN116364185B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-level directory DNA storage encoding and decoding method based on modulation technology, comprising the following steps: S100: generating a modulation code table for primer sequences and file data sequences, constructing a primer sequence data set and a file absolute path set; S200: modulating the binary file to be stored into a DNA sequence, synthesizing and storing it in vitro; S300: selecting target primers, amplifying DNA molecules under the target logical disk from the synthesis pool, and sequencing them; S400: calculating the absolute file path of each read length of the sequenced data based on the observed modulation sequence, and grouping the sequenced data based on this; S500: generating a modulation sequence for decoding each grouped data based on the encoded DNA chain length and the absolute file path; and using this modulation sequence to decode the data according to the modulation DNA storage decoding algorithm. The present invention designs an addressing sequence based on modulation technology and proposes a multi-level directory DNA storage encoding and decoding method, which can realize an efficient and flexible DNA storage file management method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of systems biology research, and in particular to a multi-level directory DNA storage encoding and decoding method based on modulation technology. Background Art

[0002] With the rapid development of technologies such as cloud computing, the Internet of Things, and big data, the total amount of global data continues to grow exponentially or even super-exponentially. However, traditional storage media (such as magnetic, optical, and solid-state storage), currently used for cloud storage, face technical bottlenecks in terms of power consumption, size, reliability, and effective storage time. Exploring new storage media and corresponding read / write technologies has become a key foundational issue for the sustainable development of information technology.

[0003] Compared with traditional storage media, DNA molecules have great advantages in data storage: (1) Ultra-high storage density. The storage density of DNA molecules can reach 10 19 bits / cm 3 , which is 6 orders of magnitude higher than traditional storage media. (2) Ultra-long service life. Data stored in DNA can be preserved for thousands of years without special human intervention. (3) Ultra-low maintenance cost. The land, resources and energy required for DNA storage are far less than those of traditional storage media, and the maintenance cost is extremely low. In addition, the biochemical reactions and operations of DNA molecules themselves have huge parallelism. In recent years, DNA storage has become a hot topic in interdisciplinary research. Nature and Science publish research papers on DNA storage every year, and governments of various countries, such as the United States and my country, have also gradually incorporated DNA storage into their national development strategies.

[0004] The construction of future DNA storage addressing systems for massive data requires two key considerations: file access and storage capacity. However, current DNA storage research struggles to simultaneously achieve both high capacity and high accessibility. This is primarily due to the crosstalk and diffusion characteristics of DNA molecules. Therefore, in addition to considering conventional biological sequence constraints such as GC balance, the absence of homopolymers, and the absence of unusual secondary structures, the design of address sequences (primer sequences) must also consider factors such as a certain orthogonal distance between address sequences and between address sequences and valid data sequences. However, the currently commonly used brute-force address sequence search method is extremely time-consuming and lacks a unified standard. Address sequences from different databases are incompatible, and the space of available primer address sequences is extremely limited. Furthermore, errors such as base insertions, deletions, and substitutions within DNA sequences during synthesis and sequencing hinder reliable data recovery, further exacerbating the future demand for real-time readout throughput and speed for massive DNA storage data.

[0005] To address these issues, the present invention proposes a multi-level directory DNA storage encoding and decoding method based on modulation technology to design addressing sequences. This method offers the following advantages: The addressing sequence design is simple, efficient, and scalable, and it meets biological constraints for addressing sequences. Addressing sequences from different databases are compatible, supporting unified management using modulation codes. This method provides an efficient and flexible DNA storage file management method similar to the computer's "hard drive - logical drive letter - folder - file" structure. Summary of the Invention

[0006] The purpose of the present invention is to provide a multi-level directory DNA storage encoding and decoding method based on modulation technology. Through the new technical concept, it can efficiently and flexibly manage DNA storage file data in a manner similar to the "hard disk-logical drive letter-folder-file" method of a computer, improve real-time data readout throughput, and solve the problems existing in the above-mentioned prior art.

[0007] The present invention provides the following technical solutions:

[0008] A multi-level directory DNA storage encoding and decoding method based on modulation technology includes the following steps:

[0009] S100: Generate a modulation code table of primer sequences and file data sequences, and then construct a primer sequence data set and a file absolute path set;

[0010] S200: Select appropriate primers from the primer sequence data set and select appropriate file absolute paths from the file absolute path set as needed, modulate the binary file to be stored into a DNA sequence of a certain length, and store it in vitro after synthesis;

[0011] S300: Select target primers, amplify the DNA molecules under the target logical disk from the synthesis pool, and sequence them;

[0012] S400: For each read length of the sequencing data, according to the observed modulation sequence and the principle of shortest edit distance, the absolute file path of the read length is determined; and the sequencing data is grouped according to the absolute file path;

[0013] S500: Generate a modulation sequence for decoding the grouped data by referring to the absolute file path and DNA coding sequence length corresponding to each grouped sequencing data; and use the modulation sequence to decode the data according to the modulation DNA storage decoding algorithm.

[0014] Preferably, the step S100 specifically includes the following steps:

[0015] 1) Generate a set of binary sequences M of a specified length n as the modulation code table for subsequent primers and file data sequences, where n>0, |M|≥3, and the modulation code table set meets the following three conditions:

[0016] Condition 1: The content of the characters '0' and '1' in any element is between 45% and 55%;

[0017] Condition 2: The number of consecutive identical characters in any element does not exceed 3;

[0018] Condition 3: The shift distance between any two elements is greater than a certain threshold d, 0 <d<n;

[0019] Assume two strings x of length n i and x j , string x i and x j The displacement distance H(x i ,x j ) is defined as where ρ k Indicates an offset of k positions, c ij For the sequence x j After shifting k positions, i The maximum sum of identical characters;

[0020] 2) Any element m in the binary sequence set M i Modulation sequence c required to generate primer sequences p , generating a length of |c p The set C of binary sequences of | b , and the set C b The minimum Hamming distance between any two elements is greater than or equal to |C p | / 2, C b Each element of the set is related to c p , modulated into a DNA sequence and placed into the primer set P;

[0021] 3) Delete m from the binary sequence set M i , the new set is expressed as M', the absolute path of the file is a binary string whose length is a multiple of n, and the substring of each n characters from left to right of the string comes from M'; in order to facilitate user file management, the absolute path of the file can be divided into two parts: directory ID and file ID. Assume that the absolute path length of the file is A dir , dynamically adjust the ratio of directory ID and file ID, obtain different numbers of file absolute paths, and put the file absolute paths that meet the conditions into set C dir middle.

[0022] More preferably, any element m in the binary sequence set M i Modulation sequence c required to generate primer sequences p , the modulation sequence is generated according to the following principles:

[0023] Principle 1: cp By m i Splicing Seconds and thirds;

[0024] Principle 2: c p The length of the primer is within the length range of conventional primer sequences.

[0025] More preferably, the directory ID is represented as a multi-layer nested directory to facilitate user file management. The specific representation method is as follows:

[0026] Assume that the directory ID length is D L , then the maximum number of nested levels that a directory of this length can represent is D L / n, each level represents |M'| directories at the same level, and the maximum number of directories that can be represented is

[0027] More preferably, the number of files represented by the file ID is in exponential form of |M'|, and the specific number represented is:

[0028] Assume that the length of the file ID is F L , then the maximum number of files that can be represented by a file ID of this length is

[0029] Preferably, the S200 specifically includes:

[0030] 1) Select a primer p from the primer sequence dataset P r Used as a logical disk identifier;

[0031] 2) From the absolute path of the file set C dir Select an absolute file path c i , repeat the absolute path of the file according to the length N of the encoding DNA sequence The second structure is used for the file to be stored f i The modulation sequence c f ;

[0032] 3) Using c f The file to be stored f i Modulated into DNA sequence set S i

[0033] 4) S i Add primer p to the head of each sequence r Constitute the DNA sequence set S' i , used for storage after DNA synthesis.

[0034] Preferably, the S400 specifically includes:

[0035] 1) For each read length r of sequencing data jAccording to the modulation rule, the corresponding observation modulation sequence o is obtained j ;

[0036] 2) According to the principle of shortest edit distance, from the file absolute path set C dir Choose one with o j Edit the absolute path of the file closest to c i , and assign the absolute path of the file to the current sequencing read length r j ;

[0037] 3) Grouping according to the absolute file path corresponding to each sequencing read length.

[0038] More preferably, according to the observed modulation sequence o j Determine the absolute path of the file to which it belongs c i The specific method is:

[0039] 1) Set the absolute path of the file to C dir For each absolute file path in The modulation sequence used to store the file is obtained again, and the new modulation sequence set is recorded as C′ dir ;

[0040] 2) Determine the observed modulation sequence o according to the principle of the shortest edit distance j The corresponding modulation sequence c′ i ∈C′ dir , and then the absolute path of the file c i ∈C dir Assigned to the observed modulation sequence.

[0041] Preferably, the S500 specifically includes:

[0042] 1) Refer to the absolute file path c corresponding to each group sequencing data i ∈C dir and DNA coding sequence length to generate the modulation sequence c' used to decode the packet data i ∈C' dir ;

[0043] 2) Apply the modulation sequence c' i According to the modulation DNA storage decoding algorithm, the absolute path of the decoding file is c i ∈C dir All sequencing read lengths.

[0044] Compared with the prior art, the present invention has the following advantages:

[0045] 1. This invention proposes a novel multi-level directory DNA storage encoding and decoding method based on modulation technology to design addressing sequences. Advantages of this method include: simple, efficient, and scalable addressing sequence design, and compliance with biological constraints; compatible addressing sequences across different databases, supporting unified management via modulation codes; and providing an efficient and flexible DNA storage file management method similar to the computer's "hard drive - logical drive letter - folder - file" model.

[0046] 2. The embodiment of the present invention does not require displayed storage address information (file path and file information are modulated and stored in DNA fragments), has high information logic density (the number of data bits that a single base can carry), and has a low write cost.

[0047] 3. The method proposed in the present invention can be deployed on the current mainstream DNA synthesis and sequencing platforms (error rate is reduced by 1% to 10%). Tasks such as data encoding and determining the absolute path of the file to which the sequencing read length belongs support computer multi-threaded parallel processing, meeting the needs of high-throughput and high-reliability reading and writing for massive DNA data storage.

[0048] 4. Experimental data shows that (the code lengths in the modulation code table are 4, 8, 12, and 16 respectively), the modulation code table meets the storage capacity requirements. Assuming that the DNA storage sequence length is 200, the modulation code table code length is 4, and the absolute file path is repeated once in the DNA coding sequence, the maximum that can be represented is 10 15 files; the absolute file path is repeated twice, and the maximum number of files that can be represented is 10 7.5 files.

[0049] 5. In addition, under different error rates, the method of the present invention can successfully identify the absolute file path to which each read in the sequencing data belongs. Subsequently, the modulated DNA storage decoding method (CN 113299347 A) can be used to successfully decode the data. The specific data results are as follows:

[0050] 1) The absolute file path is repeated once in the DNA coding sequence:

[0051] a) The error rate is less than or equal to 0.05, and the correct identification rate of the absolute path of the sequencing data file reaches more than 99%;

[0052] b) When the error rate is 0.1, the correct recognition rate of the absolute path of the sequencing data file reaches more than 90%.

[0053] c) When the error rate is 0.15, the correct recognition rate of the absolute path of the sequencing data file reaches more than 86%.

[0054] 2) The absolute file path is repeated three times in the DNA coding sequence:

[0055] a) When the error rate is less than or equal to 0.1, the correct recognition rate of the absolute path of the sequencing data file is 100%.

[0056] b) When the error rate is 0.15, the correct recognition rate of the absolute path of the sequencing data file reaches more than 98%.

[0057] These results indicate that the proposed method can be deployed on current mainstream DNA synthesis and sequencing platforms (saving error rates between 1% and 10%), meeting the needs of high-throughput and highly reliable reading and writing for massive DNA data storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.

[0059] Figure 1 This is a flow chart of an implementation of a multi-level directory DNA storage encoding and decoding method based on modulation technology according to an embodiment of the present invention;

[0060] Figure 2 This is a schematic diagram of a modulation sequence structure corresponding to a typical DNA storage sequence according to an embodiment of the present invention;

[0061] Figure 3 The algorithm flow for decoding DNA data provided by the embodiment of the present invention;

[0062] Figure 4 The accuracy of identifying the absolute file path of the sequencing data under error rates of 0.03, 0.05, 0.1, and 0.15 for different modulation code table code lengths and one repetition of the absolute file path provided by the embodiment of the present invention;

[0063] Figure 5 The code lengths of different modulation code tables provided in the embodiment of the present invention are repeated three times, and the absolute file path is used to identify the absolute file path to which the sequencing data belongs when the error rates are 0.03, 0.05, 0.1, and 0.15. DETAILED DESCRIPTION

[0064] The following is a further detailed description of the multi-level directory DNA storage encoding and decoding method based on modulation technology in conjunction with specific embodiments. These embodiments are only for comparison and explanation purposes, and the present invention is not limited to these embodiments.

[0065] Example

[0066] See also Figure 1-5 The multi-level directory DNA storage encoding and decoding method based on modulation technology proposed in an embodiment of the present invention includes the following steps:

[0067] S100: Generate a modulation code table for primer sequences and file data sequences, and then construct a primer sequence data set and a set of absolute file paths;

[0068] S200: Select appropriate primers (target logical disks) from the primer sequence data set as needed, select appropriate absolute file paths from the set of absolute file paths, modulate the binary file to be stored into a DNA sequence of a certain length, and perform in vitro storage after synthesis;

[0069] S300: Select target primers, amplify DNA molecules under the target logical disk from the synthesis pool, and sequence them;

[0070] S400: For each read length of the sequencing data, determine the absolute file path of this read length according to the observed modulation sequence based on the principle of the closest edit distance; and group the sequencing data according to the absolute file path;

[0071] S500: Generate a modulation sequence for decoding the grouped data by referring to the absolute file path corresponding to each grouped sequencing data and the DNA coding sequence length; and decode the data using this modulation sequence according to the modulation DNA storage decoding algorithm.

[0072] Preferably, the S100 specifically includes the following steps:

[0073] 1) Generate a set M of binary sequences of a specified length n as the modulation code table for subsequent primers and file data sequences, where n>0 and |M|≥3. This modulation code table set satisfies the following three conditions:

[0074] Condition 1: The content of any element character '0' and '1' is between 45% and 55%;

[0075] Condition 2: The number of consecutive identical characters in any element does not exceed 3;

[0076] Condition 3: The shift distance between any two elements is greater than a certain threshold d, where 0<d<n; usually d is taken as

[0077] Assume two strings x of length n i and x j , and the shift distance H(x i and x j ) between the strings x i and x j is defined as where ρ [[ID=4!6]] k represents an offset of k positions, and c ij is the sum of the maximum identical characters between the sequence x j offset by k positions and x i ;

[0078] 2) Any element m in the binary sequence set M i Modulation sequence c required to generate primer sequences p , generating a length of |c p The set C of binary sequences of | b , and the set C b The minimum Hamming distance between any two elements is greater than or equal to |C p | / 2, C b Each element of the set is related to c p , modulated into a DNA sequence according to Table 1 and placed into the primer set P;

[0079] Table 1 Modulated DNA sequence list

[0080] Modulation code bits 0 0 1 1 Information bits 0 1 0 1 DNA bases A T C G

[0081] 3) Delete m from the binary sequence set M i , the new set is expressed as M', the absolute path of the file is a binary string whose length is a multiple of n, and the substring of each n characters from left to right of the string comes from M'; in order to facilitate user file management, the absolute path of the file can be divided into two parts: directory ID and file ID. Assume that the absolute path length of the file is A dir , dynamically adjust the ratio of directory ID and file ID, obtain different numbers of file absolute paths, and put the file absolute paths that meet the conditions into set C dir middle.

[0082] More preferably, any element m in the binary sequence set M i Modulation sequence c required to generate primer sequences p , the modulation sequence is generated according to the following principles:

[0083] Principle 1: c p By m i Splicing Seconds and thirds;

[0084] Principle 2: c p The length of the primer is within the length range of conventional primer sequences.

[0085] More preferably, the directory ID is represented as a multi-layer nested directory to facilitate user file management. The specific representation method is as follows:

[0086] Assume that the directory ID length is D L , then the maximum number of nested levels that a directory of this length can represent is D L / n, each level represents |M'| directories at the same level, and the maximum number of directories that can be represented is

[0087] More preferably, the number of files represented by the file ID is in exponential form of |M'|, and the specific number represented is:

[0088] Assume that the length of the file ID is F L , then the maximum number of files that can be represented by a file ID of this length is

[0089] Preferably, the S200 specifically includes:

[0090] 1) Select a primer p from the primer sequence dataset P r Used as a logical disk identifier;

[0091] 2) From the absolute path of the file set C dir Select an absolute file path c i , repeat the absolute path of the file according to the length N of the encoding DNA sequence The second structure is used for the file to be stored f i The modulation sequence c f ;

[0092] 3) Using c f The file to be stored f i Modulated into DNA sequence set S i

[0093] In this embodiment, the DNA sequence set S is modulated according to the invention patent publication number (CN 113299347 A). i ;

[0094] 3) S i Add primer p to the head of each sequence r Constitute the DNA sequence set S' i , used for storage after DNA synthesis.

[0095] A typical DNA storage sequence corresponding to the modulation sequence (including primer modulation sequence and file absolute path modulation sequence) is shown in the following diagram: Figure 2 shown.

[0096] Preferably, the S400 specifically includes:

[0097] 1) For each read length (read) r of sequencing data j According to the modulation rule, the corresponding observation modulation sequence o is obtained j ;

[0098] 2) According to the principle of shortest edit distance, from the file absolute path set C dir Choose one with o j Edit the absolute path of the file closest to c i, and assign the absolute path of the file to the current sequencing read length r j ;

[0099] 3) Grouping according to the absolute file path corresponding to each sequencing read length.

[0100] More preferably, according to the observed modulation sequence o j Determine the absolute path of the file to which it belongs c i The specific method is:

[0101] 1) Set the absolute path of the file to C dir For each absolute file path in The modulation sequence used to store the file is obtained again, and the new modulation sequence set is recorded as C′ dir ; (Note: C dir and C' dir One-to-one mapping)

[0102] 2) Determine the observed modulation sequence o according to the principle of the shortest edit distance j The corresponding modulation sequence c′ i ∈C′ dir , and then the absolute path of the file c i ∈C dir Assigned to the observed modulation sequence.

[0103] Preferably, the S500 specifically includes:

[0104] 1) Refer to the absolute file path c corresponding to each group sequencing data i ∈C dir and DNA coding sequence length to generate the modulation sequence c' used to decode the packet data i ∈C' dir ;

[0105] 2) Apply the modulation sequence c' i According to the modulation DNA storage decoding algorithm, the absolute path of the decoding file is c i ∈C dir All sequencing read lengths.

[0106] In this embodiment, according to the invention patent publication number (CN 113299347 A), the absolute path of the decoded file is c i ∈C dir All sequencing read lengths.

[0107] like Figure 3The figure shows the algorithm flow for decoding DNA data (steps S300 to S500) provided by an embodiment of the present invention. When a user wants to read a file from a specified disk directory from a DNA synthesis pool: first, primers representing the disk partition (single-stranded DNA molecules) need to be added to the DNA synthesis pool for a specific PCR reaction, and the amplified DNA molecules are sequenced (step S300); second, the read length data output by the sequencing is grouped according to the file ID (step S400); finally, modulation and decoding are performed on each group of sequencing data. (For the specific scheme of step S500, please refer to the modulation and decoding patent CN 113299347 A).

[0108] The embodiments of the present invention have the following beneficial effects:

[0109] The present invention proposes a multi-level directory DNA storage encoding and decoding method based on modulation technology to design addressing sequences. This method offers the following advantages: simple, efficient, and scalable addressing sequence design, while also meeting biological constraints on addressing sequences; compatible addressing sequences across different databases, supporting unified management via modulation codes; and provides an efficient and flexible DNA storage file management method similar to the computer's "hard drive - logical drive letter - folder - file" model.

[0110] Experimental data shows that (the code lengths in the modulation code table are 4, 8, 12, and 16 respectively), the modulation code table meets the storage capacity requirements. Assuming that the DNA storage sequence length is 200, the modulation code table code length is 4, and the absolute file path is repeated once in the DNA sequence, the maximum that can be represented is 10 15 files; the absolute file path is repeated twice, and the maximum number of files that can be represented is 10 7.5 files.

[0111] 1) The absolute file path is repeated once in the DNA coding sequence, such as Figure 4 As shown:

[0112] a) The error rate is less than or equal to 0.05, and the correct identification rate of the absolute path of the sequencing data file reaches more than 99%;

[0113] b) When the error rate is 0.1, the correct recognition rate of the absolute path of the sequencing data file reaches more than 90%.

[0114] c) When the error rate is 0.15, the correct recognition rate of the absolute path of the sequencing data file reaches more than 86%.

[0115] 2) The absolute file path is repeated three times in the DNA coding sequence, such as Figure 5 Shown:

[0116] 1) When the error rate is less than or equal to 0.1, the correct recognition rate of the absolute path of the sequencing data file is 100%.

[0117] 2) When the error rate is 0.15, the correct recognition rate of the absolute path of the sequencing data file reaches more than 98%.

[0118] These results show that the method proposed in this patent can be deployed on the current mainstream DNA synthesis and sequencing platforms (error rate reduction between 1% and 10%), meeting the needs of high-throughput and high-reliability reading and writing for massive DNA data storage.

[0119] The multi-level directory DNA storage encoding and decoding method based on modulation technology provided by the above embodiment of the present invention generates a modulation code table for primer sequences and file data sequences, thereby constructing a primer sequence data set and a file absolute path set; selecting appropriate primers from the primer sequence data set and an appropriate file absolute path from the file absolute path set as needed, modulating the binary file to be stored into a DNA sequence of a certain length, and storing it in vitro after synthesis; selecting target primers, amplifying the DNA molecules under the target logical disk from the synthesis pool and sequencing them; determining the file absolute path of each read length of the sequencing data based on the observed modulation sequence and the principle of the shortest edit distance; and grouping the sequencing data according to the file absolute path; generating a modulation sequence for decoding the grouped data with reference to the file absolute path and DNA coding sequence length corresponding to each grouped sequencing data; and using this modulation sequence to decode the data according to the modulated DNA storage decoding algorithm.

[0120] Those skilled in the art will understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A multi-level directory DNA storage encoding and decoding method based on modulation technology, characterized in that: The steps include: S100: Generate a modulation code table of primer sequences and file data sequences, and then construct a primer sequence data set and a file absolute path set; S200: Select appropriate primers from the primer sequence data set and select appropriate file absolute paths from the file absolute path set as needed, modulate the binary file to be stored into a DNA sequence of a certain length, and store it in vitro after synthesis; S300: Select target primers, amplify the DNA molecules under the target logical disk from the synthesis pool, and sequence them; S400: For each read length of the sequencing data, according to the observed modulation sequence and the principle of shortest edit distance, the absolute file path of the read length is determined; and the sequencing data is grouped according to the absolute file path; S500: Generate a modulation sequence for decoding the data group by referring to the absolute file path and DNA coding sequence length corresponding to each data group; and decode the data using the modulation sequence according to the modulation DNA storage decoding algorithm; The S200 specifically includes: 1) Select a primer p from the primer sequence dataset P r Used as a logical disk identifier; 2) From the absolute path of the file set C dir Select an absolute file path c i , repeat the absolute path of the file according to the length N of the encoding DNA sequence The second structure is used for the file to be stored f i The modulation sequence c f ; 3) Using c f The file to be stored f i Modulated into DNA sequence set S i ; 4) S i Add primer p to the head of each sequence r Constitute the DNA sequence set S' i , used for storage after DNA synthesis; The S400 specifically includes: 1) For each read length r of sequencing data j According to the modulation rule, the corresponding observation modulation sequence o is obtained j ; 2) According to the principle of shortest edit distance, from the file absolute path set C dir Choose one with o j Edit the absolute path of the file closest to c i , and assign the absolute path of the file to the current sequencing read length r j ; 3) Grouping according to the absolute file path corresponding to each sequencing read length; According to the observed modulation sequence o j Determine the absolute path of the file to which it belongs c i The specific method is: 1) Set the absolute path of the file to C dir For each absolute file path in The modulation sequence used to store the file is obtained again, and the new modulation sequence set is recorded as C' dir ; 2) Determine the observed modulation sequence o according to the principle of the shortest edit distance j The corresponding modulation sequence c' i ∈C' dir , and then the absolute path of the file c i ∈C dir Assigned to the observed modulation sequence.

2. The multi-level directory DNA storage encoding and decoding method based on modulation technology according to claim 1 is characterized in that: The step S100 specifically includes the following steps: 1) Generate a set of binary sequences M of a specified length n as the modulation code table for subsequent primers and file data sequences, where n>0, |M|≥3, and the modulation code table set meets the following three conditions: Condition 1: The content of the characters '0' and '1' in any element is between 45% and 55%; Condition 2: The number of consecutive identical characters in any element does not exceed 3; Condition 3: The shift distance between any two elements is greater than a certain threshold d, 0 <d<n; Assume two strings x of length n i and x j , string x i and x j The displacement distance H(x i ,x j ) is defined as where ρ k Indicates an offset of k positions, c ij For the sequence x j After shifting k positions, i The maximum sum of identical characters; 2) Any element m in the binary sequence set M i Modulation sequence c required to generate primer sequences p , generating a length of |c p The set C of binary sequences of | b , and the set C b The minimum Hamming distance between any two elements is greater than or equal to |C p | / 2, C b Each element of the set is related to c p , modulated into a DNA sequence and placed into the primer set P; 3) Delete m from the binary sequence set M i , the new set is expressed as M', the absolute path of the file is a binary string whose length is a multiple of n, and the substring of each n characters from left to right of the string comes from M'; in order to facilitate user file management, the absolute path of the file can be divided into two parts: directory ID and file ID. Assume that the absolute path length of the file is A dir , dynamically adjust the ratio of directory ID and file ID, obtain different numbers of file absolute paths, and put the file absolute paths that meet the conditions into set C dir middle.

3. The multi-level directory DNA storage encoding and decoding method based on modulation technology according to claim 2 is characterized in that: Any element m in the binary sequence set M i Modulation sequence c required to generate primer sequences p , the modulation sequence is generated according to the following principles: Principle 1: c p By m i Splicing Seconds and thirds; Principle 2: c p The length of the primer is within the length range of conventional primer sequences.

4. The multi-level directory DNA storage encoding and decoding method based on modulation technology according to claim 2 is characterized in that: The directory ID is represented as a multi-layer nested directory to facilitate user file management. The specific representation method is as follows: Assume that the directory ID length is D L , then the maximum number of nested levels that a directory of this length can represent is D L / n, each level represents |M'| directories at the same level, and the maximum number of directories that can be represented is 5. The multi-level directory DNA storage encoding and decoding method based on modulation technology according to claim 2 is characterized in that: The number of files represented by the file ID is in exponential form |M'|, and the specific number represented is: Assume that the length of the file ID is F L , then the maximum number of files that can be represented by a file ID of this length is 6. The multi-level directory DNA storage encoding and decoding method based on modulation technology according to claim 1 is characterized in that: The step S500 specifically includes: 1) Refer to the absolute file path c corresponding to each group sequencing data i ∈C dir and DNA coding sequence length to generate the modulation sequence c' used to decode the packet data i ∈C' dir ; 2) Apply the modulation sequence c' i According to the modulation DNA storage decoding algorithm, the absolute path of the decoding file is c i ∈C dir All sequencing read lengths.

Citation Information

Patent Citations

  • DNA storage method based on modulation coding

    CN113299347A

  • DNA (deoxyribonucleic acid) storage cascade coding and decoding method for type-1 and type-2 segmented error correction inner codes

    CN114328000A