Improved DNA (Deoxyribose Nucleic Acid) compression coding method

Through the improved DNA compression encoding method, the encoding rules of the DNA tree are used to encode binary files into DNA sequences, solving the problem that traditional storage media is difficult to meet the needs of high storage density, and achieving more efficient DNA storage density and biological constraint satisfaction.

CN119943167APending Publication Date: 2025-05-06YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510033084.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional storage media are difficult to meet the rapidly growing demand for digital data storage, and DNA, as a new storage medium, has not yet fully utilized its advantages of high storage density and stability.

Method used

An improved DNA compression coding method is proposed, by converting binary files into binary strings, dividing them into binary substrings of length i, calculating the probability of each substring, building a DNA tree, and labeling them along the path of the DNA tree using four nucleotides to improve the encoding density.

Benefits of technology

It achieves a higher storage density, saves the use of nucleotides, reduces the cost of subsequent synthesis and storage, and effectively avoids the generation of homopolymers and meets the biological constraints of the DNA strand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943167A_ABST
    Figure CN119943167A_ABST
Patent Text Reader

Abstract

The invention discloses an improved DNA compressed encoding method, which is applied to the technical field of storage and aims at solving the problem that the traditional storage medium such as a hard disk and a magnetic tape has been explored to the limit in potential and is difficult to adapt to the storage demand of explosive increase of data volume. The stable chemical structure and high-density information storage capacity of DNA molecules are utilized, and digital information is coded into a nucleotide sequence. By converting the data into a specific base combination, the DNA not only can realize long-term storage of the data, but also has extremely high storage density and stability. As high cost is needed for synthesizing DNA, an efficient coding mode is needed, and original data is stored by using a smaller number of nucleotides, namely, the storage density is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of storage technology, and in particular relates to a DNA storage technology. Background Art

[0002] With the development of science and technology, the world is becoming more and more digitalized, resulting in an explosive growth in the amount of global data. According to a report by the International Data Corporation (IDC), by 2025, digital data will reach a huge amount of 175 zettabytes, which poses a challenge to data compression and storage. However, the potential of traditional storage media such as hard disks and tapes has been explored to the limit, and it is difficult to make further improvements, so it is urgent to explore new forms of storage. DNA is a long-chain polymer composed of many nucleotides. There are four types of nucleotides, including adenine (A), thymine (T), cytosine (C), and guanine (G). It is a medium for storing genetic information of organisms in nature. It has many advantages such as high stability, high storage density, easy synthesis and replication, which has been verified in the thousands of years of evolution of organisms. Summary of the invention

[0003] In order to improve the storage density, the present invention proposes a method for compressing binary files and encoding them using nucleotides to synthesize a DNA chain. During the encoding process, the biological constraints that need to be followed are fully considered, and the encoding density is improved, so that a DNA chain of the same length can store more information.

[0004] The technical solution adopted by the present invention is: an improved DNA compression encoding method, comprising:

[0005] S1, converting the original text or image into a binary string, and then dividing the binary string into binary substrings of length i;

[0006] S2, treating each binary substring as a source symbol, and calculating the probability of occurrence of each source symbol;

[0007] S3, constructing a DNA tree based on the calculated probability;

[0008] S4. Use four nucleotides to mark along the path starting from the root node to obtain the codeword corresponding to each source symbol.

[0009] Step S1 is specifically as follows:

[0010] S11. For text, use binary ASCII for conversion; for images, each pixel has eight bits, and the pixels are read one by one for conversion;

[0011] S12. If the length of the binary string is an integer multiple of i, then the binary string is directly divided according to the length of each binary substring being i; otherwise, 0 is added to the end of the binary string so that the length of the binary string after adding 0 at the end is an integer multiple of i, and then the binary string is divided according to the length of each binary substring being i.

[0012] Step S2 is specifically as follows: scan all binary substrings, calculate the probability of each binary substring appearing, each substring corresponds to a node, and its probability corresponds to the probability of the node, and obtain the node queue of the DNA tree.

[0013] Step S3 is specifically as follows:

[0014] S31. Determine whether the number of nodes is just enough to build a complete DNA tree according to the following formula

[0015] x-2*(n-1)-3=1

[0016] Among them, x is the number of substrings, n is the number of layers of the tree, and 2 means that two nodes are reduced each time encoding. Specifically, two nodes are reduced each time encoding through ternary encoding. In this process, the three smallest probabilities are selected and represented by code elements 0, 1, and 2 respectively. They are replaced with nucleotides according to the rules of round-robin encoding, and then these three probabilities are added to get a new queue and rearranged. This cycle is repeated until there are four nodes left. 3 is because the last encoding is performed when there are only four tree nodes left, and there is only one root node in the end. If the above formula cannot be satisfied, adjustments need to be made at the last layer of the tree, and virtual nodes with a corresponding frequency of 0 need to be added. By simplifying the equation, it is possible to directly determine whether the number of substrings is an even number. If it is an odd number, a virtual node with a probability of 0 is added. This is to just meet the number of nodes in the tree. If it is an even number, it remains unchanged.

[0017] S32, arranging the initial node queue in descending order, adding the three nodes with the smallest probability each time, and then rearranging them until four nodes are left;

[0018] S33. Add the probabilities corresponding to the four remaining nodes in step S32 to get 1, which is the root node of the DNA tree. At this time, the first-level node of the DNA tree is 4, and the other-level nodes are 3.

[0019] Step S4 is specifically as follows:

[0020] S41. Use ACGT to represent the four branches of the first layer in the DNA tree. According to the biological GC content constraint, A and T are assigned to the two branches at the edge, corresponding to the node with the smallest probability and the node with the largest probability, respectively, while G and C are assigned to the two branches in the middle, corresponding to the two nodes in the middle probability.

[0021] S42. After the four branches of the first layer are marked, the branches of the remaining layers are assigned values ​​according to the Goldman coding rule. The specific assignment rules are as follows:

[0022]

[0023] Beneficial effects of the present invention: In the present invention, by setting four branches in the first layer of the DNA tree and marking them in a corresponding order, the AT content and the GC content can be balanced; and in the marking of the subsequent layers according to the above method, there will be no continuous repeated nucleotides, and binary data can be effectively compressed and stored in the DNA sequence; the method of the present invention has the following advantages:

[0024] 1. Compared with the traditional binary tree, the innovative ternary tree is proposed. The same target file requires fewer nucleotides to be stored, thus improving the storage density and saving subsequent costs;

[0025] 2. By setting four branches in the first layer and marking the two middle branches with G and C, that is, the two branches with relatively middle corresponding probabilities, the GC content is effectively controlled to be close to 50%;

[0026] 3. In addition to the function of balancing GC content mentioned in 2, the appearance of four branches in the first layer can also further improve the coding density;

[0027] 4. In the subsequent process of allocating codewords, the nucleotides corresponding to the branches in the previous layer will no longer be used to mark the current layer, thus effectively avoiding the generation of homopolymers. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flow chart of DNA storage according to the present invention;

[0029] Figure 2 The coding process of establishing a DNA tree according to the present invention;

[0030] Figure 3 The structure of establishing a DNA tree according to the present invention;

[0031] Figure 4 A DNA tree corresponding to an example of the present invention;

[0032] Figure 5 The storage density effect of the present invention is exemplified;

[0033] Figure 6 The biological restraint effect exemplified by the present invention;

[0034] Figure 7 This is an example of the encoding time effect of the present invention. DETAILED DESCRIPTION

[0035] To facilitate those skilled in the art to understand the technical content of the present invention, the present invention is further explained below with reference to the accompanying drawings.

[0036] The process of DNA storage can be roughly described as follows: Figure 1 In the DNA storage process, the original information needs to be converted into binary information, and then four nucleotides are used for encoding and mapping. In this process, data compression encoding and error correction encoding need to be implemented to achieve efficient and high-quality information storage. After the base sequence is determined, it is artificially synthesized and the generated DNA chain is stored in a DNA pool. When certain information needs to be extracted, the nucleotide sequence needs to be read through DNA sequencing technology, and then decoded to finally obtain the original information.

[0037] The essence of DNA storage is to use the stable chemical structure and high-density information storage capacity of DNA molecules to encode digital information into nucleotide sequences. By converting data into specific base combinations, DNA can not only achieve long-term storage of data, but also has extremely high storage density and stability. Since synthesizing DNA requires a high cost, an efficient encoding method is needed to use fewer nucleotides to store the original data, that is, to increase storage density.

[0038] like Figure 1 As shown, the present invention provides a method for compressing binary files and mapping the compressed binary files to DNA, which has the advantage of improving the storage density of DNA. The implementation steps of the present invention are specifically introduced below in conjunction with the accompanying drawings.

[0039] A1. Figure 2 As shown in the figure, firstly, the target file such as image or text is converted into a binary file, and then the binary is scanned, where the number of scan bits is i, which means that the binary file is divided into several i-bit length binary substrings. Then, these several binary substrings are scanned and counted, the probability of occurrence is calculated, and they are sorted in order to obtain the initial node queue.

[0040] A2. The main purpose of this step is to construct a DNA tree for the node queue formed by A1, such as Figure 3 As shown, it represents the allocation rule of the DNA tree. The first-level tree has four branches, which are considered by ACGT respectively, and the following ones are all three.

[0041] A is the abbreviation of Adenine, which means adenine; C is the abbreviation of Cytosine, which means cytosine; T is the abbreviation of Thymine, which means thymine; G is the abbreviation of Guanine, which means guanine.

[0042] Two important biological constraints will be given special consideration: one is the GC content. The four nodes in the first layer are arranged in order according to their corresponding probabilities. Therefore, by assigning A and T to the nodes with the highest or lowest probabilities, the AT content can be made close to the GC content to the greatest extent, which can be verified in subsequent experiments. The second is that the maximum continuous repeat length in the coding DNA sequence is 2. The reasoning is as follows: In the second layer and subsequent branches, the Goldman coding rule is used to mark the codewords on the path from the root node to the leaf node, thereby ensuring that the same base does not appear continuously in each source symbol.

[0043] GC content: The ratio of nucleotides G and C in all nucleotide chains. The closer this ratio is to 50%, the better. If this content is unbalanced, it will not only increase the probability of sequence errors, but also affect the efficiency and accuracy of PCR amplification.

[0044]

[0045] The numerator represents the sum of the two nucleotides, and |x| represents the length of the nucleotide chain.

[0046] Homopolymer: refers to the continuous arrangement of the same nucleotides. The occurrence of homopolymers can lead to a series of errors such as insertion, deletion and substitution. For example, AGGGGCT, it is easy to read long G as short G, thereby increasing the probability of errors in the synthesis process. It is generally required to control the running length not to exceed 3.

[0047] The encoding rules are shown in Table 1 below:

[0048] Table 1 Ternary coding table when the number of nodes is greater than 4

[0049]

[0050] When the number of nodes is greater than 4, ternary encoding is used. The specific process of ternary encoding is as follows: three nodes with the smallest weights are selected each time and merged into a new node. During the encoding process, a ternary value (0, 1, 2) is assigned to each branch of the tree with reference to Table 1, and nucleotides are replaced according to the rotation rule; the ternary encoding process is repeated until four nodes are left, and then quaternary encoding is adopted. The quaternary encoding process is as follows: quaternary values ​​(0, 1, 2, 3) are assigned to the four nodes, and they are replaced according to 0-A, 1-C, 2-G, and 3-T. The path from the root node to the leaf node is the final codeword.

[0051] It can be seen that the encoding of the current bit is related to the previous nucleotide, and the same nucleotide will not appear repeatedly, so the appearance of homopolymers can be avoided. This encoding rule ensures that there are at most two consecutive identical bases in the final nucleotide chain, which can be described as follows: If the codeword of the previous source symbol ends with T, then in order to maintain the GC content constraint, no restrictions are imposed in the first layer of the DNA tree, and all four bases are allowed to be used as marker symbols, so the current source symbol may start with T, which represents the worst case, and there are two repeated nucleotides appearing consecutively in the encoded nucleotide chain. Specifically,

[0052] A3. Based on the above description, a DNA tree will be formed. An example is Figure 4 As shown, this embodiment takes A corresponding to the node with the highest probability and T corresponding to the node with the lowest probability as an example for explanation. If there are six source symbols abcdef, the probabilities of occurrence are 5 / 15, 4 / 15, 2 / 15, 2 / 15, 1 / 15, and 1 / 15, respectively. Since 6 is an even number, a DNA tree can be formed. First, the three smallest probabilities are added and then sorted, which are 5 / 15, 4 / 15, 4 / 15, and 2 / 15. These four probabilities correspond to the four nodes of the first layer, which are marked T, G, C, and A respectively. The corresponding codewords are read from the first layer downward to obtain the final code table. For example, the codeword corresponding to f is CG. According to the obtained code table, the binary sequence is read and encoded string by string.

[0053] The DNA encoding scheme proposed in this paper is to improve the storage density of DNA, thereby saving costs and improving efficiency. Under the premise of determining the storage target, the storage density is mainly related to the number of scan bits.

[0054] The following uses the encoding process of the string of letters 'adadcfbebacbaba' as an example to illustrate the implementation process of the present invention:

[0055] First, convert 'adadcfbebacbaba' directly into a binary sequence through ASCII code, and get: 0110000101100100011000010110010001100011011001100110001001100101011000100110000101100001011000010110000100110000101100001

[0057] Then set the length of the binary substring to 8 and divide the above binary sequence to get:

[0058] 01100001 01100100 01100001 01100100 01100011 01100110 0110001001100101 01100001 01100001 01100010 01100001 011000010 01100001.

[0059] By counting the above binary substrings, we get the probabilities shown in Table 2:

[0060] Table 2 Probability of occurrence of binary substrings

[0061] Binary Substring Probability a=01100001 5 / 15 b=01100010 4 / 15 c=01100011 2 / 15 d=01100100 2 / 15 e=01100101 1 / 15 f=01100110 1 / 15

[0062] The DNA tree constructed according to the probabilities shown in Table 2 is as follows Figure 4 The obtained code table is shown in Table 3. According to the code table shown in Table 3, the corresponding encoding result is TCATCAACGGCTGTAGTGT; for example, nucleotide A appears continuously, and the last A of the code word CA corresponding to the previous letter d is the same as the first A of the code word A corresponding to the current letter c.

[0063] Table 3 Code table corresponding to the letter string 'adadcfbebacbaba' in this embodiment

[0064] Binary Substring Probability Codeword a=01100001 5 / 15 T b=01100010 4 / 15 G c=01100011 2 / 15 A d=01100100 2 / 15 CA e=01100101 1 / 15 CT f=01100110 1 / 15 CG

[0065] Experimental Results

[0066] Through experiments, we found that the effect of scanning the number of bits is significantly greater when it is an even number than when it is an odd number. This is because the number of bits in a binary file is an even number. If the number of bits to be scanned is an odd number, it is necessary to add 0 after the binary string to complete this process, which will have a bad impact on the storage density, and the distribution of the signal source after scanning does not match the distribution of the source file. Therefore, we focus on observing the relationship between the storage density and the number of bits to be scanned when the number of bits to be scanned is an even number.

[0067] Storage density Figure 5 As shown, its unit is bit / nucleotide. According to the image, it can be found that when the number of scan bits is 8 or its integer multiples, the storage density is better. This can be explained as follows: for text, each letter corresponds to its unique binary code, and the ASCII code is composed of an eight-bit binary string; for images, the entire image is composed of many pixels, and the depth of a pixel is also eight bits. Therefore, when the number of binary data bits is an integer multiple of 8, the probability difference between the source symbols is closest to the difference in the original rules of the file, and the encoding effect is better. As the number of scan bits increases, the storage density tends to rise overall.

[0068] This is because increasing the number of scan bits means that when the text or image is converted into a binary sequence, the codeword encoded by the DNA tree can represent more binary string bits, so the storage density can be improved. On the other hand, as the number of binary string bits of the source symbol increases, the types of source symbols in a longer text or image will also increase (there are 2 to the power of n types of source symbols, n is the number of scan bits). Although it may not strictly conform to this law, it is definitely an increasing trend. Therefore, it is necessary to increase the code length of the codeword by increasing the depth of the tree to represent more source symbols, and the encoding complexity increases. The length of the codeword corresponding to some binary substrings with low probability of occurrence is even the same as the length of the source itself. From this perspective, the storage density will be reduced accordingly. In addition, the inability to increase the number of scan bits indefinitely is also related to the encoding time. If the number of scan bits increases, the encoding time will also show a significant increasing trend.

[0069] Compared with the method based on fixed-length coding, the encoding effect of this improved encoding method also depends on the characteristics of the input file because its entire process depends on the occurrence probability of the source symbols. If the source symbols are significantly different, the encoding effect will be better.

[0070] GC content Figure 6 As shown, its value is strictly between 45% and 55%, close to 50%, which meets the biological constraints required by the DNA chain. The biological constraints that the DNA chain in the present invention must meet also include homopolymers and repeated sequences. Homopolymers refer to sequences composed of repeated identical bases. This sequence may be normal in an organism, but may cause problems during artificial synthesis and sequencing. The first layer of the encoded DNA tree in the present invention has four branches, and the following layers are all three branches, which can effectively ensure that the same nucleotides do not appear continuously; repeated sequences may cause DNA polymerase to "slip" during replication or sequencing, resulting in insertion or deletion errors. The method of the present invention can effectively avoid the occurrence of repeated sequences. Figure 6 The upper-threshold in the figure indicates the upper limit of the threshold value, which is 55% in this embodiment; the down-threshold indicates the lower limit of the threshold value, which is 45% in this embodiment.

[0071] Encoding time Figure 7As shown, it can be seen that the time required for compressing images is significantly greater than that for text, because the image itself contains a larger amount of data. For image encoding, its running time first increases and then decreases. The initial increase is because as the length of the substring increases, the number of source symbols increases, so a deeper DNA tree needs to be built to meet this demand, and accordingly, the time will increase. However, when the length of the block becomes larger, the time will gradually decrease. This is because as the value further increases, the number of binary substrings that need to be encoded decreases, so the time will also decrease accordingly. Since both biological constraints are met here, according to the storage density and storage time, it can be concluded that when the string length is equal to 16, the storage effect will be better.

[0072] Figure 5 , 6 In 7, text means text, and image means image.

[0073] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, the present invention may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of the claims of the present invention.

Claims

1. An improved DNA compression encoding method, characterized in that: include: S1, converting the original text or image into a binary string, and then dividing the binary string into binary substrings of length i; S2, treating each binary substring as a source symbol, and calculating the probability of occurrence of each source symbol; S3, constructing a DNA tree based on the calculated probability; S4. Use four nucleotides to mark along the path starting from the root node to obtain the codeword corresponding to each source symbol.

2. An improved DNA compression encoding method according to claim 1, characterized in that: Step S1 is specifically as follows: S11. For text, use binary ASCII for conversion; for images, each pixel has eight bits, and the pixels are read one by one for conversion; S12. If the length of the binary string is an integer multiple of i, then the binary string is directly divided according to the length of each binary substring being i; otherwise, 0 is added to the end of the binary string so that the length of the binary string after adding 0 at the end is an integer multiple of i, and then the binary string is divided according to the length of each binary substring being i.

3. An improved DNA compression encoding method according to claim 2, characterized in that: Step S2 is specifically as follows: scan all binary substrings, calculate the probability of each binary substring appearing, each substring corresponds to a node, and its probability corresponds to the probability of the node, and obtain the node queue of the DNA tree.

4. An improved DNA compression encoding method according to claim 3, characterized in that: Step S3 is specifically as follows: S31. If the number of source symbols is an even number, the number of layers of the DNA tree to be constructed is calculated by the following formula. Otherwise, for nodes with a probability of 0, the number of layers of the DNA tree to be constructed is calculated by the following formula: x-2*(n-1)-3=1 Among them, x is the number of binary substrings, n is the number of tree layers, 2 means that two nodes are reduced each time the encoding is performed, and 3 means that the last encoding is performed when there are only four tree nodes left, and finally there is only one root node; S32. Arrange the initial node queue in descending order according to the probability of occurrence of the binary substring. Each time encoding is performed, add the three nodes with the smallest probability and then rearrange them until four nodes are left. S33. Add the probabilities corresponding to the four remaining nodes in step S32 to get 1, which is the root node of the DNA tree. At this time, the first-level node of the DNA tree is 4, and the nodes of other levels are 3.

5. An improved DNA compression encoding method according to claim 4, characterized in that: Step S4 is specifically as follows: S41. Use A, C, G, and T to represent the four branches of the first layer in the DNA tree. According to the biological GC content constraint, A and T are assigned to the two branches at the edge, corresponding to the node with the smallest probability and the node with the largest probability, respectively, while G and C are assigned to the two branches in the middle, corresponding to the two nodes in the middle probability. S42. After the four branches of the first layer are marked, the branches of the remaining layers are assigned values ​​according to the Goldman coding rule. The specific assignment rules are as follows: