An Optimization Method for DNA Storage Encoding Based on the RBS Algorithm
Through the DNA storage coding optimization method based on RBS algorithm, binary data is converted into base sequences and divided into different frequency files, and data block merging and word matching are performed, which solves the problem of low DNA storage density and achieves efficient and stable information storage.
Patent Information
- Application Number
- CN202210945577.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-08-08
AI Technical Summary
The existing DNA storage solutions have low storage density and cannot effectively realize the information storage potential of DNA molecules.
The DNA storage coding optimization method based on the RBS algorithm includes converting binary data into base sequences, segmenting them into base files of different frequencies, and merging data blocks, RH mixed encoding and word matching, and finally segmenting the base sequence into a group of 120nt to generate high-quality encodings.
It improves the storage density and encoding quality of DNA storage, realizes lossless storage of text and picture files, and improves the practicality and stability of DNA storage.
Smart Images

Figure CN115271069B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNA storage coding, and particularly relates to an optimization method for DNA storage coding based on the RBS algorithm. Background Art
[0002] Since the 20th century, life sciences have developed rapidly. In 1988, Davis14 et al. demonstrated the first instance of DNA storage for the first time. In 1995, Baum theoretically proposed the first DNA storage model. In recent years, with the development of DNA storage, people have been able to store encoded data on 3D printed plastic rabbits, also realized reading a specified piece of information from electrodes according to the coding method, and also completed in vivo storage. Compared with silicon-based calculators that use 0 and 1 binary for storage, DNA storage can be regarded as a carbon-based calculator that stores quaternary information through four bases: adenine (A), guanine (G), cytosine (C), and thymine (T). They can complete information archiving by arranging in a specific order. With the rapid development of DNA synthesis and sequencing methods recently, DNA storage will be a very competitive storage solution in the future. However, the actual storage density of current DNA information storage solutions is low, and the information storage potential of DNA molecules cannot be well exerted. Summary of the Invention
[0003] In view of the characteristics of DNA storage coding, the present invention provides an optimization method for DNA storage coding based on the RBS algorithm, which can generate codes with higher storage density and better coding quality, and effectively improves the efficiency, practicability, and stability of DNA storage.
[0004] To achieve the above object, the present invention proposes an optimization method for DNA storage coding based on the RBS algorithm, including the following steps:
[0005] Step 1: Convert the DNA storage coding file into a binary data stream, and then perform a random exclusive OR operation on each bit of data;
[0006] Step 2: Convert the binary data stream into a base sequence;
[0007] Step 3: Detect the frequency of each base in the base sequence, sort the base frequencies, and store the frequency sorting in a Buffer file;
[0008] Step 4: Split the base sequence, and the splitting method is as follows:
[0009] (1) First, set the position of the first frequency base to "1", and set the positions of the remaining frequency bases to "0" to generate a First_frequency file;
[0010] (2) After deleting the first-frequency bases, set the positions of the second-frequency bases to "1" and the positions of the third- and fourth-frequency bases to "0" to generate the Second_frequency file;
[0011] (3) After deleting the second-frequency bases, record the third-frequency bases and the fourth-frequency bases in the form of characters to generate the Other_frequency file;
[0012] Step 5: Perform data block merging processing on the First_frequency file;
[0013] Step 6: Perform RH hybrid encoding processing on the Second_frequency file;
[0014] Step 7: Perform character merging processing on the Other_frequency file;
[0015] Step 8: Encode all files, convert them into base sequences, and split the base sequences into short sequences.
[0016] Further, the binary data stream is converted into a base sequence according to the reference conversion (00->A, 01->G, 10->C, 11->T).
[0017] Further, the base frequencies are arranged from largest to smallest.
[0018] Further, the specific content of step 5 is as follows: Take the alternation of 0 / 1 data as the boundary, and merge the data 0 block and the data 1 block: First, set an operation bit, which represents which data block is being processed currently. Secondly, compress the number of data in this data block. The compression rule is to record the data in half in the file First_forward. When the number is even, record (even / 2) data in the file. When the number is odd, record (odd + 1) / 2 data in the file; At the same time, record whether the number is odd or even in the file First_backward. If it is even, record "0", and if it is odd, record "1".
[0019] Further, the specific content of step 6 is as follows: First, perform RLE encoding on the file to generate a data pattern of (data bit, frequency bit); Secondly, use Huffman encoding to re-encode the frequency bit; Finally, record the first data in the Buffer file, separate it from the previous data with "#", and then arrange the data of the frequency bit in sequence according to the Huffman-generated encoding to generate the change file.
[0020] Further, step 7 is specifically as follows: After merging the same characters, if they are still the same character, record the first character when different characters are merged, and record it in the file Other_forward; at the same time, record the order in the file Other_backward. If it is the merging of the same characters, record "1", if it is the merging of different characters, record "0", and if there are single bases left, store them in the buffer file Buffer.
[0021] Further, in step 8, the splitting of the base sequence into short sequences is as follows: The base sequence is split into groups of 120 nt, including a 6-nt address bit, a 110-nt data bit, and a 4-nt error correction bit.
[0022] The advantages of the above technical solutions adopted by the present invention compared with the prior art are as follows:
[0023] 1. It can store text and picture files losslessly, and the designed algorithm process can improve the storage density of the DNA storage model;
[0024] 2. The RBS algorithm splits the file and performs different encoding operations on different files according to their content sorting rules, greatly improving the encoding quality;
[0025] 3. A DNA storage coding optimization method based on the RBS algorithm proposed by the present invention can achieve lossless storage of files such as text and pictures, and at the same time can give play to the storage advantage of high DNA storage density, generate high-quality codes, and improve the practicability and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flowchart of a DNA storage coding optimization method based on the RBS algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] A DNA storage coding optimization method based on the RBS algorithm includes:
[0029] Step 1: Convert the file into a binary data stream, and then perform a random exclusive OR operation on each bit of data.
[0030] Step 2: Convert the binary data stream according to the benchmark conversion (00 -> A, 01 -> G, 10 -> C, 11 -> T) into a base sequence.
[0031] Step 3: Detect the frequencies of each base in the base sequence, arrange them from largest to smallest, and store the frequency sorting in a Buffer file.
[0032] Step 4: Split the base sequence. The splitting method is as follows:
[0033] (1) First, set the position of the base with the first frequency to "1", and the positions of the bases with the remaining frequencies to "0" to generate a First_frequency file.
[0034] (2) Then, after deleting the base with the first frequency, set the position of the base with the second frequency to "1", and the positions of the bases with the third and fourth frequencies to "0" to generate a Second_frequency file.
[0035] (3) Then, after deleting the base with the second frequency, record the bases with the third and fourth frequencies in character form to generate an Other_frequency file.
[0036] Step 5: Perform data block merging processing on the First_frequency file. Use the alternating position of 0 / 1 data as the boundary to merge the data 0 block and the data 1 block. First, set an operation bit, which represents which data block is currently being processed. Secondly, compress the number of data in this data block. The compression rule is to record the data in half in the file First_forward. When the number is even, record (even / 2) data in the file. When the number is odd, record (odd + 1) / 2 data in the file; at the same time, it is necessary to record whether it is odd or even in the file First_backward. If it is even, record "0", and if it is odd, record "1".
[0037] Step 5: Perform RH hybrid encoding processing on the Second_frequency file. First, perform RLE (run-length encoding) encoding on the file to generate a data pattern of (data bit, frequency bit); secondly, use Huffman encoding to re-encode the frequency bit; finally, record the first data in the Buffer file, separated from the previous data by "#", and then arrange the data of the frequency bit in the order generated by Huffman to generate a change file.
[0038] Step 6: Perform character merging processing on the Other_frequency file. After merging the same characters, they remain the same character. When merging different characters, record the first character in the file Other_forward; at the same time, it is necessary to record the order in the file Other_backward. If it is the merging of the same characters, record "1", and if it is the merging of different characters, record "0". If there are single bases left, store them in the buffer file Buffer.
[0039] Step 7: Encode all files using the rotation coding shown in Table 1, convert them into base sequences, and split the base sequences into groups of 120 nt, including a 6-nt address bit, an 110-nt data bit, and a 4-nt error-correction bit.
[0040] Table 1. Rotation Coding Table
[0041]
[0042] Example 1
[0043] DATA: CCACCCACCCACCCACCCACCCGACCAACGTCCGCTCGCTCGTACGTACG
[0044] In this example, the number of DNA bases is 50, and the process of random XOR and conversion of binary data into DNA codewords according to the benchmark conversion is omitted.
[0045] As Figure 1 shown, the processing method for the above base sequence is as follows:
[0046] Step 1: File splitting. The frequency word distribution of this example is: CAGT. According to this distribution, the following 3 files are split: The first frequency word file sets base C to "1" and other bases to "0"; the second frequency word file deletes base C and then sets base A to "1" and other bases to 0; the third and fourth frequency word files delete bases C and A and record bases G and T.
[0047] First_frequency: 11011101110111011101110011001001101010101000100010
[0048] Second_frequency: 1111101110000000010010
[0049] Other_frequency: GGTGTGTGTGTG
[0050] Step 2: Data block merging. Write operation bits according to the data block content, record the content halved in First_forward, and record the odd-even identification bit in First_backward.
[0051] First_forward: 1101111111111111110110010011010101010000000
[0052] First_backward:1111111
[0053] Step 3: RLE and Huffman hybrid coding. First, perform RLE coding on Second_frequency, and then perform Huffman coding on the number bits of RLE.
[0054] change:0011011110
[0055] Step 4: Character merging. Merge them in pairs, record the first character in Other_forward, then convert it into a binary file according to the reference coding, and record the same and different character identifiers in Other_backward.
[0056] Other_forward:GTTTTT
[0057] Other_change:011111111111
[0058] Other_backward:100000
[0059] Step 5: Buffer file. Record the frequency word distribution, the last character of each file, the first data bit of RLE coding, etc.
[0060] Buffer:CAGT#1#
[0061] The present invention conducts a simulation experiment on this method by means of MATLAB in the operating environment of Intel(R) Core(TM) i5-10500 3.10GHz CPU and 16.00GB memory, Windows 10. The experimental results show that the results of the method of the present invention are better than those of other algorithms.
[0062] Method comparison:
[0063] In order to compare the advantages of the algorithm in terms of storage density, other current classic DNA storage encodings are compared. The comparison results are shown in Table 2. The optimal results will be in bold.
[0064] Table 2. Comparison of DNA storage coding and experimental results
[0065]
[0066]
[0067] In order to compare the advantages of the algorithm in terms of coding quality, an encoding test is conducted on an interval color picture (pixel values 65 and 122 appear alternately), and the results are evaluated in terms of entropy change, free energy, Hamming distance, etc. The comparison results are shown in Table 3. The optimal results will be in bold.
[0068] Table 3. Interval data evaluation, including information entropy, cross entropy, relative entropy, free energy, Hamming distance.
[0069]
[0070] Comparative analysis:
[0071] Judging from the results in Table 2, the DNA storage coding optimization method based on the RBS algorithm can achieve the best storage density among all methods, which indicates that the DNA storage coding optimization method based on the RBS algorithm has good practicability in practical applications. Judging from the results in Table 3, the best results are obtained in the three criteria of relative entropy, free energy and Hamming distance, indicating that the coding quality of this method is good and the stability of the generated codewords is improved, which shows that the DNA storage coding optimization method based on the RBS algorithm has good performance.
[0072] In summary, a path planning method based on the RBS algorithm proposed by the present invention has better practicability and stability compared with other advanced methods, and can improve the storage density and coding quality of DNA storage coding.
[0073] Obviously, the above embodiments are only examples given for clear illustration, rather than limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to list all implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. An optimized method for DNA storage encoding based on the RBS algorithm, characterized in that, It includes the following steps: Step 1: Convert the DNA storage encoding file into a binary data stream, and then perform a random exclusive OR operation on each bit of data; Step 2: Convert the binary data stream into a base sequence; Step 3: Detect the frequency of each base in the base sequence, sort the base frequencies, and store the frequency sorting in a Buffer file; Step 4: Split the base sequence, and the splitting method is as follows: (1) First, set the position of the base with the first frequency to "1", and the positions of the bases with the remaining frequencies to "0" to generate a First_frequency file; (2) Then, after deleting the base with the first frequency, set the position of the base with the second frequency to "1", and the positions of the bases with the third and fourth frequencies to "0" to generate a Second_frequency file; (3) Then, after deleting the base with the second frequency, record the bases with the third and fourth frequencies in character form to generate an Other_frequency file; Step 5: Perform data block merging processing on the First_frequency file; Step 6: Perform RH hybrid encoding processing on the Second_frequency file; Step 7: Perform character merging processing on the Other_frequency file; Step 8: Encode all files, convert them into a base sequence, and split the base sequence into short sequences.
2. The DNA storage coding optimization method based on the RBS algorithm according to claim 1, wherein The binary data stream is converted into a base sequence according to the benchmark conversion (00->A, 01->G, 10->C, 11->T).
3. The DNA storage coding optimization method based on the RBS algorithm according to claim 1, wherein The base frequencies are arranged from largest to smallest.
4. A DNA storage coding optimization method based on the RBS algorithm according to claim 1, characterized in that, Specifically, Step 5 is as follows: Take the alternating position of 0 / 1 data as the boundary, and merge the data 0 block and the data 1 block: First, set an operation bit, which represents which data block is currently being processed. Secondly, compress the number of data in this data block. The compression rule is to record the data in half in the file First_forward. When the number is even, record (even / 2) data in the file. When the number is odd, record (odd + 1) / 2 data in the file; at the same time, record whether the number is odd or even in the file First_backward. If it is even, record "0", and if it is odd, record "1".
5. A method for optimizing DNA storage encoding based on the RBS algorithm according to claim 1, characterized in that, Specifically, Step 6 is as follows: First, perform RLE encoding on the file to generate a data pattern of (data bit, frequency bit); secondly, use Huffman encoding to re-encode the frequency bit; finally, record the first bit of data in the Buffer file, separated from the previous data by "#", and then arrange the data of the frequency bit in the order of the Huffman-generated encoding to generate a change file.
6. A method for optimizing DNA storage encoding based on the RBS algorithm according to claim 1, characterized in that, Specifically, Step 7 is as follows: After merging the same characters, it is still the same character. When merging different characters, record the first character and record it in the file Other_forward; at the same time, record the order in the file Other_backward. If it is the merging of the same characters, record "1", if it is the merging of different characters, record "0", and if there are single bases left, store them in the buffer file Buffer.
7. A method for optimizing DNA storage coding based on the RBS algorithm according to claim 1, characterized in that The splitting of the base sequence into short sequences described in step 8 is as follows: the base sequence is split into groups of 120 nt, including a 6-nt address bit, a 110-nt data bit, and a 4-nt error-correction bit.