DNA-based Dynamic Equilibrium System, Data Storage Method, and Decoding Method
The dynamic balancing system and encoding/decoding methods for DNA storage optimize data storage efficiency and accuracy by stabilizing DNA strands and providing efficient compression and encryption, addressing high synthesis and sequencing costs and encoding limitations.
Patent Information
- Application Number
- CN202210953717.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-08-10
AI Technical Summary
Existing DNA storage technology faces high synthesis and sequencing costs, and lacks effective coding methods, so it is impossible to maximize the storage space of DNA utilization.
A DNA-based dynamic equalization system is adopted, including a acquisition module, an encoding module, a Chinese character compression module and a decoding module. Through pseudo-random memory, exclusive or checksum dynamic equalization technology, efficient encoding and decoding of data is achieved.
It improves the efficiency and accuracy of data storage, reduces coding complexity, realizes efficient storage of Chinese characters and data stability, reduces storage redundancy, and improves data recovery rate and access speed.
Smart Images

Figure CN115423096B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage, and in particular to a DNA-based dynamic balancing system, a data storage method and a decoding method. Background Art
[0002] As the world enters the information age, the exponential increase in data volume under the background of big data has higher requirements for data storage devices. DNA, as a high-density storage medium, is the choice of natural evolution for hundreds of millions of years. At the same time, due to its excellent structural performance, mainly phosphodiester bonds and bases based on hydrogen bond complementary pairing rules, this combination structure has stability and can also complete information replication, retrieval, transcription and other tasks through enzyme catalysis. The size of a single DNA base is less than 1 nanometer. At the same time, DNA is flexible and can curl up and extend in three-dimensional space. Its information storage density is about 1019bit / cm3. Compared with the smallest unit of electronic computers, triodes, 5 nanometers are almost the limit of current silicon crystal materials. The storage density of hard disks is about 1013Bit / cm3, and that of flash memory is about 1016Bit / cm3. In terms of storage cost, DNA can be stably and long-term stored in low temperature to room temperature environments, does not require energy maintenance, and can be stored for at least 100,000 years. The main problems that currently plague DNA storage are, on the one hand, its high synthesis and sequencing costs, and on the other hand, the lack of effective encoding methods that can maximize the storage space of DNA. To this end, we proposed a DNA-based dynamic balancing system, data storage method, and decoding method. Summary of the invention
[0003] 1. Technical issues to be resolved
[0004] In view of the deficiencies in the prior art, the present invention provides a DNA-based dynamic balancing system, a data storage method and a decoding method to solve the above-mentioned problems.
[0005] (II) Technical solution
[0006] In order to achieve the above-mentioned purpose, the present invention provides the following technical solutions:
[0007] A DNA-based dynamic balancing system, comprising:
[0008] An acquisition module for acquiring first data, an encoding module for encoding the first data, a Chinese character compression module for compressing storage space and encryption, and a decoding module for decoding the encoded DNA molecular chain.
[0009] A data storage method of a DNA-based dynamically balancing system comprises the following steps:
[0010] Step 1: Obtain the first data through an acquisition module;
[0011] Step 2: Encode the first data through an encoding module to obtain a DNA molecular chain including the number of times of pseudo-random memory storage, encoded data, equalization, re-equalization, verification, and forward and reverse primers. The forward primer is located at the 3' end of the DNA molecular chain, and the reverse primer is located at the other end of the DNA molecular chain, that is, the 5' end.
[0012] Preferably, the source of the first data can be any existing electronic file stored in a computer.
[0013] Preferably, after the forward primer undergoes reverse substitution through the DNA base complementary pairing rule, it becomes the reverse primer, and the lengths of the forward primer and the reverse primer are the same;
[0014] In each of the primers, the content of guanine and cytosine accounts for a preset ratio of the total content of guanine, cytosine, adenine, and thymine contained in the primer.
[0015] Preferably, when the first data is Chinese text information in the second step, it is necessary to perform Chinese character compression, including the following steps:
[0016] S1: First, analyze the target document to obtain the total number of characters K in the document. After removing duplicate characters, count the number of characters S that appear in the document, and sort them from high to low according to the character appearance frequency to obtain a character table;
[0017] S2: Calculate the base length corresponding to the Chinese character. The calculation rule is: According to the number of characters S, calculate that a character should be encoded with n bits of bases, and n needs to ensure >= log4S;
[0018] S3: Establish a base sequence to label the text characters. The generation method of the base sequence is:
[0019] (a) Use a length of m as a seed to inject into a pseudo-random generator to randomly generate a non-repeating base sequence of n bits;
[0020] (b) Screen the generated base sequence. If there are 2 or more consecutive bases, fill them in corresponding from the bottom to the top of the character table, otherwise fill them in corresponding from the top to the bottom until all character base code correspondences are filled to obtain a filled dictionary table, which is used as the dictionary and key for DNA document statistics.
[0021] (4) According to the base sequence corresponding to the character in the dictionary table, convert all Chinese characters in the original document into DNA base sequences to obtain a compressed and encrypted DNA base sequence.
[0022] Preferably, the steps for obtaining the DNA molecular chain are as follows:
[0023] Step 1: Divide the first data into several data sub - packets;
[0024] Step 2: Generate a 0 / 1 random matrix through the pseudo - random number generator of an electronic computer;
[0025] Step 3: According to the element values and element positions in the random matrix, specify the corresponding data sub - packets for exclusive - OR encoding; perform dynamic equalization on the encoded data;
[0026] Step 4: Record the data during the process. According to a preset alphabet, map the running digital sum to guanine, cytosine, adenine, and thymine, and perform encoding to obtain and form the DNA molecular chain data.
[0027] Preferably, check for errors in the generated DNA strand. When an error occurs in the generated DNA strand, it can be recognized;
[0028] The possible errors in the DNA strand include: substitution, insertion, and deletion errors;
[0029] For insertion and deletion errors, they can be judged by the length of the strand. For substitution errors;
[0030] Adopt the error - correction method of XOR exclusive - OR check. By performing exclusive - OR bit - by - bit, finally obtain the result of exclusive - OR of all base data, and it can be judged whether there is a substitution error in the relevant DNA molecular chain.
[0031] Preferably, perform screening processing on the DNA molecular chain, screen out the folded disordered structures and / or unbounded running digital sum codes in the DNA molecular chain, and perform dynamic equalization on the strands that fail the screening.
[0032] A decoding method for a DNA - based dynamically equalizable system includes the following: According to the packing result, use the decoding module for decoding processing.
[0033] (III) Advantageous Effects
[0034] Compared with the prior art, the DNA - based dynamically equalizable system, data storage method, and decoding method provided by the present invention have the following advantageous effects:
[0035] 1. In the process of encoding the first data into a DNA molecular chain in the DNA-based dynamic balancing system, data storage method, and decoding method, various constraints are added to the first address and the second address, enabling efficient and accurate reading of the encoded data. For example, the Hamming distance between the first address and the second address is greater than or equal to half of the length of the first address, reducing the possibility of address selection errors during reading; the prefix of the first address is different from the prefix of the second address and the suffix of the second address, avoiding the possibility of matching errors during reading; the content of guanine and cytosine in the prefix of each primer accounts for a preset ratio of the total content of guanine, cytosine, adenine, and thymine contained in the primer, resulting in high accuracy when sequencing is required to read the encoded data in advance.
[0036] 2. The DNA-based dynamic balancing system, data storage method, and decoding method establish a DNA-based storage architecture, which can efficiently complete the conversion between the first data information and the DNA molecular chain, with high conversion efficiency.
[0037] 3. When encoding data on a DNA double strand in the DNA-based dynamic balancing system, data storage method, and decoding method, the absolute stability of the DNA double strand is achieved by setting a dynamic balancing method, avoiding situations such as homopolymers and uneven GC content.
[0038] 4. The DNA-based dynamic balancing system, data storage method, and decoding method design an encoding system based on a random matrix and use a multi-bit packing method, reducing the complexity of encoding and decoding, shortening the encoding time, and increasing the decoding rate; reducing the order of storage redundancy and improving the efficiency of encoding and data recovery rate in the data center, achieving highly sensitive random access and accurate rewrite addressing.
[0039] 5. The DNA-based dynamic balancing system, data storage method, and decoding method provide a new set of compression and encryption methods for Chinese character text information, achieving efficient storage of Chinese character text. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic flowchart of the steps of the data storage method in a specific embodiment of the present invention;
[0041] Figure 2 It is a schematic diagram of a DNA molecular chain in a specific embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of the generation of a matrix in a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment
[0045] See Figures 1-3 , the data storage method of the DNA-based dynamic balancing system provided in this embodiment includes the following steps:
[0046] Obtain the first data;
[0047] Encode the first data to obtain a DNA molecular chain. When the first data is Chinese text information, it is necessary to compress the Chinese character characters. Chinese is currently the most widely used language in the world. According to statistics, as of 2020, the number of people using Chinese globally exceeds 1.6 billion, accounting for more than 20% of the world's total population. China is currently the country with the largest number of Internet users in the world. (As of September 2019, China had 830 million users accessing the Internet), and the amount of Chinese document resources created is huge. Currently, the common electronic storage scheme for Chinese characters is UTF-8, and a single Chinese character character needs to occupy 3-4 bytes. Referring to the ideal DNA storage density of 2bit / nt, each Chinese character character requires an information space of 12-16 bases. However, in fact, there are about 5,000 common Chinese characters, and these 5,000 characters can be stored with 4-7 bases, which undoubtedly has great room for improvement in storage space. We have proposed a method for dynamic programming storage of Chinese characters, which can perform dynamic optimal planning according to the document information to be stored. At the same time, it has good compression ability and encryption function. The implementation method is:
[0048] (1) First, analyze the target document to obtain the total number of characters K in the document. After removing duplicate characters, count the number of characters S that appear in the document, and sort them in descending order of character appearance frequency to obtain a character table;
[0049] (2) Calculate the base length corresponding to the Chinese character. The calculation rule is: according to the number of characters S, calculate that a character should be encoded with n bits of bases, and n needs to ensure >= log4S.
[0050] (3) Establish base sequence pairs to label text characters. The generation method of the base sequence is as follows: (a) Inject a seed with a length of m into a pseudo-random generator to randomly generate a non-repeating base sequence of n bits; (b) Screen the generated base sequence. If there are two or more consecutive bases, fill them in corresponding order from the bottom to the top of the character table, otherwise fill them in corresponding order from the top to the bottom until the base coding corresponding filling of all characters is completed, obtaining a filled dictionary table, which serves as the dictionary and key for DNA document statistics;
[0051] (4) According to the base sequence corresponding to the characters in the dictionary table, convert all Chinese characters in the original document into DNA base sequences, obtaining compressed and encrypted DNA base sequences.
[0052] The coding method after obtaining the first data is as follows:
[0053] 1) Determine the coding space, which is mainly determined according to the synthesis technology and budget of the DNA synthesizer.
[0054] 2) Data segmentation: Divide L into n chunks (n = S / / P + 1) to form L1, L2, L3... Ln, and the length of each chunk is p nt;
[0055] 3) Generate a random vector. Use the random function in Python to generate an integer in the range (0, 2n), and record the number of times the random generator generates. Convert the generated integer into an n-bit binary data vector. Whether the elements from the low bit (0) to the high bit (n - 1) are 1 indicates whether the chunks of L1, L2... Ln obtained by splitting the original file L in step 2 participate in coding.
[0056] 4) XOR the chunks corresponding to the positions where the elements in the generated random vector are 1 to obtain the corresponding Droplet;
[0057] 5) Each time a random vector is generated, record the number of times the pseudo-random number generator generates and store it in the Times position. Store the Droplet obtained in step 4 in the Data payload position to form a Pack. Continuously repeat step 4 until the storage space of Times is exhausted, generating a total of 4t Packs;
[0058] 6) Equalize the obtained Packs. First, perform conditional screening, requiring homopolymers (no more than 3 consecutive bases), GC content (45%-55%), and whether the primers are repeated in the information coding space. If the DNA strand passes the screening, mark the XE bit as 0 (10 A's), and store this pack as an available strand; if the pack does not pass the screening, equalize the Packs. Use the Adapter sequence as the seed of the random generator to generate k random bases of length (t + p) nt, and perform exclusive OR on these random bases and the bases of the pack that did not pass the screening to obtain a new base sequence. Continuously repeat the equalization until the base sequence passes the screening or when k > 410, indicating that the storage space for the XE equalization bit is exhausted. For the strand that passes the screening after k equalizations, store k in the XE bit and record the base sequence of the pack that passed the equalization; for the strand that still cannot pass the screening after the XE equalization bit space is exhausted, choose to discard it.
[0059] 7) Re-equalize. Since a 10-nt Xe equalization bit is added, the Xe bit itself may also cause problems such as excessive or too low homopolymers and GC content. We introduce a 2-nt re-equalization bit. By re-equalizing the XE bit, the equalization method is the same as that for equalizing Packs in step 6). After re-equalization, the number of equalization times is stored in the XEE bit, and the Packs and XE bit are stored as the equalized base sequence. The strands that do not pass the re-equalization are directly discarded. It should be noted that the re-equalization does not change the base sequence of the Packs, and the judgment conditions for re-equalization consider the homopolymer situation of the connection part between XE and Packs and the connection part between XE and XEE;
[0060] 8) Assume that the number of DNA strands after step 7 is G, and the storage redundancy is m. To ensure the robustness of the final stored content, first randomly select (n + m) strands from G, and then randomly select n 1000 times from (n + m). Generate the corresponding degree distribution sequence according to Times (the length of this sequence vector is n), construct an (n * n)-dimensional matrix, and use the exclusive OR Gaussian elimination method to solve the obtained triangular matrix, ensuring that at least 550 times the elements on the main diagonal of the finally obtained triangular matrix are all 1.
[0061] 9) Synthesis. Starting from the packs obtained after self-equalization of the (n + m) strands finally selected in step 8), perform exclusive OR in units of 3 nt up to XE and XEE, and finally perform exclusive OR to obtain a 3-nt base sequence, which is stored as the XC bit attached after XEE.
[0062] 10) Add forward and reverse primers to the obtained strands and perform biosynthesis to complete the final DNA storage.
[0063] Such as Figure 2As shown, in this embodiment, the purpose is to achieve highly sensitive random access and accurate rewrite addressing. The principle of the proposed method is that in random access, each block in the system must be equipped with an address sequence primer that allows unique selection and amplification through DNA. Due to information encoding, the length of the data information is uncertain. The data chain may be a complete data message, or it may be a segment or extremely scattered chain information. Therefore, primers are added to the front and back ends of the DNA molecular chain to identify different chain information. In this embodiment, the length of the DNA molecular chain is taken as an example of 700 bps for illustration. In other embodiments, it may be other lengths. The forward primer and the reverse primer are synthesized at both ends of the block sequence, and the forward primer and the reverse primer are respectively stored in short blocks with a length of 20 bps. The block sequence is used to store data. The encoding process is to encode the first data of the block sequence to obtain encoded data, so as to obtain the encoded DNA molecular chain.
[0064] In this embodiment, the first data is divided into a number of second data. Each second data includes a primer, a pseudo-random memory storage times bit, an encoded data bit, an equalization bit, a reequalization bit, and a check bit. The primer is a code that describes the DNA prefix synchronization. The number of selected co-encoded words is controlled by the target to make the rewrite as simple as possible and avoid error propagation due to variable code lengths. The "word encoding" operation is as follows: First, the words in different text data messages are counted and tabulated in a dictionary. Each word in the dictionary is transformed into a binary sequence long enough to allow encoding of the dictionary. Second, a quaternary model is used (for example, 00, 01, 10, 11 can be encoded one-to-one with the bases A (adenine), T (thymine), C (cytosine), G (guanine) in DNA).
[0065] The content of G (guanine) and C (cytosine) in the prefix of each primer accounts for a preset ratio of the total content of G (guanine), C (cytosine), A (adenine), and T (thymine) contained in the primer. The preset ratio includes but is not limited to 45% - 55%.
[0066] The reason is that since DNA stores information through G (guanine), A (adenine), C (cytosine), T (thymine), and A pairs with T, and C pairs with G to form a stable double-stranded structure. Whether it is single-stranded DNA or double-stranded DNA, it can store information in the form of binary encoding. In double-stranded DNA, a prefix with a certain GC content is required (so that the GC content accounts for about 45% - 55% of the total). Because a DNA double strand with a 50% GC content is more stable than a DNA double strand with a lower or higher GC content and can have better coverage during sequencing. Since the user information in the encoding is achieved through prefix synchronization, it is important to impose GC content constraints on the addresses and their prefixes, because the latter requirement also needs to ensure that all segments of the encoded data block can have the GC content removed;
[0067] The steps for decoding are as follows:
[0068] 1) Extract the DNA strand obtained by sequencing according to the forward and reverse Adapters, and judge and analyze the length of the extracted DNA strand. If the length is lower than the target length or higher than the target length, it indicates that a deletion / insertion error has occurred, and this strand is discarded;
[0069] 2) Perform exclusive OR reduction on the data according to the length specified by the XC bit. If the reduced data is inconsistent with the data stored in the XC, otherwise it indicates that a substitution error has occurred, and this strand is discarded;
[0070] 3) If the number of strands K completed by sequencing >= n, the formal decoding process can be started. First is information recovery. Inject the primer information as the seed of the random number generator into the pseudo-random number generator. According to the number of times the pseudo-random number generator generated recorded at the XEE position, recover the information of the XEE, and then according to the information of the number of times the pseudo-random number generator generated after recovery of the XEE, recover the data of the Times and data payload bits to obtain the original data information.
[0071] 4) Generate the corresponding random integers from 0 to 2**(n) according to the number of times the pseudo-random number generator generated corresponding to the Times, and convert them into n-bit binary vectors to form a K*n-dimensional binary distribution sequence.
[0072] 5) Convert the data in the Data payload from base to decimal data. The degree distribution sequence matrix forms an augmented matrix, and the XOR Gaussian elimination method is used to solve the matrix. (The solution rules are as follows: First, combine the K-order matrix D with the Data matrix of K rows and 1 column to construct an augmented matrix. Next, judge along the matrix diagonal (i from 0 - k). If D[i][i] = 1, then judge all the sequences below it along the column. If D[j][i] = 1, then XOR all the data in the i-th row with all the data in the j-th row. If D[i][i] = 0, then look down along the column until D[j][i] = 1 is found, swap the two rows, and then continue to look down. If there is still D[j][i] = 1, then XOR the i-th row with the j-th row to ensure that an upper triangular matrix is constructed, and all areas below the matrix diagonal are 0. Then, according to the previous step, perform the reverse operation to eliminate all 1s above the diagonal to 0, and the unique S1...Sk and Data1...Datak can be obtained. Complete the decoding process.) Obtain the data solutions of S1, S2...Sn, convert this part of the solutions to binary sequences, and then convert them into base sequences to obtain ChunksL1, L2...Ln.
[0073] 6) Complete the decoding of the stored text data according to Table 2.
[0074] The DNA-based dynamic balancing system includes:
[0075] An acquisition module for acquiring first data;
[0076] An encoding module for encoding the first data to obtain a DNA molecular chain, where the DNA molecular chain includes encoded data, a forward primer, and a reverse primer. The forward primer is located at one end of the encoded data, and the reverse primer is located at the other end of the encoded data. The encoded data includes the pseudo-random memory storage times bit, the encoded data bit, the balancing bit, the re-balancing bit, and the check bit;
[0077] A Chinese character compression module: used for special encoding of Chinese characters when the first data is Chinese text information, which plays a role in compressing the storage space and encrypting.
[0078] A decoding module for decoding the encoded DNA molecular chain and decoding it back to the complete first data.
[0079] An embodiment of the present invention also provides a storage medium storing a program, and the program is executed by a processor to complete the DNA-based data storage method.
[0080] This embodiment also provides a decoding method for the DNA-based dynamic balancing system, including the following steps: perform decoding processing according to the packaging result.
[0081] When the DNA is obtained by using the above-mentioned packaging method, when decoding is required to read the DNA, only the packaging result needs to be utilized, which is equivalent to only sending the row positions and column positions of 1s in the generated matrix G, instead of sending the entire generated matrix G. Then, only decoding needs to be performed according to the received row positions and column positions to restore the generated matrix to translate the original data. At this stage, the encoding and decoding process and application of the LT code are to encapsulate and transmit unit original data, while for unit data packet transmission, in the case of a large amount of data, problems such as occupying more memory and bandwidth will occur, and the effectiveness and reliability will decline when the data is relatively large. Through the above-mentioned processing method, that is, some bits after encapsulation encoding transmission are used to replace the original unit number transmission, which greatly reduces the data volume in the storage space, reduces the storage redundancy, and improves the decoding success rate.
[0082] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data storage method for a DNA-based dynamic balancing system, characterized in that, Including the following steps: Step 1: Obtain first data through an acquisition module; Step 2: Encode the first data through an encoding module to obtain a DNA molecular chain including the storage times of a pseudo-random memory, encoded data, equalization, re-equalization, verification, and forward and reverse primers. The forward primer is located at the 3' end of the DNA molecular chain, and the reverse primer is located at the other end of the DNA molecular chain, that is, the 5' end; After the forward primer undergoes reverse substitution through the DNA base complementary pairing rule, it becomes the reverse primer, and the forward primer and the reverse primer have the same length; In each of the primers, the content of guanine and cytosine accounts for a preset ratio of the total content of guanine, cytosine, adenine, and thymine contained in the primer; In the second step, when the first data is Chinese text information, it is necessary to perform Chinese character compression, including the following steps: S1: First, analyze the target document to obtain the total number of characters K in the document. After removing duplicate characters, count the number of characters S that appear in the document, and sort them in descending order according to the character occurrence frequency to obtain a character table; S2: Calculate the base length corresponding to the Chinese characters. The calculation rule is: According to the number of characters S, calculate that a character should be encoded with n bits of bases, and n needs to ensure >= log4S; S3: Establish a base sequence to label the text characters. The generation method of the base sequence is: (a) Inject a seed with a length of m into a pseudo-random generator to randomly generate a non-repeating base sequence of n bits; (b) Screen the generated base sequence. If there are 2 or more consecutive bases, fill them in corresponding from the bottom to the top of the character table, otherwise fill them in corresponding from the top to the bottom until all character base code correspondences are filled to obtain a filled dictionary table, which is used as the dictionary and key for DNA document statistics; (4) According to the base sequence corresponding to the characters in the dictionary table, convert all the Chinese characters in the original document into DNA base sequences to obtain a compressed and encrypted DNA base sequence; The steps for obtaining the DNA molecular chain are as follows: Step 1: Divide the first data into several data sub-packets; Step 2: Generate a 0 / 1 random matrix through the pseudo-random number generator of an electronic computer; Step 3: According to the element values and element positions in the random matrix, specify the corresponding data sub-packets for exclusive OR encoding; perform dynamic equalization on the encoded data; Step 4: Record the data in the process. According to a preset alphabet, map the running numbers and combine them into guanine, cytosine, adenine, and thymine, and perform encoding to obtain and form DNA molecular chain data.
2. The data storage method of the DNA-based dynamic balancing system according to claim 1, characterized in that: The source of the first data can be any existing electronic file stored in a computer.
3. The data storage method of the DNA-based dynamic balancing system according to claim 1, characterized in that: When checking for errors in the generated DNA strand, errors in the generated DNA strand can be recognized; Possible errors in the DNA strand include: substitution, insertion, and deletion errors; For insertion and deletion errors, they can be judged by the length of the strand. For substitution errors; An error correction method using XOR checksum is adopted. By performing XOR operation bit by bit, the result of XOR of all base data is finally obtained, which can be used to determine whether there are substitution errors in the relevant DNA molecular chain.
4. The data storage method of the DNA-based dynamic balancing system according to claim 1, characterized in that: The DNA molecular chain is screened, and the folded disordered structures and / or unbounded running digital sums in the DNA molecular chain are screened out. For the chains that fail the screening, dynamic equalization is performed.
5. A DNA-based dynamic balancing system, characterized in that, A data storage method applied to a DNA-based dynamic equalization system as described in claim 1, comprising: An acquisition module for acquiring the first data, an encoding module for encoding the first data, a Chinese character compression module for compressing the storage space and encrypting, and a decoding module for decoding the encoded DNA molecular chain.
6. A decoding method for a DNA-based dynamic balancing system, characterized in that, A data storage method applied to a DNA-based dynamic equalization system as described in claim 1, including the following: According to the packing result, decoding processing is performed using the decoding module.
Citation Information
Patent Citations
DNA storage coding method for optimizing Chinese storage
CN111600609A
Molecular data storage systems and methods
WO2021105974A1