A DNA storage dual coding method, device and readable storage medium
Through the dual encoding method of DNA storage, DNA code word dictionary and redundant segmentation are defined, and error correction codes and primers are added, which solves the problems of capacity, fault tolerance and address index in DNA encoding storage, and realizes efficient information storage and error correction.
Patent Information
- Application Number
- CN202210815586.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-09
AI Technical Summary
The existing DNA encoding storage methods have shortcomings in data sequence conversion, DNA capacity, fault tolerance and address indexing, resulting in problems such as low information density, high computational complexity and limited capacity.
The dual encoding method of DNA storage is adopted. By defining the n-base length DNA code word dictionary, using double-encoded characters and Arabic numerals, combining redundant segmentation and error correction codes, adding primers, information segmentation and sequencing, improving capacity and error tolerance.
It increases the capacity of a single group of DNA, simplifies the maintenance of address indexes, enhances error correction capabilities, improves error tolerance, and reduces the cost of DNA synthesis.
Smart Images

Figure CN115188422B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a DNA storage dual encoding method, device and readable storage medium, belonging to the fields of biotechnology and information technology. Background Art
[0002] DNA storage refers to using DNA as a medium to store information data. DNA molecules are extremely stable, do not require additional energy consumption for maintenance, can be preserved for millions of years in a low-temperature and dry environment, and are extremely tiny (a single base plus a phosphoribose backbone, totaling only thirty to forty atoms), with extremely high storage density. In theory, one gram of DNA can store all the movies ever filmed by humans to date, or all the books and paintings and other information materials in human history. These two advantages far exceed all other current information storage media, such as paper, optical discs, magnetic disks, magnetic tapes, etc. DNA molecules do not rely on specific reading devices, which is also different from current electronic devices. For example, the most popular floppy disks thirty years ago had very troublesome data reading due to the discontinuation of the production of reading devices. However, DNA is the genetic material of almost all organisms on Earth. No matter how future technology develops, humans will always have various methods to read DNA data, and no matter how the instrument equipment changes, it will not affect the reading of DNA sequence information. As a storage medium, the disadvantages of DNA are that data cannot be arbitrarily modified, the reading and writing time is slow, and the cost is very high. Nevertheless, using DNA to long-term backup and store archival materials and other inert information, that is, high-value data information materials that are rarely used but very important, still has broad prospects.
[0003] An important research direction of DNA storage is the encoding method, that is, how to convert binary digital information data into DNA sequences. In addition to maximizing the information density as much as possible, the following four problems also need to be solved: First, sequence limitations, that is, the GC content of the DNA sequence should be reasonable, generally 40% to 60%, and the single-base repeat sequences should be as few as possible; Second, the address indexing problem of each DNA molecule, that is, enough address indexes can be encoded; Third, how to correct errors when random errors occur during the synthesis and sequencing of the sequence; Fourth, how to redundantly recover when some sequences are lost. Currently, there are various studies solving these problems according to different ideas, and various related encoding methods have emerged, but these encoding methods still have some defects, mainly including limited capacity of a single group of DNA due to insufficient address index quantity, the need to record and maintain address index information, low information density and error tolerance rate, and high computational complexity in the decoding stage. Summary of the Invention
[0004] Technical Problems to be Solved by the Invention
[0005] In view of the problems existing in the existing DNA encoding storage method in terms of data sequence conversion, DNA capacity, and error tolerance, the present invention proposes a DNA storage dual encoding method, device, and readable storage medium.
[0006] Technical solution
[0007] To achieve the above object, the technical solution provided by the present invention is as follows:
[0008] A DNA storage dual encoding method includes the following steps:
[0009] Step 1, dictionary definition: Define a DNA codeword dictionary with an n-base length. Each DNA codeword dual-encodes a character and an Arabic numeral. The DNA codewords in the dictionary satisfy: the content of G and C bases in the codeword is between 40% and 60%, and the Hamming distance between any two codewords encoding the same Arabic numeral in the dictionary is ≥2;
[0010] Step 2, information segmentation: Segment the binary information to be stored according to an m-character length;
[0011] Step 3, generate redundant segments: Generate a number of redundant segments for a group of segmented information according to a certain redundancy ratio. The length of the redundant segments is the same as the information segmentation length in Step 2. Each character of the redundant segments is generated according to the error correction code generation rule by the corresponding characters at the same position in all information segments in the group. The characters at the same position in the redundant segments and the original information segments together form a basic error correction unit;
[0012] Step 4, segment the scale numbers according to an m-Arabic numeral length;
[0013] Step 5, according to the dictionary defined in Step 1, use DNA codewords to dual-encode the segmented characters generated in Steps 2 and 3 and the segmented Arabic numerals generated in Step 4;
[0014] Step 6, add error correction code: Add an error correction code with a certain base length to the DNA sequence generated in Step 5;
[0015] Step 7, add primers at both ends: Add primers with a certain base length to both the head and tail ends of the DNA sequence obtained in Step 6;
[0016] Step 8, DNA synthesis and sequencing: Synthesize the DNA sequence obtained in Step 7 and sequence the DNA fragments.
[0017] Furthermore, in Step 1, the base length n of the DNA codeword is ≤20.
[0018] Furthermore, the error correction code based on which the redundant segments are generated in Step 3 is Reed-Solomon code, RS erasure code, RS error correction code.
[0019] Further, the scale numbers in step 4 are normal numbers, approximate normal numbers, and their derived numbers.
[0020] Further, the error correction code in step 6 is an RS erasure code, an RS error correction code, a Hamming code, and a check code. The error correction code can correct one or more base errors in combination with the error detection mechanism of the codeword itself.
[0021] A DNA storage dual coding device includes:
[0022] A memory for storing a computer program;
[0023] A processor for implementing the steps of the above DNA storage dual coding method when executing the computer program.
[0024] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above DNA storage dual coding method are implemented.
[0025] Beneficial effects
[0026] The DNA storage dual coding method of the present invention simultaneously solves the problems of single-base information density and address indexing, increases the capacity of a single set of DNA to exceed the requirements of practical applications, makes the maintenance of address indexing information extremely simple and convenient, endows the codeword with error correction ability, can greatly improve the error tolerance rate, and will have important guiding significance for the technical upgrade of reducing the cost of DNA synthesis. Description of the drawings
[0027] Figure 1 It is a step diagram of the DNA storage dual coding method of the present invention. Specific implementation manners
[0028] To further understand the content of the present invention, the present invention will be described in detail in combination with the drawings and specific implementation manners.
[0029] Dictionary generation: A DNA codeword refers to a short segment of DNA sequence with a fixed length, such as 7 bases, 8 bases, 10 bases, 14 bases, etc. Here, the present invention is described by taking 8 bases as an example. There are four types of DNA bases: A / C / G / T. The total number of 8-base sequences is 4 8 ^8, totaling more than 60,000 different sequences. Remove the codewords whose GC content of the sequence does not meet 40% to 60%, mainly including repetitive sequences, such as AAAAAAAA; and sequences with uneven GC content, such as ATTAATTA (GC content is 0%). Then divide the selected qualified codewords into 10 groups, and the Hamming distance between any two codewords in each group is at least 2.
[0030] For the sake of simplicity and convenience, two types of code words composed of different base combinations are directly selected here. These two types of code words contain 1A-2C-2G-3T, such as AGCTTCTG; and 2A-1C-3G-2T, such as GATGGACT. The GC content of each code word in the above two types of code words is 50%, and the number of single bases does not exceed 3, so at most 3 consecutive bases can be repeated inside the code word. And among all the code words in these two categories, if two code words are randomly selected for comparison, the bases in at least 2 of the 8 positions in the sequence are different, that is, the Hamming distance is at least 2. The advantage of this setting is that when any single base of a code word has a mutation error, the erroneous invalid code word can be directly judged according to the base composition of the code word.
[0031] According to the permutation and combination formula, it can be concluded through simple calculation that the number of different sequences of these two types of code words is 1680 (8*7*6*5*4 / 2 / 2). 1280 of each of the two types of code words are selected to form a code word library with a total of 2560 members. The two types of code words are arbitrarily divided into 5 groups, each with 256 code words, for a total of 10 groups, respectively called Group 0, Group 1, Group 2, ..., Group 9, each code word encodes an 8-bit binary sequence, and each character in the 256 binary sequences (00000000 to 11111111) has 10 code words to encode, and these 10 code words belong to one of the above groups 0 to 9. In other words, there are 256 binary 8-bit characters and 10 Arabic numerals in any combination, a total of 2560, each of which is encoded using a code word, and each combination corresponds to a code word, thus forming a double encoding dictionary. Any code word in the dictionary has a dual meaning. A single code word encodes a binary byte (00000000-111111111) and also encodes an Arabic numeral (0-9) of the group to which it belongs.
[0032] The generation of the above dictionary selects 8 bases to encode an 8-bit binary sequence and an Arabic numeral. This is only for the convenience of briefly explaining the implementation steps of the present invention. It does not mean that this is the only choice, nor does it mean that this choice is efficient. The codeword length and basic information characters can be flexibly adjusted according to the actual demand for coding efficiency.
[0033] Segmentation of information and scale numbers: The binary information to be stored in DNA is segmented according to characters (such as 8-bit bytes). In this embodiment, it is segmented every 25 characters in length. At the same time, redundant segments are added according to a ratio (such as 10%). For example, 100 redundant segments are added after every 900 data segments. The specific method is to arrange 900 segments into a matrix with 25 characters per row and 900 characters per column. The 900 characters in each column are calculated according to the rules of RS erasure code to generate 100 characters of redundant segments. Finally, a matrix with 1000 rows and 25 characters per row is obtained. Each row is an information segment or a redundant information segment, and each column contains 900 data characters and 100 redundant characters. The scale numbers are selected as the value of pi (π). Ignoring the decimal point, it is regarded as a string of numbers and is also segmented every 25 digits until a number string with the same number and length as the information segments (including redundancy) is finally obtained.
[0034] RS erasure code is one of the most commonly used data security backup methods in the field of computer information technology. Its data recovery ability is described as follows: for any n data blocks, after adding m data blocks through the RS error correction code calculation rules, among the total n + m data blocks, if any no more than m data blocks are lost, all the lost data blocks can be correctly recovered through calculation. The difference from the RS error correction code is that the prerequisite for the erasure code is that all known data blocks must be correct.
[0035] Double coding: According to the dictionary generated in the previous steps, accurately select codewords to double code the corresponding information segments and the number strings of scale numbers in sequence. According to the previous example, the codeword length is 8. Therefore, 25 codewords (DNA sequences with a length of 200 bases) double code 25 bytes of information and a number string of 25 Arabic numerals. After all the information segments and redundant segments are encoded, a DNA sequence library with a length of 200 bases is obtained.
[0036] Add error-correcting codes and primers: Add an 8-base error-correcting code after each 200-base DNA, and add primer sequences of 21 bases in length at both ends to obtain a DNA sequence library with a length of 250 bases. In most cases, the 8 bases of the error-correcting code generated here will not be repetitive sequences or have uneven GC content. Even if this situation occurs, considering the worst-case scenario, for example, 8 consecutive Gs or Cs with an out-of-range GC content, since it is only 8 bases in length, it will not cause too much impact. In the conventional DNA sequences in organisms, repetitions of this length also occur. Such codewords are not allowed to appear in the previous codeword selection because when storing information, multiple similar codewords may be juxtaposed, resulting in DNA sequences with a length of dozens to hundreds of bases and severely uneven GC content. There is at most one such error-correcting code in each sequence, and within the length of the entire sequence, these eight bases will not cause too much change in the overall GC content, so it is acceptable. At the same time, redundancy is considered in the coding segments, and the complete loss of a very small number of segments will not affect the complete decoding and reading of information.
[0037] The purpose of adding an error-correcting code is to correct possible base errors in this sequence. The error-correcting code can be a parity check code, Hamming code, RS error-correcting code, erasure code, etc. In principle, adding a sequence error-correcting code is not a necessary condition in the present invention because the codeword itself has a certain error-detecting ability, and adding a sequence error-correcting code only enhances the error-correcting ability.
[0038] Synthesize DNA sequences: Synthesize all DNA sequences by an array method, which can simultaneously synthesize 10,000 to 1 million DNA fragments (i.e., oligos, oligonucleotides, single-stranded DNA). This is the most commonly used sequence synthesis method for DNA storage at present, and there are different ways such as chemical synthesis or photochemical synthesis. The synthesized oligonucleotides store information and are all mixed together and cannot be directly separated individually, so it is called an oligo pool. Synthesizing array oligonucleotides is a mature commercial service, and currently, the longest can reach 300 bases in length. However, considering the actual situation, it is generally mainly 120 to 250 bases in length. The starting number of molecules for synthesizing a single sequence is approximately 10^6 to 10^7, that is, about one million to ten million molecules. Calculated according to a synthesis efficiency of 99% for each base in each round, the number of finally successfully synthesized molecules that meet the length is only about 1%, that is, about ten thousand to one hundred thousand for each sequence. This quantity is extremely scarce, and the quality of the finally obtained single sequence is about pg level (10^-12 grams).
[0039] DNA Fragment Sequencing: After more than half a century of development, DNA sequencing has been widely used in the field of life sciences and is a relatively mature commercial activity. The principles and processes of sequencing will not be elaborated here. After sequencing the oligonucleotide pool, generally through next-generation sequencing (NGS) or third-generation sequencing, the sequence data of these oligonucleotides (oligos) will be obtained. Although some companies claim that the synthesis error rate is less than one-thousandth, literature research shows that for oligonucleotide pools synthesized by different companies, the actual error rate of individual bases after sequencing totals around 0.5%-1%. In addition, the copy numbers of different oligonucleotides are also different, so the sequencing depth (coverage multiple) also determines the range of oligonucleotides that the sequencing results can cover. Here, it means that the deeper the sequencing depth, the higher the proportion of the number of oligonucleotide sequences obtained to the total number of theoretical sequences. Generally speaking, however, there will always be some sequences lost for various reasons. Generally, at least 90% of the oligonucleotide sequences will have a copy number difference within a range of 10 times. This means that with a 10-fold depth (each sequence is sequenced 10 molecules on average), it is empirically estimated that more than 98% of the sequences can be obtained, and about 95% of the sequences are completely correct after error correction.
[0040] Decoding Data: DNA fragments can be sequenced at a depth of more than a thousand times to detect almost all sequences. Generally, 10-fold depth sequencing can obtain about 98% of the complete oligonucleotide sequences, with a base error rate of about 1%. However, most errors can be corrected through the alignment of multiple molecular copies and sequence error correction codes, and about 95% of the completely correct sequences can be obtained.
[0041] Each single sequence contains 25 codewords. If all errors are corrected, the correct character string and digital string can be decoded. Even if only a single molecule of the sequence is sequenced, on average, each molecular sequence has two base errors. According to the dictionary, at least 23 codewords of the correct encoded numbers can be decoded. Based on this digital string containing two blanks or incorrect numbers, it is sufficient to accurately locate this sequence precisely.
[0042] Suppose the number of sequence fragments of a group of oligonucleotides reaches 1020, and each encodes 25 bytes. Therefore, such a group of DNA can encode 2.5 * 10 21 bytes (ZB). For comparison, the total capacity of all electronic storage devices (including various magnetic disks, tapes, optical discs, memory cards, etc.) produced worldwide each year currently does not reach the ZB level.
[0043] The reason for the ability to accurately locate based on an incomplete digital string is that although the decimal value of pi (π) has not been proven to be a normal number in the strict mathematical sense, existing data indicates that it should be a normal number (the binary value of π has been proven to be a normal number). The characteristic of a normal number is that the ten Arabic numerals from 0 to 9 tend to be randomly and evenly distributed, and the longer the digital length, the closer it is to average randomness. For π, in terms of a length exceeding 10^8 digits, for the application of the present invention, this randomness is already good enough. Therefore, when segmented into 25 - digit (or other digit numbers such as 18, 20, etc.) segments, with each segment in a row, among the 25 columns thus formed, the numbers from 0 to 9 in each column are also randomly distributed, and the quantity of each number from 0 to 9 is basically the same, each accounting for about one - tenth. Therefore, if the decoded string is incomplete, it can be compared and located one by one. Each digit can narrow down the range by one - tenth. So theoretically, 19 digits can determine a unique position within the range of a total of 10^20 sequences. On the premise of having 2 missing or incorrect codewords, there will be 23 digits in the above - mentioned string. The total number of these 10^20 strings far exceeds the actual demand. Currently, it is generally believed that a single - group DNA fragment capacity of 10^12 bytes can meet the actual application requirements, and the number of digital strings required for this capacity reaches 10 11 levels, that is, 10 - digit numbers can be used for precise positioning. Moreover, the value of pi has currently only been calculated to approximately 6 * 10 13 digits.
[0044] After 10 - fold deep sequencing, all measured sequences, whether they are one molecular copy or multiple molecular copies, and whether they contain a small number of errors or have all been corrected, can be accurately located. Therefore, according to the 10% redundancy design of the original RS erasure code, only 90% of the correct codewords need to be obtained to completely recover 100% of the codewords. With about 95% of the correct codewords, of course, the correct data can be recovered.
[0045] Suppose the RS error - correcting code has a 10% redundancy, and among the sequences with 98% coverage obtained by sequencing, about 95% are completely correct without errors, 3% of the sequences contain errors, and 2% of the sequences are missing. According to the accurate positioning of the address index, after filling the missing codewords with random numbers and examining a single RS error - correcting code block, the maximum proportion of error code elements totals 2.24% (2% + 3% * 1% * 8), while the RS code redundancy is 10%, which can correct up to 5% of the total code element errors within a single code block. Obviously, it is sufficient to correct 2.24% of the code element errors.
[0046] After using the erasure code or error - correcting code to recover all data, each information data is segmented and connected in the order of the digital strings of their respective accurate address indexes, which is the original binary data stored in DNA.
[0047] Method for constructing a codeword library: In addition to selecting codewords according to the base composition and dividing them into groups, there are also various methods for constructing a DNA codeword library. As long as two selected principles are followed to ensure compliance with the conditions, that is, the Hamming distance between any two codewords in a group of codewords is at least 2, which can be achieved in at least three different ways. The first is the remainder of the sum of base numbers modulo 4. According to the remainder of the sum of the base numbers of a codeword (A / C / G / T are regarded as the four numbers 0, 1, 2, 3 respectively) modulo 4, any DNA sequence can be divided into four categories, namely the four categories with remainders of 0, 1, 2, and 3. For example, for the codeword AGCTTCTG, calculated according to the classification, 0 + 2 + 1 + 3 + 3 + 1 + 3 + 2, the sum is 14, and the remainder modulo 4 is 2; for the codeword GATGGACT, 2 + 0 + 3 + 2 + 2 + 0 + 1 + 3, the sum is 13, and the remainder modulo 4 is 1. For two codewords with the same remainder, any single-base error will cause the remainder of the sum of the base numbers of this incorrect codeword modulo 4 to change; the second is the different base compositions of the codewords, with at least three of the four bases ACGT having different quantities. For example, the compositions of two 10-base codewords, one group is 2A - 2C - 3G - 3T; the other group is 2A - 3C - 1G - 4T. The quantities of CGT in these two groups are different. Any single-base error that mutates into another base will not cause the codeword to become another group. Therefore, these two groups of codewords can be encoded together; the third is that the difference in the number of the same base is at least 2. For example, for two codewords with different base compositions, the number of A bases in one group of codewords is 2, and the number of A bases in the other group of codewords is 4. Then the Hamming distance between these two groups of codewords is also at least 2, because any single-base error in the same codeword will not allow it to become another group of codewords.
[0048] After the above method classifies the codewords in the dictionary, the Hamming distance between any two codewords within a class is at least 2, which is to ensure that a single error in a single codeword can be detected. If the Hamming distance is at least 3, then a single-base error within the codeword can be directly corrected. This can be achieved by using Hamming codes or RS codes for a single codeword. The cost is a reduction in storage efficiency, and the advantage is a significant increase in the error tolerance rate.
[0049] Another method for constructing a codeword library is the brute-force method. For example, for 10-base length codewords, all codewords are examined one by one. For codewords that meet the conditions of uniform GC content and few repeat sequences, they are placed in one of the groups from group 0 to group 9. Then the next codeword is examined. If it meets the conditions (uniform GC content, few repeat sequences), before throwing this codeword into a group, it is examined whether the Hamming distance between it and any other codeword in this group is at least 2. If not, change to another group. If none of them are suitable, it is directly discarded. During this process, try to keep the number of codewords in these ten groups uniform. Using a computer program, the screening of codewords can be quickly completed according to the selection method.
[0050] A DNA codeword library with a Hamming distance of at least 2 or 3, and DNA codewords with several bases are mainly considered in terms of cost or storage efficiency, as well as the technical specifications of synthetic oligonucleotides. When the error rate is relatively high, a Hamming code or RS codeword with a high error tolerance and the ability to correct single errors can be used to construct the codeword library. And shorter codewords can greatly increase the capacity of a single group of DNA and are more sensitive to errors. For example, compared with 8-base codewords, 6-base codewords can hold 33 in a 200-base length sequence, 8 more than the 25 codewords of 8-base. The length of the digital string can reach 33 digits, so the capacity of a single group of DNA is increased by one hundred million times (10^8), and the probability of 2-base mutations occurring within a single codeword is also obviously smaller than that of 8-base codewords.
[0051] Generally speaking, different from DNA synthesis in the field of life sciences, DNA storage does not require that every sequence can synthesize a copy of the completely correct sequence, because error-correcting codes can be used to correct it. Reducing the requirement for the correct rate can prompt researchers to adopt low-cost DNA synthesis technologies, thereby reducing the economic cost of DNA storage.
[0052] Segmented redundancy code: As mentioned above, the segmented redundancy code uses RS erasure code, and RS error-correcting code can also be used here. Which one to choose depends on the specific situation and requirements. Generally speaking, under the same conditions in the present invention, the RS erasure code is more efficient than the error-correcting code. After redundancy of the error-correcting code, its error-correcting ability is at most only 50% of the redundancy, and after redundancy of the erasure code, the recovery ability is 100% of the redundancy. The codewords of the present invention can initially detect errors themselves. Therefore, when combined with the RS erasure code, the efficiency will be higher than that of the RS error-correcting code. Here is just a simple comparison. In principle, both of these two redundancy codes can be used. The codewords of the present invention can detect errors by themselves. Whether there are a small number of molecules sequenced for a single sequence or multiple molecules sequenced, it can be determined whether the codeword is in error. Strictly speaking, when it is determined that there is an error, it must be in error. When it is determined that there is no error, there is a very small probability that two or more base errors occur within the codeword, and by chance, the incorrect codeword and the correct codeword encode the same Arabic numeral. Due to uneven sequencing, when only a small number of, such as two or three molecules, of a target sequence are sequenced, the codeword can detect errors, and most of the errors can be excluded, but the correct determination cannot be made for errors using the majority voting principle. Therefore, whether it is the RS error-correcting code or the erasure code, the error detection of a single codeword of the present invention can greatly improve the efficiency and function of the RS code.
[0053] Sequence error-correcting code: The sequence error-correcting code is an optional item in the present invention. For a single sequence, there are four error-correcting codes that can be used, namely RS erasure code, RS error-correcting code, Hamming code, and remainder error-correcting code. In the present invention, the RS error-correcting code is still slightly more efficient.
[0054] Since each codeword of the present invention can detect errors of at least one base by comparing with the scale numbers itself, it can be used in combination with error-correcting codes and improve the efficiency of error-correcting codes. For example, under normal circumstances, for Hamming codes, 200 data bits require 8 parity bits to correct a single base error. However, since a single codeword of the present invention can detect a single error, the Hamming code (12, 8) using four base parity codes can be used. Four parity bit bases check 8 data, and these 8 data respectively correspond to the sum of the first bases of all codewords, to the respective sums of the eighth bases, and check the remainder of modulo 4, which is equivalent to the parity check in quaternary. The results of the check are four types: 0 / 1 / 2 / 3, rather than two types: 0 / 1 in binary, and the others are the same. If a single base error occurs, the codeword itself or the comparison with the scale numbers gives the position of the error codeword, and the Hamming code gives which base is in error, so as to determine the position of the error and correct it. Similarly for correcting a single base error, the method of the present invention only requires four bases instead of eight bases.
[0055] In summary, error-correcting codes with 4 additional bases, RS erasure codes, Hamming codes or parity codes can all correct the errors of a single codeword (including error-correcting codes), and can detect two base errors of a single codeword hidden in the sequence. The above analysis is only based on the sequencing results of a single molecule. If deep sequencing obtains two or more molecule sequencings of a single sequence and refers to each other, it is easier to correct errors and obtain the correct sequence.
[0056] Scale numbers: Besides the pi (π), other irrational numbers such as Any irrational number that may be a normal number and has not been disproven, such as the natural constant e or the golden ratio, but pi is the simplest and most convenient scale number. Scale numbers have two functions. One is to serve as the address index of the sequence to accurately locate the encoded information. The other is to serve as a scale to check whether an error has occurred in the codeword. Even if a codeword looks like a legitimate codeword and exists in the dictionary, if the encoded number does not match the scale number, it will also be treated as an incorrect codeword. The error detection ability of the scale number is closely related to the codeword length and error rate of the sequence. In principle, the natural number sequence can also be directly used as the scale number. Using a sufficiently large number P that is relatively prime to 10 as the counting unit, such as 57, and counting by adding P each time while ignoring the hundreds digit, the sequence 0, 57, 14, 71, 28, 85,... can be obtained, and all numbers can eventually be traversed. The purpose of doing this is that if the index has a dozen digits, when using the natural sequence, the high digits of the first 1000 numbers are all 0. If there may be problems after converting such a group of numbers into DNA sequences due to index repeated digits, then these problems will be concentrated together, and these numbers are all in one redundant code block, so it may bring difficulties in error correction. By using the method of counting by adding P, the indexes of these repeated digits can be dispersed into different error correction code blocks, so the possibility of error correction difficulties will be greatly reduced. This problem of the present invention has actually been well avoided through the coding method. Even if the index numbers are all repeated, however, because the codewords are double-coded, the codeword sequence will not repeat due to the repetition of the index numbers. The codeword is determined by two factors, the encoded information and the index number. Unless both are repeated, the possibility of a repeated sequence is extremely small. Even if both are repeated, it is only a 10-base repeat of the codeword, and a certain proportion of sequence loss is allowed in the sequence redundancy. Nevertheless, using the natural sequence with counting by adding P is better than using the natural sequence because it can better assist in error correction, but obviously, using a sequence of normal numbers will have a better effect.
[0057] The present invention requires that for all codewords in the dictionary that encode the same basic number (such as Arabic numerals), the Hamming distance between any two codewords is at least 2, and it does not require that the Hamming distance between two codewords encoding different basic numbers must be at least 2. Obviously, after a single-base mutation of an incorrect base, even if this codeword is a legitimate codeword and can be found in the dictionary, the basic number it encodes will definitely not be the same as the scale number, and thus it will be identified as an incorrect codeword.
[0058] Whether it is required that the Hamming distance between any two codewords of all legal codewords is ≥2, and without this requirement, only requiring that the Hamming distance between any two codewords of all codewords encoding the same basic digit is ≥2, the difference lies in the total number of available codewords, so there is a slight difference in coding efficiency. For example, the total number of 10-base sequences exceeds 1 million. If they are divided into four groups according to the remainder of the sum of base digits modulo 4, and only one of the groups is used as the codeword library, compared with using all four groups as the codeword library, obviously the number of available codewords differs by 4 times. Therefore, the number of bits that can be encoded by a single codeword, theoretically, the latter can increase by 2 bits compared to the former. When using four groups of codewords, it is impossible to determine whether there is an error only from the codewords themselves and it is necessary to compare with the scale digits to make a judgment. But when using a single group of codewords, the four groups of codewords can be used to encode different types of data respectively for convenient identification, such as storing different types of files such as text, audio and video, pictures, digital data, etc. separately.
[0059] In principle, it is not a necessary condition of the present invention that the Hamming code distance between all codewords encoding the same digit is at least 2. Even without this requirement, it is still effective, but the error detection efficiency is a bit lower. In this case, most single-base errors can be corrected through multi-copy comparison of deep sequencing, and other errors and sequence losses can be corrected through the sequence redundancy RS code. Requiring that the Hamming distance between any two codewords encoding the same digit is at least 2 increases the error correction method. Even if there are only two copies of the sequence, it is possible to distinguish the right and wrong of the codewords, so the requirement for sequencing depth is reduced and the sensitivity of error detection is improved.
[0060] The scale digits only need to be approximately normal numbers, and so far, only about 10^13 digits of pi have been calculated. Therefore, derivative numbers of pi can be used, such as n times pi, as long as n is not a multiple of 10. It can be inferred by simple direct verification that nπ also has the characteristics of a normal number. Or use (n + 0.0123456789)π, where n ranges from 1 to 10^12 and 10^12 digits of pi are taken, which can generate enough digital strings for use. With enough scale digits, the present invention can store data in DNA with a high enough information density, and different primer combinations can be used to randomly read different files.
[0061] Sequence length: The lengths of oligonucleotides currently commercially available are generally within 300 bases, and stepped quotations are based on different lengths. The lengths commonly ordered by actual customers generally do not exceed 200 bases. Since the synthesis error rate is generally relatively fixed, the longer the length, the more incorrect bases will ultimately be contained. After increasing the length, the number of errors to be processed increases. Since none of the current other coding methods use a method similar to the fixed-length sequence codewords of the present invention, single-base error detection and correction cannot be performed, which will increase the difficulty of data processing. For the method of the present invention, increasing the length can not only increase the number of codewords in a single sequence, thereby increasing the capacity of a single set of DNA, but also has a higher actual cost performance. For example, the commercial quotation for the synthesis of an oligonucleotide pool with a length of 200 bases is only less than 5% higher than that with a length of 150 bases. However, after removing the primer sequences at both ends (a total of about 50 bases), the length of DNA that can carry information can be increased by 50%. The quotation for a length of 250 bases is about less than one-fourth higher than that for a length of 200 bases, and the length that can carry information can be increased by about one-third.
[0062] Using longer DNA sequences, combined with the codeword length and the number of information bits carried, appropriate adjustment can be made to achieve the best effect. If the DNA length reaches several thousand bases or longer and can accommodate hundreds of codewords, then the numbers encoded by the codewords can be the binary 0 and 1. Each codeword encodes a number of 0 or 1 and simultaneously encodes a character. For 100 codewords, there are 2^100 binary digit strings, and this address index number exceeds 10^30. The number of bits of the encoded character can be increased to improve the information density.
[0063] High-density coding: The codeword length determines the range of the number of codewords that can be screened. The total number of sequences with a length of 10 bases exceeds 10 6 ^6. If each single codeword encodes 16 bits, there are a total of 65,536 characters. After combining with 10 Arabic numerals, 655,360 codewords are required. Therefore, encoding 16-bit binary numbers with a 10-base sequence requires using two-thirds of all sequences as codewords.
[0064] Among all sequences with a length of 10 bases, through calculation, it can be known that the total number of sequences satisfying the condition of a GC content of 40% to 60% is 688,128, and the GC content of each codeword is between 40% and 60%. And for each of the 16-bit characters, only 655,360 codewords are needed when each is assigned 10 codewords. Therefore, by selecting about 655,360 from more than 680,000 sequences as codewords and through fine adjustment, it can be ensured that the Hamming distance between any two codewords encoding the same number is ≥2.
[0065] First, all 10-base codewords with a GC content of 40% to 60% are divided into four groups according to the remainder of the sum of base numbers modulo 4, namely remainder 0, remainder 1, remainder 2, and remainder 3. After careful analysis (calculating A / C / G / T as 0 / 1 / 2 / 3 respectively, and it is basically similar for other corresponding relationships), the number of codewords in these four groups is as follows: the remainder 0 group has a total of 172032, the remainder 1 group has a total of 169344, the remainder 2 group has a total of 172032, and the remainder 3 group has a total of 174720. The base compositions of each group are shown in Table 1. Each of these four groups encodes two Arabic numerals, and each group will use 131072 (65536 * 2), and a total of 8 Arabic numerals can be encoded. Additionally, a total of 65536 codewords are extracted from the remainder 0 group and the remainder 1 group to encode 1 Arabic numeral, and a total of 65536 codewords are also extracted from the remainder 2 group and the remainder 3 group to encode 1 Arabic numeral. In this way, for all the codewords encoding each Arabic numeral, the Hamming distance between any two codewords is at least 2. When extracting codewords from each group, according to the codeword composition and the number of bases, the extracted codewords can ensure that the Hamming distance is at least 2. Table 1 lists the base compositions of some 10-base codewords that meet the GC content of 40% to 60%. From these four groups of base compositions modulo 4, 32768 codewords can be extracted from each group.
[0066] Table 1 Base Compositions of Some 10-Base Codewords (GC% 40 - 60)
[0067]
[0068]
[0069] In addition, since the total number of codewords is more than 680,000, and actually only more than 650,000 are needed, all codewords with the same first four bases or the same last four bases can be removed. After calculation, there are 9248 such codewords (1156 * 8). This can ensure that for any arrangement of the codewords of the present invention, at most 6 consecutive single-base repeats can occur. Therefore, the sequence restriction screening conditions of the present invention are a GC content of 40% to 60% and no more than 6 single-base repeats. Currently, the standard of other methods is generally no more than 3 consecutive single-base repeats, but this standard is a man-made rigid regulation and not a necessary condition. In the gene DNA sequences of normal organisms, it is quite common for the continuous repeat of a single base to exceed 6, and there are also many cases where it exceeds 10. The most important thing for sequence restriction is the GC content. The single-base repeats of the present invention will not significantly affect the synthesis and sequencing of DNA sequences. Research also shows that there is no obvious difference in the error rate between synthetic genes and DNA storage when using array synthesis of oligonucleotide fragments.
[0070] High-fault-tolerant coding: When the Hamming distance between codewords is ≥3, it can correct errors by itself, so the fault tolerance rate of this coding is very high. As mentioned before, when encoding 16 bits of data with 10 bases, if 4-base error-correcting codes, RS error-correcting codes, or Hamming codes are added to each codeword, single-base errors can be corrected. When using 200-base-length oligonucleotides, with 40 bases for the primers at both ends, there are 160 bases available for encoding, which can accommodate 11 14-base codewords, totaling 154 bases. In theory, each sequence can correct up to 11 base errors, encode 176 bits of data and an 11-digit Arabic numeral string. After about 10% sequence redundancy and calculated according to the full length of 200 bases, it can ultimately achieve a single-base density of approximately 0.8 bits and a maximum capacity of 20 TB for a single group of DNA.
[0071] The maximum fault tolerance rate of high-fault-tolerant coding is approximately 7% (11 / 154). Currently, the error rate of array-synthesized oligonucleotides is less than 1%, and the best measured data is approximately around 0.7%. Therefore, combined with deep sequencing, the present invention can use a DNA synthesis method with a higher error rate to store data, thereby further reducing costs.
[0072] High-capacity coding: By using codewords with a compliant GC content, the sequence limitation problem is solved. Instead of using dual coding to solve the address indexing problem, a small number of codewords can be directly used to encode the address index, and other codewords can encode segmented information data.
[0073] For example, for a 200-base-length DNA fragment, when using 10-base codewords, two codewords are used to encode the address index, and other codewords encode the information. For 10-base codewords with a GC content of 40%-60%, after grouping according to the remainder of the sum of base numbers modulo 4, each group has more than 170,000 codewords. Select 100,000 codewords from one group to encode 5-digit decimal numbers from 00000 to 99999. Then two codewords can encode a total of 10 11 address indexes, and the codewords themselves can detect errors. The other 14 codewords directly encode the information without error detection, and the RS error-correcting code with segmented redundancy is used for unified error correction at the end. There are approximately 680,000 codewords encoding this information, and each codeword encodes a 19-bit character. Only less than 530,000, that is, 524,288 codewords need to be selected for use (2 19 ). The total number of bits encoded by a single sequence then reaches 266 bits (19 bits * 14 codewords). The upper limit of the capacity of a single group of DNA after segmented redundancy is approximately around 3 TB. If 4 codewords are selected for the address index, the number of address indexes that can be encoded can reach 10 21 address indexes, and the capacity upper limit is approximately 30 ZB, and a single sequence encodes 228 bits. The capacity upper limit can be flexibly adjusted according to actual needs. For each additional codeword used for the address index, the capacity upper limit can be increased by 100,000 times.
[0074] Since there are four groups of codewords available for the address index, the group of the address index codeword can be specified to distinguish different data. When encoding the address index with two codewords, there are 16 different group information. When using modulo 4 remainders 0, 1, 2, and 3 to represent the groups respectively, for example, the address index codeword group 00 encodes text, 11 encodes pictures, 22 encodes videos, 33 encodes other data, etc. Here only the non-dual encoding of 10-base codewords is simply described. Similar ideas can also be adopted for codewords of other lengths, and generally good results can be obtained, so they will not be elaborated one by one.
[0075] Information density: Currently, the information density of a single base is generally used as one of the quality criteria for measuring DNA storage encoding methods, and it is best to approach the Shannon limit. However, from a commercial perspective, the price density ratio of a single base, that is, the cost performance, is a more important criterion. Even so, when simply examining the single-base information density of various methods, the objective standard is to use the total length of oligonucleotides including primer sequences, error-correcting codes, and sequence redundancy as the denominator for measurement. The single-base density of the fountain code at Columbia University is approximately 1.17 bits. The hybrid radix at the University of Washington uses a method of converting DNA sequences in 6-bit blocks, and its density is about 0.8 bits. Due to the high proportion of sequence redundancy in the yin-yang code, it is slightly lower than the fountain code, at 1.02 bits.
[0076] If codewords with a length of 10 bases are used, and each codeword encodes 16 bits of information, and oligonucleotides with a length of 250 bases are used, 20 codewords can be used, plus 4-base Hamming code error correction and 5% redundant segmentation. Under this condition, the actual information density is approximately 1.21 bits. Without relying on the use of primer combinations to increase the DNA capacity of a single group, the upper limit of the DNA capacity of a single group can reach 3.8 ZB.
[0077] Adjustment of the encoding scheme: According to specific needs, the encoding scheme can be flexibly adjusted. For example, if the characters encoded by the 10-base sequence are reduced to 14 bits instead of 16 bits, then a group of codewords using modulo 4 remainders will be sufficient. Therefore, four different groups of codewords can be used to encode different data types, such as text, picture and video, audio data, and other data, etc. After such adjustment, error detection is very simple. After a single-base mutation occurs in the codeword, it can be determined as an illegal codeword without comparing with the scale numbers.
[0078] If address index growth is required, short-length codewords can be selected. Just as described in the encoding process before, 8 bases are used to encode 8 bits. If enhanced error correction ability is needed, error correction codes can be added to individual codewords. For example, 4 bases are added to a 10-base codeword to make the codeword a Hamming code (14, 10) or RS code. Such a 14-base Hamming code can directly correct single-base errors within the codeword. If a 200-base DNA fragment is used, in addition to the 40-base primer, 11 codewords (154 bases) can be accommodated. A single fragment can still encode 176 bits. The actual single-base information density after redundancy still exceeds 0.8 bits. At the same time, a single set of DNA can still reach more than 20 TB. The sequencing result of a single molecule can theoretically correct up to 11 base errors by itself.
[0079] Another way to enhance error correction ability is to reduce the codeword length. For example, 6 bases are used to encode 7 bits. There are 1280 6-base codewords that meet the restriction condition of 50% GC content, which can exactly encode the combination of 7 bits and 10 Arabic numerals. For a 200-base DNA fragment, after removing the 40-base primer length, at least 26 codewords (156 bases) can be encoded. Therefore, even if an average of 10 base errors occur in a single fragment, at least 16 codewords can still be correct. According to the 16-bit index number encoded by it, when the data of a single set of DNA is at the TB level, it is sufficient to correctly locate this fragment. Through deep sequencing, the sequencing of multiple fragment molecules can basically restore the correct sequence. Coupled with sequence redundancy, the data can still be fully restored.
[0080] Therefore, the present invention can flexibly and appropriately adjust and balance the ratio among the encoding information density, the number of address indexes, and the error correction ability to meet different requirements.
[0081] Direct text encoding: The characters encoded by the present invention can not only be binary characters, but also directly and efficiently encode text information, such as Chinese text and English text.
[0082] 65536 16-bit characters can be used to encode Chinese characters. Other methods and steps are basically the same. When generating sequence redundancy, Chinese characters are still regarded as 16-bit numbers and redundancy is also performed. Since there are only about less than 7,000 commonly used legal simplified Chinese characters in Chinese, high-quality code words can be selected to encode commonly used Chinese characters, such as 50% GC content and as few single base repeats as possible. As for more uncommon or rare Chinese characters, it is only necessary to select a few code words as prefixes and form a double code sub combination with other code words to encode rare Chinese characters. For example, only 60,000 code words are used to encode 60,000 Chinese characters, and 10 code words are taken as prefixes only, and combined with these 60,000 code words into double code words, with a total of 600,000 combinations, which can encode another 600,000 rare Chinese characters. At the same time, since the number of code words is sufficient, some code words can also be allocated to encode punctuation marks, Arabic numerals (here are Arabic numerals as characters, which are not the same as Arabic numerals used for address indexing), uppercase and lowercase English letters, Latin characters, etc. This can efficiently encode various required Chinese texts, such as ancient literary works, modern scientific papers, etc.
[0083] The encoding of English text is similar. The encoding can be directly targeted at words instead of individual letters. Although the number of English words can be as high as hundreds of thousands or even millions, the frequency of use of words is completely different. According to statistical results, the 1,000 most frequently used English words have a total frequency of use exceeding 80%, and the top 10,000 words exceed 90%. Then use 600,000 code words to double encode the 60,000 most frequently used words, and then use 100 code words as word heads. The double code word method can encode another 6 million uncommon words, which is enough. Unlike Chinese, English has various forms of verbs and adjectives. For the past tense and past participle of verbs, irregular ones are treated as different words, and regular ones are treated as different words. Different forms of adjectives are also treated similarly, including prefixes and suffixes. The first letter of the first word in each sentence is capitalized, which is just a format. This rule can be used when recovering data. Individual letters are also assigned code words and treated as words. In this way, since there are more than 650,000 available code words, it can be processed according to this idea in general. Moreover, this process actually compresses the English text. The average English word is about 5.6 letters (including spaces), which is about 45 bits in total, calculated as 8 bits per character. Using 10-base code words to encode a word has compressed the data to nearly one-third of the original size.
[0084] Since Chinese characters are very concise, Chinese text and English text can be processed in the same way. According to the frequency statistics of Chinese characters, the total frequency of use of the 2,400 most frequently used Chinese characters is as high as 99%, which means that 99% of the Chinese characters in the current Chinese text are these 2,400 characters. Therefore, 24,500 code words are set aside to encode the combination of these 2,400 Chinese characters and address index Arabic numerals, as well as the combination of 50 prefix code words and address index Arabic numerals, which can encode 2,400 commonly used Chinese characters and 120,000 uncommon Chinese characters. The remaining code words are used to encode English words, which has basically no significant impact, because there are more than 650,000 code words in total, and Chinese characters only need more than 20,000 code words.
[0085] Redundancy error correction algorithm: When performing segmented redundancy of binary data, multiple algorithms can be selected. The present invention can use RS error correction code or RS erasure code. Other algorithms such as LDPC low-density parity check code can also be used in principle, but in terms of comprehensive efficiency, RS erasure code is slightly more efficient with the help of single codeword error detection.
[0086] Since RS error correction code and RS erasure code have been developed very maturely and are widely used in the current communication and storage fields, their detailed principles will not be elaborated here, only the reasons for selection and efficiency will be explained. RS error correction code is based on the finite field of linear algebra, and the calculation of each data is converted into the power of polynomial coefficients for calculation. According to the set redundancy ratio, m code data (such as 8-bit bytes) generate k check code data (also 8-bit bytes). These code data constitute an RS code block, which can correct up to k / 2 arbitrary code data errors through calculation.
[0087] RS erasure code calculates n original data (such as 256 different data of 8 bits) with appropriate matrices to obtain m check data, and randomly selects at least n correct data from n+m data (each of which can come from the original data and the check data), that is, all data can be restored through calculation. Ready-made matrices such as Vandermonde matrix and Cauchy matrix can be used.
[0088] As can be seen from the above, RS erasure codes are highly efficient, but they need to ensure that the data used for calculations are correct. The present invention can detect errors in the codewords of a single data. After deep sequencing and single codeword error detection confirmation, the data codewords encoded by these codewords can be used in RS erasure codes to recover missing data. Another difference between RS error correction codes and RS erasure codes is that RS error correction codes can use bit characters of any length, such as 3 bits, 4 bits, 8 bits, 13 bits, 16 bits or more, while RS erasure codes can currently only run on data codewords of 8 bits or 16 bits in length.
[0089] Coding efficiency: Coding efficiency refers to the proportion of the information utilized in coding to the total information. The information capacity limit of a single base is 2 bits, and the proportion of the utilized bits is the coding efficiency. Both storing information and address indexing require the utilization of information. The total number of information bits that a coding method can encode includes these two data, and both are important. The amount of encoded information determines the storage density, and the amount of encoded address indexing determines the capacity of a single group of DNA. Sequence error-correcting codes and sequence error-correcting redundancy are used to solve the random error problem in sequence synthesis, which can determine the ability to recover data.
[0090] Among the current coding methods, the fountain code has the highest obvious efficiency. Using 128 bases to encode 256 bits of information, 16 bases can encode up to approximately 16M address indexes (this value is derived based on its stated highest capacity of 500M and 32 encoding 256 bits for a single fragment), which is equivalent to approximately 24 bits (2^24 is approximately equal to 16M). There is also an 8-base error-correcting code. Therefore, the total is 280 bits, and the total number of bits for 152 bases is 304. Thus, its coding efficiency is 92% (280 / 304). According to this calculation principle, the comprehensive comparison data of each coding method using address indexing is shown in Table 2. From the comparison data, it can be seen that the coding efficiency of the present invention is much higher than other coding methods. Even considering the actual coding efficiency after taking into account the redundant sequences, it exceeds 90%. Moreover, the upper limit of the capacity of a single group of DNA using this coding method reaches the EB level, exceeding other methods by a billion times and also exceeding the decimal coding method of another patent application of the present inventor by ten thousand times. Using a DNA sequence with a length of 160 bases, it can contain at most 320 bits of data, but the present invention can use at most 312 bits of them to encode data information and address indexing information respectively, with a coding efficiency as high as 97.5%. Considering the limiting factors of the sequence, this is very close to the efficiency upper limit.
[0091] Table 2 Performance Comparison of Main Coding Methods
[0092] Mixed radix Fountain code Yin-yang code Decimal system The present invention DNA length (excluding primers) 110 152 160 160 160 Base length of encoded data 96 128 128 128 160 Number of information bits stored 144 256 256 208 256 Base length of address index 14 16 16 24 0 Maximum number of address indexes <![CDATA[2x10 6 > <![CDATA[1.6x10 7 > <![CDATA[6x10 4 > <![CDATA[10 13 > <![CDATA[10 17 > Equivalent number of bits of address index 21 24 16 40 56 Base length of error correction code 0 8 16 8 0 Total number of encoded bits 165 280 272 248 312 Coding efficiency 75% 92.1% 85% 77.5% 97.5% Redundancy sequence ratio 15% 7% 20% 10% 5% Actual coding efficiency after redundancy 65% 86.1% 68% 69.8% 92.6% Whether 100% of the data can be restored Yes Yes No Yes Yes Upper limit of single-group DNA capacity in GB 0.06 0.5 0.0008 220000 3000000000
[0093] DNA Storage Dual Coding Device and Readable Storage Medium
[0094] The present invention provides a DNA storage dual coding device, including:
[0095] A memory for storing a computer program;
[0096] A processor for implementing the steps of the above-mentioned DNA storage dual coding method when executing the computer program.
[0097] The steps of the DNA storage dual coding method described above can all be implemented by the structure of this device.
[0098] Corresponding to the foregoing method embodiments, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the foregoing DNA storage dual encoding method are implemented.
[0099] The above schematically describes the present invention and its implementation manners. The description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to the technical solution without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A DNA storage dual encoding method, characterized in that, It includes the following steps: Step S1, dictionary definition: Define a DNA codeword dictionary with a length of n bases. Each DNA codeword double-encodes a character and an Arabic numeral. The DNA codewords in the dictionary satisfy that the content of G and C bases in the codeword is between 40% and 60%, and the Hamming distance between any two codewords encoding the same Arabic numeral in the dictionary is ≥2; Step S2, information segmentation: Segment the binary information to be stored according to a length of m characters; Step S3, generate redundant segments: Generate several redundant segments for a group of segmented information according to a certain redundancy ratio. The length of the redundant segments is the same as the length of the information segmentation in Step S2. Each character of the redundant segment is generated according to the error correction code generation rule from the corresponding characters at the same position of all information segments in the group. The characters at the same position of the redundant segment and the original information segment together form a basic error correction unit; Step S4, segment the scale number according to a length of m Arabic numerals; Step S5, according to the dictionary defined in Step S1, use DNA codewords to double-encode the segmented characters generated in Steps S2 and S3 and the segmented Arabic numerals generated in Step S4; Step S6, add error correction code: Add an error correction code with a certain base length to the DNA sequence generated in Step S5; Step S7, add primers at both ends: Add primers with a certain base length to both the head and the tail of the DNA sequence obtained in Step S6; Step S8, DNA synthesis and sequencing: Synthesize the DNA sequence obtained in Step S7 and sequence the DNA fragment.
2. The DNA storage dual encoding method according to claim 1, wherein, In Step S1, the base length n of the DNA codeword is n≤20.
3. A DNA storage dual coding method as described in claim 1, characterized in that, The error correction code based on which the redundant segments are generated in Step S3 is Reed-Solomon code, RS erasure code, RS error correction code.
4. A DNA storage dual encoding method as claimed in claim 1, wherein, The scale number in Step S4 is a normal number, an approximate normal number, and their derivatives.
5. A DNA storage dual encoding method according to claim 1, characterized in that, The error correction code in Step S6 is RS erasure code, RS error correction code, Hamming code, and parity check code. The error correction code can correct one or more base errors in combination with the error detection mechanism of the codeword itself.
6. A DNA storage dual encoding device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the DNA storage double-encoding method as described in any one of claims 1-5 when executing the computer program.
7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the DNA storage double-encoding method as described in any one of claims 1-5 are implemented.
Citation Information
Patent Citations
DNA data storage and coding method
CN111737956A
Data storage method, decoding method, system and device and storage medium
CN113314187A