Information encoding and decoding method, device, storage medium, and information storage and interpretation method
By converting two binary information into coding sequences with intersections into output symbols, the problems of low storage density and difficult sequencing in existing DNA data storage technologies are solved, and high encoding density and data fidelity are achieved.
Patent Information
- Application Number
- CN201980099954.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2039-09-24
AI Technical Summary
Existing DNA data storage technologies are difficult to effectively improve storage density, and they cannot completely avoid continuous GC or AT situations in DNA sequences, resulting in difficulty in reading sequence information during sequencing.
By obtaining two binary information and applying two encoding rules, it converts them into an encoding sequence with intersections into output symbols, the information capacity limit and high encoding density are achieved.
The information encoding density is improved, the continuous GC or AT situation in the DNA sequence is avoided, and the fidelity and storage density of data reading are improved.
Smart Images

Figure CN114730616B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information storage technology, and in particular to an information encoding and decoding method, device, storage medium, and information storage and interpretation method. Background Art
[0002] With the development of modern technology, especially the Internet, the global data is showing an exponential growth trend. The ever-increasing amount of data places higher and higher demands on storage technology. Traditional storage technologies, such as magnetic tape and optical disc storage, are increasingly unable to meet current data storage needs due to limited storage density and time. In recent years, the development of DNA storage technology has provided a new way to solve these problems. Compared with traditional storage media, DNA as a medium for information storage has the advantages of long storage time (can be stored for more than thousands of years, more than a hundred times that of existing magnetic tape and optical disc media), high storage density (reaching about 10 9 Gb / mm 3 , which is more than 10 million times that of tape and CD media) and good storage security.
[0003] DNA data storage generally includes the following steps: 1) Encoding: converting the binary 0 / 1 code of computer information into DNA sequence information of A / T / C / G bases; 2) Synthesis: synthesizing the corresponding DNA sequence using DNA synthesis technology, and storing the obtained DNA molecules in an in vitro medium or living cells; 3) Sequencing: reading the DNA sequence of the stored DNA molecules using sequencing technology; 4) Decoding: using the method corresponding to the encoding process in step 1), converting the DNA sequence obtained by sequencing into a binary 0 / 1 code, and further converting it into computer information.
[0004] In order to achieve effective DNA data storage, it is necessary to develop technologies for the above steps. Among them, the encoding and decoding technologies involved in step 1) and step 4) are the most critical technologies for DNA data storage. The most critical part of this technology is: 1) How to maximize the density of DNA-encoded 0 / 1 binary information. The improvement of DNA storage density is crucial to saving the cost of DNA synthesis for storing information in step 2). 2) When 0 / 1 binary information is converted into A / T / C / G base sequences, single base repetitions, high GC and high AT between sequences are avoided to the greatest extent. Generally speaking, continuous single base repetitions, high GC and high AT in DNA sequences will make it difficult to read sequence information during sequencing. The way 0 / 1 binary information is converted to A / T / C / G DNA sequences directly determines the difficulty of interpreting the DNA sequence during sequencing, thereby determining the fidelity of the data during reading.
[0005] The most classic DNA data storage methods currently include the data encoding methods proposed by George Church, Goldman and others. In 2012, George Church and others proposed the most primitive method of converting binary 0 / 1 information to A / T / C / G base information, that is, 0 represents A / T, 1 represents C / G, and each nucleotide encodes 1 binary data. This method can avoid continuous single-base repetitions to a certain extent, but it cannot avoid the high GC or high AT in the data. When a large number of 0 or 1 repetitions appear in the data, the encoded DNA sequence will produce continuous GC or AT. At the same time, in this encoding method, each nucleotide only encodes 1 bit of information, which is very limited.
[0006] In 2013, Goldman et al. proposed a new DNA encoding method, which improved the storage density of DNA data to a certain extent. It first converts binary data into ternary 0 / 1 / 2 data through Huffman coding, and then converts 0 / 1 / 2 data into "quaternary" A / T / C / G sequence through a designed rule. This encoding method also avoids continuous single base repetition, but it is still limited in DNA data storage density, and its maximum encoding density is 1.6 bits (bits) / base (nt). At the same time, this encoding method cannot completely solve the situation where continuous GC or AT appears in the coding sequence. Although some other encoding methods are different from the above two methods, they cannot completely avoid continuous GC and AT situations or the storage density is still limited.
[0007] In order to more effectively utilize DNA for binary data storage, improve the density of DNA data storage, and avoid continuous GC and AT situations, it is particularly important to develop more efficient DNA data encoding methods. Summary of the invention
[0008] The present invention provides an information encoding and decoding method, device, storage medium and information storage and interpretation method, which convert and integrate two input binary information into one information, reach the information capacity limit and have high information encoding density.
[0009] According to the first aspect, an embodiment provides an information encoding method, including:
[0010] Acquire first binary information and second binary information, as well as a first encoding rule and a second encoding rule, wherein the first encoding rule is used to encode the first binary information, and the second encoding rule is used to encode the second binary information;
[0011] Obtain a first candidate output symbol corresponding to a current input of the first binary information according to a first coding rule, and obtain a second candidate output symbol corresponding to a current input of the second binary information according to a second coding rule, and take the intersection of the first candidate output symbol and the second candidate output symbol as the output symbol corresponding to the current input;
[0012] The output symbol corresponding to each binary bit of the first binary information and the second binary information is determined in sequence by the first coding rule and the second coding rule to obtain a coding sequence composed of a plurality of output symbols.
[0013] In a preferred embodiment, the first output candidate symbol is two symbols among four symbols, the second output candidate symbol is two symbols among four symbols, and the first output candidate symbol and the second output candidate symbol have a same symbol;
[0014] The first coding rule is that, under the support bit of the first coding rule, two symbols are selected from the four symbols as the first output candidate symbols when the current input of the first binary information is 0, and the other two symbols are selected as the first output candidate symbols when the current input is 1, wherein the support bit of the first coding rule is any one of the four symbols, and the support bit of the first coding rule has a first corresponding relationship with the current output bit;
[0015] The above-mentioned second coding rule is that, under the support bits of the first coding rule and the support bits of the second coding rule, two symbols are selected from the four symbols as the second output candidate symbols when the current input of the second binary information is 0, and the other two are selected as the second output candidate symbols when the current input is 1, wherein the support bits of the above-mentioned second coding rule are any one of the four symbols, and the support bits of the above-mentioned second coding rule have a second corresponding relationship with the current output bits, and each support bit of the first coding rule corresponds to the support bits of four second coding rules.
[0016] In a preferred embodiment, the length of the first binary information is equal to the length of the second binary information.
[0017] In a preferred embodiment, the first binary information and the second binary information are split from the same binary information.
[0018] In a preferred embodiment, the first corresponding relationship is a first predetermined number of bits before the current output bit; and the second corresponding relationship is a second predetermined number of bits before the current output bit.
[0019] In a preferred embodiment, the method further comprises: obtaining a starting base sequence, which provides support bits for the first encoding rule and the second encoding rule before generating the output base.
[0020] In a preferred embodiment, the first output candidate symbol is two base symbols among the four bases A, T, C, and G, the second output candidate symbol is two base symbols among the four bases A, T, C, and G, and the supporting bit is one base symbol among the four bases A, T, C, and G; and,
[0021] The above coding sequence is a nucleic acid sequence containing the above four bases A, T, C, and G.
[0022] In a preferred embodiment, the method further comprises: obtaining a starting sequence, which provides a starting support bit for the first encoding rule and the second encoding rule before generating the first bit of the encoding sequence.
[0023] In a preferred embodiment, the method further comprises: before acquiring the first binary information and the second binary information, extracting the first binary information and the second binary information from a computer storage device.
[0024] According to the second aspect, an embodiment provides an information encoding device, including:
[0025] An information acquisition unit, used to acquire first binary information and second binary information, as well as a first encoding rule and a second encoding rule, wherein the first encoding rule is used to encode the first binary information and the second encoding rule is used to encode the second binary information;
[0026] An information encoding unit, configured to obtain a first candidate output symbol corresponding to a current input of the first binary information according to a first encoding rule, and obtain a second candidate output symbol corresponding to a current input of the second binary information according to a second encoding rule, and take an intersection of the first candidate output symbol and the second candidate output symbol as the output symbol corresponding to the current input;
[0027] The result generating unit is used to determine the output symbol corresponding to each binary bit of the first binary information and the second binary information in sequence through the first coding rule and the second coding rule to obtain a coding sequence composed of a plurality of the above output symbols.
[0028] In a preferred embodiment, the first output candidate symbol is two symbols among four symbols, the second output candidate symbol is two symbols among four symbols, and the first output candidate symbol and the second output candidate symbol have a same symbol;
[0029] The first coding rule is that, under the support bit of the first coding rule, two symbols are selected from the four symbols as the first output candidate symbols when the current input of the first binary information is 0, and the other two symbols are selected as the first output candidate symbols when the current input is 1, wherein the support bit of the first coding rule is any one of the four symbols, and the support bit of the first coding rule has a first corresponding relationship with the current output bit;
[0030] The above-mentioned second coding rule is that, under the support bits of the first coding rule and the support bits of the second coding rule, two symbols are selected from the four symbols as the second output candidate symbols when the current input of the second binary information is 0, and the other two are selected as the second output candidate symbols when the current input is 1, wherein the support bits of the above-mentioned second coding rule are any one of the four symbols, and the support bits of the above-mentioned second coding rule have a second corresponding relationship with the current output bits, and each support bit of the first coding rule corresponds to the support bits of four second coding rules.
[0031] In a preferred embodiment, the first corresponding relationship is a first predetermined number of bits before the current output bit; and the second corresponding relationship is a second predetermined number of bits before the current output bit.
[0032] In a preferred embodiment, the first output candidate symbol is two base symbols among the four bases A, T, C, and G, the second output candidate symbol is two base symbols among the four bases A, T, C, and G, the supporting position is one base symbol among the four bases A, T, C, and G; and the coding sequence is a nucleic acid sequence containing the four bases A, T, C, and G.
[0033] According to a third aspect, an embodiment provides a computer-readable storage medium, comprising a program, wherein the program can be executed by a processor to implement the method of the first aspect.
[0034] According to a fourth aspect, an embodiment provides a method for storing information using a DNA sequence, comprising:
[0035] The binary information to be stored is converted into DNA sequence information by the information encoding method as described above in the first aspect, wherein the DNA sequence information includes a base sequence formed by four bases: A, T, C, and G;
[0036] synthesizing a corresponding DNA sequence according to the above DNA sequence information;
[0037] The above DNA sequence is saved to realize the storage of information.
[0038] In a preferred embodiment, the method further comprises: splitting the DNA sequence information into multiple pieces of DNA short sequence information, and adding a DNA index sequence identifier to each of the split DNA short sequence information, wherein the DNA index sequence identifier includes position order information of the DNA short sequence information;
[0039] synthesizing a corresponding DNA sequence according to the above-mentioned DNA short sequence information;
[0040] The above DNA sequence is saved to realize the storage of information.
[0041] In a preferred embodiment, the above DNA sequence is stored in the form of dry powder, or embedded in an embedding material.
[0042] In a preferred embodiment, the above DNA sequences are transferred into living cells for preservation.
[0043] In a preferred embodiment, the living cells are microbial cells.
[0044] In a preferred embodiment, the living cell is Escherichia coli or Saccharomyces cerevisiae.
[0045] According to the fifth aspect, an embodiment provides an information decoding method, including:
[0046] Obtain a coding sequence generated by the above-mentioned coding method of the first aspect, as well as a first coding rule and a second coding rule, wherein the first coding rule is used to encode the first binary information, and the second coding rule is used to encode the second binary information;
[0047] Reading a current symbol of the coding sequence, and converting the current symbol into binary bits of the first binary information and the second binary information according to the correspondence between the four different symbols and the binary information in the first coding rule and the second coding rule;
[0048] Through the correspondence between different symbols and binary information in the first coding rule and the second coding rule, each binary bit of the first binary information and the second binary information corresponding to each symbol bit of the above coding sequence is determined in turn to obtain the first binary information and the second binary information with a determined binary bit order.
[0049] In a preferred embodiment, the above symbols are four base symbols: A, T, C, and G.
[0050] In a preferred embodiment, the above coding sequence is obtained by the steps in (1) or (2):
[0051] (1) sequencing each DNA sequence synthesized by the method of the fourth aspect to obtain the above coding sequence; or
[0052] (2) Sequencing each DNA sequence synthesized by the method of the fourth aspect to obtain information of each DNA short sequence; obtaining positional order information of each DNA short sequence based on the DNA index sequence identifier; and combining each DNA short sequence into the coding sequence based on the positional order information.
[0053] In a preferred embodiment, the above decoding method further comprises:
[0054] The first binary information and the second binary information are transcoded into corresponding information.
[0055] In a preferred embodiment, the corresponding information is text information, image information, audio information and / or video information.
[0056] According to the sixth aspect, an embodiment provides an information decoding device, including:
[0057] An information acquisition unit, used to acquire a coding sequence generated by the coding device of the second aspect, as well as a first coding rule and a second coding rule, wherein the first coding rule is used to encode the first binary information, and the second coding rule is used to encode the second binary information;
[0058] An information decoding unit, used for reading a current symbol of the coding sequence, and converting the current symbol into binary bits of the first binary information and the second binary information according to the correspondence between different symbols and binary information in the first coding rule and the second coding rule;
[0059] The result generating unit is used to determine each binary bit of the first binary information and the second binary information corresponding to each symbol bit of the above coding sequence in turn through the correspondence between the four different symbols and the binary information in the first coding rule and the second coding rule, and obtain the first binary information and the second binary information with a determined binary bit order.
[0060] In a preferred embodiment, the above different symbols are four base symbols: A, T, C, and G.
[0061] In a preferred embodiment, the above coding sequence is obtained by the following unit (1) or (2):
[0062] (1) a sequencing unit, used to sequence each DNA sequence synthesized by the method of the fourth aspect to obtain the above coding sequence; or
[0063] (2) a sequencing unit, used to sequence each DNA sequence synthesized by the method described in the fourth aspect to obtain information of each DNA short sequence;
[0064] An index unit, used to obtain the position order information of each of the above-mentioned DNA short sequences according to the above-mentioned DNA index sequence identifier;
[0065] The combining unit is used to combine the above-mentioned DNA short sequences into the above-mentioned coding sequence according to the above-mentioned positional sequence information.
[0066] In a preferred embodiment, the above-mentioned device further comprises: a transcoding unit, which is used to transcode the above-mentioned first binary information and second binary information into corresponding information.
[0067] In a preferred embodiment, the corresponding information is text information, image information, audio information and / or video information.
[0068] According to the seventh aspect, an embodiment provides a computer-readable storage medium, including a program, which can be executed by a processor to implement the decoding method as in the fifth aspect.
[0069] The information encoding and decoding method provided by the present invention can convert and integrate two input binary information into one piece of information, reach the information capacity limit, and have high information encoding density. In a preferred embodiment, four base symbols A, T, C, and G are used as four different symbols to cleverly convert and integrate two input binary information into a DNA sequence, so that the information capacity limit of a single DNA base reaches 2 bits / base, and the information encoding density is high.
[0070] In practical applications, the method of the present invention can be well combined with long-fragment gene storage to improve DNA storage density. The method of the present invention uses two encoding rules to generate 5566277615616 encoding systems, providing a rich compilation method rule library for DNA storage applications, greatly expanding the selection space of DNA storage compilation methods.
[0071] At the same time, the coding rules in the method of the present invention depend on two supporting bits, so that the coding method has a higher tolerance for repeated input, and can effectively avoid or reduce the impact caused by continuous single base and double base repetition. In addition, the coding rule library generated by the method of the present invention uses different coding rule combinations for encrypted storage, which increases the difficulty of cracking DNA storage information, greatly improves the security of DNA storage information, and improves the guarantee for future data security storage applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 A general flow chart of DNA storage and interpretation of information in an embodiment of the present invention;
[0073] Figure 2 It is a flow chart of the encoding method for converting binary information into DNA sequence in an embodiment of the present invention;
[0074] Figure 3 Schematic diagram of a general representation method of encoding rules in an embodiment of the present invention;
[0075] Figure 4 A schematic diagram of a case representation method of encoding rules in an embodiment of the present invention;
[0076] Figure 5 It is a schematic diagram of the principle of converting binary information into a DNA sequence in an embodiment of the present invention;
[0077] Figure 6 It is a structural block diagram of an encoding device for converting binary information into a DNA sequence in an embodiment of the present invention;
[0078] Figure 7 It is a flow chart of a decoding method for converting a DNA sequence into binary information in an embodiment of the present invention;
[0079] Figure 8 It is a structural block diagram of a decoding device for converting a DNA sequence into binary information in an embodiment of the present invention;
[0080] Fig. 9 This is a flow chart of a method for storing information using a DNA sequence in an embodiment of the present invention;
[0081] Fig.10 Flow chart of a method for interpreting information stored in the form of DNA sequences in an embodiment of the present invention. DETAILED DESCRIPTION
[0082] The present invention is further described in detail below by specific embodiments in conjunction with the accompanying drawings. In the following embodiments, many details are described to enable the present invention to be better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other materials or methods.
[0083] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various implementations. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a required sequence, unless otherwise specified that a certain sequence must be followed.
[0084] In this article, the serial numbers assigned to the features, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning.
[0085] In this document, the first encoding rule is also referred to as "encoding rule 1", and the two have equivalent meanings; the second encoding rule is also referred to as "encoding rule 2", and the two have equivalent meanings.
[0086] like Figure 1 As shown, the data encoding, storage and interpretation process of the present invention includes the following steps:
[0087] (1) Data encoding and storage process:
[0088] Step 1: Use any program that comes with the computer operating system or a program specially written for extracting binary 0 / 1 codes to extract the binary "0 / 1" computer information to be stored.
[0089] Step 2: Use the coding rules in a set of coding and decoding methods in the present invention to convert two binary "0 / 1" information into a DNA sequence represented by A / T / C / G bases. Generally speaking, the two binary "0 / 1" information have the same length, which is greater than or equal to 1.
[0090] Step 3: Split the DNA sequence obtained from the 0 / 1 binary computer information conversion into short fragments of a certain length to facilitate DNA synthesis in the next step.
[0091] Step 4: Add an index sequence (index 1 and index 2 in the figure) to each short DNA fragment. The index sequence can encode the sequence information of the short fragments obtained in step 3.
[0092] Step 5: Use DNA synthesis technology to synthesize the DNA sequence fragment obtained in step 4 and store it in a corresponding medium. Generally speaking, the DNA sequence can be stored in a sample tube in the form of dry powder, or embedded in an embedding material such as amber or silica sphere for preservation, or transferred into living cells for preservation. The living cells can be microbial cells, more preferably Escherichia coli or Saccharomyces cerevisiae.
[0093] (2) DNA data interpretation process:
[0094] Step 6: Sequence the stored DNA fragments containing the data information using sequencing technology (eg, Sanger sequencing or high-throughput sequencing) to obtain the DNA sequences of these fragments.
[0095] Step 7: Decipher the order of the DNA fragment encoding information according to the pre-set index sequence information of the DNA sequence, and sort the DNA fragments in sequence.
[0096] Step 8: According to the sequence of the DNA fragment encoding information region obtained by sorting, the DNA information is converted into 0 / 1 binary information by using the decoding rule corresponding to the encoding rule in step 2.
[0097] Step 9: Use any program that comes with the operating system or a program specially written to convert 0 / 1 into data information to convert the 0 / 1 binary information obtained in step 8 into storage information (i.e., a file, such as text, image, audio or video, etc.).
[0098] like Figure 2 As shown, the encoding method for converting binary information into DNA sequence in one embodiment of the present invention comprises the following steps:
[0099] S201: Acquire first binary information and second binary information, as well as a first encoding rule and a second encoding rule, wherein the first encoding rule is used to encode the first binary information, and the second encoding rule is used to encode the second binary information.
[0100] In an embodiment of the present invention, the first binary information and the second binary information are 0 / 1 binary information to be encoded. The first binary information and the second binary information may have the same or different sources, that is, the two may be associated 0 / 1 binary information or unassociated 0 / 1 binary information. An example of associated 0 / 1 binary information, such as these two 0 / 1 binary information are binary information split from the same binary information. It is generally required that these two 0 / 1 binary information have the same length, because in the method of the present invention, each time a binary bit (0 / 1) of each 0 / 1 binary information is read, a pair of binary bits (0 / 1) from these two 0 / 1 binary information is converted into a symbol by the method of the present invention, for example, base symbol (A, T, G or C) information. An example of unassociated 0 / 1 binary information, such as two 0 / 1 binary information derived from text information and pattern information, respectively.
[0101] In the embodiment of the present invention, the first coding rule is a coding rule for the correspondence between binary symbols 0 / 1 and output symbols (e.g., base symbols A, T, G, or C). As a typical but non-limiting example, one case of the first coding rule is that, under the support bit of the first coding rule, two bases of A, T, C, and G are selected as the first output candidate bases when the current input of the first binary information is 0, and the other two are selected as the first output candidate bases when the current input is 1.
[0102] like Figure 3As shown, the vertical A, T, G, C represent the supporting positions of four different bases, N1, N2, N3, N4 represent one of the four bases A, T, C, G respectively; a1, a2, a3, a4 are 0 or 1, and a1+a2+a3+a4=2; b1, b2, b3, b4 are 0 or 1, and b1+b2+b3+b4=2; c1, c2, c3, c4 are 0 or 1, and c1+c2+c3+c4=2; d1, d2, d3, d4 are 0 or 1, and d1+d2+d3+d4=2. In other words, in the a series, there are two 1s and two 0s. Similarly, in the b, c, and d series, there are two 1s and two 0s respectively. Therefore, the encoding rule 1 (first encoding rule) has a total of [4-choose-2 combinations, that is, "choose any 2 of the 4 bases, and there are 6 possible combinations in each support position"] 4, that is, 6^4 or 1296, because the base selection between the 4 different support positions is independent of each other. This encoding rule is interpreted as, when the support position is Nx (x = 1, 2, 3, 4), when the input bit is 0 / 1, the N corresponding to the two output bits of 0 / 1 is the output possibility.
[0103] like Figure 4 As shown, an example of coding rule 1 is shown. In this example, under different support bits of coding rule 1, the input bit of the first binary information corresponds to the corresponding output bit in the manner shown in the figure.
[0104] In the embodiment of the present invention, the input bit refers to the currently input binary bit of the first binary information (or the second binary information), which is 0 / 1; the output bit refers to the base bit to be output after conversion according to the first encoding rule (or the second encoding rule) corresponding to the input bit.
[0105] In an embodiment of the present invention, the support bits of the first coding rule refer to the support information required for the first coding rule to select the correct output symbol according to different inputs of the first binary information (the so-called "input" refers to the binary bits of the current input of the first binary information). For example, when the output symbol is a base symbol A, T, G or C, the support bits of the first coding rule are also base symbols. Specifically, the support bits of the first coding rule are known bases that have a first corresponding relationship with the current output bit. Generally speaking, the support bits are the base information that has been converted before, for example, the number of bits before the data bit (current output bit) currently being converted, such as the previous 3rd or 6th base information. Therefore, in one embodiment of the present invention, the so-called "first corresponding relationship" refers to the first predetermined number of bits before the current output bit (for example, the 3rd or 6th bit before the current input bit, etc., an arbitrarily set number of bits). Of course, the support bit can also be virtual randomly generated information, which has an artificially set first correspondence with the current output bit. In the embodiment of the present invention, the so-called "random generation" refers to generating random numbers through various random methods, and then corresponding to ATCG according to certain rules. Examples of random methods include but are not limited to: Monte Carlo random numbers or U (0,1) random numbers, etc. In addition, the support bit can also come from a specific selection method of the reference sequence, which has a specific first correspondence with the current output bit. For example, each output bit corresponds to a known base on the reference sequence in a specific mapping method, such as the base sequence of the reference sequence corresponds to each output bit in turn. In the embodiment of the present invention, when the support bits are different, the same input will have different base outputs.
[0106] In the embodiment of the present invention, the second coding rule is a coding rule for the correspondence between binary symbols 0 / 1 and output symbols (e.g., base symbols A, T, G, or C). As a typical but non-limiting example, one case of the second coding rule is that, under the support bits of the first coding rule and the support bits of the second coding rule, two bases are selected from the four bases as the second output candidate bases when the input of the second binary information is 0, and the other two are selected as the second output candidate bases when the input is 1; wherein the support bits of the second coding rule are any one of the four bases, and the support bits of the second coding rule have a second correspondence with the current output bit, and each support bit of the first coding rule corresponds to four support bits of the second coding rule.
[0107] like Figure 3As shown, N01 and N02 are two possible bases corresponding to the input of 0 in coding rule 1, and N11 and N12 are two possible bases corresponding to the input of 1 in coding rule 1. Since different supporting bits in coding rule 1 have different output possibilities, when the supporting bits in coding rule 1 are different, the corresponding coding rule 2 (second coding rule) may be different. Among them, xn+yn=1, xn*yn=0, n=1,2,3,4,5,6,7,8. That is to say, there is one 0 and one 1 in each x,y group. Therefore, corresponding to each possibility given in coding rule 1, coding rule 2 has 2 to the power of 8. This is because the base selection between different supporting bits of coding rule 2 is independent. Each supporting bit of each coding rule 1 corresponds to (2^2)^4 possibilities of coding rule 2, that is, 256 possibilities. This encoding rule is explained as, when the support bits are Nx (x=01, 02, 11, 12), when the input bits are 0 / 1, the N corresponding to the two output bits of 0 / 1 is the output possibility.
[0108] like Figure 4 As shown, a case of coding rule 2 is shown. In this case, under different support bits of coding rule 1, it is shown that the input bit of the second binary information corresponds to the corresponding output bit in the manner shown in the figure under different support bits of coding rule 2. Specifically, the coding rule 2 corresponding to the support bit of coding rule 1 is A, the coding rule 2 corresponding to the support bit of coding rule 1 is C, the coding rule 2 corresponding to the support bit of coding rule 1 is G, and the coding rule 2 corresponding to the support bit of coding rule 1 is T are shown.
[0109] In an embodiment of the present invention, similar to the support bits of the first coding rule, the support bits of the second coding rule refer to the support information required for the second coding rule to select the correct output symbol according to different inputs of the second binary information (the so-called "input" refers to the binary bits of the current input of the second binary information). For example, when the output symbol is a base symbol A, T, G or C, the support bits of the second coding rule are also base symbols. Specifically, the support bits of the second coding rule are known bases that have a second corresponding relationship with the current output bit. Generally speaking, the support bits are bases that have been converted before, for example, the number of bits before the data bit (current output bit) currently being converted, such as the base information of the 4th or 8th bit before. Therefore, in one embodiment of the present invention, the so-called "second corresponding relationship" refers to the second predetermined number of bits before the current output bit (for example, the number of bits arbitrarily set before the 4th or 8th bit). Of course, the support bit can also be virtual randomly generated information, which has an artificially set second corresponding relationship with the current output bit. In the embodiment of the present invention, the so-called "random generation" refers to generating random numbers through various random methods, and then corresponding to ATCG according to certain rules. Examples of random methods include but are not limited to: Monte Carlo random numbers or U (0,1) random numbers. In addition, the support bit can also come from a specific selection method of the reference sequence, which has a specific second corresponding relationship with the current output bit. For example, each output bit corresponds to a known base on the reference sequence according to a specific mapping method, such as the base sequence of the reference sequence corresponds to each output bit in turn. In a preferred embodiment of the present invention, the first corresponding relationship is different from the second corresponding relationship, that is, the support bit of the first encoding rule and the support bit of the second encoding rule take different base positions. For example, in one embodiment of the present invention, the known base at the 6th position before the current output bit is selected as the support bit of the first encoding rule, and the known base at the 1st position before the current output bit is selected as the support bit of the second encoding rule.
[0110] Since the base selection of different supporting positions in encoding rule 1 is independent, and the base selection of different supporting positions in encoding rule 2 is also independent, the number of types of encoding rule 2 corresponding to each encoding rule 1 is 256^4, that is, 4294967296 types. Therefore, the total number of types of this dual encoding rule system is 1296*4294967296, that is, about 5.6×10^12 types out of 5566277615616.
[0111] In this way, a binary input bit (0 / 1) from the first binary information is encoded through rule 1 to obtain two possible lists of outputs, and a binary input bit (0 / 1) from the second binary information is encoded through rule 2 to obtain two possible lists of outputs. The intersection of the two output lists is taken to determine the output bases corresponding to the two input binary bits.
[0112] It should be noted that, when the previously converted base information is used as the support bit, when the conversion is initially started, since the previously converted base information does not exist as the support bit, it is necessary to solve the initial support bit problem through appropriate methods. In one embodiment of the present invention, the encoding method of the present invention also includes: obtaining a starting base sequence, which provides support bits for the first encoding rule and the second encoding rule before generating the output base. In other cases, such as when virtual randomly generated information or a specific selection method from a reference sequence is used as the support bit, there is no such problem.
[0113] S202: Obtain a first output candidate symbol corresponding to a current input of the first binary information according to a first coding rule, and obtain a second output candidate symbol corresponding to a current input of the second binary information according to a second coding rule, and take the intersection of the first output candidate symbol and the second output candidate symbol as the output symbol corresponding to the current input.
[0114] In the embodiment of the present invention, as a preferred example, the first output candidate symbol and the second output candidate symbol are respectively two base symbols among the four bases of A, T, C, and G. Therefore, the first output candidate symbol and the second output candidate symbol refer to the first output candidate base and the second output candidate base respectively.
[0115] In the embodiment of the present invention, the so-called "current input" refers to the 0 / 1 binary bit currently being read from the first binary information or the second binary information, which respectively represents one bit of binary data. The present invention reads the 0 / 1 binary bit in the first binary information and the second binary information simultaneously. When reading a 0 / 1 binary bit of the first binary information, a 0 / 1 binary bit of the second binary information is read simultaneously. A 0 / 1 binary bit of the read first binary information is converted into a first output candidate base (including two bases) through a first coding rule, and a 0 / 1 binary bit of the read second binary information is converted into a second output candidate base (including two bases) through a second coding rule. There is a common base between the above-mentioned first output candidate base and the second output candidate base, that is, the intersection of the first output candidate base and the second output candidate base, and the intersection is the output base corresponding to the two current inputs (the current input of the first binary information and the current input of the second binary information).
[0116] S203: Determine the output symbol corresponding to each binary bit of the first binary information and the second binary information in sequence according to the first encoding rule and the second encoding rule, and obtain a coding sequence composed of a plurality of the above output symbols.
[0117] Since the first binary information and the second binary information each have a certain length, such as tens, hundreds, thousands or tens of thousands of bits (0 / 1 binary bits), and each conversion operation of the above step S202 can only convert one bit (0 / 1 binary bit) of the first binary information and the second binary information into an output base, it is necessary to continuously perform the conversion operation until all bits (0 / 1 binary bits) of the first binary information and the second binary information are converted into corresponding output bases, and such multiple output bases form a corresponding DNA sequence. At this point, the conversion of two 0 / 1 binary information (the first binary information and the second binary information) into one DNA sequence information is completed.
[0118] In order to make the present invention more easily understood, as a typical but non-limiting example, Figure 5 As shown: There are two binary information a and b, and their 0 / 1 binary information encoding sequence is as follows Figure 5 As shown in the figure, a virtual sequence is shown as the starting base support bit, and the ninth binary bit 1 of the binary information a is converted into a G / C base according to the first encoding rule (Table 1) when the support bit is A; when the support bit of the ninth binary bit 1 of the binary information a is A, and when the support bit of the ninth binary bit 1 of the binary information b is A, the ninth binary bit 1 of the binary information b is converted into an A / C base according to the second encoding rule (Table 2) when the support bit is A, and the intersection of the G / C base and the A / C base, that is, the C base, is taken as the output base.
[0119] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above-mentioned embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above-mentioned functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above-mentioned functions can be implemented. In addition, when all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and can be downloaded or copied and saved in the memory of the local device, or the system of the local device is updated, and when the program in the memory is executed by the processor, all or part of the functions in the above-mentioned embodiments can be implemented.
[0120] Therefore, corresponding to the information encoding method of the present invention, the embodiment of the present invention also provides an information encoding device, such as Figure 6As shown, the device includes: an information acquisition unit 601, used to acquire first binary information and second binary information, as well as a first encoding rule and a second encoding rule, wherein the first encoding rule is used to encode the first binary information, and the second encoding rule is used to encode the second binary information; an information encoding unit 602, used to acquire a first output candidate symbol corresponding to a current input of the first binary information according to the first encoding rule, and acquire a second output candidate symbol corresponding to a current input of the second binary information according to the second encoding rule, and take the intersection of the first output candidate symbol and the second output candidate symbol as the output symbol corresponding to the current input; a result generating unit 603, used to determine the output symbol corresponding to each binary bit of the first binary information and the second binary information in turn through the first encoding rule and the second encoding rule, and obtain a coding sequence composed of a plurality of the above output symbols.
[0121] In a preferred embodiment of the present invention, the first output candidate symbol and the second output candidate symbol are two base symbols among the four bases A, T, C, and G respectively; the first coding rule is that, under the support bit of the first coding rule, two bases are selected from the four bases A, T, C, and G as the first output candidate base when the current input of the first binary information is 0, and the other two are used as the first output candidate base when the current input is 1; the second coding rule is that, under the support bit of the first coding rule and the support bit of the second coding rule, two bases are selected from the four bases as the second output candidate base when the current input of the second binary information is 0, and the other two are used as the second output candidate base when the current input is 1; wherein the support bit of the first coding rule is a known base having a first corresponding relationship with the current output bit, the support bit of the second coding rule is a known base having a second corresponding relationship with the current output bit, and each support bit of the first coding rule corresponds to the support bits of four second coding rules.
[0122] Corresponding to the information encoding method of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, including a program, which can be executed by a processor to implement the information encoding method of the present invention.
[0123] As the reverse process of the encoding method of the present invention, the embodiment of the present invention also provides an information decoding method, such as Figure 7 As shown, the method includes:
[0124] S701: Obtain a coding sequence generated by the above coding method, as well as a first coding rule and a second coding rule, wherein the first coding rule is used to encode the first binary information, and the second coding rule is used to encode the second binary information.
[0125] In the embodiment of the present invention, the coding sequence generated by the coding method can be, for example, a DNA sequence information. Accordingly, as a typical but non-limiting example, the first coding rule is: under the support position of the first coding rule, two bases are selected from the four bases of A, T, C, and G as the first output candidate bases when the input of the first binary information is 0, and the other two are selected as the first output candidate bases when the input is 1.
[0126] As a typical but non-limiting example, in an embodiment of the present invention, the second encoding rule is: under the support bits of the first encoding rule and the support bits of the second encoding rule, two bases are selected from the four bases as the second output candidate bases when the input of the second binary information is 0, and the other two are selected as the second output candidate bases when the input is 1.
[0127] In an embodiment of the present invention, the support bits of the first encoding rule are known bases having a first corresponding relationship with the current output bits, the support bits of the second encoding rule are known bases having a second corresponding relationship with the current output bits, and each support bit of the first encoding rule corresponds to four support bits of the second encoding rules.
[0128] S702: Read the current symbol of the information encoded by the four different symbols, and convert the current symbol into binary bits of the first binary information and the second binary information according to the correspondence between the four different symbols and the binary information in the first encoding rule and the second encoding rule.
[0129] As a typical but non-limiting example, in the embodiment of the present invention, four bases A, T, C, and G are used as four different encoding symbols, and accordingly, the current symbol is also referred to as the "current base". The so-called "current base" refers to the base on the DNA sequence that is being read and converted into a binary bit. Since there are dozens, hundreds, thousands, or even tens of thousands of bases on the DNA sequence, each base that is being read and converted is the so-called "current base".
[0130] S703: Determine in sequence each binary bit of the first binary information and the second binary information corresponding to each symbol bit of the coding sequence through the correspondence between different symbols and binary information in the first coding rule and the second coding rule, and obtain the first binary information and the second binary information with a determined binary bit order.
[0131] In one embodiment of the present invention, the above coding sequence can be obtained by the following steps (1) or (2):
[0132] (1) sequencing each DNA sequence synthesized by the method of the present invention to obtain the above coding sequence; or
[0133] (2) Sequencing each DNA sequence synthesized by the method of the present invention to obtain information of each DNA short sequence; obtaining positional order information of each DNA short sequence based on the DNA index sequence identifier; and combining each DNA short sequence into the coding sequence based on the positional order information.
[0134] In one embodiment of the present invention, the decoding method further comprises: transcoding the first binary information and the second binary information into corresponding information, such as text information, image information, audio information and / or video information.
[0135] It should be noted that many details in the above-mentioned decoding method, especially the details involving technical features such as the first binary information, the second binary information, the first encoding rule, the second encoding rule, the input bit, the output bit, and the support bit, are the same as the details of such technical features in the above-mentioned encoding method, and therefore will not be repeated herein.
[0136] Corresponding to the information decoding method of the present invention, the embodiment of the present invention also provides an information decoding device, such as Figure 8 As shown, the device includes: an information acquisition unit 801, which is used to acquire a coding sequence generated by the above-mentioned coding device, as well as a first coding rule and a second coding rule, wherein the first coding rule is used to encode the first binary information, and the second coding rule is used to encode the second binary information; an information decoding unit 802, which reads the current symbol of the above-mentioned coding sequence, and converts the above-mentioned current symbol into binary bits of the first binary information and the second binary information according to the correspondence between different symbols and binary information in the first coding rule and the second coding rule; a result generating unit 803, which is used to determine each binary bit of the first binary information and the second binary information corresponding to each symbol bit of the above-mentioned coding sequence in turn through the correspondence between four different symbols and binary information in the first coding rule and the second coding rule, and obtain the first binary information and the second binary information with a determined binary bit order.
[0137] In one embodiment of the present invention, the above coding sequence is obtained by the following unit (1) or (2):
[0138] (1) a sequencing unit, used to sequence each DNA sequence synthesized by the method of the present invention to obtain the above coding sequence; or
[0139] (2) a sequencing unit, used to sequence each DNA sequence synthesized by the method of the present invention to obtain information of each DNA short sequence; an indexing unit, used to obtain positional order information of each DNA short sequence based on the DNA index sequence identifier; and a combining unit, used to combine each DNA short sequence into the coding sequence based on the positional order information.
[0140] In one embodiment of the present invention, the decoding device further comprises: a transcoding unit, configured to transcode the first binary information and the second binary information into corresponding information, such as text, image, audio or video information.
[0141] In one embodiment of the present invention, the above-mentioned different symbols are four base symbols A, T, C, and G; similarly, many details in the above-mentioned decoding device, especially the details of technical features such as the first binary information, the second binary information, the first encoding rule, the second encoding rule, the input bit, the output bit, and the supporting bit, are the same as the details of such technical features in the above-mentioned encoding method, and therefore will not be repeated herein.
[0142] Corresponding to the information decoding method of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, including a program, which can be executed by a processor to implement the information decoding method of the present invention.
[0143] In one embodiment of the present invention, a method for storing information using DNA sequences is also provided. Fig. 9 As shown, the method includes:
[0144] S901: The binary information to be stored is converted into DNA sequence information by the information encoding method of the present invention. The DNA sequence information includes a base sequence formed by four bases: A, T, C, and G.
[0145] S902: Synthesize the corresponding DNA sequence according to the above DNA sequence information.
[0146] S903: Save the above DNA sequence to achieve information storage.
[0147] In one embodiment of the present invention, a method for decoding information stored in the form of DNA sequences is also provided. Fig.10 As shown, the method includes:
[0148] S1001: Obtaining DNA fragments storing data information;
[0149] S1002: obtaining the DNA sequence of the DNA fragment by sequencing;
[0150] S1003: Convert the DNA sequence into binary information through a decoding method for converting the DNA sequence into binary information.
[0151] In some embodiments, the DNA sequence information converted from binary information is long and not conducive to direct synthesis. Therefore, as a preferred method, the DNA sequence information is split into multiple DNA short sequence information, and a DNA index sequence identifier is added to each of the split DNA short sequence information, and the DNA index sequence identifier contains the position order information of the DNA short sequence information; then, the corresponding DNA sequence is synthesized according to the DNA short sequence information; finally, the DNA sequence is saved to achieve information storage.
[0152] Generally speaking, DNA sequences can be stored in sample tubes in the form of dry powder, or embedded in amber, silica balls and other embedding materials for preservation. DNA sequences can also be transferred into living cells for preservation. The living cells can be microbial cells, more preferably Escherichia coli or Saccharomyces cerevisiae.
[0153] In some embodiments, there are multiple DNA fragments storing data information, and each DNA fragment carries a corresponding index sequence as the coding order information of the DNA fragment. In this case, in order to obtain complete binary information, it is necessary to first sort the DNA sequence obtained by sequencing according to the index sequence to obtain a sorted complete DNA sequence, and then convert the complete DNA sequence into binary information through the decoding method of converting a DNA sequence into binary information of the present invention.
[0154] The technical solutions of the present invention are described in detail below through examples. It should be understood that the examples are merely exemplary and cannot be construed as limiting the protection scope of the present invention.
[0155] Example
[0156] In this embodiment, the following encoding rule 1 in Table 1 and encoding rule 2 in Table 2 are selected:
[0157] Table 1 Coding rules 1
[0158] Supporting base Output candidate base when input bit is 0 Output candidate base when input bit is 1 A T / A G / C T G / C T / A C T / A G / C G G / C T / A
[0159] Table 2 Coding rules 2
[0160]
[0161]
[0162] In this embodiment, the supporting bit selection method is that the sixth bit before the current output bit is used as the supporting bit of encoding rule 1, and the first bit before the current output bit is used as the supporting bit of encoding rule 2.
[0163] In this embodiment, the data selected is to convert Li Bai's "Viewing the Waterfall at Mount Lu" into DNA code.
[0164] Wanglushan Waterfall
[0165] --Li Bai, Tang Dynasty
[0166] The sun shines on the incense burner, producing purple smoke, and in the distance I can see a waterfall hanging in front of the river.
[0167] The waterfall drops three thousand feet, as if the Milky Way is falling from the sky.
[0168] The specific process is as follows:
[0169] 1. Encoding and storage
[0170] (1) Extract the binary code corresponding to “Viewing the Waterfall at Mount Lu”, as shown in the following binary code: 1110011010011100100110111110010110101001000011100101101100011011000111 100111100000001001000111100101101110001000001100000101001001011010010 110111100101100101001001001000011000011010101101100011 1011100111100110011011110100001010111001101001011110011110000101 10100111111010011010011010011001111001111000001010011110011110010010010 0111111110011110110100101010111110011110000011100111111101111001000 11001110100110000001101001011110011110011100100010111110011110000000100100 01111001011011100010000011111001101000110001001100011001 111001011011101111001111001110001110000000100000010000001010111010011010001110 01111011100110101011000000111100111100110110100100101110001000 1011111001001011100010001001111001011000011110010110110000101110 10111011111011110010001100111001110011110011010100100011110011010011000101011111110100110010011101101101100110101010100101100111101000100100001011110111 1001001011100110011101111001011010010010101001111000111000000010000001000001010
[0171] (2) Split the above binary code into the following two binary codes of equal length, namely, binary code 1 and binary code 2:
[0172] Binary code 1: 1110011010011100100110111110010110101001000011100101101100011000111 10011110000000100100001111001011011100010000011001000001001001011010010 11011110010110010100100100100100001100001101011011000001110011010011 10111001111001100110111101000010101110011010010111101001011110011110000101 10100111111010011010011010011001111001111000001010001001111001111001010010 011111111001111011010010101011111001111000001110011111110111110010001100110011000000110100101111001111001000101111001111000000010010001
[0174] Binary code 2: 11100101101110001000001111100110100011001000001011100101100010011000110111 10010110110111001110111001110111000111000000010000010000010101110100100111001 111011100110101101011000000111100111100110110100100101110001000100010 11111001001011100010001001111001011000110110000011111001011010 111011111011110010001100111001110011010010001111001101001000111100110100110001010111111 1010011001001110110110111001101010101001111101000100100001011110111100100101011100110011101111001011010010010101001111000111000000010000001000001010
[0176] (3) Select the starting base sequence as: "ATCAGTGCTA". The starting base sequence is used to provide the starting support bit information when the base has not been output. The starting base sequence is an agreed virtual sequence that is only reflected during conversion and does not appear in the final DNA sequence. It is also used during decoding.
[0177] (4) Convert binary code 1 and binary code 2 into DNA coding sequences according to coding rules 1 and 2, as shown below:
[0178] GTACAAGGAGCATATGCACGCTTACCCTTGACTTACTGAAAGGCCGTTACCTCG
[0179] AATTCTGCATGTTTGAGAATATGATAAATTGGACTTTCAACAAGCTGAAACGAG
[0180] TTGATCTACGCGACTGAACAATAAGACCCACACGATGAGAACTCCCTTAAAAG
[0181] GGGAAGGTCAAAGCCGAAACTTTAACCCCTAAGCGCATGTTTGGGATATCACTG
[0182] GCTATGCCCTAAGCATGTCAGGTATTTGTCGTTCCGAATATCGGGCTTGCAGAC
[0183] TATATACGCAGAGCATGATTGCACGCTTCGCACAGCATCGCGAGGTTTTCAGCG
[0184] GGGAGCTATACTAGGGTTTCAGGGACCCTTATCATACTTCCTGAGACCCCCACG
[0185] GGCTTACCTCGCATCTATACTCTAACCAGACAGACTCAGGAAATGAAGGCAGTC
[0186] ATTAGACTGCGCCCAACAGTCCCAAAGGAAAATGCCACTATATCTCGCAACTAA
[0187] CATCTGCAGAGGCGCTAACCTC
[0188] (5) The DNA coding sequence is synthesized using chemical synthesis methods.
[0189] (6) Freeze-dry the synthesized DNA into powder and store it.
[0190] 2. Read the stored DNA sequence
[0191] (1) The stored DNA dry powder is used to construct a library, and then its sequence is obtained by high-throughput sequencing technology, as shown below:
[0192] GTACAAGGAGCATATGCACGCTTACCCTTGACTTACTGAAAGGCCGTTACCTCG
[0193] AATTCTGCATGTTTGAGAATATGATAAATTGGACTTTCAACAAGCTGAAACGAG
[0194] TTGATCTACGCGACTGAACAATAAGACCCACACGATGAGAACTCCCTTAAAAG
[0195] GGGAAGGTCAAAGCCGAAACTTTAACCCCTAAGCGCATGTTTGGGATATCACTG
[0196] GCTATGCCCTAAGCATGTCAGGTATTTGTCGTTCCGAATATCGGGCTTGCAGAC
[0197] TATATACGCAGAGCATGATTGCACGCTTCGCACAGCATCGCGAGGTTTTCAGCG
[0198] GGGAGCTATACTAGGGTTTCAGGGACCCTTATCATACTTCCTGAGACCCCCACG
[0199] GGCTTACCTCGCATCTATACTCTAACCAGACAGACTCAGGAAATGAAGGCAGTC
[0200] ATTAGACTGCGCCCAACAGTCCCAAAGGAAAATGCCACTATATCTCGCAACTAA
[0201] CATCTGCAGAGGCGCTAACCTC
[0202] (2) By using encoding rules 1 and 2, the above DNA sequence is decoded according to the reverse process of the encoding process to obtain two binary codes, namely binary code 1 and binary code 2:
[0203] Binary code 1: 1110011010011100100110111110010110101001000011100101101100011000111 10011110000000100100001111001011011100010000011001000001001001011010010 11011110010110010100100100100100001100001101011011000001110011010011 10111001111001100110111101000010101110011010010111101001011110011110000101 10100111111010011010011010011001111001111000001010001001111001111001010010 011111111001111011010010101011111001111000001110011111110111110010001100110011000000110100101111001111001000101111001111000000010010001
[0205] Binary code 2: 11100101101110001000001111100110100011001000001011100101100010011000110111 10010110110111001110111001110111000111000000010000010000010101110100100111001 111011100110101101011000000111100111100110110100100101110001000100010 11111001001011100010001001111001011000110110000011111001011010 111011111011110010001100111001110011010010001111001101001000111100110100110001010111111 1010011001001110110110111001101010101001111101000100100001011110111100100101011100110011101111001011010010010101001111000111000000010000001000001010
[0207] (3) Convert binary code 1 and binary code 2 into text information as shown below:
[0208] Wanglushan Waterfall
[0209] --Li Bai, Tang Dynasty
[0210] The sun shines on the incense burner, producing purple smoke, and in the distance I can see a waterfall hanging in front of the river.
[0211] The waterfall drops three thousand feet, as if the Milky Way is falling from the sky.
[0212] The above specific examples are used to illustrate the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art, according to the concept of the present invention, some simple deductions, modifications or substitutions can be made.
Claims
1. An information encoding method, characterized in that: The method comprises: Acquire first binary information and second binary information, as well as a first encoding rule and a second encoding rule, wherein the first encoding rule is used to encode the first binary information, and the second encoding rule is used to encode the second binary information; Obtaining a first candidate output symbol corresponding to a current input of the first binary information according to the first coding rule, and obtaining a second candidate output symbol corresponding to the current input of the second binary information according to the second coding rule, and taking the intersection of the first candidate output symbol and the second candidate output symbol as the output symbol corresponding to the current input; The output symbol corresponding to each binary bit of the first binary information and the second binary information is determined in sequence by the first encoding rule and the second encoding rule to obtain a coding sequence consisting of a plurality of the output symbols. The first output candidate symbols are two symbols among four symbols, the second output candidate symbols are two symbols among four symbols, and the first output candidate symbols and the second output candidate symbols have a same symbol; The first coding rule is that, under the support bit of the first coding rule, two symbols are selected from four symbols as the first output candidate symbols when the current input of the first binary information is 0, and the other two symbols are selected as the first output candidate symbols when the current input is 1, wherein the support bit of the first coding rule is any one of the four symbols, and the support bit of the first coding rule has a first corresponding relationship with the current output bit; The second coding rule is that, under the support bits of the first coding rule and the support bits of the second coding rule, two symbols are selected from four symbols as the second output candidate symbols when the current input of the second binary information is 0, and the other two are selected as the second output candidate symbols when the current input is 1, wherein the support bits of the second coding rule are any one of the four symbols, and the support bits of the second coding rule have a second corresponding relationship with the current output bits, and each support bit of the first coding rule corresponds to four support bits of the second coding rule.
2. The method according to claim 1, characterized in that The length of the first binary information is equal to the length of the second binary information.
3. The method according to claim 2, characterized in that The first binary information and the second binary information are obtained by splitting the same piece of binary information.
4. The method according to claim 1, characterized in that: The first corresponding relationship is a first predetermined number of bits before the current output bit; and the second corresponding relationship is a second predetermined number of bits before the current output bit.
5. The method according to claim 1, characterized in that The first output candidate symbol is two base symbols among the four bases A, T, C, and G, the second output candidate symbol is two base symbols among the four bases A, T, C, and G, and the supporting bit is one base symbol among the four bases A, T, C, and G; and, The coding sequence is a nucleic acid sequence comprising the four bases A, T, C, and G.
6. The method according to claim 1, characterized in that The method further comprises: A starting sequence is obtained, which provides a starting support bit for the first encoding rule and the second encoding rule before generating a first bit of the encoding sequence.
7. The method according to claim 1, characterized in that The method further comprises: Before acquiring the first binary information and the second binary information, the first binary information and the second binary information are extracted from a computer storage device.
8. An information encoding device, characterized in that: The device comprises: An information acquisition unit, used to acquire first binary information and second binary information, as well as a first encoding rule and a second encoding rule, wherein the first encoding rule is used to encode the first binary information, and the second encoding rule is used to encode the second binary information; an information encoding unit, configured to obtain a first candidate output symbol corresponding to a current input of the first binary information according to the first encoding rule, and obtain a second candidate output symbol corresponding to the current input of the second binary information according to the second encoding rule, and take an intersection of the first candidate output symbol and the second candidate output symbol as an output symbol corresponding to the current input; A result generating unit is used to determine the output symbol corresponding to each binary bit of the first binary information and the second binary information in sequence according to the first coding rule and the second coding rule, and obtain a coding sequence composed of a plurality of the output symbols. The first output candidate symbols are two symbols among four symbols, the second output candidate symbols are two symbols among four symbols, and the first output candidate symbols and the second output candidate symbols have a same symbol; The first coding rule is that, under the support bit of the first coding rule, two symbols are selected from four symbols as the first output candidate symbols when the current input of the first binary information is 0, and the other two symbols are selected as the first output candidate symbols when the current input is 1, wherein the support bit of the first coding rule is any one of the four symbols, and the support bit of the first coding rule has a first corresponding relationship with the current output bit; The second coding rule is that, under the support bits of the first coding rule and the support bits of the second coding rule, two symbols are selected from four symbols as the second output candidate symbols when the current input of the second binary information is 0, and the other two are selected as the second output candidate symbols when the current input is 1, wherein the support bits of the second coding rule are any one of the four symbols, and the support bits of the second coding rule have a second corresponding relationship with the current output bits, and each support bit of the first coding rule corresponds to four support bits of the second coding rule.
9. The device according to claim 8, characterized in that The first corresponding relationship is a first predetermined number of bits before the current output bit; and the second corresponding relationship is a second predetermined number of bits before the current output bit.
10. The device according to claim 8, characterized in that The first output candidate symbol is two base symbols among the four bases A, T, C, and G, the second output candidate symbol is two base symbols among the four bases A, T, C, and G, and the supporting bit is one base symbol among the four bases A, T, C, and G; and, The coding sequence is a nucleic acid sequence comprising the four bases A, T, C, and G.
11. A computer-readable storage medium, characterized in that: The method comprises a program which can be executed by a processor to implement the method according to any one of claims 1 to 7.
12. A method for storing information using DNA sequences, characterized in that: The method comprises: The binary information to be stored is converted into DNA sequence information by the information encoding method according to any one of claims 1 to 7, wherein the DNA sequence information includes a base sequence formed by four bases: A, T, C, and G; synthesizing a corresponding DNA sequence according to the DNA sequence information; The DNA sequence is stored to achieve information storage.
13. The method according to claim 12, characterized in that The method further comprises: Splitting the DNA sequence information into multiple pieces of DNA short sequence information, and adding a DNA index sequence identifier to each of the split DNA short sequence information, wherein the DNA index sequence identifier includes position order information of the DNA short sequence information; synthesizing a corresponding DNA sequence according to the DNA short sequence information; The DNA sequence is stored to achieve information storage.
14. The method according to claim 12 or 13, characterized in that The DNA sequence is stored in the form of dry powder or embedded in an embedding material.
15. The method according to claim 12 or 13, characterized in that The DNA sequence is transferred into living cells and preserved.
16. The method according to claim 15, characterized in that The living cells are microbial cells.
17. The method according to claim 15, characterized in that The living cells are Escherichia coli or Saccharomyces cerevisiae.
18. An information decoding method, characterized in that: The method comprises: Obtain a coding sequence generated by the coding method according to any one of claims 1 to 7, as well as a first coding rule and a second coding rule, wherein the first coding rule is used to encode first binary information, and the second coding rule is used to encode second binary information; Reading a current symbol of the coding sequence, and converting the current symbol into binary bits of the first binary information and the second binary information according to the correspondence between different symbols and binary information in the first coding rule and the second coding rule; Through the correspondence between different symbols and binary information in the first coding rule and the second coding rule, each binary bit of the first binary information and the second binary information corresponding to each symbol bit of the coding sequence is determined in turn to obtain the first binary information and the second binary information with a determined binary bit order.
19. The method according to claim 18, characterized in that The symbols are four base symbols: A, T, C, and G.
20. The method according to claim 19, characterized in that The coding sequence is obtained by the steps in (1) or (2): (1) sequencing each DNA sequence synthesized by the method according to claim 12 to obtain the coding sequence; or (2) sequencing each DNA sequence synthesized according to the method of claim 13 to obtain information of each DNA short sequence; According to the DNA index sequence identifier, obtaining the position order information of each of the DNA short sequences; According to the positional sequence information, the short DNA sequences are combined into the coding sequence.
21. The method according to claim 18, characterized in that The method further comprises: The first binary information and the second binary information are transcoded into corresponding information.
22. The method according to claim 21, characterized in that The corresponding information is text information, image information, audio information and / or video information.
23. An information decoding device, characterized in that: The device comprises: An information acquisition unit, used to acquire a coding sequence generated by the coding device according to any one of claims 8 to 10, as well as a first coding rule and a second coding rule, wherein the first coding rule is used to encode the first binary information, and the second coding rule is used to encode the second binary information; an information decoding unit, configured to read a current symbol of the coding sequence, and convert the current symbol into binary bits of the first binary information and the second binary information according to the correspondence between different symbols and binary information in the first coding rule and the second coding rule; The result generating unit is used to determine each binary bit of the first binary information and the second binary information corresponding to each symbol bit of the coding sequence in turn through the correspondence between different symbols and binary information in the first coding rule and the second coding rule, and obtain the first binary information and the second binary information with a determined binary bit order.
24. The device according to claim 23, characterized in that The symbols are four base symbols: A, T, C, and G.
25. The device according to claim 24, characterized in that The coding sequence is obtained by the following unit (1) or (2): (1) a sequencing unit, used to sequence each DNA sequence synthesized by the method according to claim 12 to obtain the coding sequence; or (2) a sequencing unit, used to sequence each DNA sequence synthesized according to the method of claim 13 to obtain information of each DNA short sequence; An index unit, used to obtain the position order information of each short DNA sequence according to the DNA index sequence identifier; A combining unit is used to combine the short DNA sequences into the coding sequence according to the positional sequence information.
26. The device according to claim 23, characterized in that The device further comprises: A transcoding unit is used to transcode the first binary information and the second binary information into corresponding information.
27. The device according to claim 26, characterized in that The corresponding information is text information, image information, audio information and / or video information.
28. A computer-readable storage medium, characterized in that: The medium includes a program, and the program can be executed by a processor to implement the decoding method according to any one of claims 18 to 22.
Citation Information
Patent Citations
Artificially synthesized DNA storage medium with coding information, storage reading method for information, and applications
CN104850760A
Encoding method and decoding method for performing information storage by means of DNA
CN105022935A