DNA data storage method, reading method and terminal based on coding optimization
By adopting encoding optimization and module mapping coding technologies in DNA data storage, the problems of high cost and low efficiency of DNA data storage in the prior art are solved, and more efficient data writing and recovery are achieved.
Patent Information
- Application Number
- CN202510139117.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The existing DNA data storage technology has high cost, time-consuming and error-prone problems in the data writing and sequencing process, and the synthetic DNA cannot be reused, which increases storage costs.
The DNA data storage method based on coding optimization is adopted to reduce the number of ‘bits 1’ by precoding, thereby reducing the usage of molecular modules and achieving cost and speed optimization. Use module mapping encoding to convert binary data into DNA module combinations and connect them to DNA molecular chains through enzyme catalytic reactions.
It realizes the reduction of the cost of DNA data storage and improves the writing speed, ensures the reliability and recovery capabilities of data, and solves the inefficiency of data storage and recovery in the prior art.
Smart Images

Figure CN119580800B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data storage technology, and in particular relates to a DNA-based data storage method, a data reading method, a data storage device, a data reading device, a terminal device and a computer-readable storage medium, in particular a DNA data storage method, a reading method and a terminal based on coding optimization. Background Art
[0002] In an era of explosive digital data growth, DNA is being explored as a next-generation molecular storage medium. Data can be encoded in vitro by using four natural nucleotides (A, T, G, and C), synthesizing data in DNA molecules, and retrieving data from DNA by sequencing. DNA data storage achieves extremely high data density due to the manipulation of molecules at the atomic level. DNA materials are very stable both in liquids and at relatively high temperatures, providing high durability (high retention) compared to many existing media materials. Data in DNA can be easily generated in hundreds of millions of copies by simple PCR reactions while maintaining low energy. DNA data retrieval has greatly benefited from advances in sequencing technology, including Illumina Next Generation Sequencing (NGS) and Nanopore 3. RD produces sequencing that can rapidly sequence human and other genomes at ever-decreasing prices.
[0003] However, current strategies for DNA data storage also face challenges. The storage of any dataset requires the template-free synthesis of specific long DNA molecules by chemical or enzymatic methods, which is still very expensive, time-consuming, labor-intensive and error-prone. The synthesized DNA can no longer be used to store other datasets, further increasing storage costs. For data retrieval, synthesis-based Illumina sequencing must cut long data DNA molecules into short fragments (<300 bases) to maintain a low error rate (<0.1%), and complex post-sequencing bioinformatics analysis is required to assemble the fragmented data.
[0004] For these reasons, efforts have been made to write data into universal DNA sequences, including natural DNA. There are still many problems to be solved in the research of its practical application. One of the great advantages of DNA as a storage medium is the stability of DNA molecules, which can be preserved for up to a hundred years without human intervention.
[0005] The applicant disclosed DNA data storage and reading solutions based on coding optimization in the previous applications with application number 202210114165.6, etc. During the implementation of these solutions, it was found that improvements in the solutions, including cost control, were needed. Summary of the invention
[0006] The purpose of the present invention is to provide a DNA data storage method, reading method, data storage device, data recovery device, terminal equipment and computer-readable storage medium based on coding optimization in the embodiments of the present application, which can solve the problem of cost reduction due to the inability to ensure the existence of the preprocessing algorithm adopted.
[0007] In a first aspect, this embodiment provides a DNA data storage method based on coding optimization, comprising:
[0008] S1: obtaining binary data of data to be stored, and pre-encoding the binary data according to an optimization target of DNA storage coding, wherein the optimization target of DNA storage coding includes that the total number of modules consumed for information expression meets a preset condition;
[0009] S2: After obtaining a set of DNA molecular chains represented by a module combination through module mapping coding, the corresponding DNA molecular chains are obtained by assembly or synthesis.
[0010] Pre-coding the binary data according to the optimization goal of DNA storage coding further includes:
[0011] Dividing the binary data into a plurality of segments;
[0012] The N binary data in each segment are converted into a symbol sequence according to a preset interval distance, and the probability information of each symbol in the sequence is counted. According to the optimization goal of minimizing the total module consumption, a precoding rule for converting the symbol encoding into new symbol information is set according to the probability information, and the binary data in the segment are re-encoded according to the precoding rule, where N is the number of binary numbers in the segment.
[0013] Setting a precoding rule for converting the symbol code into new symbol information according to the probability information further includes:
[0014] The optimization goal is to minimize the total number of bit 1s in the recoded binary sequence, and the set precoding rule is to replace characters with high probability with symbols marked with more bit 0s, and replace characters with low probability with symbols marked with more bit 1s.
[0015] Setting a precoding rule for converting the symbol code into new symbol information according to the probability information further includes:
[0016] The symbol with the highest probability is replaced with 00…0; the remaining symbols are assigned 1, 01, 001, … in descending order of probability, and through the variable-length recoding rule, each character after replacement has at most one “bit 1” in it.
[0017] Pre-coding the binary data according to the optimization goal of DNA storage coding further includes:
[0018] Dividing the binary data into a plurality of segments;
[0019] The N binary numbers in each fragment use bit flip information to extract the position where the flip occurs, and obtain a new sequence with the same length as the N binary data. In the new sequence, "bit 1" is assigned to the position where each flip occurs. The first bit information of the new sequence is consistent with the first bit information of the original data, and the binary data in the fragment is re-encoded, where N is the number of binary numbers in the fragment.
[0020] Pre-coding the binary data according to the optimization goal of DNA storage coding further includes:
[0021] Dividing the binary data into a plurality of segments;
[0022] In each segment, according to the pre-set run information, N binary numbers are converted into a character sequence represented by the run length, the probability information of each run symbol in the sequence is counted, and a recoding rule including long recoding and variable-length recoding is set for the run symbol according to the probability information; the N binary data are recoded according to the pre-coding rule, where N is the number of binary numbers in the segment.
[0023] After obtaining the set of DNA molecular chains represented by the module combination form, obtaining the corresponding DNA molecular chain by assembling means further includes:
[0024] Selecting suitable modules from a preset DNA module library and determining the order of the modules to obtain a corresponding module combination;
[0025] The modules in the module combination are connected into corresponding DNA molecular chains to complete the data storage.
[0026] Selecting the adapted module from the preset DNA module library through the module mapping code further includes:
[0027] Setting meta information and content information for each pre-divided segment, and the meta information further includes at least one piece of information including an interval distance and a precoding rule;
[0028] The meta information and the content information are mapped and encoded respectively using a meta information DNA module library and a content information DNA module library, and the meta information DNA module library and the content information DNA module library correspond to the rules of size and address recoding.
[0029] Using the content information DNA module library to map and encode the content information further includes:
[0030] Divide the binary information into m bits to obtain several short messages, each of which has corresponding address information;
[0031] Re-encoding the address information according to a-bit b-binary system;
[0032] Get the "data-address pair" corresponding to each re-encoded short message. When m=1, save the "data-address pair" with data 1 or data 0; when m>1, save any 2 m - 1 case of "data-address pair";
[0033] The "data-address pair" information is subjected to module mapping using the content information DNA module library to obtain a module combination adapted to the "data-address pair" of each short message and corresponding encoding information of each module.
[0034] The meta-information mapping is encoded using a meta-information DNA module library, which further comprises:
[0035] Each piece of meta-information corresponds to a module combination. The meta-information DNA module library is used to perform module mapping on each bit of information of each piece of meta-information. Each group of modules corresponds to a bit of the re-encoded information. Different modules in each group represent the content of the bit, so as to obtain the module combination adapted to the meta-information of each short message and the corresponding encoding information of each module therein.
[0036] Determining the module sequence to obtain a corresponding module combination further includes:
[0037] The meta information and content information of each segment are mapped and encoded to form the current segment combination unit.
[0038] The order of the corresponding fragment combination units is obtained according to the order of the fragments, thereby obtaining one or more module combinations of the binary data and the order between the modules in each module combination.
[0039] Connecting the modules in the module combination into corresponding DNA molecular chains to complete the data storage further comprises:
[0040] Find the module combination of the "data-address pair" of the meta information and content information of each fragment of the binary data and the corresponding encoding information of each module, and compose them into a fragment combination unit corresponding to the fragment;
[0041] Each module of the fragment assembly unit is a DNA module containing a specific base sequence. According to the pre-agreed module order, the corresponding DNA modules are connected into corresponding DNA molecular chains through enzyme catalysis reaction, and the DNA molecular chains are one or more.
[0042] According to the pre-agreed module sequence, connecting the corresponding DNA modules into corresponding DNA molecular chains through enzyme catalysis reaction further includes:
[0043] The terminal module of the meta information of the segment combination unit is connected to the first module of the content information, and the modules of the content information in the same segment combination unit are connected in sequence;
[0044] The modules between the fragment assembly units are connected by matching in a pre-agreed order: the first module of the rear-end fragment assembly unit is connected to the terminal module of the current fragment assembly unit, or the terminal module of the current fragment assembly unit is the tail DNA module of the current DNA molecular chain, and the first module of the rear-end fragment assembly unit is independently the first DNA module of a DNA molecular chain, thereby forming one or more DNA molecular chains for centralized storage.
[0045] Pre-coding the binary data according to the optimization goal of DNA storage coding further includes:
[0046] Dividing the binary data into a plurality of segments;
[0047] Meta information and content information are respectively set in each fragment, and the binary data in the fragment is divided into "data-address pairs" through a preset interval distance, and the "data-address pairs" include a content module and an address module; an optimization calculation module is set, and the content module, the address module and the optimization calculation module are represented as a module combination form of the content information according to the set optimization rules.
[0048] Setting an optimization calculation module, and expressing the content module, the address module and the optimization calculation module into a module combination form of the content information according to the set optimization rule further includes:
[0049] Setting the optimization calculation module further includes setting a run character quantity module representing the number of times a character appears;
[0050] Representing the content module, the address module and the optimization calculation module in a module combination form according to the set optimization rule further includes: optimizing according to the stored content module, and using a combination of the address module and the run module to represent the module combination of the content information.
[0051] Selecting the adapted module from the preset DNA module library through the module mapping code further includes:
[0052] Meta information is set for each pre-divided segment, which further includes at least one information including an interval distance, a recoding rule, and an optimization rule;
[0053] The meta information and the content information are mapped and encoded respectively using a meta information DNA module library and a content information DNA module library, and the meta information DNA module library and the content information DNA module library are in regular correspondence of size and address.
[0054] Setting an optimization calculation module, and expressing the content module, the address module and the optimization calculation module into a module combination form according to the set optimization rules further includes:
[0055] Setting the optimization calculation module further includes setting a run character quantity module representing the number of characters that appear;
[0056] The content module, the address module and the optimization calculation module are expressed as a module combination according to the set optimization rules, further including: after a single "data-address pair", multiple run modules are connected to directly express a longer set of data through a longer DNA molecular chain.
[0057] Obtaining the corresponding DNA molecule chain by synthesis further includes:
[0058] The corresponding module combination is mapped to the corresponding base sequence according to the pre-designed DNA module, converted into DNA molecule long chain information, and further directly synthesized into the molecule long chain.
[0059] DNA data reading method based on coding optimization, including:
[0060] Perform module recognition on DNA molecular chains;
[0061] Decode the module mapping information according to the module information to obtain the corresponding binary data;
[0062] The original data content is reconstructed by decoding the binary data according to the decoding operation corresponding to the pre-coding rules to realize data reading.
[0063] The present invention introduces precoding and module mapping cascade coding, based on a pre-synthesized module library, and physically stores electronic data through DNA assembly. The precoding in the present invention mainly refers to reducing the number of "bit 1", thereby reducing the amount of molecular modules, achieving the effect of reducing costs and increasing speed. By introducing module mapping coding, DNA data storage based on assembly technology is realized, and a specific module combination is selected from a preset DNA module library for reaction connection into a DNA molecular chain, and its speed and cost can be optimized compared to synthesis technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0065] Figure 1 A schematic diagram of a system architecture for an application scenario provided in an embodiment of the present application;
[0066] Figure 2 This is a flow chart of a DNA data storage method based on coding optimization according to an embodiment of the present application;
[0067] Figure 3 A flowchart of DNA data reading based on coding optimization;
[0068] Figure 4 The following is an example diagram of the complete process of writing and reading data. DETAILED DESCRIPTION
[0069] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0070] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0071] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0072] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0073] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0074] With the widespread application of digital information and the rapid development of big data science, the information data generated by people every day is growing exponentially, and the existing traditional storage media have gradually failed to meet the demand. As a new type of storage medium, DNA molecular chain has attracted widespread attention due to its advantages such as high storage density, long preservation time, low maintenance cost and strong stability.
[0075] At present, research on DNA storage is more focused on improving storage efficiency through efficient encoding, decoding and improving fault tolerance, thereby reducing the cost of synthetic sequencing; and on DNA-based storage media, more emphasis is placed on the exploration of theory and material structure.
[0076] The general process of DNA storage is to encode digital data into a DNA base sequence, synthesize DNA according to the encoded base sequence, and store it in a storage medium inside or outside the body. Among them, the synthesis of DNA can be achieved by writing the nucleotide base sequence through a synthesizer, and then stored in a pooled liquid as a medium. When reading data, it can be read and sequenced by a sequencer, and the data can be restored through subsequent decoding processing.
[0077] The existing DNA storage direction mainly focuses on efficient encoding and decoding methods, related data modeling, and research on DNA storage medium materials and structures; in the process of DNA-based data storage, digital data will be preprocessed before it is stored, such as compression, deletion, encryption, encoding, etc. If the data can still be restored after decades or hundreds of years of data storage, it is necessary to know the data and its corresponding preprocessing algorithm; For example, the applicant disclosed a DNA data storage scheme and reading scheme based on coding optimization in the previous application of application number 202210114165.6, etc. However, in the implementation process, how to reduce the cost of synthetic sequencing has become a big problem in the research of this industry. The expected hope of the proposal of this application is to reduce the number of DNA base sequences used in synthesis, thereby achieving the effect of reducing the cost of synthetic sequencing.
[0078] See also Figure 1 , is a schematic diagram of the system architecture of an application scenario provided by an embodiment of the present application. Figure 1 The embodiment of the present application shows an end-to-end full-process DNA storage system architecture that realizes data self-storage and self-recovery, such as Figure 1 As shown, before storing the data, digital data (raw data) of different formats are pre-encoded to obtain binary data to be encoded; through module mapping encoding, an adapted module combination is selected from a preset DNA module library, and the binary data is converted into module coding information required for DNA assembly; these module coding information are reacted and connected into DNA molecular chains, and stored in in vivo and in vitro storage media. In addition to the above-mentioned assembly form, the present invention can also obtain the corresponding DNA molecular chains by synthesis after obtaining a set of DNA molecular chains represented by a module combination.
[0079] During the storage process, the layout of data in the DNA storage medium is optimized according to the characteristics of different DNA storage media and in combination with traditional silicon-based storage media. When analysis is required, the DNA molecular chain is subjected to module identification, and module matching is performed according to a pre-set module library to obtain module mapping coding information; the module mapping coding information is decoded and converted into the binary data; the binary data is reconstructed into the original binary data content according to the decoding operation corresponding to the pre-coding to realize data reading. Module identification can be performed by sequencing to obtain the base sequence; and then module matching is performed according to the pre-set module library based on the base sequence results. At the same time, if the corresponding DNA molecular chain is obtained by synthesis, the base sequence can be obtained by sequencing; and then module matching is performed according to the pre-set module library based on the base sequence results.
[0080] Even if the DNA data is lost during storage, the data can still be fully restored to all data, and even the complete original data can be restored through very little DNA information; in the application scenario of large-scale and complex data storage, it is to realize the self-containment and self-analysis of the data stored in DNA; the reliability of large-scale data storage in a long-term uncertain environment and the integrity of data recovery are guaranteed. More importantly, through module mapping coding, an adapted module combination is selected from a preset DNA module library, and the binary data is converted into the module coding information required for DNA module assembly; these module coding information are reacted and connected into a DNA molecular chain, which makes the DNA molecular chain grow into a fully modular connection and quickly connected, solving the problem of slow writing speed and high cost when DNA storage is combined with biosynthesis technology for data storage. The DNA storage of the present invention has a fast writing speed, modular writing and low cost, and when the data of DNA storage is recovered, the reliability is high and the data recovery is fast.
[0081] A complete data operation process includes the process of writing data into a DNA molecule chain and reading or restoring data from a DNA molecule chain. However, within the protection scope of the present invention, whether data is read out of a DNA molecule chain or data is written into a DNA molecule chain, as long as it is implemented using coding optimization technology, it should be within the protection scope of the present invention and is one of the embodiments of the present invention.
[0082] First embodiment
[0083] See also Figure 2 , which is a flow chart of a DNA data storage method based on coding optimization. It includes:
[0084] S110: Obtain binary data of data to be stored, and pre-encode the binary data according to the optimization target of DNA storage coding, wherein the optimization target of DNA storage coding includes that the total number of modules consumed for information expression meets a preset condition;
[0085] S120: After obtaining a set of DNA molecular chains represented by a module combination through module mapping coding, the corresponding DNA molecular chains are obtained through assembly or synthesis.
[0086] Obtaining the corresponding DNA molecular chain by assembly further includes: selecting the adapted modules from the preset DNA module library by module mapping coding, and determining the module sequence to obtain the corresponding module combination; connecting the modules in the module combination into the corresponding DNA molecular chain to complete the data storage.
[0087] Obtaining the corresponding DNA molecular chain by synthesis further includes: mapping the corresponding module combination to the corresponding base sequence according to the pre-designed DNA module, converting it into DNA molecular long chain information, and further directly synthesizing the molecular long chain.
[0088] See also Figure 3 , which is a flowchart of DNA data reading based on coding optimization. It includes:
[0089] S210: performing module identification on the DNA molecular chain (for example, by sequencing the DNA and matching the base sequences obtained by sequencing to obtain module information);
[0090] S220: Decode the module mapping information according to the module information to obtain corresponding binary data;
[0091] S230: Reconstruct the original data content by performing the decoding operation corresponding to the binary data pre-coding to realize data reading.
[0092] Module Identification Module identification through DNA sequencing is only one means of identification.
[0093] The complete process of data writing and reading mainly includes the following steps: (1) pre-coding; (2) module mapping coding; (3) DNA assembly and storage. Figure 4 shown.
[0094] S1: Precoding
[0095] If the data to be stored is non-binary data, it is first converted into binary data and divided into k segments according to a certain length of bits.
[0096] The precoding in the present invention mainly refers to reducing the number of "bits 1", thereby reducing the amount of molecular modules, achieving the effect of reducing costs and increasing speed. There are many precoding algorithms, and the following are just examples.
[0097] The first implementation method of precoding (the method of precoding DNA data using information entropy) can use the precoding methods mentioned in this article to reduce the number of "bits 1", thereby reducing the amount of molecular modules, achieving the effect of reducing costs and increasing speed. "Precoding the binary data according to the optimization goal of DNA storage coding" further includes:
[0098] Divide the binary data into a number of segments (e.g., k segments);
[0099] The N binary data in each fragment are converted into a symbol sequence according to a preset interval distance, and the probability information of each symbol in the sequence is counted. According to the optimization goal of minimizing the total module consumption, the precoding rule for converting the symbol encoding into new symbol information is set according to the probability information, and the binary data in the fragment is re-encoded according to the precoding rule, where N is the number of binary numbers in the fragment. It should be noted that this step can form a complete processing process with the subsequent step S120. Generally speaking, the processing of each fragment can be a separate process, and in this separate processing process, this step, module mapping encoding, etc. become a continuous process. However, while the previous fragment is being processed, the subsequent one or more fragments can be one or more parallel processing processes.
[0100] Let me give you an example:
[0101] S11: Assume that the binary data of a segment input is 0110110011011110;
[0102] S12: Convert the binary data into a symbol sequence using a custom interval distance (assuming it is 2). In this example, there are four symbols (01, 10, 11, 00), as shown in the first row of Table 1. Statistical information on the probability of the four symbols in the sequence is shown in the second row of Table 1.
[0103] Further, according to the optimization goal of the subsequent DNA storage coding (i.e., the total number of modules consumed for information expression is minimized), the above symbols are re-encoded. For example, in this example, the optimization goal is to minimize the total number of bits 1 in the re-encoded binary sequence, and the following re-encoding is performed, as shown in the third row of Table 1:
[0104] Table 1 Probability information of four symbols in the sequence and fixed-length recoding example table
[0105] Original symbol 01 10 11 00 Probability 0.25 0.25 0.375 0.125 New symbols 01 10 00 11
[0106] The binary data of the input segment of this embodiment is re-encoded as shown in Table 2:
[0107] Table 2 Example of fixed-length re-encoding of input binary data in this embodiment
[0108] Original data 0110110011011110 Encoded data 0110001100010010
[0109] The precoding method described in Table 2 is: replace the characters with high probability with symbols with "more 0 bits", and replace the characters with low probability with symbols with "more 1 bits". It is worth noting that the method described in Table 2 is a fixed-length recoding. In practice, variable-length recoding can also be used. The probability information of the four symbols in the sequence and the variable-length recoding are shown in Table 3:
[0110] Table 3 Probability information of four symbols in the sequence and variable length recoding example table
[0111] Original symbol 01 10 11 00 Probability 0.25 0.25 0.375 0.125 New symbols 1 01 000 001
[0112] The variable length recoding of binary data of the input segment in this embodiment is shown in Table 4:
[0113] Table 4 Example of variable length re-encoding of input binary data in this embodiment
[0114] Original data 0110110011011110 Encoded data 101000001000100001
[0115] This variable-length recoding method replaces the symbol with the highest probability with 00…0 (a total of N-1 zeros), where N is the total number of symbols. The remaining symbols are assigned 1, 01, 001, … in descending order of probability. This variable-length recoding method ensures that each character after replacement has at most one “bit 1” inside, but the total length of the recoded data will also increase.
[0116] "Setting precoding rules for converting the symbol code into new symbol information based on the probability information" further includes: the optimization goal is to minimize the total number of bit 1s in the re-encoded binary sequence, and the set precoding rules are to replace characters with high probability with symbols with "more bit 0s" and characters with low probability with symbols with "more bit 1s".
[0117] S13: In the subsequent DNA module data encoding of S120 (the module image encoding process is that after data division, content address pairs are formed; address recoding is represented by corresponding module connection; content corresponding module is connected with address corresponding module; block address is represented by sticky end), the interval distance of symbol sequence division and recoding rules and other meta information can be stored in the first several bits of the data document according to the set rules. Through this method, in one example, the information of "bit 1" in the data obtained after encoding is effectively reduced, and the total number of DNA modules involved in the DNA module encoding process is also reduced, thereby playing a role in optimizing storage.
[0118] The second implementation method of precoding: a method of precoding DNA data using flipping information.
[0119] Wherein, “pre-encoding the binary data according to the optimization target of DNA storage coding” further includes:
[0120] Divide binary data into several segments;
[0121] The N binary numbers in each fragment use the bit flip information to extract the position where the flip occurs, and obtain a new sequence with the same length as the N binary data. In the new sequence, "bit 1" is assigned to the position where each flip occurs. The first bit information of the new sequence is consistent with the first bit information of the original data, and the binary data in the fragment is re-encoded, where N is the number of binary numbers in the fragment.
[0122] Let me give you an example:
[0123] S21: Assume that the input binary data is: 1111000111110011110;
[0124] S22: Use its bit flip information to extract the position where the flip occurs, that is: [5,8,13,15,19], as shown in the second row of Table 5. Take a new sequence of the same length as the original data, and assign "bit 1" to each position where the flip occurs in the new sequence. The first bit information of the new sequence is consistent with the first bit information of the original data. In this example, since the first bit of the original data is 1, the first bit of the encoded new sequence is also 1. If the first bit of the original data is 0, the first bit of the encoded new sequence is also 0. The data obtained after encoding is shown in the third row of Table 5:
[0125] Table 5 Bit flip encoding example table
[0126] Original data 1111000111110011110 Flip Information [5,8,13,15,19] After encoding, the data is obtained 1000100100001010001
[0127] (S23) After the DNA module data is encoded in the subsequent step S120, the basic module image encoding process is as follows: after data division, content address pairs are formed; addresses are re-encoded and represented by corresponding module connections; content corresponding modules are connected with address corresponding modules; block addresses are represented by sticky ends.
[0128] In the obtained data, the information of "bit 1" is effectively reduced, and the total number of DNA modules involved in the DNA module encoding process is also reduced, thereby playing a role in optimizing storage.
[0129] The third implementation method of precoding: a method of precoding DNA data using run-length information.
[0130] Wherein, “pre-encoding the binary data according to the optimization target of DNA storage coding” further includes:
[0131] Divide binary data into several segments;
[0132] In each segment, according to the pre-set run information, N binary numbers are converted into a character sequence represented by the run length, the probability information of each run symbol in the sequence is counted, and a recoding rule including long recoding and variable-length recoding is set for the run symbol according to the probability information; the N binary data are recoded according to the pre-coding rule, where N is the number of binary numbers in the segment.
[0133] Let me give you an example:
[0134] S31: Assume that the input binary data is 0110110011011110;
[0135] S32: Using its run information, convert it into a character sequence represented by the run length, that is, [3, 2, 2, 3, 1, 4, 2, 1, 3, 3]. In actual use, the number of run characters can be limited. For example, when the run is limited to 10, the run length 17 exceeding 10 can be converted into [10, 0, 7]. The probability information of the run symbol in the sequence is statistically analyzed, as shown in the second row of Table 6.
[0136] The run symbols can be re-encoded according to the probability information, including equal-length re-encoding and variable-length re-encoding as shown in the third row of Table 6 and the fourth row of Table 6 respectively.
[0137] Table 6 Probability information of run symbols in sequence and recoding example table
[0138] symbol 1 2 3 4 Probability 0.2 0.3 0.4 0.1 Equal length recoding 10 01 00 11 Variable length recoding 01 1 000 001
[0139] The data obtained after equal-length recoding and variable-length recoding of the binary data input in this embodiment are shown in Table 7:
[0140] Table 7 Example of equal-length recoding and variable-length recoding of binary data input in this embodiment
[0141] Original data 111001100010000110111000 The data is obtained after fixed-length re-encoding 00010100101101100000 Data obtained after variable length re-encoding 0001100001001101000000
[0142] S33: The basic module image encoding process in the DNA module data encoding in the subsequent step S120 is as follows: after data division, content address pairs are formed; addresses are re-encoded and represented by corresponding module connections; content corresponding modules are connected to address corresponding modules; block addresses are represented by sticky ends.
[0143] The first bit information (0 or 1), re-encoding rules, and the set value of the longest run, etc., can be stored in the first few bits of the data file according to the set rules. Through this method, the information of "bit 1" in the encoded data is effectively reduced, and the total number of DNA modules involved in the DNA module encoding process is also reduced, thereby playing a role in optimizing storage.
[0144] S120: Module mapping encoding (the first embodiment first introduces assembly).
[0145] After obtaining n new data fragments, the binary data information in each fragment is converted into the module coding information required for DNA assembly through module mapping coding.
[0146] Specifically, each new segment information contains two parts of information: meta information and content information.
[0147] Meta information records the information of the information fragment. For example, when precoding, the meta information records the address information and / or encoding rules of the information fragment in the original data. The meta information record may also include the number of K, the value of N, and information about the fragment combination. The content information includes the data information and redundant information of the fragment. The redundant information can be generated by any method based on the meta information and content information, and is used to restore information when a read error occurs within the fragment.
[0148] For meta information and content information, two different module libraries can be used for mapping encoding, or only one module library can be used to complete the mapping encoding. The meta information module library and the content information module library are logically divided into two libraries, which can be merged into the same module library. The modules used to distinguish between the modules representing meta information and the modules used to represent content information are mapped and matched separately only for the convenience of expression. In the specific implementation, only one module library needs to be set to complete the image encoding.
[0149] That is, meta information and content information are set for each fragment respectively; the meta information and content information are mapped and encoded respectively using a prefabricated meta information DNA module library and a content information DNA module library, the meta information DNA module library and the content information DNA module library correspond to the rules of size and address recoding, if a-bit b-base recoding is adopted, the module library is set to a*b modules, each group of modules corresponds to a bit of the recoded information, and different modules in each group represent the content of the bit.
[0150] "Use the content information DNA module library to map and encode the content information", which further includes: dividing the binary information into m bits to obtain a number of short messages, each short message has corresponding address information; re-encoding the address information according to the a-bit b-base; obtaining the "data-address pair" corresponding to each re-encoded short message, when m=1, saving the "data-address pair" with either data 1 or data 0; when m>1, saving any 2 m -1 case of "data-address pair"; use the content information DNA module library to perform module mapping on the "data-address pair" information to obtain the module combination adapted to the "data-address pair" of each short message and the corresponding encoding information of each module.
[0151] The above content information is processed and encoded through the following three steps, as shown in Table 8. The specific example process is as follows:
[0152] 1) Information splitting: Split the binary information into m bits to obtain several short messages, each of which has its own address information. For example, assuming the data is 10011010, the information splitting example when m=1 is shown in the second row of Table 8.
[0153] 2) Information reconstruction: Re-encode the address. Assuming the re-encoding rule is 3-bit binary, the information reconstruction is shown in the third row of Table 8.
[0154] 3) Information mapping: According to the above two steps, the "data-address pair" is obtained as shown in the fourth row of Table 8:
[0155] Table 8 Example table of information splitting, reconstruction and mapping
[0156] data 1 0 0 1 1 0 1 0 address 7 6 5 4 3 2 1 0 Address recoding 111 110 101 100 011 010 001 000 Data-Address Pair 1-111 0-110 0-101 1-100 1-011 0-010 1-001 0-000
[0157] It is worth noting that when m=1, it is only necessary to save the "data-address pair" of either data 1 or data 0; when m>1, it is necessary to save any 2 m -1 case of "data-address pair". For this example, assume that the "data-address pair" with data 1 is saved as shown in Table 9:
[0158] Table 9: Example table of "data-address pairs" assuming the stored data is 1
[0159] data 1 0 0 1 1 0 1 0 Data-address pairs to be saved 111 / / 100 011 / 001 /
[0160] Furthermore, the "data-address pair" information that needs to be actually stored is mapped to modules. Prepare a DNA module library, whose size corresponds to the address recoding rule, that is, a-bit b-binary recoding corresponds to (a*b modules). In the above example, the module library contains a total of 6 modules (3 groups, 2 modules in each group). Each group of modules corresponds to a bit of the recoded information, and different modules in each group represent the content of the bit. That is , where a=0,1,2, b=0,1. For example, for the actually stored information "011", the corresponding module combination mapping is: For the data in the above example, it is converted into the module mapping code shown in Table 10.
[0161] Table 10 Module mapping encoding conversion example table
[0162]
[0163] For the meta-information, directly perform the information mapping of the third step above to obtain the module combination result. "Use the meta-information DNA module library to map and encode the meta-information", which further includes: each meta-information corresponds to a module combination, and the meta-information DNA module library is used to perform module mapping on each bit of information of each meta-information, each group of modules corresponds to a bit of the re-encoded information, and different modules in each group represent the content of the bit, so as to obtain the module combination adapted to the meta-information of each short message and the corresponding encoding information of each module therein. That is, each meta-information corresponds to a module combination, and each content information corresponds to multiple module combinations. At this point, the meta-information and content information in each information segment are converted into two groups of module combination information after the module mapping encoding in the above steps. For example, when the meta-information is "110" and the content information is as in the above example, the encoding shown in Table 11 is obtained:
[0164] Table 11 Example of mapping encoding for meta information "110"
[0165]
[0166] That is, find the module combination of the "data-address pair" of the meta information and content information of each fragment of the binary data and the corresponding encoding information of each module, and compose them into the fragment combination unit corresponding to the fragment. In other words, the meta information and content information of each fragment are respectively mapped and encoded to form the current fragment combination unit, and the order of the corresponding fragment combination units is obtained according to the order of the fragments, so as to obtain one or more module combinations of binary data and the order between each module in the module combination (such as the partial module combination shown in Table 11). The DNA molecular chain formed by the DNA module can be one or more. The meta information of each module in the DNA molecular chain contains its corresponding fragment information. Therefore, when the data is read out, the position information of the corresponding fragment information can also be obtained by parsing the meta information, and the corresponding original binary data can also be read out at the same time.
[0167] S3: DNA assembly and storage
[0168] The module combination of the "data-address pair" of the meta-information and content information of each fragment of the binary data and the corresponding encoding information of each module are found, and they are assembled into a fragment combination unit corresponding to the fragment; each module of the fragment combination unit is a DNA module containing a specific base sequence, and the corresponding DNA modules are connected into corresponding DNA molecular chains through enzyme-catalyzed reactions according to the pre-agreed module order.
[0169] According to the pre-agreed module sequence, connecting the corresponding DNA modules into corresponding DNA molecular chains through enzyme catalysis reaction further includes:
[0170] The terminal module of the meta information of the segment combination unit is connected to the first module of the content information, and the modules of the content information in the same segment combination unit are connected in sequence;
[0171] The modules between the fragment assembly units are connected according to the pre-agreed sequence matching between the modules: the first module of the rear-end fragment assembly unit is connected to the terminal module of the current fragment assembly unit, or the terminal module of the current fragment assembly unit is the tail DNA module of the current DNA molecular chain, and the first module of the rear-end fragment assembly unit is independently the first DNA module of a DNA molecular chain, thereby forming one or more DNA molecular chains for centralized storage, and the connection refers to the design of different DNA modules through sticky ends so that they can perform chemical reactions under an enzyme catalytic environment to achieve DNA assembly. For example, when the terminal module of the meta-information and the first module of the content information generate a DNA molecular chain, they are connected by covalent bonds to achieve assembly. After the DNA modules in the fragment assembly unit are connected to form a DNA molecular chain, they can be stored independently to form multiple DNA molecular chains, or several DNA molecular chains can be further connected into longer molecular chains for storage.
[0172] For example: In the DNA module library, each module is a DNA module containing a specific base sequence, which is distinguished by the difference in base sequence. The modules between specific groups can be assembled through chemical reactions under enzyme catalysis by designing sticky ends. Specifically, in the above example, the module representing meta-information and Can connect, and , Can be connected, the terminal module of meta information The first module with content information It can be connected, and the connection of other modules is similar. According to the results of erasure code and module mapping cascade coding, DNA assembly reaction is carried out to form several DNA molecular chains for centralized storage.
[0173] When the fragment combination unit is a DNA module including 10 bases, the edit distance between the fragment combination units "ATCGTAGCCA" and "TTCGTAGCCA" is 1, and the edit distance between the fragment combination units "ATCGTAGCCA" and "TAGCATCGGT" is 10. In order to facilitate reading, when designing the fragment combination unit, only multiple fragment combination units whose edit distances between each other are greater than or equal to the preset distance threshold can be selected. In this way, in the process of reading information, even if errors occur in individual bases or other fragments in the fragment combination unit, as long as it can be determined that the edit distance is less than the preset distance threshold, the read fragment combination unit can still be mapped to a specific code, thereby improving the fault tolerance of the fragment combination unit storage.
[0174] In the method for storing information in molecules disclosed in the present invention, information is stored according to content-address pairs, and by re-encoding the address and / or content of the information, fragment combination units in a prefabricated module library are repeatedly used for large-scale parallel assembly to achieve information storage. Compared with storing information by synthesizing DNA by growing nucleotides one by one, the number of types of fragment combination units required is greatly reduced, and parallel assembly greatly improves the efficiency of the combination, thereby reducing the difficulty of storage and improving storage efficiency.
[0175] Second embodiment
[0176] Precoding can set an optimization calculation module, which can have many algorithms. The content information can include a content module, an address module and an optimization calculation module, and optimization rules are given to perform optimization.
[0177] In other words, the binary data is pre-encoded according to the optimization target of the DNA storage coding. In addition to the direct pre-coding of the first example, the pre-coding of this example can also be to first divide the binary data into several segments, set meta information and content information for each segment, set an optimization calculation module for the content information, and give optimization rules to optimize the DNA coding. This pre-coding is a coding in a broad sense. The pre-coding of the present invention, that is, the setting of reducing the total number of module consumption for information expression, belongs to the protection scope of the present invention.
[0178] Pre-coding the binary data according to the optimization goal of DNA storage coding further includes:
[0179] Divide the binary data into a number of segments (e.g., K);
[0180] Meta information and content information are respectively set in each fragment, and the binary data in the fragment is divided into "data-address pairs" through a preset interval distance, and the "data-address pairs" include a content module and an address module; an optimization calculation module is set, and the content module, the address module and the optimization calculation module are represented as a module combination form of the content information according to the set optimization rules.
[0181] "Setting an optimization calculation module, and representing the content module, the address module and the optimization calculation module as a module combination form of the content information according to the set optimization rules" further includes: "Setting an optimization calculation module" further includes setting a run character quantity module representing the number of times a character appears; "Representing the content module, the address module and the optimization calculation module as a module combination form according to the set optimization rules" further includes: optimizing according to the stored content modules, and using a combination of the address module and the run module to represent the module combination of the content information.
[0182] “Selecting a suitable module from a preset DNA module library through module mapping encoding” further includes:
[0183] Meta information is set for each pre-divided segment, which further includes at least one information including an interval distance, a recoding rule, and an optimization rule;
[0184] The meta information and the content information are mapped and encoded respectively using a meta information DNA module library and a content information DNA module library, and the meta information DNA module library and the content information DNA module library are in regular correspondence of size and address.
[0185] In another example, “setting an optimization calculation module, and expressing the content module, the address module and the optimization calculation module into a module combination form according to a set optimization rule” further includes:
[0186] "Setting the optimization calculation module" further includes setting a run character quantity module that represents the number of characters that appear; "representing the content module, the address module and the optimization calculation module into a module combination form according to the set optimization rules" further includes: after a single "data-address pair", connecting multiple run modules to directly represent a longer set of data through a longer DNA molecule chain.
[0187] An example is:
[0188] S41: Assume that the input binary data is 0110110011011110;
[0189] S42: Divide the original data sequence into data address pairs by a custom interval distance K (assuming K=1), as shown in Table 12.
[0190] Table 12 Example of dividing the original data sequence to generate address pairs Table 1
[0191] content 0 1 1 0 1 1 0 0 1 1 0 1 1 1 1 0 address 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
[0192] The above division method forms data address pairs: 0-0, 1-1, 1-2, 0-3, 1-4, ...
[0193] In this example, the DNA modules corresponding to the preset addresses and contents are designed, and there are 2+16=18 modules. The design of the DNA module is a short chain of DNA molecules composed of a certain number of bases of a custom length.
[0194] Each data address pair can be represented by a module connection, such as "content module 0-address module 3", "content module 1-address module 4". In practice, when the interval distance K=1, only the data address pair corresponding to content module 0 or content module 1 can be stored, and the other one can be represented by a whole vacancy.
[0195] Set a run character number module, for example: use "run module 1" to indicate that the number of consecutive characters is 1, use "run module 2" to indicate that the number of consecutive characters is 2, and so on. Connect the content module, address module and run module to form a representation of a segment of consecutive repeated characters. In actual use, the number of run characters can be limited. For example, when the upper limit of the run is set to 10, the run length exceeding 17 can be converted into two DNA molecular chains containing address module I and address module I+10, namely "content module X-address module I-run module 10" and "content module Y-address module I+10-run module 7".
[0196] Applying this method to the above example, the resulting module combination is shown in Table 13:
[0197] Table 13 Module combination example Table 1
[0198] Serial number Content Module Address module Run Module 1 Content Module 1 Address module 1 Travel module 2 2 Content Module 1 Address module 4 Travel module 2 3 Content Module 1 Address module 8 Travel module 2 4 Content Module 1 Address module 11 Travel module 4
[0199] When the interval distance K=1, since all the stored content modules are content modules 1 (or all are content modules 0), the connection of the content modules can be omitted, and only the combination of the address module and the run module is used to represent each DNA chain, as shown in Table 14:
[0200] Table 14 Example of using the combination of address module and run module Table 1
[0201] Serial number Address module-travel module 1 Address module 1-Run module 2 2 Address module 4 - Run module 2 3 Address module 8-Run module 2 4 Address module 11-travel module 4
[0202] Furthermore, multiple run modules can be connected after a single content address module pair, and a longer set of data can be directly represented by a longer DNA molecule chain. For example, the data segment "011011" can be represented by "content module 1-address module 1-run module 2-run module 1-run module 2". The second run module represents the number of consecutive bits 0 after consecutive bits 1, the third run module represents the number of consecutive bits 1 after consecutive bits 0, and so on.
[0203] S43: After encoding a set of DNA molecular chains represented by a combination of modules, the corresponding DNA molecular chains can be obtained by assembly or synthesis. The assembly method is to prepare the corresponding short DNA molecular chains in advance, add sticky ends to the beginning and end of the corresponding modules, mix the modules according to the coding combination, and add the corresponding DNA ligase to achieve the assembly reaction. The synthesis method is to map the corresponding module combination to the corresponding base sequence according to the pre-designed DNA module, convert it into DNA molecular long chain information, and further directly synthesize the molecular long chain.
[0204] Another example of the above method: the interval distance K=2, comprising:
[0205] Assume the input binary data is 0110011111001010;
[0206] The interval distance K=2, divide the data address pairs, as shown in Table 15:
[0207] Table 15 Example of dividing the original data sequence to generate address pairs Table 2
[0208] content 01 10 01 11 11 00 10 10 address 0 1 2 3 4 5 6 7
[0209] Next, the content and content module are mapped, as shown in Table 16:
[0210] Table 16 Example of content and content module mapping
[0211] content Content Module 01 Content Module 1 10 Content Module 2 11 Content Module 3 00 Content Module 4
[0212] Only the module combinations corresponding to three of the content modules can be stored. In practice, the module combinations corresponding to the most frequent content can be represented by blanks. For example, in this example, only the module combinations corresponding to content module 1, content module 3, and content module 4 can be stored, and the module combination corresponding to content module 2 can be represented by blanks. Then there are module combinations as shown in Table 17:
[0213] Table 17 Module combination example Table 2
[0214] Serial number Content Module Address module Run Module 1 Content Module 1 Address module 0 Journey Module 1 2 Content Module 1 Address module 2 Journey Module 1 3 Content Module 3 Address module 3 Travel module 2 5 Content Module 4 Address module 5 Journey Module 1
[0215] S53: Similar to the above method, after encoding, the corresponding DNA molecule chain is obtained by synthesis or assembly.
[0216] It should be noted that the "data-address pair" in the first embodiment can be combined into modules using the "content-address-run" in the second embodiment, and an optimization algorithm can also be added.
[0217] Third embodiment
[0218] S4: Data decoding
[0219] Data decoding is the reverse process of encoding. Here we only explain the process and will not elaborate on it in detail. Sequence the DNA molecular chain to obtain the base sequence. Perform module matching based on the base sequence results to obtain module mapping encoding information. Further decode the module encoding information to obtain meta information and content information. When the content information contains redundant information, the redundant information can be used to perform error correction operations on the meta information and data information, so that separate fragment information can be obtained. After reading and obtaining a sufficient number (m≥k) of fragment information, the original binary data content can be reconstructed through the decoding operation of the erasure code to achieve data reading.
[0220] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0221] Fourth embodiment
[0222] A DNA data storage device based on coding optimization, comprising:
[0223] Precoding unit: obtains the data to be stored, and obtains the binary data to be encoded by correcting the precoding;
[0224] DNA module library: used to pre-store modules corresponding to the rules of size and address recoding. If a-bit b-base recoding is adopted, the module library is set to a*b modules, each group of modules corresponds to a bit of the recoded information, and different modules in each group represent the content of the bit;
[0225] Module mapping encoding unit: used to select an adapted module combination from the DNA module library through module mapping encoding, and convert the binary data into module encoding information required for DNA assembly;
[0226] DNA information writing unit: determines the corresponding module combination through the module encoding step, and connects it into the corresponding DNA molecular chain through enzyme catalysis reaction to complete the data storage.
[0227] The module mapping encoding unit further includes a meta information module mapping unit, a content information module mapping unit and a segment combining unit.
[0228] Meta information module mapping unit: used for dividing binary information into m bits to obtain a number of short messages, each short message has corresponding address information, and re-encoding the address information according to a-bit b-binary system; obtaining the "data-address pair" corresponding to each re-encoded short message, and performing module mapping on the "data-address pair" information using the content information DNA module library to obtain the module combination adapted to the "data-address pair" of each short message and the corresponding encoding information of each module;
[0229] The content information module mapping unit is used to use the meta information DNA module library to perform module mapping on each bit of each meta information, each group of modules corresponds to a bit of the re-encoded information, and different modules in each group represent the content of the bit, so as to obtain the module combination adapted to the meta information of each short message and the corresponding encoding information of each module therein.
[0230] Fragment combination unit: used to find the module combination of the "data-address pair" of the meta information and content information of each fragment of the binary data and the corresponding encoding information of each module, and form them into the fragment combination unit corresponding to the fragment; each module of the fragment combination unit is a DNA molecular chain containing a specific base sequence, the modules of the meta information are connected in sequence, the terminal module of the meta information is connected to the first module of the content information, and the modules of the content information are connected in sequence; the modules between the fragment combination units are connected according to the pre-agreed sequence matching between the modules: the first module of the rear-end fragment combination unit is connected to the terminal module of the current fragment combination unit, or the terminal module of the current fragment combination unit is the tail DNA module of the current DNA molecular chain, and the first module of the rear-end fragment combination unit is independently the first DNA module of a DNA molecular chain, thereby forming one or more DNA molecular chains for centralized storage.
[0231] Fifth embodiment
[0232] A DNA data reading device based on coding optimization, comprising:
[0233] Module recognition unit: performs module recognition on DNA molecular chains;
[0234] Decoding unit: used to decode module mapping information according to module information to obtain corresponding binary data;
[0235] Decoding unit: used to reconstruct the original data content according to the pre-encoding decoding operation of binary data to realize data reading.
[0236] The module recognition unit may be to sequence the DNA molecule chain to obtain the base sequence; the module mapping encoding unit is to perform module matching according to the base sequence result and the pre-set module library.
[0237] The structure of a terminal device provided in an embodiment of the present application. The terminal device of this embodiment includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and when the processor executes the computer program, the steps in any of the above-mentioned method embodiments are implemented.
[0238] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the terminal may have more or fewer components, or a combination of certain components, or different components, for example, it may also include input and output devices, network access devices, etc. The processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0239] In some embodiments, the memory may be an internal storage unit of the terminal device, such as a hard disk or memory of the terminal device. In other embodiments, the memory may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal device. Further, the memory may also include both an internal storage unit of the terminal device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory may also be used to temporarily store data that has been output or is to be output.
[0240] The present application also provides a computer-readable storage medium storing a computer program, which can implement the steps in the above-mentioned method embodiments when the computer program is executed by a processor. The present application also provides a computer program product, which can implement the steps in the above-mentioned method embodiments when the computer program product is executed on a mobile terminal.
[0241] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0242] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0243] Sixth embodiment
[0244] Basically similar to the first and second examples, the DNA module is not only deoxyribonucleic acid (DNA), but can also be expanded to molecular modules, including ribonucleic acid (RNA), peptides, organic polymers, organic small molecules, carbon nanomaterials, inorganic substances, etc. When storing information, it involves the combination of different modules representing content coding and address coding. These modules can be combined together in the form of covalent bonds, ionic bonds, hydrogen bonds, intermolecular forces, hydrophobic forces, base complementary pairing, etc.
[0245] In a fifth aspect, a molecular module data access method based on coding optimization is provided, wherein the storage process further comprises:
[0246] S1: Obtain data to be stored, and obtain binary data to be encoded by pre-encoding;
[0247] S2: selecting the adapted modules from the preset molecular module library through module mapping coding, and determining the module sequence to obtain the corresponding module combination;
[0248] S3: connecting the modules in the module combination into corresponding molecular modules to complete the data storage;
[0249] The reading process further includes:
[0250] S4: module identification of molecular modules;
[0251] S5: Decode the module mapping information according to the module information to obtain corresponding binary data;
[0252] S6: Reconstruct the original data content according to the pre-encoding decoding operation of the binary data to realize data reading.
[0253] That is, coding optimization can be applied to larger molecular modules to achieve the same function. Similarly, there are many schemes for module identification. Obtaining the corresponding base sequence through DNA sequencing and then obtaining module information through base sequence matching is only one implementation scheme.
[0254] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0255] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0256] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0257] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A DNA data storage method based on coding optimization, characterized in that: include: S1: Obtain binary data of data to be stored, and pre-encode the binary data according to the optimization target of DNA storage coding, wherein the optimization target of DNA storage coding includes that the total number of modules consumed for information expression meets a preset condition; Wherein, pre-coding the binary data according to the optimization target of DNA storage coding includes: Dividing the binary data into a plurality of segments; Meta information and content information are respectively set in each segment, and binary data in the segment is divided into data-address pairs by a preset interval distance, wherein the data-address pairs include a content module and an address module; an optimization calculation module is set, and the content module, the address module and the optimization calculation module are expressed as a module combination form of the content information according to a set optimization rule, further comprising: Setting the optimization calculation module further includes setting a run character quantity module representing the number of times a character appears; The content module, the address module and the optimization calculation module are expressed as a module combination according to the set optimization rule, further comprising: optimizing according to the stored content module, using a combination of the address module and the run module to express the module combination of the content information; S2: After obtaining a set of DNA molecular chains represented by a combination of modules through module mapping coding, the corresponding DNA molecular chains are obtained through assembly or synthesis.
2. The method according to claim 1, characterized in that Pre-coding the binary data according to the optimization goal of DNA storage coding further includes: Dividing the binary data into a plurality of segments; The N binary data in each segment are converted into a symbol sequence according to a preset interval distance, and the probability information of each symbol in the sequence is counted. According to the optimization goal of minimizing the total module consumption, a precoding rule for converting the symbol encoding into new symbol information is set according to the probability information, and the binary data in the segment are re-encoded according to the precoding rule, where N is the number of binary numbers in the segment.
3. The method according to claim 2, characterized in that Setting a precoding rule for converting the symbol code into new symbol information according to the probability information further includes: The optimization goal is to minimize the total number of bit 1s in the recoded binary sequence, and the set precoding rule is to replace characters with high probability with symbols marked with more bit 0s, and replace characters with low probability with symbols marked with more bit 1s.
4. The method according to claim 2, characterized in that Setting a precoding rule for converting the symbol code into new symbol information according to the probability information further includes: The symbol with the highest probability is replaced with 00…0; the remaining symbols are assigned 1, 01, 001, … in descending order of probability, and a variable-length recoding rule is used to ensure that each character after replacement has at most one 1 bit.
5. The method according to claim 1, characterized in that Pre-coding the binary data according to the optimization goal of DNA storage coding further includes: Dividing the binary data into a plurality of segments; The N binary data in each fragment use the bit flipping information to extract the position where the flipping occurs, and obtain a new sequence with the same length as the N binary data. In the new sequence, bit 1 is assigned to the position where each flipping occurs. The first bit information of the new sequence is consistent with the first bit information of the original data, and the binary data in the fragment is re-encoded, where N is the number of binary numbers in the fragment.
6. The method according to claim 2, characterized in that Pre-coding the binary data according to the optimization goal of DNA storage coding further includes: Dividing the binary data into a plurality of segments; In each segment, according to the pre-set run information, N binary numbers are converted into a character sequence represented by the run length, the probability information of each run symbol in the sequence is counted, and a recoding rule including long recoding and variable-length recoding is set for the run symbol according to the probability information; the N binary data are recoded according to the pre-coding rule, where N is the number of binary numbers in the segment.
7. The method according to claim 1, characterized in that After obtaining the set of DNA molecular chains represented by the module combination form, obtaining the corresponding DNA molecular chain by assembling means further includes: Selecting suitable modules from a preset DNA module library and determining the order of the modules to obtain a corresponding module combination; The modules in the module combination are connected into corresponding DNA molecular chains to complete the data storage.
8. The method according to claim 1, characterized in that Through the module mapping encoding, the DNA molecular chain set represented by the module combination form further includes: Setting meta information and content information for each pre-divided segment, and the meta information further includes at least one piece of information including an interval distance and a precoding rule; The meta information and the content information are mapped and encoded respectively using a meta information DNA module library and a content information DNA module library, and the meta information DNA module library and the content information DNA module library correspond to the rules of size and address recoding.
9. The method according to claim 8, characterized in that Using the content information DNA module library to map and encode the content information further includes: Divide the binary information into m bits to obtain several short messages, each of which has corresponding address information; Re-encoding the address information according to a-bit b-binary system; Get the data-address pair corresponding to each recoded short message. When m=1, save the data-address pair with data 1 or data 0; when m>1, save any (2 m -1) Data-address pair of the case; The data-address pair information is subjected to module mapping using the content information DNA module library to obtain a module combination adapted to the data-address pair of each short message and corresponding encoding information of each module.
10. The method according to claim 8 or 9, characterized in that The meta-information mapping is encoded using a meta-information DNA module library, which further comprises: Each piece of meta-information corresponds to a module combination. The meta-information DNA module library is used to perform module mapping on each bit of information of each piece of meta-information. Each group of modules corresponds to a bit of the re-encoded information. Different modules in each group represent the content of the bit, so as to obtain the module combination adapted to the meta-information of each short message and the corresponding encoding information of each module therein.
11. The method according to claim 8, characterized in that Through the module mapping encoding, the DNA molecular chain set represented by the module combination form further includes: The meta information and content information of each segment are mapped and encoded to form the current segment combination unit. The order of the corresponding fragment combination units is obtained according to the order of the fragments, thereby obtaining one or more module combinations of the binary data and the order between the modules in each module combination.
12. The method according to claim 8, characterized in that Obtaining the corresponding DNA molecule chain by assembly or synthesis further includes: Find the module combination of the data-address pair of the meta information and content information of each fragment of the binary data and the corresponding encoding information of each module, and compose them into a fragment combination unit corresponding to the fragment; Each module of the fragment assembly unit is a DNA module containing a specific base sequence. According to the pre-agreed module order, the corresponding DNA modules are connected into corresponding DNA molecular chains through enzyme catalysis reaction, and the DNA molecular chains are one or more.
13. The method according to claim 12, characterized in that According to the pre-agreed module sequence, connecting the corresponding DNA modules into corresponding DNA molecular chains through enzyme catalysis reaction further includes: The terminal module of the meta information of the segment combination unit is connected to the first module of the content information, and the modules of the content information in the same segment combination unit are connected in sequence; The modules between the fragment assembly units are connected by matching in a pre-agreed order: the first module of the rear-end fragment assembly unit is connected to the last module of the current fragment assembly unit, or the last module of the current fragment assembly unit is the tail DNA module of the current DNA molecular chain, and the first module of the rear-end fragment assembly unit is independently the first DNA module of a DNA molecular chain, thereby forming one or more DNA molecular chains for centralized storage.
14. The method according to claim 1, wherein: Through the module mapping encoding, the DNA molecular chain set represented by the module combination form further includes: Meta information is set for each pre-divided segment, which further includes at least one information including an interval distance, a recoding rule, and an optimization rule; The meta information and the content information are mapped and encoded respectively using a meta information DNA module library and a content information DNA module library, and the meta information DNA module library and the content information DNA module library are in regular correspondence of size and address.
15. The method according to claim 1, wherein: Setting an optimization calculation module, and expressing the content module, the address module and the optimization calculation module into a module combination form according to the set optimization rules further includes: Setting the optimization calculation module further includes setting a run character quantity module representing the number of characters that appear; The module combination form of representing the content information by the content module, the address module and the optimization calculation module according to the set optimization rules further includes: connecting multiple run modules after a single data-address pair to directly represent a longer set of data through a longer DNA molecular chain.
16. The method according to claim 1, wherein: Obtaining the corresponding DNA molecule chain by synthesis further includes: The corresponding module combination is mapped to the corresponding base sequence according to the pre-designed DNA module, converted into DNA molecule long chain information, and further directly synthesized into the molecule long chain.
17. A DNA data reading method based on coding optimization, characterized in that: include: S1: obtaining binary data of data to be stored, and pre-encoding the binary data according to the optimization target of DNA storage coding, further comprising: the optimization target of DNA storage coding includes that the total number of modules consumed for information expression meets a preset condition; Wherein, pre-coding the binary data according to the optimization target of DNA storage coding includes: Dividing the binary data into a plurality of segments; Meta information and content information are respectively set in each segment, and binary data in the segment is divided into data-address pairs by a preset interval distance, wherein the data-address pairs include a content module and an address module; an optimization calculation module is set, and the content module, the address module and the optimization calculation module are expressed as a module combination form of the content information according to a set optimization rule, further comprising: Setting the optimization calculation module further includes setting a run character quantity module representing the number of times a character appears; The content module, the address module and the optimization calculation module are expressed as a module combination according to the set optimization rule, further comprising: optimizing according to the stored content module, using a combination of the address module and the run module to express the module combination of the content information; S2: After obtaining a set of DNA molecular chains represented by a module combination through module mapping coding, the corresponding DNA molecular chains are obtained through assembly or synthesis; Perform module recognition on DNA molecular chains; Decode the module mapping information according to the module information to obtain the corresponding binary data; The original data content is reconstructed by decoding the binary data according to the decoding operation corresponding to the pre-coding rules to realize data reading.
Citation Information
Patent Citations
Deoxyribose nucleic acid (DNA) data storage method and reading method based on erasure code and assembly technology and terminal
CN116564424A
Method, apparatus and system for storing information in molecule
CN118412026A
Video compression using analytical entropy coding
WO2003067763A2