DNA-based data classified storage method and application
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
Smart Images

Figure CN122090936A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of DNA data storage technology, and in particular to a DNA-based data classification and storage method and application. Background Technology
[0002] DNA storage technology is an information storage method that leverages the high density, long-term stability, and enormous information capacity of DNA molecules to address the growing challenges of digital data. In DNA data storage, digital information is first encoded into DNA sequences, which are then stored and retrieved using chemical synthesis and high-throughput sequencing technologies. This technology not only far surpasses traditional storage media in information density but also enables long-term data preservation, thanks to the stability of DNA molecules under suitable conditions, which can last for thousands of years.
[0003] In related technologies, data is typically categorized into three types based on its access frequency and importance: "cold data," "hot data," and "warm data." "Hot data" refers to data that is frequently accessed and has high requirements for business immediacy; it usually requires real-time processing and access, such as transaction data and real-time analytics data. "Warm data" falls between hot and cold data; although accessed less frequently, it still requires relatively timely responses, such as some semi-real-time reports or historical data. "Cold data" refers to data that is rarely accessed but needs long-term storage, such as archived data and backup data; it typically does not have high requirements for access speed but requires reliable storage and low-cost solutions. Different data types require different storage strategies to meet their respective performance, cost, and reliability requirements.
[0004] Current DNA storage technologies typically target specific data layers (such as focusing on "cold" data layers) and lack methods for classifying and managing data stored in DNA. Existing methods for expanding the dimensionality of information stored in DNA usually involve DNA assembly, such as assembling artificial chromosomes to introduce higher-dimensional information, or introducing base modifications to the DNA sequence to store additional information. However, these methods only effectively demonstrate DNA's ability to store multi-dimensional information; they fail to classify and manage data of different dimensions according to different levels of data classification. This lack of ability to classify and manage data at different levels, as opposed to existing information storage systems, limits the development of this technology.
[0005] Therefore, there is an urgent need for a DNA-based data classification and storage method to achieve data classification and storage, and improve the flexibility of information storage and retrieval. Summary of the Invention
[0006] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a DNA-based data classification and storage method and its application. The core of this DNA-based data classification and storage method is to divide DNA into multi-dimensional storage spaces, and then map these multi-dimensional storage spaces to different levels of data in a data classification system according to their characteristics, thereby achieving data classification and storage. Employing this DNA-based data classification and storage method enables efficient data management and helps improve the flexibility of data storage and retrieval.
[0007] A first aspect of the present invention provides a DNA-based data classification and storage method, comprising:
[0008] The information to be stored is divided into one-dimensional information, two-dimensional information, and three-dimensional information according to the data access frequency. The one-dimensional information is low read / write frequency information, the two-dimensional information is high read / write frequency information, and the three-dimensional information is medium read / write frequency information.
[0009] Based on the principle of binary encoding, the one-dimensional information is converted into a nucleic acid sequence and fixed on a first storage medium containing a flexible substrate for storage.
[0010] Based on the high-low bit encoding principle, the two-dimensional information is converted into two-dimensional dot matrix information, and then the two-dimensional dot matrix information is written into a second storage carrier containing a flexible substrate using DNA as raw material for storage.
[0011] Based on the voxel encoding principle or pixel encoding principle, the three-dimensional information is converted into dot matrix information. At the same time, according to the preset mapping relationship, the voxels or pixels are converted into different base types and arranged in order according to the preset direction to obtain a nucleic acid sequence. The nucleic acid sequence is then fixed on a third storage carrier containing a flexible substrate for storage.
[0012] The DNA-based data classification and storage method according to embodiments of the present invention has at least the following beneficial effects: The present invention utilizes the combination of information-carrying DNA sequences and micro / nano structures to achieve multi-level data management and hierarchical storage in a DNA storage system. For DNA data with high access frequency, the DNA-based data classification and storage method of the present invention enables rapid reading, greatly improving the reading efficiency of DNA data; for DNA data with low access frequency, it enables high-density storage and, under appropriate conditions, can be preserved for thousands of years, greatly improving the flexibility of DNA data storage and retrieval.
[0013] In some embodiments of the present invention, converting the one-dimensional information into a nucleic acid sequence based on the binary encoding principle includes:
[0014] The one-dimensional information is converted into a binary sequence using an encoding algorithm, and the binary sequence is converted into a nucleic acid fragment according to a preset mapping relationship.
[0015] In some embodiments of the present invention, the bases of the nucleic acid fragment are selected from natural bases and / or non-natural bases.
[0016] In some embodiments of the present invention, the natural bases include adenine (A), guanine (G), cytosine (C), and thymine (T); the non-natural bases include at least one of the following: methylcytosine (mC), methyladenine (mA), methylguanine (mG), methylthymine (mT), hydroxymethylcytosine (hmC), carboxycytosine (caC), formylcytosine (fC), isocytosine (isoC), inosine (I), nitroindole, nitropyrrole, artificial base Z, artificial base P, artificial base S, and artificial base B.
[0017] In some embodiments of the present invention, the preset mapping relationship can be: 00 is mapped to A, 01 is mapped to T, 11 is mapped to C, and 10 is mapped to G;
[0018] Alternatively, it can be a combination of different binary codes and base codes, such as 11 mapping to A, 10 mapping to T, 01 mapping to C, and 00 mapping to G.
[0019] In some embodiments of the present invention, the high-low coding principle includes using two states, namely the presence or absence of DNA, to represent two-dimensional information.
[0020] In some embodiments of the present invention, the method of writing the two-dimensional dot matrix information using DNA as raw material includes micro-nano lithography, electrochemical writing, or inkjet printing.
[0021] In some embodiments of the present invention, the conversion of the three-dimensional information into raster information based on the voxel encoding principle includes:
[0022] The three-dimensional information is sliced according to a preset direction, and the slices are mapped to a voxel arrangement to obtain two-dimensional dot matrix information containing the slice information.
[0023] In some embodiments of the present invention, the preset direction can be any specified direction.
[0024] In some preferred embodiments of the present invention, the preset direction includes a preset X-axis, Y-axis, or Z-axis direction.
[0025] In some embodiments of the present invention, the conversion of the three-dimensional information into dot matrix information based on the pixel encoding principle includes:
[0026] The three-dimensional information is sliced according to a preset direction, and the slices are mapped to a voxel arrangement to obtain two-dimensional dot matrix information containing the slice information.
[0027] In some embodiments of the present invention, the preset direction includes a preset time axis direction.
[0028] In some embodiments of the present invention, the information to be stored includes at least one of images, text, programs, audio, video, and three-dimensional objects.
[0029] In some embodiments of the present invention, the first storage carrier containing a flexible substrate, the second storage carrier containing a flexible substrate, and the third storage carrier containing a flexible substrate may be the same or different.
[0030] In some embodiments of the present invention, the first storage carrier containing a flexible substrate, the second storage carrier containing a flexible substrate, and the third storage carrier containing a flexible substrate are each independently a microfluidic chip containing a flexible substrate.
[0031] In some embodiments of the present invention, the flexible substrate is independently selected from any one of nylon film, modified nylon film, flexible gold film, and flexible platinum film.
[0032] In some embodiments of the present invention, the modified nylon film includes a nylon film modified with chemical groups.
[0033] In some embodiments of the present invention, the chemical group modification includes carboxyl group modification and / or aldehyde group modification.
[0034] A second aspect of the present invention provides a DNA-based classification data reading method, comprising:
[0035] Obtain a storage medium carrying the one-dimensional, two-dimensional, or three-dimensional information;
[0036] When reading the one-dimensional information, the corresponding nucleic acid sequence in the storage vector is read using in situ sequencing, and the one-dimensional information is decoded to obtain the one-dimensional information.
[0037] When reading the two-dimensional information, a reading method corresponding to the storage method of the two-dimensional information is determined according to the storage method of the two-dimensional information, and the reading method is used to read the corresponding two-dimensional dot matrix information in the storage carrier and decode to obtain the two-dimensional information;
[0038] When reading the two-dimensional information, a reading method corresponding to the three-dimensional information storage method is determined according to the three-dimensional information storage method. The reading method is used to read the dot matrix information of the corresponding slice in the storage carrier, and then the slices are arranged in a preset direction order to obtain spatial information. After decoding, the three-dimensional information is obtained.
[0039] The multi-level DNA data reading method according to embodiments of the present invention has at least the following beneficial effects: the reading method of the present invention can read corresponding data quickly and in a targeted manner, thereby improving information reading efficiency.
[0040] In some embodiments of the present invention, when reading the two-dimensional information and / or the three-dimensional information, the reading method includes optical microscopy reading, scanning electron microscopy reading, atomic force microscopy reading, or super-resolution microscopy reading.
[0041] In some embodiments of the present invention, when the writing method is micro-nano lithography, the reading method is scanning electron microscopy or atomic force microscopy.
[0042] In some embodiments of the present invention, when the writing method is electrochemical writing or inkjet printing, the reading method is fluorescence imaging reading or optical microscopy reading.
[0043] A third aspect of the present invention provides a storage system including a storage unit and a retrieval unit, wherein the storage unit operates the DNA-based data classification and storage method described in any of the first aspects.
[0044] In some embodiments of the present invention, the storage unit is provided with three storage sub-units, wherein the first storage sub-unit is used to store the one-dimensional information, the second storage sub-unit is used to store the two-dimensional information, and the third storage sub-unit is used to store the three-dimensional information.
[0045] In some embodiments of the present invention, the reading unit operates a DNA-based data reading method as described in any of the second aspects.
[0046] In some embodiments of the present invention, the reading unit further includes a recognizer, which identifies the index pattern on the DNA storage medium, thereby controlling a specific area on the DNA storage medium to react with the reaction solution to complete operations such as writing, reading, and erasing.
[0047] A fourth aspect of the present invention provides the application of the DNA-based data classification and storage method as described in any of the first aspects and / or the DNA-based classification data reading method as described in any of the second aspects in DNA data classification management or storage. Attached Figure Description
[0048] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:
[0049] Figure 1 A schematic diagram illustrating the principle of using carboxyl-modified nylon porous membranes to store DNA data;
[0050] Figure 2 The results validate the feasibility of using carboxyl-modified nylon porous membranes to store DNA data.
[0051] Figure 3 The statistical results of DNA loading on nylon porous membranes before and after activation treatment;
[0052] Figure 4 This is a flowchart illustrating the fabrication process of the PDMS patterned nylon microfluidic chip of the present invention.
[0053] Figure 5 This is the functional verification result of the PDMS patterned nylon microfluidic chip of the present invention;
[0054] Figure 6 This is a schematic diagram illustrating the conversion of a four-color Mario pattern based on the high-low bit encoding principle of the present invention;
[0055] Figure 7 This is a schematic diagram illustrating the principle of converting four-stranded helical proteins based on the voxel encoding principle of the present invention;
[0056] Figure 8 This is a schematic diagram illustrating the spatial information of a four-stranded helical protein obtained through in situ sequencing according to the present invention.
[0057] Figure 9 This is a flowchart illustrating in situ sequencing using a PDMS patterned nylon microfluidic chip, as described in this invention.
[0058] Figure 10 This is a schematic diagram illustrating the mapping relationship between bases and fluorescence signals during in-situ sequencing using a PDMS-patterned nylon microfluidic chip, as described in this invention.
[0059] Figure 11 The results of in-situ sequencing using a PDMS patterned nylon microfluidic chip are presented in this invention.
[0060] Figure 12 This is a schematic diagram of the Mario GIF image slice conversion of the present invention;
[0061] Figure 13 This is a schematic diagram illustrating the DNA information encryption principle based on three-dimensional information storage of the present invention;
[0062] Figure 14 This is a schematic diagram of the encrypted DNA information decoding based on three-dimensional information storage according to the present invention. Detailed Implementation
[0063] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.
[0064] The terms "preferred," "more preferably," etc., used in this invention refer to embodiments of the invention that provide certain beneficial effects under certain circumstances. However, other embodiments may also be preferred under the same or other circumstances. Furthermore, the description of one or more preferred embodiments does not imply that other embodiments are unavailable, nor is it intended to exclude other embodiments from the scope of this invention.
[0065] In the description of this invention, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0066] In the description of this invention, high read / write frequency information ("hot" data) refers to data accessed frequently. This data typically requires fast and efficient access and processing, and therefore needs to be stored on high-performance, low-latency storage devices to meet the needs of rapid reading. Medium read / write frequency information ("warm" data) refers to data with moderate access frequency and a certain degree of importance. This data does not require the rapid access and processing of "hot" data, but still needs to be reliably stored and accessed within a certain period of time. Low read / write frequency information ("cold" data) refers to data accessed less frequently. This data typically needs to be stored for a long time, but does not require frequent access and processing, and is therefore more suitable for storage on lower-cost, higher-capacity storage devices. Specifically, the criteria for classifying data access frequency in this invention can be flexibly adjusted according to specific business needs. Typically, an access frequency threshold is set by monitoring the access frequency. For example, the number of accesses within a time window can be set: if the access frequency of data within this time window exceeds a certain value, it is defined as "hot" data; if the access frequency is below the threshold but still regular, it is "warm" data; data that is rarely or almost never accessed is "cold" data. Subsequently, the entire data is converted into DNA data for storage.
[0067] Unless otherwise specified in the examples, the procedures should be performed under standard conditions or conditions recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.
[0068] This invention proposes a DNA-based data classification and storage method. Its core principle is to divide DNA information into multi-dimensional storage spaces based on DNA data access frequency. These multi-dimensional storage spaces are then mapped to different levels of data in a data classification system according to their characteristics, thereby achieving data classification and storage. Specifically, this application's scheme divides the information to be stored into one-dimensional, two-dimensional, and three-dimensional information based on data access frequency. One-dimensional information represents low read / write frequency information, two-dimensional information represents high read / write frequency information, and three-dimensional information represents medium read / write frequency information. The classification and storage methods correspond to one-dimensional information storage, two-dimensional information storage, and three-dimensional information storage, respectively.
[0069] 1. One-dimensional information storage:
[0070] The core of one-dimensional information storage is to utilize high-surface-area materials to achieve high-density loading of DNA information, corresponding to "cold" data storage in the data hierarchy. Its information storage process includes: first, using an encoding algorithm to convert the data to be stored into a binary sequence, and then, according to a pre-defined mapping relationship, converting the binary sequence into nucleic acid fragments; then, fixing the nucleic acid fragments onto a storage carrier containing a flexible substrate, thereby achieving one-dimensional information storage.
[0071] In some specific implementations, the data to be stored may include one or more of the data information that can exist on the computer, such as images, text, programs, audio, and video.
[0072] In some specific implementations, when obtaining the binary sequence corresponding to the data to be stored, the encoding information corresponding to the data to be stored can be obtained, and the corresponding encoding information can be converted into binary encoding information to obtain the corresponding binary sequence. For example, the text in the text information can be converted into the corresponding ASCII (American Standard Code for Information Interchange) encoding or UNICODE (Universal Character Set) encoding, and then the encoding information can be converted into a binary sequence. In some specific implementations, the binary sequence is converted into a base sequence according to a preset mapping relationship. Since the bases in DNA include four types of natural bases: adenine (A), guanine (G), cytosine (C), and thymine (T), the preset mapping relationship can be a binary-to-quaternary mapping relationship.
[0073] In some specific implementations, the mapping relationship can be: 00 maps to A, 01 maps to T, 11 maps to C, and 10 maps to G. This mapping relationship can also be a combination of different binary codes and base codes, for example, 11 maps to A, 10 maps to T, 01 maps to C, and 00 maps to G.
[0074] In some specific embodiments, the bases in the DNA may also be non-natural bases. Non-natural bases refer to optional bases or base analogues other than the four bases A, G, C, and T mentioned above. In some embodiments, non-natural bases include at least one of the following: methylcytosine (mC), methyladenine (mA), methylguanine (mG), methylthymidine (mT), hydroxymethylcytosine (hmC), carboxycytosine (caC), formylcytosine (fC), isocytosine (isoC), inosine (I), nitroindole, nitropyrrole, artificial base Z, artificial base P, artificial base S, and artificial base B. In some embodiments, non-natural bases include at least one of m6A, m5C, hm5C, f5C, ca5C, 4-nitroindole, 5-nitroindole, 6-nitroindole, and 3-nitropyrrole.
[0075] In some specific implementations, if the data to be stored is large, the file containing the data can be split into multiple sub-files, and the order of the sub-base sequences corresponding to each sub-file can be recorded using an index sequence. After splitting the data or file to be stored, the resulting sub-files can be recorded using an index sequence.
[0076] In some specific embodiments, to immobilize nucleic acid fragments, the flexible substrate is preferably a flexible substrate capable of binding to nucleic acid chemical bonds, such as nylon membranes, modified nylon membranes (e.g., carboxyl groups), flexible gold membranes, flexible platinum membranes, and other flexible substrate materials that can covalently bind to nucleic acid fragments. In some preferred embodiments, the flexible substrate can obtain more binding sites through activation, for example, by activation with ethyldimethylaminopropylcarbodiimide (EDAC) or EDC / NHS, thereby enabling the binding of more nucleic acid fragments within the same surface area, and further enabling the storage of more nucleic acid fragments.
[0077] In some specific embodiments, nucleic acid fragments are chemically bonded to a flexible substrate on the storage carrier. The nucleic acid fragments can be chemically modified nucleic acid molecules and / or unmodified nucleic acid molecules. Unmodified nucleic acid molecules directly bind to the flexible substrate via chemical bonds, while modified nucleic acid molecules can bind to the flexible substrate via modified groups or atoms. Further, the chemical bonding methods include, but are not limited to, at least one of amino-carboxyl group bonding, metal-thiol group bonding, amino-aldehyde group bonding, and coordination bonds between metal and nucleic acid. In some preferred embodiments, the nucleic acid fragments are immobilized by covalent bonds formed between them and the flexible substrate through amino, carboxyl, or other groups. Preferably, a solution containing nucleic acid fragments is spotted onto the flexible substrate, and under certain conditions, a reaction occurs to covalently bind the flexible substrate and the nucleic acid fragments. After the reaction, the flexible substrate can be washed with a typical washing solution composed of SSC (Saline Sodium Citrate) or SSPE (Saline Sodium Phosphate EDTA) and SDS (Sodium Dodecyl Sulfate), or other washing solutions known in the art capable of washing nucleic acid fragments, to remove unbound nucleic acid fragments.
[0078] In some specific implementations, the nucleic acid fragment can be an oligonucleotide sequence or a double-stranded nucleic acid fragment.
[0079] In some specific embodiments, when the nucleic acid fragment is an oligonucleotide sequence, it is a linear polynucleotide fragment consisting of 2 to 300 nucleotide residues linked by phosphodiester bonds. Preferably, the number of nucleotide residues in the oligonucleotide is 2 to 250, 2 to 200, 2 to 160, 2 to 130, or 2 to 100. Preferably, the oligonucleotide sequence can be a chemically modified nucleic acid molecule and / or an unmodified nucleic acid molecule. Unmodified nucleic acid molecules directly bind to the flexible substrate of the storage carrier through chemical bonds, while modified nucleic acid molecules can bind to the strip substrate through modified groups or atoms.
[0080] In some specific embodiments, when the nucleic acid fragment is a double-stranded nucleic acid fragment, it includes the original nucleic acid fragment to be stored and its complementary fragment. The purpose of storing the nucleic acid fragment is to store the nucleic acid fragment itself or the nucleic acid file represented by the nucleic acid fragment, for example, encoding the original stored information into a nucleic acid sequence through a specific method. In some embodiments, during one-dimensional information reading, the nucleic acid fragment to be read can be reacted with a denaturing solution to cause the nucleic acid fragment to unwind and be released into the denaturing solution, the denaturing solution is recovered, and sequencing is performed. In some specific embodiments, the denaturing solution is an alkaline solution, specifically including but not limited to NaOH, KOH, etc. Nucleic acids can be denatured by acid, alkali, heat, etc., but thermal denaturation needs to be carried out at low nucleic acid concentration and low salt concentration, while strong acid will degrade nucleic acid and cause information loss. Therefore, it is preferable to use an alkaline solution to denature it during reading to avoid nucleic acid degradation. In some embodiments, the sequencing method includes mixing the recovered denaturing solution with dNTPs, buffer components, primers, and polymerase, using the complementary fragment in the denaturing solution as a template to reconstruct double strands before sequencing. In some implementations, after sequencing is complete, the base sequence in the sequencing results is re-decoded into the original storage file for reading.
[0081] 2. Two-dimensional information storage:
[0082] The core of two-dimensional information storage is to integrate DNA with nanostructures, transforming DNA libraries, which are traditionally stored with disordered DNA strands, into ordered ones. Information is stored and transmitted using the two-dimensional matrix formed by the ordered arrangement of DNA strands.
[0083] In some specific embodiments, the two-dimensional information storage process includes: first, converting the data to be stored into two-dimensional dot matrix information (array code) based on the high-low bit encoding principle; then, writing the two-dimensional dot matrix information onto a flexible substrate of the storage carrier using DNA as the raw material, thereby realizing data storage. In some preferred embodiments, the method for writing two-dimensional dot matrix information using DNA as the raw material includes micro-nano lithography, electrochemical writing, or inkjet printing. When the writing method is inkjet printing, the corresponding reading method can be optical microscopy; when the writing method is micro-nano lithography, the reading method is determined to be scanning electron microscopy or atomic force microscopy, that is, observing the arrangement of the dot matrix in the storage carrier using scanning electron microscopy or atomic force microscopy to obtain the two-dimensional dot matrix information.
[0084] In some specific implementations, the high- and low-bit encoding principle is based on storing or transmitting information in two different states of DNA. Taking the conversion of a four-color image into a two-color image for storage as an example, the presence of DNA is represented by black, and its absence by gray. In the high-bit encoding scheme, red pixels in the original four-color image are replaced with black, and orange pixels are replaced with gray. Conversely, in the low-bit encoding scheme, red pixels are replaced with gray, and orange pixels are replaced with black. During image retrieval, if both the high- and low-bit bits of a pixel are gray, it is interpreted as gray; if the high-bit is black and the low-bit is gray, it is red; if the high-bit is gray and the low-bit is black, it is orange; and if both the high- and low-bit bits are black, it is black.
[0085] In some specific implementations, the data to be stored may include one or more of the data information that can exist on the computer, such as images, text, programs, audio, and video.
[0086] In some specific embodiments, the storage medium includes a microfluidic chip; specifically, it can be a PDMS patterned nylon microfluidic chip.
[0087] In some specific implementations, writing two-dimensional dot matrix information using DNA as raw material includes writing it using DNA as printing ink via inkjet printing.
[0088] In some specific implementations, the reading of two-dimensional information only requires a temperature cycle through which a template primer is paired with DNA in the printing ink. The reading speed of the two-dimensional pattern information is only three minutes, which is suitable for thermal data storage.
[0089] 3. Three-dimensional information storage:
[0090] The core of three-dimensional information storage is to introduce voxel encoding on the basis of one-dimensional sequence encoding and two-dimensional pixel encoding. Voxel-encoded data can be read out through in-situ sequencing on a PDMS patterned nylon microfluidic chip, which is suitable for "warm" data storage.
[0091] In some specific implementations, taking the storage of three-dimensional object information as an example, the information storage process includes: first, the data to be stored (such as a three-dimensional object) is sliced along a preset direction, and the slices are mapped to voxel array codes. Then, the slices are arranged in order according to a specific axis (such as the z-axis). By mapping the voxels on the slices to base types according to certain rules (such as mapping gray values to specific bases), the coding information on each slice is obtained in this way. Then, the slices are arranged in order according to the preset specific axis direction to obtain the nucleic acid sequence containing the information of the three-dimensional object to be stored. Finally, the sequence is fixed on a flexible substrate of the storage carrier to realize the storage of three-dimensional information.
[0092] In some specific implementations, taking video information storage as an example, the information storage process includes: first, the video to be stored is sliced along the time axis, and the slices are mapped to a two-dimensional pixel arrangement. Then, the slices are arranged in order according to a specific axis (such as the time axis). By mapping the pixels on the slices to base types according to certain rules, the encoding information on each slice is obtained in this way. Then, the slices are arranged in order according to the preset time axis direction to obtain the nucleic acid sequence containing the video information. Finally, the sequence is fixed on a flexible substrate of the storage carrier to achieve three-dimensional information storage.
[0093] In some specific implementations, the overall algorithm complexity is high when storing three-dimensional information due to the simultaneous dual encoding of spatial information of the sequence. A simpler approach is to align the sequence containing the spatial information of the three-dimensional object to be stored with a large DNA database, select DNA strands with identical sequences, store them on a two-dimensional matrix, and then recover the data through in-situ sequencing. This method also allows for DNA information encryption and more diverse information transmission methods.
[0094] This invention also proposes a DNA-based classification data reading method, comprising:
[0095] Obtain a storage medium carrying the one-dimensional, two-dimensional, or three-dimensional information;
[0096] When reading the one-dimensional information, the corresponding nucleic acid sequence in the storage vector is read by in situ sequencing and decoded to obtain the one-dimensional information.
[0097] When reading the two-dimensional information, a reading method corresponding to the two-dimensional information storage method is determined according to the two-dimensional information storage method. The reading method is used to read the corresponding two-dimensional dot matrix information in the storage carrier and decode to obtain the two-dimensional information.
[0098] When reading the two-dimensional information, a reading method corresponding to the three-dimensional information storage method is determined according to the three-dimensional information storage method. The reading method is used to read the dot matrix information of the corresponding slice in the storage carrier, and then the slices are arranged in a preset direction order to obtain spatial information. After decoding, the three-dimensional information is obtained.
[0099] In some specific implementations, the reading methods for reading two-dimensional and / or three-dimensional information include optical microscopy, scanning electron microscopy, atomic force microscopy, or super-resolution microscopy.
[0100] The present invention also proposes a storage system, including a storage unit and a retrieval unit, wherein the storage unit runs the aforementioned DNA-based data classification and storage method.
[0101] In some specific implementations, the storage unit is provided with three storage sub-units, wherein the first storage sub-unit is used for storing one-dimensional information, the second storage sub-unit is used for storing two-dimensional information, and the third storage sub-unit is used for storing three-dimensional information.
[0102] In some specific implementations, the reading unit performs the aforementioned DNA-based data reading method.
[0103] In some specific embodiments, the reading unit further includes a recognizer, which identifies the index pattern on the DNA storage medium, thereby controlling the reaction of specific areas on the DNA storage medium with the reaction solution to complete operations such as writing, reading, and erasing.
[0104] In some specific implementations, the functional units (such as storage units and retrieval units) can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0105] The DNA-based data classification and storage method of this application will be explained in detail below with specific examples.
[0106] Example 1: One-dimensional information storage
[0107] This embodiment uses a nylon porous filter membrane as a flexible substrate to verify the feasibility of one-dimensional information storage, wherein... Figure 1 The reaction principle is illustrated, involving the modification of a nylon porous membrane with carboxyl groups and the modification of a DNA template sequence with amino groups. With the assistance of EDAC activation, covalent amide bonds are formed to densely immobilize the DNA sequence on the nylon porous membrane, thereby achieving information storage. Specific verification methods are as follows:
[0108] S1. Place the prepared nylon porous filter membrane with barcode pattern into 10% w / v ethyl dimethylaminopropyl carbodiimide (EDAC) (the EDAC can be adjusted to 5-20% w / v depending on the actual situation), activate for 20 minutes, and then rinse with deionized water to obtain carboxyl-modified nylon porous filter membrane.
[0109] S2. Add 500 pmol of DNA template containing amino groups and Cy3 (ACCTAGAAGTAAAGTAAGCGATCAGTCAGTCTTAGCGGCTATGGCTACGA) to a carboxyl-modified nylon porous membrane and react at room temperature for 5–30 minutes. Then, wash away the unreacted DNA template by placing it in 2×SSPE / 0.1% SDS at 50°C for 5–10 minutes to obtain a nylon porous membrane with DNA template immobilized.
[0110] S3. Contact the DNA primer with Cy5.5 (TCGTAGCCATAGCCGCTAAG) with a nylon porous filter membrane immobilized with DNA template. At a certain reaction temperature, allow the DNA primer with Cy5.5 to perform complementary pairing with the DNA template on the nylon porous filter membrane, and then detect the fluorescence signal on the nylon porous filter membrane.
[0111] As a comparison, this embodiment also provides control group 1, which uses a nylon porous filter membrane that has not been activated by EDAC as a flexible substrate, and control group 2, which does not set a reaction temperature in step S3 (the other conditions are the same).
[0112] The results are as follows Figure 2 and Figure 3 As shown, using carboxyl-modified nylon porous membranes as flexible substrates, a high information storage density can be achieved under suitable conditions. The DNA loading per unit volume of the nylon membrane was calculated using a conversion of two bits per base, and the results indicate an information storage density of 568 TB / mm². 3 It is evident that it is suitable for "cold" data storage with low access frequency and large data storage volume.
[0113] Furthermore, based on the above detection results, it can be concluded that the EDAC activation treatment and reaction temperature affect the amount of DNA loaded per unit volume of nylon filter membrane.
[0114] Example 2: Two-dimensional information storage
[0115] This embodiment uses a PDMS patterned nylon microfluidic chip to verify the feasibility of storing two-dimensional information (such as a four-color Mario pattern).
[0116] 1. Fabrication and Verification of PDMS Patterned Nylon Microfluidic Chips
[0117] By using PDMS as ink, hydrophobic spaces are introduced into a nylon membrane through screen printing to create a PDMS-patterned nylon microfluidic chip. The fabrication process is as follows: Figure 4 As shown, the specific steps include:
[0118] A mixture of PDMS and curing agent (10:1 ratio) was centrifuged to remove air bubbles and used as ink for screen printing. PDMS patterns were then screen printed onto a carboxyl-modified porous nylon membrane. After baking at 80°C for 20 minutes, a nylon microfluidic chip with the PDMS pattern was obtained.
[0119] Based on the PDMS patterned nylon microfluidic chip prepared above, the function of the hydrophobic compartments created based on PDMS was further verified using pigment dyes. The specific method is as follows:
[0120] 0.3 μL of green pigment solution was added to the center of the PDMS patterned 3*3 dot matrix and the center of the unpatterned nylon film, and the changes over time were observed.
[0121] Functional verification results are as follows Figure 5 As shown, the results indicate that pigments and dyes can be confined within hydrophobic compartments, proving that it is feasible to create two-dimensional lattice-type hydrophobic compartments on nylon membranes.
[0122] 2. Four-color Mario pattern storage
[0123] Based on the aforementioned PDMS patterned nylon microfluidic chip, this embodiment further uses a four-color Mario pattern as an example to store two-dimensional information using the high-low bit encoding principle. The specific storage steps are as follows:
[0124] refer to Figure 6 This paper describes a method for converting a four-color Mario pattern into two-color image data based on high- and low-bit encoding principles. The original four colors of the four-color Mario pattern are red, orange, gray, and black. In the high-bit encoding scheme, red pixels are replaced with black, and orange pixels with gray. Conversely, in the low-bit encoding scheme, red pixels are replaced with gray, and orange pixels are replaced with black. During image retrieval, if both the high and low bits of a pixel are gray, it is interpreted as gray; if the high bit is black and the low bit is gray, it is red; if the high bit is gray and the low bit is black, it is orange; and if both the high and low bits are black, it is black. Based on this encoding principle, a two-dimensional dot matrix arrangement data of the two-color image converted from the four-color Mario pattern is obtained.
[0125] Furthermore, the storage is achieved by converting between the presence and absence of DNA, where the presence of DNA represents black and the absence represents gray. Based on this principle, a two-dimensional dot matrix arrangement is written onto a PDMS patterned nylon microfluidic chip. The specific method is as follows:
[0126] The nylon microfluidic chip with PDMS pattern was incubated in 918 mM EDAC solution for 10 minutes, followed by washing three times with DNase / RNase-free water. Residual surface moisture was carefully removed. Then, 100 mM DNA solution and 750 mM sodium bicarbonate solution were mixed at a 1:2 ratio, and 0.3 μL of the mixture was added to each well of the chip array. After incubation at room temperature for 2 hours, the chip was washed three times with 1×TE solution to complete the dot matrix writing.
[0127] Based on the above method, four-color Mario patterns can be stored on PDMS patterned nylon microfluidic chips. Since the two-dimensional pattern information can be read after only one temperature cycle with template primer pairing, the reading speed is only three minutes, making it suitable for storing "hot" data.
[0128] Example 3: Three-dimensional information storage
[0129] This embodiment uses a PDMS patterned nylon microfluidic chip to verify the feasibility of storing three-dimensional information (such as quadruple helical proteins and GIFs of Mario).
[0130] 1. Storage of 3D object information
[0131] refer to Figure 7 The quadrature helical protein to be stored is sliced along the z-axis, and the slices are mapped to a voxel arrangement. The slices are then arranged in z-axis order. By mapping the voxels on the slices to base types according to certain rules, a DNA sequence containing the spatial information of the quadrature helical protein can be obtained. This DNA sequence is then immobilized on a PDMS patterned nylon microfluidic chip for storage.
[0132] During storage, the overall algorithm complexity is high due to the simultaneous dual encoding of spatial information of the sequence. A simpler approach is to align the sequence containing the spatial information of the four-helix protein with the large DNA library to be stored, selecting identical sequences for storage on a two-dimensional dot matrix. During retrieval, the spatial information of the four-helix protein can be read through in situ sequencing; the specific principle can be found in [reference needed]. Figure 8 Specifically, in situ sequencing can be performed using PDMS patterned nylon microfluidic chips.
[0133] In this embodiment, four nucleotides were labeled with different fluorescent groups to verify the feasibility of in situ sequencing using a PDMS-patterned nylon microfluidic chip. A schematic diagram of the principle is provided below. Figure 9Thymine was labeled with FITC, cytosine with Cy5, adenine with TRITC, and guanine with Dylight 800. Base identification was performed using the fluorescence intensity of four channels. The in situ sequencing cycle includes polymerization extension, signal capture, and removal of fluorescent / blocking groups. Through repeated cycles, DNA molecules are decoded sequentially from one end of the DNA. The fluorescence signals corresponding to the merged bases on the chip are captured using a confocal microscope and gel imaging system (see [link to detailed principle]). Figure 10 (Detection is performed by utilizing the fact that each base carries a different fluorescent group). The specific experimental method for the in situ sequencing cycle is as follows:
[0134] Sequencing was performed using the MGISEQ-2000RS high-throughput rapid sequencing kit. Specific methods were described in the kit's instruction manual. The basic steps are as follows:
[0135] S1. Immerse the chip to be tested in the fluorescent sequencing reagent, which is prepared by fluorescently labeled dNTPs, DNA polymerase and reagent No. 1 in a ratio of 1:1:43.75, and incubate at 65°C for 30 seconds.
[0136] S2. Next, the chip is transferred to another set of sequencing reagents (dNTPs, DNA polymerase, and reagent No. 1 in a 1:1:40 ratio) and incubated again at 65°C for 30 seconds.
[0137] S3. Thoroughly clean the chip three times with cleaning solution (2×SSPE, 0.1% SDS, 1% Triton X-100), then clean three times with TE buffer, and finally clean twice with DNase / RNase anhydrous solution.
[0138] S4. Chip imaging can be performed using a gel imaging system (BIO-RAD) or a confocal microscope (Nikon A1R confocal microscope and FLIM).
[0139] S5. After imaging, treat the chip with regeneration reagent No.9 at 55°C for 60 seconds, and then clean the chip again according to the above cleaning steps.
[0140] S6. Repeat steps S1 to S5 to perform subsequent cycles in sequence.
[0141] The results are as follows Figure 11 As shown, this method can accurately read DNA sequence information.
[0142] 2. Video information storage
[0143] Based on the above principles, the three-dimensional information storage method of this invention can also be used to store video information. Unlike the sequential arrangement of three-dimensional objects along the z-axis, the slices are arranged sequentially along the time axis during video information storage. (Refer to...) Figure 12 As shown, a GIF of Mario is used as a verification. The GIF is sliced in chronological order, and different pixels in a single slice can be mapped to different base types, thereby realizing the storage of the GIF. For the specific storage and retrieval principle, please refer to the "3D Object Information Storage" section above.
[0144] 3. DNA information encryption
[0145] By using site-selective sequencing, base combinations from different spatial locations can be read out on a single plane, thus achieving DNA information encryption in three-dimensional information storage and enabling more diverse information transmission methods. The principle and verification of the method are as follows: Figure 13 and Figure 14 As shown, each site can be sequenced or not in each cycle, which means that the base combinations at different Z-axis positions can be displayed on a plane. Different information can be transmitted through selective sequencing of each site.
[0146] In summary, this invention provides a DNA-based data classification and storage method and application. The method involves dividing the information to be stored into one-dimensional, two-dimensional, and three-dimensional information based on data access frequency. One-dimensional information is low-frequency data (i.e., "cold" data), which is converted and stored based on binary encoding principles. Two-dimensional information is high-frequency data (i.e., "hot" data), which is converted into two-dimensional dot matrix information based on high-low bit encoding principles for storage. Three-dimensional information is medium-frequency data (i.e., "warm" data), which is converted into dot matrix information based on voxel encoding principles or pixel encoding principles for storage. Using this DNA-based data classification and storage method enables efficient data management and helps improve the flexibility of DNA data storage and retrieval.
[0147] The embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention. Furthermore, the embodiments of the present invention and the features thereof can be combined with each other unless otherwise specified.
Claims
1. A DNA-based data classification and storage method, characterized in that, include: The information to be stored is divided into one-dimensional information, two-dimensional information, and three-dimensional information according to the data access frequency. The one-dimensional information is low read / write frequency information, the two-dimensional information is high read / write frequency information, and the three-dimensional information is medium read / write frequency information. Based on the principle of binary encoding, the one-dimensional information is converted into a nucleic acid sequence and fixed on a first storage medium containing a flexible substrate for storage. Based on the high-low bit encoding principle, the two-dimensional information is converted into two-dimensional dot matrix information, and then the two-dimensional dot matrix information is written into a second storage carrier containing a flexible substrate using DNA as raw material for storage. Based on the voxel encoding principle or pixel encoding principle, the three-dimensional information is converted into dot matrix information. At the same time, according to the preset mapping relationship, the voxels or pixels are converted into different base types and arranged in order according to the preset direction to obtain a nucleic acid sequence. The nucleic acid sequence is then fixed on a third storage carrier containing a flexible substrate for storage.
2. The DNA-based data classification and storage method according to claim 1, characterized in that, The process of converting the one-dimensional information into a nucleic acid sequence based on the binary encoding principle includes: The one-dimensional information is converted into a binary sequence using an encoding algorithm, and the binary sequence is converted into a nucleic acid fragment according to a preset mapping relationship.
3. The DNA-based data classification and storage method according to claim 1, characterized in that, The high-low coding principle includes using two states of DNA presence or absence to represent two-dimensional information; Preferably, the method for writing the two-dimensional dot matrix information using DNA as raw material includes micro-nano lithography, electrochemical writing, or inkjet printing.
4. The DNA-based data classification and storage method according to claim 1, characterized in that, The process of converting the three-dimensional information into raster information based on the voxel encoding principle includes: The three-dimensional information is sliced according to a preset direction, and the slices are mapped to a voxel arrangement to obtain two-dimensional dot matrix information containing slice information. Preferably, the preset direction includes a preset X-axis, Y-axis, or Z-axis direction; Preferably, the step of converting the three-dimensional information into dot matrix information based on the pixel encoding principle includes: The three-dimensional information is sliced according to a preset direction, and the slices are mapped to a voxel arrangement to obtain two-dimensional dot matrix information containing slice information, wherein the preset direction includes the time axis direction.
5. The DNA-based data classification and storage method according to any one of claims 1 to 4, characterized in that, The information to be stored includes at least one of the following: images, text, programs, audio, video, and three-dimensional objects.
6. A DNA-based classification data reading method, characterized in that, include: Obtain a storage medium carrying the one-dimensional, two-dimensional, or three-dimensional information; When reading the one-dimensional information, the corresponding nucleic acid sequence in the storage vector is read using in situ sequencing, and the one-dimensional information is decoded to obtain the one-dimensional information. When reading the two-dimensional information, a reading method corresponding to the storage method of the two-dimensional information is determined according to the storage method of the two-dimensional information, and the reading method is used to read the corresponding two-dimensional dot matrix information in the storage carrier and decode to obtain the two-dimensional information; When reading the two-dimensional information, a reading method corresponding to the three-dimensional information storage method is determined according to the three-dimensional information storage method. The reading method is used to read the dot matrix information of the corresponding slice in the storage carrier, and then the slices are arranged in a preset direction order to obtain spatial information. After decoding, the three-dimensional information is obtained.
7. The classification data reading method according to claim 6, characterized in that, When reading the two-dimensional information and / or the three-dimensional information, the reading method includes fluorescence imaging reading, optical microscopy reading, scanning electron microscopy reading, atomic force microscopy reading, or super-resolution microscopy reading.
8. A storage system, characterized in that, It includes a storage unit and a retrieval unit, wherein the storage unit operates the DNA-based data classification and storage method as described in any one of claims 1 to 5.
9. The storage system according to claim 8, characterized in that, The reading unit operates the DNA-based classification data reading method as described in claim 6 or 7.
10. The application of the DNA-based data classification and storage method as described in any one of claims 1 to 5 and / or the DNA-based classification data reading method as described in any one of claims 6 to 7 in DNA data classification management or storage.