DNA Data Hierarchical Storage and Retrieval Methods
By classifying DNA information into primary and secondary information and combining DNA base sequence codes and array codes for storage, the high cost and operational complexity of existing DNA information storage methods have been solved, enabling efficient and automated information storage and retrieval.
Patent Information
- Application Number
- CN202211217725.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Existing DNA information storage technologies are costly, slow, complex to operate, and difficult to automate, especially when it is necessary to read or modify a small amount of key information, which consumes a lot of time and effort.
The information to be stored is classified into primary information and secondary information, which are stored using DNA base sequence codes and array codes respectively. By utilizing DNA molecular synthesis and sequencing technology and storage unit array arrangement technology, the hierarchical storage and retrieval of information can be achieved.
It achieves simple and highly automated operation for DNA data storage, is suitable for long-term storage of large amounts of cold information, and is highly efficient when frequently reading and writing key information, reducing the cost and complexity of information retrieval.
Smart Images

Figure CN115421669B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage technology, and in particular to a method for hierarchical storage and retrieval of DNA data. Background Technology
[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.
[0003] DNA information storage technology, as a novel information storage method, has seen significant development in recent years. DNA, as a natural information storage medium within living organisms, possesses extremely high information storage density and lifespan. DNA information storage technology utilizes artificial DNA synthesis to create DNA strands with specific base sequences for information storage. Then, DNA sequencing technology detects the base sequence arrangement of unknown DNA strands, enabling information retrieval. Its biggest drawbacks are high information reading and writing costs, slow speeds, and reliance on biochemical reactions, making the operation complex and difficult to automate. Especially for a DNA repository containing a large amount of digital information, reading or modifying even a small amount of critical information requires specialized personnel to conduct manual chemical experiments, consuming considerable time and effort in sequencing and synthesizing the entire DNA information, making its application quite inconvenient. Summary of the Invention
[0004] This invention provides a method for hierarchical storage of DNA data, which is simple to operate, highly automated, convenient to apply, and highly efficient. The method includes:
[0005] The overall information to be stored is divided into primary information and secondary information. Primary information includes information that needs to be read and written frequently, while secondary information is the overall information.
[0006] The secondary information is used to form DNA base sequence codes; the primary information is used to form array codes in the corresponding base.
[0007] Synthesize a set of DNA strands that contain the DNA base sequence code;
[0008] According to the preset method for writing primary information, the DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector through the writing method to obtain a storage vector.
[0009] This invention provides another method for hierarchical reading of DNA data, which enables hierarchical reading of stored DNA data. The method is simple to operate, highly automated, convenient to apply, and highly efficient. The method includes:
[0010] A storage medium carrying primary and secondary information is obtained, wherein the storage medium is obtained by writing the array code corresponding to the primary information into a DNA vector, and the DNA vector contains the DNA base sequence code corresponding to the secondary information.
[0011] When it is necessary to read first-level information, the reading method corresponding to the writing method of the first-level information is determined, the array code of the storage medium is read using the reading method, the read array code is decoded, and the first-level information is obtained.
[0012] When secondary information needs to be read, the DNA strand set in the storage medium is obtained, the DNA base sequence code in the DNA strand set is read, and the read DNA base sequence code is decoded to obtain the secondary information;
[0013] The first-level information includes information that needs to be read and written frequently in the overall information, while the second-level information is the overall information.
[0014] In this invention, DNA information storage technology based on DNA molecule synthesis and sequencing is combined with information storage technology based on storage unit array arrangement to perform hierarchical storage of DNA information. Different levels of information are stored using these two technologies respectively, fully leveraging their advantages and overcoming the shortcomings of existing DNA information storage methods, such as high cost and cumbersome operation. This approach is suitable for applications requiring long-term storage of large amounts of cold information, where frequent reading and writing of key information is not necessary, but rather frequent reading and writing of all information is required. It is convenient and highly efficient. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0016] Figure 1 This is a flowchart of the DNA information hierarchical storage method in an embodiment of the present invention;
[0017] Figure 2 This is a schematic diagram illustrating the preparation of DNA ink in an embodiment of the present invention;
[0018] Figure 3 and Figure 4 These are schematic diagrams illustrating the principles of inkjet printing and hollow probe printing in embodiments of the present invention.
[0019] Figure 5 This is a schematic diagram illustrating the preparation of the DNA membrane in an embodiment of the present invention;
[0020] Figure 6 This is a schematic diagram illustrating the principle of scanning probe photolithography printing in an embodiment of the present invention;
[0021] Figure 7 This is a flowchart of the DNA information hierarchical reading method in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0023] The inventors discovered that a method for information storage using atomic force microscopy (AFM) to etch micro / nano arrays had previously been proposed. This technique, similar to optical discs, uses raised / recessed dots etched onto a flat substrate to represent "0" and "1" respectively, arranging them in a specific sequence to achieve information storage. The precision of this AFM lithography method is significantly higher than that of optical disc laser heads, thus achieving higher information storage density. Compared to DNA information storage, this method is lower in cost, faster, simpler to operate, and easier to automate. Therefore, this invention proposes a hierarchical DNA data storage method that combines DNA information storage technology based on DNA molecule synthesis and sequencing with optical disc / hard disk information storage technology based on storage unit arrays. This allows for hierarchical storage of DNA information, utilizing both technologies to store different levels of information, fully leveraging the advantages of both technologies, and overcoming the drawbacks of existing DNA information storage methods, such as high cost and cumbersome operation. This method is particularly suitable for applications requiring long-term storage of large amounts of cold information, where frequent reading and writing of key information is required, rather than reading and writing all information frequently.
[0024] The principle of this invention is as follows: using DNA chemical synthesis technology, based on the DNA base sequence code encoded by digital information, a DNA strand storing corresponding information is synthesized as secondary information storage. The DNA strand containing secondary information is formed into a DNA carrier. As will be seen later, the carrier form includes, but is not limited to, dissolving DNA to form DNA ink, or forming a DNA membrane on a substrate. Through appropriate processing technology, these carriers are used to form a specific information storage matrix, with the pattern distribution of the matrix serving as primary information storage.
[0025] Secondary information, stored in DNA, boasts extremely high information storage density. However, its synthesis relies on chemical processes such as sequencing and synthesis, leading to high automation difficulty and inconvenient application. This invention utilizes secondary information as a carrier, forming primary information through an array. While primary information has a lower storage density than secondary information, its information reading and writing processes are simple array detection and printing etching processes, offering high efficiency, speed, and ease of implementation. Hierarchical storage of secondary and primary information fully leverages the high storage density of secondary information while also incorporating the convenient reading and writing advantages of primary information. The combination of these two elements forms the method of this invention, which will be described in detail below.
[0026] Figure 1 The flowchart of the DNA information hierarchical storage method in this embodiment of the invention includes:
[0027] Step 101: Divide the overall information to be stored into primary information and secondary information. Primary information includes information that needs to be frequently read and written in the overall information, while secondary information is the overall information.
[0028] Step 102: Form DNA base sequence codes from secondary information; form array codes of the corresponding base from primary information;
[0029] Step 103: Synthesize a set of DNA strands containing DNA base sequence codes;
[0030] Step 104: According to the preset method for writing primary information, the DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector through the writing method to obtain a storage vector.
[0031] In this embodiment of the invention, DNA information storage technology based on DNA molecular synthesis and sequencing is combined with information storage technology based on storage unit array arrangement to perform hierarchical storage of DNA information. Different levels of information are stored using both technologies, fully leveraging their advantages and overcoming the shortcomings of existing DNA information storage methods, such as high cost and cumbersome operation. This approach is suitable for applications requiring long-term storage of large amounts of cold information, where frequent reading and writing of key information is necessary, rather than frequent reading and writing of all information.
[0032] In step 101, the overall information to be stored is divided into primary information and secondary information. The primary information includes information in the overall information that needs to be read and written frequently, and the secondary information is the overall information.
[0033] The overall information refers to all the information to be stored. For example, if you want to store information about the National Library, you can treat it as the overall information.
[0034] In one embodiment, the primary information includes the index, summary, key information, physical location of the secondary information in the storage medium, and primer information in the DNA strand.
[0035] Among them, key information refers to information that requires frequent reading and writing operations. For example, if the information of the entire National Library is to be stored, the indexes, abstracts, and full texts of commonly used books (key information) can be compiled as primary information, and the information of the entire library can be compiled as secondary information.
[0036] In daily use, users do not need to read or modify all the massive secondary information. When they only care about the key information, they can easily read and write the primary information without introducing a complex DNA synthesis and sequencing process.
[0037] In step 102, the secondary information is converted into DNA base sequence codes; the primary information is converted into array codes of the corresponding base.
[0038] In this embodiment of the invention, since the specific reading and writing methods of secondary information and primary information are different, different methods are used to encode them respectively.
[0039] In one embodiment, the secondary information is formed into a DNA base sequence code, including:
[0040] The secondary information is encoded as the arrangement sequence of DNA bases A (adenine), T (thymine), C (cytosine), and G (guanine);
[0041] The arrangement sequence is divided into base sequence codes for different DNA strands, and the base sequence codes in each DNA strand include primer sequence, position sequence, information sequence, and error correction sequence.
[0042] In one embodiment, forming the primary information into an array code of a corresponding base includes:
[0043] According to the preset method for writing first-level information, the first-level information is encoded into a number system arrangement sequence corresponding to the writing method;
[0044] The array code is formed by arranging the number system into the array code, which includes a positioning flag code, a position code, an information code, and an error correction code.
[0045] For Level 1 information, the natural information is encoded into a corresponding binary arrangement sequence based on the available information states in the selected reading and writing method. For example, if the preset writing method for Level 1 information is micro-nano lithography, and it forms three information states: raised "2", flat "1", and recessed "0", then the information is encoded in ternary form and arranged accordingly.
[0046] In step 103, a set of DNA strands containing DNA base sequence codes is synthesized;
[0047] Specifically, DNA strands can be synthesized artificially. Based on the characteristics of the synthesis technology, multiple DNA strands are obtained, forming a DNA strand set. Each DNA strand contains primer information of the DNA molecule in the DNA base sequence code, and each DNA strand has a large number of copies, forming multiple backup information.
[0048] Each DNA strand's base sequence also includes an address code and a check code. For example, in one available encoding method, each DNA strand's base sequence is a 132nt fragment, into which a 14nt address code and a 12nt check code are added, forming a 158nt long DNA strand. The address code is used for numbering, and the check code is used to verify if the fragment is corrupted.
[0049] In step 104, according to the preset method for writing primary information, the DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector through the writing method to obtain a storage vector.
[0050] In this embodiment of the invention, there are two methods for writing the preset first-level information.
[0051] In one embodiment, according to a preset method for writing primary information, a DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector using the writing method, including:
[0052] When the preset method for writing primary information is droplet printing, the DNA strands are dissolved in water to form DNA ink, which is the DNA carrier.
[0053] DNA ink is printed using a droplet printing method to form a dot matrix corresponding to the array code on a preset substrate, thereby obtaining a storage medium and realizing primary information writing.
[0054] Figure 2 This is a schematic diagram illustrating the preparation of DNA ink in an embodiment of the present invention. DNA strands are dissolved in water to form DNA ink, resulting in a DNA carrier that can be used for droplet printing methods such as inkjet printing and hollow probe printing. Specifically, droplet printing methods include inkjet printing or AFM hollow probe (FluidFM) printing, and other droplet printing methods are also possible; no limitation is made here. After the DNA ink is used to form a dot matrix corresponding to the array code on a preset substrate using a droplet printing method, this dot matrix becomes a hierarchical DNA storage carrier containing secondary and primary information. The arrangement of the dots represents primary information, and the DNA molecules within the dot matrix represent secondary information. Figure 3 and Figure 4 These are schematic diagrams illustrating the principles of inkjet printing and hollow probe printing, respectively. The array is a DNA droplet array on a substrate.
[0055] In one embodiment, according to a preset method for writing primary information, a DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector using the writing method, including:
[0056] When the preset method for writing primary information is micro-nano lithography, the DNA strands are dissolved in a solvent to prepare a DNA solution; the solvent can be composed of water and methanol.
[0057] A DNA solution is dropped onto a specific substrate and then spin-coated to prepare a DNA membrane, which serves as a DNA carrier.
[0058] The DNA membrane is etched using micro-nano lithography to form a dot matrix corresponding to the array code on the DNA membrane, thereby obtaining a storage carrier and realizing primary information writing.
[0059] Figure 5 This is a schematic diagram illustrating the preparation of a DNA membrane in an embodiment of the present invention. DNA is dissolved in a 1:1 solution of water and methanol to prepare a DNA solution. This solution is then formed on a flat substrate such as silicon / gold using a spin-coating method to obtain a DNA carrier suitable for micro / nano lithography methods such as scanning probe lithography. Specifically, the micro / nano lithography method can be scanning probe lithography. After forming a dot matrix corresponding to the array code on the DNA membrane, this dot matrix becomes a hierarchical DNA storage carrier containing both secondary and primary information. The arrangement of the dots on the DNA membrane represents primary information, and the DNA molecules within the DNA membrane represent secondary information. Figure 6 This is a schematic diagram of the principle of scanning probe photolithography printing in an embodiment of the present invention. The array is a DNA membrane on the substrate, and the DNA membrane has a dot matrix formed by probe etching points.
[0060] After DNA information is stored in a hierarchical manner, it can be retrieved according to specific needs. Figure 7 The flowchart of the DNA information hierarchical reading method in this embodiment of the invention includes:
[0061] Step 701: Obtain a storage carrier carrying primary information and secondary information. The storage carrier is obtained by writing the array code corresponding to the primary information into a DNA carrier. The DNA carrier contains the DNA base sequence code corresponding to the secondary information.
[0062] Step 702: When it is necessary to read first-level information, determine the reading method corresponding to the writing method of first-level information, read the array code of the storage medium using the reading method, decode the read array code, and obtain the first-level information.
[0063] Step 703: When it is necessary to read secondary information, obtain the DNA strand set in the storage medium, read the DNA base sequence code in the DNA strand set, decode the read DNA base sequence code, and obtain the secondary information.
[0064] The first-level information includes information that needs to be read and written frequently in the overall information, while the second-level information is the overall information.
[0065] In this embodiment of the invention, since the primary information stores information that requires frequent reading and writing, only the primary information needs to be read in normal use. Primary information is typically presented in a dot matrix format, allowing for rapid reading. This process is significantly faster than reading the DNA molecule base sequence and is simpler and easier to automate. Secondary information stores the overall information; when comprehensive reading is required, the secondary information can be read entirely. Secondary information is presented in the form of DNA molecule base sequences. After extracting the corresponding DNA molecules from the carrier of the primary information, DNA sequencing is performed. This hierarchical reading method achieves rapid hierarchical reading of DNA information.
[0066] In step 701, a storage carrier carrying primary information and secondary information is obtained. The storage carrier is obtained by writing the array code corresponding to the primary information into a DNA carrier. The DNA carrier contains the DNA base sequence code corresponding to the secondary information.
[0067] The aforementioned storage medium was obtained using the aforementioned DNA information hierarchical storage method and can be used to read DNA information on demand.
[0068] In step 702, when it is necessary to read first-level information, the reading method corresponding to the writing method of the first-level information is determined, the array code of the storage medium is read using the reading method, the read array code is decoded, and the first-level information is obtained.
[0069] In this embodiment of the invention, two methods for writing first-level information are given. Here, two methods for reading first-level information are also given, which will be described below.
[0070] In one embodiment, based on the writing method of the first-level information, a reading method corresponding to the writing method is determined, and the array code of the storage medium is read using the reading method, including:
[0071] When the writing method is a droplet printing method, the reading method is determined to be optical microscope reading;
[0072] This allows the DNA on a pre-set substrate in the storage medium to become colored or fluorescent; specifically, this can be achieved using DNA dyes or corresponding DNA probes.
[0073] The array code is obtained by observing the arrangement of dots in the storage medium using an optical microscope.
[0074] In one embodiment, based on the writing method of the first-level information, a reading method corresponding to the writing method is determined, and the array code of the storage medium is read using the reading method, including:
[0075] When the writing method is a micro-nano lithography method, the reading method is determined to be a scanning electron microscope or an atomic force microscope;
[0076] The array code is obtained by observing the arrangement of dots in the storage medium using a scanning electron microscope or an atomic force microscope.
[0077] Specifically, there are two scenarios for reading secondary information: reading all secondary information and reading a portion of secondary information, which will be described below.
[0078] In one embodiment, obtaining a set of DNA strands in a storage medium and reading the DNA base sequence codes in the set of DNA strands includes:
[0079] When all secondary information needs to be read, the storage medium is dissolved in a solvent to obtain a DNA strand set. The DNA base sequence codes of all DNA molecules in the DNA strand set are then read through DNA molecule sequencing.
[0080] When it is necessary to read some secondary information, first-level information is obtained; based on the summary of the secondary information to be read, the index of the secondary information to be read is obtained from the first-level information; the physical location of the secondary information corresponding to the index in the storage medium and the primer information in the DNA strand are determined; the DNA vector corresponding to the physical location is extracted from the storage medium; the obtained DNA vector is dissolved in a solvent to obtain a DNA strand set; the DNA molecules to be read are selected from the DNA strand set according to the primer information, and DNA molecule sequencing is performed to read the DNA base sequence code of the corresponding DNA molecule.
[0081] In this process, DNA molecules to be read are selected from the DNA strand set based on the primer information, and DNA sequencing is performed. Specifically, specific PCR amplification is conducted based on the primer information to further index the required DNA information for targeted sequencing. This avoids sequencing all DNA molecules to access only partial information, significantly improving the indexing and access speed of secondary information. Compared to directly using primer information for indexing, the hierarchical storage index combines physical location with primer information, improving overall efficiency, reducing the difficulty of random access primer design, and storing primer information in the primary information, avoiding the need for additional databases for storage and retrieval.
[0082] The following is a specific embodiment to illustrate the specific application of the method proposed in this invention.
[0083] The information from local gazetteers is stored and retrieved using the method proposed in the embodiments of this invention.
[0084] First, information classification
[0085] The information in the local gazetteer is classified as secondary information, while the index information, summary, annual major events and catalog (key information) of the local gazetteer are classified as primary information.
[0086] Second, encoding
[0087] The information in the entire local gazetteer is secondary information. Based on its encoding format in existing computer storage, its binary encoded sequence is obtained. Following traditional DNA information storage encoding methods, the binary encoded sequence is converted into a DNA base sequence. In one available encoding method, the DNA base sequence is divided into multiple 132nt fragments. A 14nt address code and a 12nt check code are added to these fragments, ultimately forming multiple 158nt long DNA strands. Furthermore, according to a pre-defined primary information writing method, the primary information is encoded into a radix-based sequence corresponding to the writing method. The physical location of the secondary information in the storage medium and the primer information in the DNA strands are considered as primary information.
[0088] Third, writing secondary information.
[0089] Based on the encoded DNA base sequence, a set of DNA strands is synthesized artificially, specifically, multiple 158nt DNA strands are obtained, each with primer information and a large number of copies of each strand.
[0090] Fourth, writing first-level information.
[0091] Example A: The large number of 158nt DNA strands obtained in the previous step are dissolved in water to form DNA ink, which is the DNA carrier. This example includes, but is not limited to, the following two methods: ① Loading the DNA ink into an inkjet printer. According to the array code corresponding to the primary information, ink is printed at a certain point on a preset substrate to represent "1", and no ink is printed at a certain point to represent "0". The preset substrate can be a glass slide, paper, nylon film, etc. The printed information is an array code, thus achieving primary information writing. The primary information written in this method has an information storage unit size of approximately 50µm. ② Loading the DNA ink into the ink storage chamber of an AFM hollow probe (FluidFM). Appropriate air pressure is applied behind it, forcing the DNA ink to the front end of the hollow probe nozzle. When it contacts the substrate with appropriate force and then leaves, ink printing is achieved. According to the array code corresponding to the primary information, ink is printed at a certain point on a preset substrate to represent "1", and no ink is printed at a certain point to represent "0". The printed information is an array code, thus achieving primary information writing. The first-level information written using this method has an information storage unit size of approximately 1µm.
[0092] Example B: The large number of 158nt DNA strands obtained in the previous step were dissolved in a 1:1 mixture of water and methanol to prepare a DNA solution. The DNA solution was dropped onto a flat, specific substrate such as a gold substrate / highly doped silicon substrate, and a DNA film with a thickness of approximately 10nm was obtained by spin coating at 2000 rpm for 30 seconds. This DNA film served as a DNA carrier. The specific substrate was used as a sample for atomic force microscopy (AFM). A voltage of 10-30V was applied to the specific substrate, and the conductive probe of the AFM was grounded, creating an electric field between the probe and the specific substrate. When the probe was less than 10nm away from the substrate, local anodizing of the DNA film was achieved. Oxidized points represented "1" and non-oxidized points represented "0". Based on the array code corresponding to the primary information, an array code was printed on the DNA film, thus achieving primary information writing. The primary information written using this method has an information storage unit size of approximately 10nm-100nm.
[0093] Fifth, Level 1 Information Reading
[0094] Since the primary information stores frequently used key information, only the primary information needs to be read during normal use. The primary information is presented in the form of an array.
[0095] In Example A, the storage cell size is in the micrometer range, falling within the observation range of an optical microscope. DNA on the substrate is colored or fluoresced using a DNA dye, and the arrangement of the dots in the storage medium is observed using an optical microscope to obtain the array code.
[0096] For Example B, the array code can be obtained by observing the arrangement of the dot matrix in the storage medium using a scanning electron microscope or an atomic force microscope.
[0097] This process is significantly faster than reading the base sequence of a DNA molecule, and it is simpler to manipulate and easier to automate.
[0098] (6) Secondary information retrieval: The secondary information stores all information in the system. When it is necessary to fully retrieve information (such as consulting all local chronicles), the secondary information can be retrieved. The secondary information is presented in the form of DNA molecular base sequences and is realized through DNA molecular sequencing.
[0099] When it is necessary to read all the secondary information, the storage medium is dissolved in a solvent to obtain a DNA strand set, which includes two cases:
[0100] For Example A, the storage vector is dissolved in a solvent (which can be water) to obtain a DNA solution (here, the DNA solution is DNA ink), that is, the DNA strand set on the preset substrate is transferred into the DNA aqueous solution. After purification and PCR amplification, the DNA strand set is obtained, and high-throughput sequencing is performed to read the DNA base sequence code of all DNA molecules in the DNA strand set;
[0101] For Example B, the storage vector is dissolved in a solvent (such as water) to obtain a DNA solution, that is, the DNA strand set on the substrate is transferred to the DNA aqueous solution. After purification and PCR amplification, the DNA strand set is obtained, and high-throughput sequencing is performed to read the DNA base sequence code of all DNA molecules in the DNA strand set.
[0102] When it is necessary to read some secondary information, such as reading local history information of a city in 2018, the index of the secondary information to be read is obtained from the primary information based on the abstract of the secondary information to be read (e.g., the city in 2018); the physical location of the secondary information corresponding to the index in the storage medium and the primer information in the DNA strand are determined; the DNA vector corresponding to the physical location is extracted from the storage medium according to the physical location; the extraction method is as described above, that is, the obtained DNA vector is dissolved in a solvent to obtain a set of DNA strands, and then the DNA molecules to be read are screened from the set of DNA strands according to the primer information, and DNA molecules are sequenced to read the DNA base sequence code of the corresponding DNA molecules.
[0103] (7) Decoding: Based on the array code corresponding to the primary information or the DNA base sequence code corresponding to the secondary information, the read specific array is decoded into the required information, i.e., primary information or secondary information, according to the encoding rules.
[0104] Figure 3This is a schematic diagram illustrating the principle of storing and retrieving primary and secondary information using micro-nano lithography in this embodiment of the invention. First, based on the secondary information, DNA is synthesized, and the secondary information is written into a DNA strand set. The DNA strand set is dissolved in a 1:1 mixture of water and methanol, and a DNA membrane is obtained by homogenization on a specific substrate. Primary information is then written onto the DNA membrane using scanning probe lithography. In daily use, commonly used information can be obtained through AFM scanning / electron microscopy. When partial secondary information needs to be retrieved, a portion of the DNA membrane is extracted, and then partially dissolved in water to obtain a DNA solution. After purification and PCR processing, the DNA strand set is obtained, and sequencing yields the complete information, i.e., the secondary information.
[0105] In summary, the method proposed in this invention combines DNA information storage technology based on DNA molecule synthesis and sequencing with information storage technology based on storage cell array arrangement to perform hierarchical storage of DNA information. It utilizes both technologies to store different levels of information, fully leveraging the advantages of each and overcoming the shortcomings of existing DNA information storage methods, such as high cost and cumbersome operation. This method is suitable for applications requiring long-term storage of large amounts of cold information, where frequent reading and writing of key information is necessary, rather than frequent reading and writing of all information.
[0106] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for hierarchical storage of DNA information, characterized in that, include: The overall information to be stored is divided into primary information and secondary information. Primary information includes information that needs to be read and written frequently, while secondary information is the overall information. The secondary information is used to form DNA base sequence codes; the primary information is used to form array codes in the corresponding base. Synthesize a set of DNA strands that contain the DNA base sequence code; According to the preset method for writing primary information, the DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector through the writing method to obtain a storage vector. The primary information includes the index, summary, key information, physical location of the secondary information in the storage medium, and primer information in the DNA strand; Forming primary information into an array code of a corresponding base includes: encoding primary information into a base arrangement sequence corresponding to a preset primary information writing method; forming the array code from the base arrangement sequence, wherein the array code includes a positioning flag code, a position code, an information code, and an error correction code.
2. The method as described in claim 1, characterized in that, The secondary information is used to form a DNA base sequence code, including: The secondary information is encoded as the arrangement sequence of DNA bases A, T, C, and G; The arrangement sequence is divided into base sequence codes for different DNA strands, and the base sequence codes in each DNA strand include primer sequence, position sequence, information sequence, and error correction sequence.
3. The method as described in claim 1, characterized in that, According to a preset method for writing primary information, a DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector using the writing method, including: When the preset method for writing primary information is droplet printing, the DNA strands are dissolved in water to form DNA ink, which is the DNA carrier. DNA ink is printed using a droplet printing method to form a dot matrix corresponding to the array code on a preset substrate, thereby obtaining a storage medium and realizing primary information writing.
4. The method as described in claim 1, characterized in that, According to a preset method for writing primary information, a DNA strand set is prepared into a DNA vector corresponding to the writing method, and an array code is written into the DNA vector using the writing method, including: When the preset method for writing primary information is micro-nano lithography, the DNA strands are dissolved in a solvent to prepare a DNA solution. A DNA solution is dropped onto a specific substrate and then spin-coated to prepare a DNA membrane, which serves as a DNA carrier. The DNA membrane is etched using micro-nano lithography to form a dot matrix corresponding to the array code on the DNA membrane, thereby obtaining a storage carrier and realizing primary information writing.
5. A method for hierarchical reading of DNA information, characterized in that, include: A storage medium carrying primary and secondary information is obtained, wherein the storage medium is obtained by writing the array code corresponding to the primary information into a DNA vector, and the DNA vector contains the DNA base sequence code corresponding to the secondary information. When it is necessary to read first-level information, the reading method corresponding to the writing method of the first-level information is determined, the array code of the storage medium is read using the reading method, the read array code is decoded, and the first-level information is obtained. When secondary information needs to be read, the DNA strand set in the storage medium is obtained, the DNA base sequence code in the DNA strand set is read, and the read DNA base sequence code is decoded to obtain the secondary information; The primary information includes information that needs to be frequently read and written within the overall information, while the secondary information is the overall information. The primary information includes the index, summary, key information, physical location of the secondary information in the storage medium, and primer information in the DNA strand. The array code includes a positioning marker code, a location code, an information code, and an error correction code.
6. The method as described in claim 5, characterized in that, Based on the writing method of the first-level information, determine the reading method corresponding to the writing method, and use the reading method to read the array code of the storage medium, including: When the writing method is a droplet printing method, the reading method is determined to be optical microscope reading; To make the DNA on a preset substrate in the storage medium show color or fluorescence; The array code is obtained by observing the arrangement of dots in the storage medium using an optical microscope.
7. The method as described in claim 5, characterized in that, Based on the writing method of the first-level information, determine the reading method corresponding to the writing method, and use the reading method to read the array code of the storage medium, including: When the writing method is a micro-nano lithography method, the reading method is determined to be a scanning electron microscope or an atomic force microscope; The array code is obtained by observing the arrangement of dots in the storage medium using a scanning electron microscope or an atomic force microscope.
8. The method as described in claim 5, characterized in that, Obtain the DNA strand set in the storage medium, and read the DNA base sequence code in the DNA strand set, including: When all secondary information needs to be read, the storage medium is dissolved in a solvent to obtain a DNA strand set. The DNA base sequence codes of all DNA molecules in the DNA strand set are then read through DNA molecule sequencing. When it is necessary to read some secondary information, first-level information is obtained; based on the summary of the secondary information to be read, the index of the secondary information to be read is obtained from the first-level information; the physical location of the secondary information corresponding to the index in the storage medium and the primer information in the DNA strand are determined; the DNA vector corresponding to the physical location is extracted from the storage medium; the obtained DNA vector is dissolved in a solvent to obtain a DNA strand set; the DNA molecules to be read are selected from the DNA strand set according to the primer information, and DNA molecule sequencing is performed to read the DNA base sequence code of the corresponding DNA molecule.
Citation Information
Patent Citations
Distributed array storage and microorganism-based high-capacity error correction DNA storage technology (Bio-RAID)
CN114927169A