A DNA storage method and data information storage entity capable of realizing data random lossless reading
By immobilizing DNA double-stranded structures on silica microspheres or magnetic beads and encapsulating them with hydrogel materials, combined with fluorescent probes and primer indexes, the problems of long-term preservation and random, lossless reading in DNA storage are solved, achieving efficient data access and preservation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2026-03-31
AI Technical Summary
Existing DNA storage methods cannot simultaneously achieve long-term preservation and random lossless reading in systems with extremely high file volumes. Existing indexing methods suffer from problems such as low information utilization and high risk of file loss.
DNA double-stranded structures are immobilized on silica microspheres or magnetic beads, encapsulated with hydrogel materials, and random access is achieved using fluorescent probes and primer indexing. Data is then read through PCR amplification, ensuring that the data strand is preserved in situ.
It enables random, lossless data reading in systems with extremely high file volumes, maintaining high information utilization and preventing data loss, while supporting integrated and miniaturized operations.
Smart Images

Figure CN116364194B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of DNA storage, and more specifically to a DNA storage method and a data storage entity that enables random and lossless reading of data. Background Technology
[0002] DNA, due to its ultra-high storage density and ultra-long storage time, is considered a highly promising future storage medium, capable of alleviating the storage pressure brought about by the exponential growth of modern data. As a data storage device, it must be able to perform the functions of reading, writing, and storing information. Therefore, the process of DNA data storage involves encoding and synthesizing data information and writing it into DNA, various methods of storage and extraction in between, and finally sequencing and decoding the extracted DNA to read out the data information.
[0003] Current preservation methods include lyophilizing DNA into powder, ligating it to plasmids and encapsulating it with protective substances, and ligating it to microspheres and encapsulating it with protective substances. Lyophilizing DNA into powder allows for long-term preservation, as seen in fossil DNA. However, this method is complex to read and not conducive to integration and miniaturization. Compared to ligating DNA to plasmids and then encapsulating it in silica microspheres, directly ligating the DNA data strand to silica microspheres offers higher information utilization and simplifies amplification and recovery.
[0004] Currently, random access methods include conventional primer-indexed PCR, fluorescent primer-indexed droplet PCR sorting, and PCR after DNA strand immobilization. Achieving random access requires indexing for specific data access. Therefore, the initial approach used primers as indexes. When a file needs to be read, it is amplified using primers specific to that file, followed by sequencing to obtain the desired file data. However, this method has a problem: the entire file system is contiguous, so amplification increases the proportion of files read and may even cause other files to lose stored information. Therefore, droplet PCR was proposed to achieve near-lossless reading and recovery. The method involves using fluorescent probes to bind to the file to be read, then using fluorescently sorted droplets, and finally amplifying the selected file. This avoids affecting other files. However, this method reduces information utilization and ultimately leads to the loss of oligonucleotide data. This is because unpaired single-stranded DNA requires careful design to prevent specific binding by other probes, which significantly reduces information utilization—a counterproductive approach—and cannot completely prevent certain probes from binding, resulting in data loss. Therefore, double-stranded DNA preservation has been further proposed, where one strand is short and the other is long. The short strand pairs with the long strand to protect the data and prevent other probes from specifically binding to it. The remaining portion of the long strand serves as an index, allowing random access via specific binding of fluorescent probes. This is the protection method during access, but its long-term preservation method is not yet well understood. Other methods utilize combinations of fluorescent bands or QR codes as indexes, abandoning the advantages of primers as indexes and reducing the number of indexable files.
[0005] Therefore, the full functionality of storage and random access in the current storage process is not effectively achieved. Relatively speaking, the advantageous storage method does not fully utilize the available information, and the advantageous indexing random access method is not deeply integrated with the complementary storage method. As of now, the existing technologies are not conducive to the effective real-time performance and integration of the entire storage process. Summary of the Invention
[0006] The purpose of this invention is to provide a DNA storage method and a data storage entity that enables random and lossless reading of data, thereby solving the problem that existing DNA storage methods cannot simultaneously achieve long-term preservation and random and lossless reading in systems with extremely high file counts.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0008] According to a first aspect of the present invention, a DNA storage method enabling random and lossless data reading is provided, comprising the following steps: S1, preparation of a DNA double-stranded structure: encoding a binary file into a DNA base sequence using fountain codes, determining the base sequences of primers and probes, preparing the DNA base sequence, primers, and probes, amplifying the DNA base sequence using primers to obtain a DNA double-stranded structure modified with NH2 at the 5' end; S2, preparation of a data storage entity: reacting the DNA double-stranded structure prepared in step S1 with silica microspheres or magnetic beads, and then encapsulating it with a hydrogel material to form a storage entity. A data storage entity is prepared; S3, random and non-destructive reading of DNA data: the hydrogel material of the data storage entity is removed, and fluorescent probes are used to specifically bind to the long chain head for fluorescence screening to separate the file to be read. Then, PCR is performed using primers to obtain the base sequence of the read file, and sequencing is used to obtain the DNA base sequence to be read; S4, recovery of the data storage entity: since the original DNA base sequence is still preserved on silica microspheres or magnetic beads, it is encapsulated again with hydrogel material and recovered to the file system to ensure that the data information is fixed in place and not lost.
[0009] In step S1, the prepared DNA double-stranded structure consists of a short strand paired with a long strand, and is divided into a fluorescent probe single-stranded segment, a primer index double-stranded segment, and a data carrier double-stranded segment. The fluorescent probe single-stranded segment is located at the 3' end of the long strand and is used for fluorescent sorting to achieve searching. The primer index double-stranded segment is used to achieve primer indexing. The data carrier double-stranded segment is used to complete random access. The NH2 modified at the 5' end of the long strand is used to fix the DNA double-stranded structure onto the silica microspheres or magnetic beads.
[0010] In step S2, functional groups that can react with amino groups to form chemical bonds are pre-grown on the silica microspheres or magnetic beads.
[0011] Preferably, in steps S2 and S4, the hydrogel material can be selected from: agarose, polyacrylamide, or polyethylene glycol.
[0012] Preferably, in steps S2 and S4, the encapsulation of the hydrogel material can be achieved by jet microfluidics or vortex centrifugation purification.
[0013] Preferably, in step S3, the hydrogel material on the outer layer of the data information storage entity can be removed by heating, chemical dissolution, or photoresponsive dissolution.
[0014] Preferably, in step S3, fluorescence screening, separation of the required read files, further PCR, and recovery of the data storage entity can all be achieved using a microfluidic device.
[0015] According to a second aspect of the present invention, a data information storage entity is provided, comprising the following three parts: a silica microsphere or magnetic bead located at the center; a plurality of DNA double-stranded structures fixed on the surface of the silica microsphere or magnetic bead; and a hydrogel material encapsulating the silica microsphere or magnetic bead and the DNA double-stranded structures therein; wherein the DNA double-stranded structure is composed of a short strand paired with a long strand, and is divided into a fluorescent probe single-stranded segment, a primer index double-stranded segment, and a data carrier double-stranded segment; the fluorescent probe single-stranded segment is located at the 3' end of the long strand and is used for fluorescent sorting to achieve searching; the primer index double-stranded segment is used to achieve primer indexing; the data carrier double-stranded segment is used to complete random access; and the 5' end of the long strand is modified with NH2 to fix the DNA double-stranded structure onto the silica microsphere or magnetic bead.
[0016] The silica microspheres or magnetic beads are pre-grown with functional groups that can react with amino groups to form chemical bonds. The hydrogel material can be selected from: agarose, polyacrylamide, or polyethylene glycol.
[0017] The encapsulation of the hydrogel material can be achieved by jet microfluidics or vortex centrifugation purification.
[0018] The inventiveness of this invention lies primarily in its novel method for random access and long-term DNA preservation based on primer indexing of an extremely high number of data strands. This method involves immobilizing the DNA double-stranded structure onto silica microspheres or magnetic beads, encapsulating and protecting it with hydrogel materials, and performing random access using PCR. First, the method chemically binds the DNA double-stranded structure to silica microspheres or magnetic beads. Then, it encapsulates the data storage entity with hydrogel materials, which, together with the carefully designed DNA double strands, complete the functions of preservation and random access. During random access, the hydrogel material is first removed, followed by direct PCR amplification using specific primers. The amplified DNA is then washed, collected, and sequenced to complete the random access. Since the immobilized effective data strands are preserved in situ, the long strands containing all the information are not lost, thus this invention successfully achieves lossless reading.
[0019] Secondly, the inventiveness of this invention lies in the first-ever fabrication of a data storage entity, which comprises three parts: silica microspheres or magnetic beads, a DNA double-stranded structure, and a hydrogel material. The silica microspheres or magnetic beads are used to fix the DNA double-stranded structure at the center, the DNA double-stranded structure is used to store data and serve as an index, and the hydrogel material provides outer protection. According to this data storage entity, directly connecting the DNA data strand to the silica microspheres results in higher information utilization, and amplification and recovery are simpler. The DNA double-stranded structure enables random access, while the outer hydrogel material provides long-term preservation. This allows for random reading in file systems with extremely high file counts, and the original data strand remains in place after data reading, ensuring that the original data is not lost.
[0020] In summary, the DNA storage method and data storage entity provided by the present invention, which enable random and lossless reading of data, not only utilize the DNA double-stranded structure to achieve random access, but also utilize the outer layer of hydrogel material to give it long-term preservation function. This enables random reading in a file system with an extremely high number of files and ensures that the original data is not lost. Therefore, the present invention has broad application prospects in the field of DNA storage. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of a DNA storage method that enables random and lossless data reading according to the present invention.
[0022] Figure 2 This is a schematic diagram of the structure of a DNA double strand prepared according to the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of a data information storage entity provided by the present invention;
[0024] Figure 4 A schematic diagram of the encapsulation process for hydrogel materials. Detailed Implementation
[0025] The present invention will be further described below with reference to specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0026] This invention provides a DNA storage method that enables random and lossless data reading, the overall process of which is as follows: Figure 1 As shown, this method utilizes double-stranded DNA for random access and a hydrogel material for long-term preservation. The method specifically includes the following steps:
[0027] S1: Preparation of DNA double-stranded structure
[0028] This step includes: encoding a binary file into a DNA base sequence using fountain codes; determining the base sequences of primers and probes; purchasing or preparing PCR primers, probes, and DNA base sequences, wherein the 5' end of the primers is modified with NH2; and amplifying the DNA base sequence using the primers to obtain a double-stranded DNA structure with NH2 modified at the 5' end. It should be understood that encoding a binary file into a DNA base sequence using fountain codes is a conventional technique well-known to those skilled in the art.
[0029] The DNA double-stranded structure prepared in step S1 is as follows: Figure 2 As shown, it consists of a pair of short and long chains, divided into a fluorescent probe single-stranded segment, a primer index double-stranded segment, and a data carrier double-stranded segment. The fluorescent probe single-stranded segment is located at the 3' end of the long chain and is used for fluorescent sorting to achieve searching. The primer index double-stranded segment is used for primer indexing, and the data carrier double-stranded segment is used to complete random access. The 5' end of the long chain is modified with NH2 to facilitate subsequent immobilization and attachment to silica microspheres or magnetic beads.
[0030] S2: Preparation of Data Information Storage Entities
[0031] This step includes: providing silica microspheres or magnetic beads, on which functional groups, such as carboxyl groups, that can react with amino groups to form chemical bonds are pre-grown; reacting the DNA double-stranded structure prepared in step S1 with the silica microspheres or magnetic beads, thereby immobilizing the DNA onto the microspheres or magnetic beads through the reaction of amino groups with specific functional groups; and encapsulating the DNA double-stranded structure with a hydrogel material to obtain a product such as... Figure 3 The data storage entity shown. It should be understood that the hydrogel material can be agarose or other substances that can form a gel.
[0032] According to a preferred embodiment of the present invention, agarose is used as the hydrogel material. Microspheres are first added to a warm agarose solution (2% wt, Sigma-Aldrich, Cat#A9414) to a final concentration of approximately 200 microspheres / μL. The microsphere suspension is placed on a heating block (70°C) to prevent gelation and is thoroughly vortexed before emulsification. Encapsulation of this hydrogel material can be achieved via jet microfluidics or purification via vortex centrifugation, such as... Figure 4 As shown.
[0033] 1) Jetting Microfluidics: To use jet microfluidics to emulsify the suspension, it was transferred to a 1 mL syringe. The syringe and another 2 mL syringe containing the oil phase were further connected to the microfluidic flow focusing device via plastic tubing. During emulsification, the syringe and device were kept warm by an indoor hot air heater positioned 20 cm away. The flow rates of the aqueous and oil phases were set to 2.6 mL / h and 26 mL / h, respectively, to produce a stable jet. The resulting droplets were collected in centrifuge tubes placed on ice and further incubated at 4 °C for 30 minutes to ensure microgel coagulation.
[0034] 2) Vortex centrifugation purification method: To use vortexing for emulsification, 500 μL of the prepared bead suspension was transferred into a 1.5 mL centrifuge tube pre-filled with 500 μL of oil. The tube was then fixed on a vortex mixer (DLAB, MX-S) and vortexed at 2300 rpm for 10 minutes. The resulting emulsion was then kept at 4°C for 30 minutes.
[0035] After encapsulation using the above method, the agarose forms a solidified gel that protects the inner DNA double strands.
[0036] According to another preferred embodiment of the invention, if the heating procedure for melting the agarose gel is not required, the agarose can be replaced with a covalently formed and biodegradable hydrogel, such as polyacrylamide (PAAM) and polyethylene glycol (PEG), so that the gel layer can be chemically or photoresponsively dissolved to protect the data information.
[0037] S3: Random, lossless reading of DNA data
[0038] When randomly reading files, the encapsulating hydrogel material needs to be removed first. Then, fluorescent probes are used to specifically bind to the 3' end of the long chain. After fluorescence screening, the files to be read are separated, and PCR is performed. The stored file data can be recovered by sequencing and decoding.
[0039] According to a preferred embodiment of the invention, the agarose gel layer can be melted by simply heating the droplets to 60°C without affecting the stability of the droplets.
[0040] S4: Recycling of data information storage entities
[0041] After the random and lossless reading of DNA data is completed in step S3, the original DNA base sequence is still preserved on the silica microspheres. The data information is protected by encapsulating the hydrogel material again and is completely and losslessly recycled to the file system.
[0042] Therefore, the data storage entity prepared according to the present invention can fully perform the functions of information protection and random access, and the microspheres are small in size, simple to operate, and the sugar-protected encapsulation is simple and reliable to unseal.
[0043] Preferably, the present invention can utilize microfluidic devices to achieve integrated and miniaturized encapsulation, preservation, unsealing, fluorescence sorting, and further PCR and recovery of hydrogel materials.
[0044] Furthermore, because the DNA data strands are immobilized on the microspheres, a sorting-free method can be used to directly perform random access PCR readings using primer specificity. In this way, other data strands that conflict with the primers in this reading file are also amplified. However, because the immobilized valid data strands are preserved in situ, at least the long strands containing all the information are not lost. The amplified DNA strands are washed away for sequencing, and the influence of short strand information can be removed by the decoding algorithm, and the file data can be completely recovered.
[0045] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention. Various variations can be made to the above embodiments of the present invention. All simple and equivalent changes and modifications made in accordance with the claims and description of this application fall within the protection scope of the claims of this patent. All aspects not described in detail in this invention are conventional technical content.
Claims
1. A DNA storage method that enables data random lossless read, characterized in that, The method comprises the following steps: S1, preparation of DNA double-stranded structure: encoding a binary file into a DNA base sequence by fountain code, determining the base sequence of primers and probes, preparing the DNA base sequence, primers and probes, amplifying the DNA base sequence by primers to obtain a DNA double-stranded structure with NH2 modified at the 5' end; S2, preparation of data information storage entity: reacting the DNA double-stranded structure prepared in step S1 with silica microspheres or magnetic beads, and then encapsulating by hydrogel material to prepare a data information storage entity; S3, random and lossless reading of DNA data: removing the hydrogel material outside the data information storage entity, using a fluorescent probe to specifically bind to the head of a long chain, and then performing fluorescence screening to separate the required reading file, and then using primers to perform PCR to obtain the base sequence of the reading file, and sequencing to obtain the required reading DNA base sequence; S4, recycling of data information storage entity: since the original DNA base sequence is still saved on the silica microspheres or magnetic beads, it is encapsulated by hydrogel material again and recycled to the file system to ensure that the data information is fixed in place and not lost; In step S1, the DNA double-stranded structure prepared comprises a short chain and a long chain, which are paired and divided into a fluorescent probe single-stranded segment, a primer index double-stranded segment and a data carrier double-stranded segment. The fluorescent probe single-stranded segment is located at the 3' end of the long chain and is used for fluorescence sorting to achieve searching. The primer index double-stranded segment is used to realize primer indexing. The data carrier double-stranded segment is used to complete random access. The NH2 modified at the 5' end of the long chain is used to fix the DNA double-stranded structure to the silica microspheres or magnetic beads. In step S3, the hydrogel material outside the data information storage entity can be removed by heating dissolution, chemical dissolution or light-responsive dissolution. In steps S2 and S4, the encapsulation by hydrogel material can be realized by jet microfluidic method or vortex centrifugal purification method.
2. The DNA storage method of claim 1, wherein, In step S2, the silica microspheres or magnetic beads have a functional group pre-grown thereon which can react with amino group to form a chemical bond.
3. The DNA storage method of claim 1, wherein, In steps S2 and S4, the hydrogel material can be selected from agarose, polyacrylamide and polyethylene glycol.
4. The DNA storage method of claim 1, wherein, In step S3, fluorescence screening, separation of the required reading file, further PCR and recycling of the data information storage entity can be realized by a microfluidic device.
5. A data information storage entity using the DNA storage method according to any one of claims 1-4 for data information storage, characterized in that, It comprises the following three parts: Silica microspheres or magnetic beads located in the center; A plurality of DNA double-stranded structures fixed on the surface of the silica microspheres or magnetic beads; Hydrogel material encapsulating the silica microspheres or magnetic beads and the DNA double-stranded structures; The DNA double-stranded structure comprises a short chain and a long chain, which are paired and divided into a fluorescent probe single-stranded segment, a primer index double-stranded segment and a data carrier double-stranded segment. The fluorescent probe single-stranded segment is located at the 3' end of the long chain and is used for fluorescence sorting to achieve searching. The primer index double-stranded segment is used to realize primer indexing. The data carrier double-stranded segment is used to complete random access. The long chain is modified with NH2 at the 5' end, which is used to fix the DNA double-stranded structure to the silica microspheres or magnetic beads.
6. The data information storage entity of claim 5, wherein, The silica microspheres or magnetic beads are pre-grown with functional groups that can react with amino groups to form chemical bonds, and the hydrogel material can be selected from the group consisting of agarose, polyacrylamide, and polyethylene glycol.
7. The data information storage entity of claim 5, wherein, The encapsulation of the hydrogel material can be achieved by a jet microfluidic method or a vortex centrifugal purification method.
Citation Information
Patent Citations
Magnetic DNA hydrogel and preparation method thereof
CN110272982A
Information storage method based on DNA short chain hybridization
CN112768003A