An information storage technique based on Z-DNA
By utilizing Z-DNA structure and nanopore sequencing technology, the problem of easy replication of DNA information storage in existing technologies has been solved, achieving simple and fast information writing and highly secure key protection, making it suitable for the storage and retrieval of sensitive information.
Patent Information
- Application Number
- CN202211067651.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-09-01
AI Technical Summary
Existing data copy protection technologies rely on cryptographic algorithms or hardware layer protection for media reading, which cannot fundamentally solve the data security problem of sensitive information itself. Furthermore, DNA information storage is complex and time-consuming in the writing process and is easily copied by PCR.
Information is stored using a Z-DNA structure. The properties of Z and T bases are utilized to erase information during PCR. Information is then read using nanopore sequencing technology. By splitting the key information and storing it in multiple DNA libraries and automatically updating it, security is improved.
It significantly enhances information security, the writing process is simple and fast, the cost is low, and it is difficult to copy and read via PCR, thus improving the protection strength of key information.
Smart Images

Figure BDA0003828525130000141 
Figure BDA0003828525130000151 
Figure BDA0003828525130000161
Abstract
Description
Invention Field
[0001] This application generally relates to the field of bioinformatics, and more specifically, to a sensitive information protection and storage technology. Background of the Invention
[0002] In the digital information age, information security is of paramount importance. While mature encryption algorithms can theoretically guarantee data security, key theft frequently occurs in practice, threatening information security. Traditional silicon-based storage is simple to read and write, making it easily stolen through copying. DNA information storage offers some confidentiality, but its writing process is complex and time-consuming, and it is easily copied and read using polymerase chain reaction (PCR), posing significant security risks.
[0003] Existing data copy protection technologies all rely on cryptographic principles or hardware-level access control for media reading. Each data reading process requires a correct password or fingerprint to authorize legitimate data access. While protecting data, these technologies also generate new information, such as keys, thus failing to fundamentally solve the data security problem of sensitive information itself. Invention Overview
[0004] To address the aforementioned technical problems, the inventors of this application store relevant information, such as key information, using Z-DNA structural information, which is then erased during the PCR process, significantly enhancing information security. Furthermore, the writing process is simple, fast, and low-cost.
[0005] In a first aspect, this application provides an information protection storage method, comprising: storing the information as a DNA library using one or more DNA fragments, wherein the information includes multiple bits.
[0006] Each of the one or more DNA fragments includes one or more information recording units, each of the information recording units including an index portion and a data recording portion.
[0007] The data recording portion may contain a Z base or not contain a Z base, and the data recording portion containing a Z base is set to correspond to information bit 1, while the DNA data recording portion not containing a Z base is set to correspond to information bit 0, or vice versa.
[0008] In some implementations, one or more of the information recording units may be absent. This absence can be used as an additional information recording pattern, which is set to correspond to one bit in a ternary information bit, forming a ternary information bit recording pattern with the presence of Z-bases and the absence of Z-bases.
[0009] In a second aspect, the present application provides a method for splitting and storing and updating key information, comprising:
[0010] storing the key information in the form of a plurality of DNA libraries;
[0011] the splitting method comprises splitting of different information segments, or splitting based on differences or calculation results between different segments; wherein the key information is generated by superimposed calculation of the information of the plurality of DNA libraries, each of which stores a plurality of password bits;
[0012] a part of the plurality of DNA libraries is stored in the form of a physical or virtual server, and another part of the plurality of DNA libraries is stored in the user end; and
[0013] the key of the user end and the key of the server end are invalidated immediately after use, and the key of the server end and the key information of the user end are automatically updated, so as to update the key information without re-encryption.
[0014] In a third aspect, the present application provides a method for reading information, comprising: sequencing one or more DNA fragments in a DNA library storing the information, the information comprising a plurality of password bits,
[0015] wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprises an index part and a data recording part,
[0016] wherein the data recording part contains Z base or does not contain Z base, and wherein the data recording part containing Z base is set to correspond to information bit 1, and the DNA data recording part not containing Z base is set to correspond to information bit 0, or vice versa.
[0017] In some embodiments, one or more of the information recording units can be absent, which can be set as an additional information recording mode corresponding to one of the ternary information bits, forming a ternary information bit recording mode with Z base and Z base.
[0018] In a fourth aspect, the present application provides the use of DNA fragments in the preparation of a DNA library for storing information, the information comprising a plurality of password bits,
[0019] wherein the DNA fragments comprise one or more DNA fragments,
[0020] wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprises an index part and a data recording part,
[0021] wherein the data record part comprises or does not comprise a Z-base, and wherein a data record part comprising a Z-base is set to correspond to an information bit 1, and a DNA data record part not comprising a Z-base is set to correspond to an information bit 0, or vice versa.
[0022] In some embodiments, one or more of the information record units can be absent, which absence can be set to correspond to one of the ternary information bits, forming a ternary information bit recording pattern with Z-base containing and Z-base not containing.
[0023] In a fifth aspect, the present application provides a use of a DNA library for storing information, the information comprising multi-bit password bits,
[0024] wherein the DNA library comprises one or more DNA fragments,
[0025] wherein each of the one or more DNA fragments comprises one or more information record units, each of the information record units comprising an index part and a data record part,
[0026] wherein the data record part comprises or does not comprise a Z-base, and wherein a data record part comprising a Z-base is set to correspond to an information bit 1, and a DNA data record part not comprising a Z-base is set to correspond to an information bit 0, or vice versa.
[0027] In some embodiments, one or more of the information record units can be absent, which absence can be set to correspond to one of the ternary information bits, forming a ternary information bit recording pattern with Z-base containing and Z-base not containing.
[0028] In a sixth aspect, the present application provides an information storage system comprising a DNA library obtained by the method of any one of the first to fourth aspects. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 XmaI enzyme digestion and T4 DNA ligase ligation results of dA-DNA and dZ-DNA are shown, wherein lane 1: A-XmaI-LR; lane 2: XmaI enzyme digestion results of A-XmaI-LR; lane 3: Z-XmaI-LR; lane 4: XmaI enzyme digestion results of Z-XmaI-LR; lanes 5-7: T4 DNA ligase ligation results of A-XmaI-L / A-XmaI-R, A-XmaI-L / Z-XmaI-R and Z-XmaI-L / Z-XmaI-R after XmaI enzyme digestion.
[0030] Figure 2 Bgl I cleavage and T4 DNA ligase ligation results of dA-DNA and dZ-DNA are shown, where lane 1 and 5: A-Bgl I-LR; lane 2 and 6: Bgl I cleavage results of A-Bgl I-LR 1 h and 5 h; lane 3, 7 and 9: Z-Bgl I-LR; lane 4 and 8: Bgl I cleavage results of Z-Bgl I-LR 1 h and 5 h; lane 10-12: T4 DNA ligase ligation results after Bgl I cleavage of A-Bgl I-L / A-Bgl I-R, A-Bgl I-L / Z-Bgl I-R and Z-Bgl I-L / Z-Bgl I-R.
[0031] Figure 3 Electrophoresis results of 32 upstream DNAs and 2 downstream DNAs are shown, where lane 1-32: 32 upstream dA-DNAs; lane 33: downstream dA-DNA; lane 34: downstream dZ-DNA.
[0032] Figure 4 Electrophoresis results of 32 obtained dA-A DNA fragments are shown.
[0033] Figure 5 Electrophoresis results of 26 obtained dA-Z DNA fragments are shown.
[0034] Figure 6 Electrophoresis results of three constructed DNA libraries are shown. DETAILED DESCRIPTION
[0035] The terms used in this application have the meanings commonly understood by a person of ordinary skill in the art, unless otherwise indicated.
[0036] Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation; nucleotides can be represented by commonly accepted single-letter codes.
[0037] Current copy protection technologies are based on cryptography algorithms or medium reading hardware layer protection principles, rather than underlying (medium layer) copy protection technologies. Current data disc protection technologies, such as CD-Cops and Dummy files protection mechanisms, respectively add a protective cover outside the installation program, verify the password (usually 8 bits) during installation; and modify the ISO code related to file size to create a large fake file, making the total capacity of the disc show as the maximum 2 GB, causing the common 650 MB, 700 MB CDR to be unable to backup the illusion. Current U disk data protection technologies include, but are not limited to, key encryption type, fingerprint encryption type, software encryption type, hardware encryption type, and software and hardware combined encryption type.
[0038] Key encryption type
[0039] The professionally produced copy-protected U disk can be built-in with physical digital keys, and the user can manually input the pre-set password to realize the encryption and decryption functions. The password of the key encryption type copy-protected U disk is saved in the encryption chip. In addition, the key encryption type copy-protected U disk can realize data encryption and decryption without a computer, and support setting multiple accounts and multiple permissions.
[0040] Fingerprint encryption type
[0041] The technically precise copy-protected U disk can be built-in with a fingerprint collector. Since everyone's fingerprint is unique and unchangeable for a lifetime, the fingerprint encryption type copy-protected U disk can correspond the user's fingerprint according to the uniqueness and stability of the fingerprint, thereby verifying the user's true identity, and thus realizing the encryption and decryption functions of file data through this method.
[0042] Software encryption type
[0043] According to the understanding of the product, the software encryption type copy-protected U disk refers to the encryption of the U disk content through built-in and attached software. Generally, it uses ASP encryption, so the software encryption type copy-protected U disk can partially avoid the disadvantage that the original U disk files can be read through password cracking tools on other PCB boards.
[0044] Hardware encryption type
[0045] The decryption of such technology when reading files is carried out in the U disk, and the USB port transmits decrypted files. Using USB packet capture tools can also easily obtain the disk files. Hardware encryption technology is not suitable for the scene of "copy protection" in fact.
[0046] Software and hardware combined encryption type
[0047] This kind of technology uses a unique VNAS encryption scheme, which can realize active authorization control and achieve security without vulnerabilities. It does not use any non-standard Hacker technology, and there is no risk of being mistakenly killed by antivirus software. The file is highly encrypted and almost impossible to decrypt.
[0048] However, the above protection methods all need password / fingerprint as the key for information reading, that is, these methods cannot solve the protection problem of the key itself.
[0049] DNA modifications differ in form and function but generally do not change Watson-Crick base pairing. Diaminopurine (Z base) is a unique exception because in cyanophage it completely replaces adenine and forms three hydrogen bonds with thymine. During in vitro DNA synthesis, deoxyribonucleotides with Z base can be added to in vitro amplification systems such as PCR systems, thereby replacing A and T pairing in conventional DNA to incorporate Z base in the synthesized DNA.
[0050] In the present application, DNA containing Z base is referred to as Z-DNA, and conventional DNA not containing Z base is referred to as A-DNA for distinction. The inventors of the present application ingeniously developed a password storage system based on the property of T being able to pair with both A and Z, which is not only safe and reliable, but also not easy to be stolen, because it is difficult to amplify DNA containing Z base through in vitro amplification systems without prior knowledge of the specific sequence of the DNA, and thus it is impossible to decipher the information implied in the DNA sequence. Therefore, the present applicant pioneered the application of biological means to the field of electronics, specifically, to the field of informatics.
[0051] Specifically, the present application provides the following technical solutions:
[0052] In a first aspect, the present application provides an information protection storage method, comprising: storing the information as a DNA library in the form of one or more DNA fragments, the information comprising multi-bit bits,
[0053] wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprising an index part and a data recording part,
[0054] wherein the data recording part contains Z base or does not contain Z base, and wherein the data recording part containing Z base is set to correspond to information bit 1, and the DNA data recording part not containing Z base is set to correspond to information bit 0, or vice versa.
[0055] In some embodiments, one or more of the information recording units can be absent, which can be set as an additional information recording mode corresponding to one bit in ternary information bits, forming a ternary information bit recording mode with Z base and Z base-free.
[0056] In a second aspect, the present application provides a key information splitting storage updating method, comprising:
[0057] splitting and saving the key information in the form of multiple DNA libraries;
[0058] The method of splitting includes splitting of different information segments, or splitting based on differences or calculation results between different segments; wherein the key information is generated by superimposed calculation of information of the plurality of DNA libraries, each of which stores multi-bit password bits;
[0059] Part of the plurality of DNA libraries is stored in the form of physical or virtual on the server side, and another part of the plurality of DNA libraries is stored on the user side; and
[0060] The key on the user side and the key on the server side are invalidated immediately after use, and the key information on the server side and the key information on the user side are automatically updated, so as to update the key information without re-encryption.
[0061] In some embodiments, the key information is encryption key information for digital assets.
[0062] Here, the key composed of 32-bit password bits is taken as an example for illustration.
[0063] Key A (binary): 10101110001111011000010100110110
[0064] Key B (binary): 10011101011111101010101101110101
[0065] Key A and Key B are stored as a first DNA library and a second DNA library respectively, and DNA containing Z bases in the DNA library represents password bit 1, and DNA not containing Z bases represents password bit 0, in other words, conventional DNA containing A bases represents password bit 0. Those skilled in the art can understand that the first DNA library can represent Key A, and the second DNA library can represent Key B, and by sequencing the first DNA library and the second DNA library, such as by nanopore sequencing, Key A and Key B can be obtained. However, if the DNA library is amplified, since Z bases and A bases will pair with T bases, the amplified product will not be able to distinguish A bases and Z bases, and thus the specific information of Key A and Key B cannot be obtained.
[0066] A person skilled in the art can generate a new key according to a certain encoding rule based on the superposition calculation of the key A and the key B. For example, if the DNA indicated by the first index part in the first DNA library contains Z base, and the DNA indicated by the same first index part in the second DNA library also contains Z base, then the DNA indicated by the first index part can represent the code bit 0; if the DNA indicated by the first index part in the first DNA library contains Z base, and the DNA indicated by the same first index part in the second DNA library does not contain Z base, then the DNA indicated by the first index part can represent the code bit 1; if the DNA indicated by the first index part in the first DNA library does not contain Z base, and the DNA indicated by the same first index part in the second DNA library does not contain Z base, then the DNA indicated by the first index part can represent the code bit 0; if the DNA indicated by the first index part in the first DNA library does not contain Z base, and the DNA indicated by the same first index part in the second DNA library contains Z base, then the DNA indicated by the first index part can represent the code bit 1. Through such superposition algorithm, the key A and the key B can generate a new key C as follows:
[0067] Key C (binary): 00110011010000110010111001000011.
[0068] A new key C can be generated from the key A and the key B through different superposition algorithms. In this case, the first DNA library and the second DNA library can be stored in different places, which further improves the security of information storage.
[0069] Although a new key is obtained through the superposition calculation of two keys here, a person skilled in the art will understand that more keys can be superimposed to generate a new key different from the original key through a certain encoding rule, thereby further improving the security of key storage.
[0070] In some embodiments, a special application scenario of the information storage method of the present application is the protection of digital assets (NFT). The NFT digital assets are encrypted by a specific password. Then the specific password is divided into two passwords, one is an online password and the other is an offline password. When the digital assets are traded, the online and offline passwords are updated synchronously, but the specific password remains unchanged, thereby solving the problem of the original owner of the digital assets holding the decryption password after the transaction.
[0071] In particular, the information is information for a digital asset, such as key information, and the first DNA library is stored on a server side, and the second DNA library is stored on a client side, and after each use of the information, the first DNA library is automatically updated, and the second DNA library is updated accordingly, so as to achieve the effect of updating the information without re-encrypting the digital asset. The information, such as key mode, is fixed and does not change with the update of the first DNA library and the second DNA library.
[0072] In a third aspect, the present application provides an information reading method, comprising: sequencing one or more DNA fragments in a DNA library storing the information, the information comprising a plurality of bits of information,
[0073] wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprising an index part and a data recording part,
[0074] wherein the data recording part contains Z bases or does not contain Z bases, and wherein the data recording part containing Z bases is set to correspond to information bit 1, and the DNA data recording part not containing Z bases is set to correspond to information bit 0, or vice versa.
[0075] In some embodiments, one or more of the information recording units can be absent, and this absence can be set as an additional information recording mode, corresponding to one bit of ternary information bits, forming a ternary information bit recording mode with Z bases and Z-free bases.
[0076] In some specific embodiments, the sequencing can be performed by a method selected from the group consisting of massively parallel sequencing, ion semiconductor sequencing and nanopore sequencing.
[0077] Preferably, the sequencing is performed by nanopore sequencing. Using nanopore sequencing, a single DNA or RNA molecule can be sequenced without the need for PCR amplification or chemical labeling of the sample. In any of the previously developed sequencing methods, at least one of the above steps is necessary. Nanopore sequencing has the potential to provide relatively low-cost genotyping, high test migration rate and rapid processing of samples, and can display results in real time.
[0078] In some embodiments of the method of any of the above aspects, the information can be any important private information that needs to be stored, such as a key.
[0079] In some embodiments of the method of any of the above aspects, the index portion is a DNA fragment having a specific sequence, preferably, the index portion does not contain Z-base. Those skilled in the art will understand that the index portion is used to distinguish various DNAs in the DNA library, and therefore can be located at any position of the DNA, for example, at the 5' end, the 3' end or the middle position.
[0080] In addition, since the index portion is used to distinguish various DNAs in the DNA library, any length of DNA fragment that can achieve this purpose can be used as the index portion. Preferably, the index portion can be a DNA fragment of at least 4 nucleotides in length, for example, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides or any length longer than 4 nucleotides.
[0081] In some embodiments of the method of any of the above aspects, the information includes n-bit cipher bits, wherein n is greater than or equal to 2. Preferably, n is greater than or equal to 16. Since the DNA library has a very large capacity, the information can include a very large number of cipher bits, and it is almost impossible to decipher the information.
[0082] In some embodiments of the method of any of the above aspects, the one or more DNAs each have a length of 20 bp-2000 bp, preferably 600 bp-1000 bp. The length and composition of DNA are not critical for information storage, and those skilled in the art can select a suitable length of DNA to produce a DNA library for storing information as needed.
[0083] In some embodiments of the method of any of the above aspects, one or more of the information recording units can be absent, which can be used as an additional information recording mode, set to correspond to one bit in the ternary information bit, forming a ternary information bit recording mode with Z-base and without Z-base, which can further increase the new cipher bits in the original binary key, thereby converting the binary key to a ternary key, further increasing the complexity of the information and bringing greater challenges to information deciphering.
[0084] In a fourth aspect, the present application provides use of a DNA fragment in the preparation of a DNA library for storing information, the information including a plurality of cipher bits,
[0085] wherein the DNA fragment includes one or more DNA fragments,
[0086] wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprising an index portion and a data recording portion,
[0087] wherein the data recording portion comprises or does not comprise a Z-base, and wherein a data recording portion comprising a Z-base is set to correspond to an information bit of 1, and a DNA data recording portion not comprising a Z-base is set to correspond to an information bit of 0, or vice versa.
[0088] In some embodiments, one or more of the information recording units can be absent, which absence can be set to correspond to one of the ternary information bits, forming a ternary information bit recording pattern with Z-base containing and Z-base not containing.
[0089] In a fifth aspect, the present application provides use of a DNA library for storing information, the information comprising a plurality of bits of cryptographic information,
[0090] wherein the DNA library comprises one or more DNA fragments,
[0091] wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprising an index portion and a data recording portion,
[0092] wherein the data recording portion comprises or does not comprise a Z-base, and wherein a data recording portion comprising a Z-base is set to correspond to an information bit of 1, and a DNA data recording portion not comprising a Z-base is set to correspond to an information bit of 0, or vice versa.
[0093] In some embodiments, one or more of the information recording units can be absent, which absence can be set to correspond to one of the ternary information bits, forming a ternary information bit recording pattern with Z-base containing and Z-base not containing.
[0094] In some embodiments of the use according to any of the above aspects, the information can be any important private information that needs to be stored, for example can be a cryptographic key.
[0095] In some embodiments of the use according to any of the above aspects, the index portion is a DNA fragment with a specific sequence, preferably the index portion does not contain a Z-base. It will be understood by the skilled person that the index portion is used to distinguish the various DNAs in the DNA library, and thus can be located anywhere in the DNA, for example at the 5’ end, the 3’ end or an intermediate position.
[0096] In addition, since the index part is used to distinguish various DNAs in the DNA library, any length of DNA fragment that can achieve this purpose can be used as the index part. Preferably, the index part can be a DNA fragment of at least 4 nucleotides in length, for example, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, or any length of DNA fragment longer than 4 nucleotides.
[0097] In some embodiments of the use of any of the above aspects, the information comprises n-bit cipher bits, wherein n is greater than or equal to 2. Preferably, n is greater than or equal to 16. Since the DNA library has a very large capacity, the information can include a very large number of cipher bits, and it is almost impossible to decipher the information.
[0098] In some embodiments of the use of any of the above aspects, the one or more DNAs each have a length of 20 bp to 2000 bp, preferably 600 bp to 1000 bp. The length and composition of the DNA are not critical for information storage, and one skilled in the art can select a suitable length of DNA to produce a DNA library for storing information as needed.
[0099] In some embodiments of the use of any of the above aspects, one or more of the information recording units can be absent, which absence can be set as an additional information recording mode corresponding to one bit in the ternary information bit, forming a ternary information bit recording mode with Z-base and without Z-base. This can further increase the new cipher bits in the original binary key, thereby converting the binary key to a ternary key, further increasing the complexity of the information and bringing greater challenges to information deciphering.
[0100] In a sixth aspect, the present application provides an information storage system comprising the DNA library obtained by the method of any of the first aspect to the second aspect.
[0101] In specific embodiments, the information can be any important private information that needs to be stored, for example, it can be a key. Examples
[0102] The following examples are illustrative only and are not intended to limit the scope of the embodiments of the present application or the scope of the claims appended hereto.
[0103] Materials and Methods
[0104] Experimental Materials
[0105] dZTP used in the following examples was purchased from Trilink (California, USA). Other nucleotides (dTTP, dGTP, dCTP) used were purchased from New England Biolabs (NEB). Q5 High-Fidelity DNA Polymerase, T4 DNA Ligase, and restriction enzymes Xmal and Bgll used in the experiment were all purchased from NEB. DNA Clean & Concentrator-5 was purchased from Zymo Research (California, USA). Bacterial genomic DNA extraction kit and plasmid extraction kit were purchased from Tiangen Biotech (Beijing) Co., Ltd. ClonExpress II One Step Cloning Kit was purchased from Vazyme (Jiangsu, China). Magnetic bead DNA purification and recovery kit Agencourt AMPure XP was purchased from Beckman Coulter (California, USA).
[0106] Extraction of E. coli BL21(DE3) genomic DNA
[0107] The E. coli BL21(DE3) genome used in the following examples was extracted from 4 mL of overnight cultured bacteria solution by bacterial genomic DNA extraction kit and used as the template for PCR experiments.
[0108] PCR amplification of dA-DNA and dZ-DNA
[0109] All PCR amplification experiments in the following examples were catalyzed by Q5 High-Fidelity DNA Polymerase. dZ-DNA was obtained by mixing dZTP with other nucleotides (dTTP, dGTP, dCTP) into the reaction system. dA-DNA was obtained by mixing dATP with other nucleotides (dTTP, dGTP, dCTP) into the reaction system. DNA fragments Xmal-L (260 bp), Xmal-R (420 bp), and Xmal-LR (640 bp) were obtained by primers Xmal-L-for / Xmal-L-rev, Xmal-R-for / Xmal-R-rev, and Xmal-L-for / Xmal-R-rev, respectively. According to the difference between dA-DNA and dZ-DNA, we finally obtained 6 DNA fragments containing Xmal enzyme cutting site (5'-CCCGGG-3').
[0110] Restriction enzyme digestion and ligase ligation of dA-DNA and dZ-DNA
[0111] DNA fragments of 0.5 μg of A-XmaI-LR and Z-XmaI-LR were digested by restriction enzyme XmaI in 20 μL of enzyme digestion reaction system. The electrophoresis map shows that A-XmaI-LR and Z-XmaI-LR are both digested into two short fragments of DNA. By XmaI digestion, four DNA fragments containing sticky ends can be obtained: XmaI-L(A or Z) and XmaI-R(A or Z). By different combination ways (LA-RA, LA-RZ, LZ-RZ), these fragments are connected by T4 DNA ligase. The electrophoresis map shows that, because the sticky end site generated after XmaI digestion is a palindromic sequence, the fragments themselves will also be connected to each other, and there are three (L-L L-RR-R) connection cases( Figure 1 ).
[0112] Restriction enzyme BglI (5'-GCCNNNNNGGC-3') is selected to replace XmaI. The cutting site of BglI is located at the fourth nucleotide from both ends. In order to avoid the generation of palindromic sequences, we design the sticky end as 5'-GCA-3'. The BglI endonuclease recognition site does not contain Z, while the enzyme cutting site can contain Z. DNA fragments BglI-L (260 bp) and BglI-R (420 bp) are obtained by primers XmaI-L-for / BglI-L-rev and BglI-R-for / XmaI-R-rev, respectively. BglI-LR is obtained by overlap extension PCR (OE-PCR) of BglI-L and BglI-R. According to the difference between dA-DNA and dZ-DNA, we obtain six DNA fragments containing BglI enzyme cutting sites. DNA fragments of 0.5 μg of A-BglI-LR and Z-BglI-LR are digested by restriction enzyme BglI in 20 μL of enzyme digestion reaction system. The electrophoresis map shows that A-BglI-LR and Z-BglI-LR are both digested into two short fragments of DNA, but the efficiency of enzyme digestion of dZ-DNA is lower. After 1 h of enzyme digestion, A-BglI-LR fragments are completely digested into two short fragments, while Z-BglI-LR needs 5 h. By different combination ways (LA-RA, LA-RZ, LZ-RZ), these fragments are connected by T4 DNA ligase. The electrophoresis map shows that the two short fragments are successfully connected into a long fragment( Figure 2 ).
[0113] Example 1: Design of 32-bit password storage
[0114] Design of 32 sequences
[0115] In this example, we designed 32 DNA sequences. Each sequence was divided into an upstream (L, 220 bp) and a downstream (R, 480 bp) part. In the 32 sequences, the order of the bases in the upstream sequence was different for each sequence, serving as a characteristic to distinguish the 32 sequences. The upstream sequence was composed of dA-DNA only, while the downstream sequence was either composed of dA-DNA or dZ-DNA (see Tables 3 and 4 for the sequence information of the primers used for amplification and the final 32 sequences). The design was performed according to the design principle (Table 1), where 1 indicates a downstream sequence with dZ-DNA and 0 indicates a downstream sequence with dA-DNA. We designed 4 libraries, each containing different combinations of dA-A DNA and dA-Z DNA.
[0116] Table 1. Design principle for 32 dA-A DNA and dA-Z DNA
[0117]
[0118] Separate acquisition of upstream DNA and downstream DNA
[0119] When designing the primers, a restriction enzyme site for Bgll was added to the 3' end of the upstream sequence and the 5' end of the downstream sequence, and a barcode sequence (24 bp) for nanopore sequencing recognition was also added to the 5' end of the upstream sequence. The barcode of each sequence was different from each other, for subsequent data separation and processing of nanopore sequencing. The 32 upstream dA-DNA and 2 downstream dA-DNA and dZ-DNA (although the sequences are different, but are divided into 2 according to whether there is Z in the sequence) were obtained by Q5 high-fidelity DNA polymerase amplification with E. coli BL21(DE3) genome as template. Figure 3
[0120] BglI enzyme digestion and T4 DNA ligase ligation of 32 fragments
[0121] The enzyme digestion reaction system used for the dA-DNA and dZ-DNA fragments of the upstream (L) and downstream (R) has been listed in Table 2. Since the efficiency of Bgll digestion of dZ-DNA is low, the amount of dZ-DNA used is lower than that of dA-DNA under the same system. T4 DNA ligase was used to ligate the upstream L with the downstream dA-DNA R and dZ-DNA R. The total amount of DNA added in the 10 μL ligation reaction system was 200 ng, and the molar ratio of L and R fragments was 1:1. Considering that the 5' end of L contains a barcode for sequence recognition, the unligated L fragments may interfere with the results of nanopore sequencing, so an excess of R fragments was added to the ligation reaction system (Table 2).
[0122] Table 2 Bgll enzyme digestion and T4 DNA ligase reaction system
[0123]
[0124] Obtaining of dA-A DNA fragments
[0125] The concentration of DNA recovered from the ligation reaction system is usually very low. In order to meet the needs of sequencing, a large amount of DNA needs to be added to the ligation reaction system. At the same time, the DNA sample for nanopore sequencing cannot be destroyed by agarose gel electrophoresis and ultraviolet light. For dA-A DNA fragments, 32 dA-DNA L fragments and dA-DNA R fragments were connected by T4 DNA ligase, and the ligation product was purified and recovered using DNA Clean & Concentrator-5. The pUC19 plasmid was linearized by PCR to obtain a pUC19 linearized vector with corresponding fragment homologous arms. The vector and the fragment were connected by ClonExpress II OneStep Cloning Kit and transformed into Escherichia coli DH5α. The correct transformants were screened, and the required 32 plasmids were extracted from the correct transformants using a plasmid extraction kit. The 32 dA-A DNA fragments were amplified using the 32 plasmids as templates, and the obtained DNA fragments had no band interference and met the sequencing requirements Figure 4 ) in quality.
[0126] Obtaining of dA-Z DNA fragments
[0127] dA-Z DNA fragments were obtained by connecting dA-DNA L and dZ-DNA R from the ligation system. According to the design (Table 1), 26 dA-Z DNA fragments (except 2, 9, 18, 20, 25 and 29) need to be obtained. In order to remove as much unconnected upstream dA-DNA as possible, magnetic beads were used for DNA purification. According to the instructions of Agencourt AMPure XP, when the volume of AMPure is 0.6 times the volume of the sample, the recovered product is basically free of DNA fragments less than 300 bp in length. Therefore, 0.6x AMPure was used to purify 26 dA-Z DNA fragments from the ligation reaction system, and the electrophoretogram showed that it was basically free of unconnected upstream fragments Figure 5 ).
[0128] Construction of DNA library
[0129] According to different combination modes of 32 dA-A DNA and dA-Z DNA, we designed four libraries, and constructed three of them, Pass A, Pass B and Pass D (Table 1). The combination modes of dA-A DNA and dA-Z DNA in each library are different. When constructing the library, the mass of each DNA fragment added to the library is 250 ng, and then concentrated. The volume of the obtained DNA library is 50 μL, and the concentration is higher than 100 ng / μL. Take 1 μL from each DNA library for agarose gel electrophoresis experiment. Figure 6 The brightest band in the DNA Marker contains 120 ng of DNA.
[0130] Table 3 Sequence information of 32 sequences used in the examples (Z in Z-DNA sequence is also represented by A)
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138] Table 4 Sequence information of primers used to amplify target sequences in the examples
[0139]
[0140]
[0141]
[0142] The above describes exemplary embodiments of the various inventions of the present application, but those skilled in the art can modify or improve the exemplary embodiments of the present application described above without departing from the essence and scope of the present application, and the resulting variations or equivalent schemes also belong to the scope of the present application.
Claims
1. An information protection storage method, comprising: storing said information as a DNA library in the form of one or more DNA fragments, said information comprising multi-bit bits, wherein each of said one or more DNA fragments comprises one or more information recording units, each of said information recording units comprising an index portion and a data recording portion, wherein said data recording portion comprises or does not comprise a Z-base, and wherein a data recording portion comprising a Z-base is set to correspond to an information bit 1, and a DNA data recording portion not comprising a Z-base is set to correspond to an information bit 0, or vice versa, optionally, one or more of said information recording units is absent, such absence being set as an additional information recording mode, corresponding to one of the ternary information bits, forming a ternary information bit recording mode with the Z-base containing and Z-base not containing.
2. An information reading method comprising: sequencing one or more DNA fragments of a DNA library holding said information, said information comprising multi-bit code bits, wherein each of said one or more DNA fragments comprises one or more information recording units, each of said information recording units comprising an index portion and a data recording portion, wherein said data recording portion comprises or does not comprise a Z-base, and wherein a data recording portion comprising a Z-base is set to correspond to an information bit 1, and a DNA data recording portion not comprising a Z-base is set to correspond to an information bit 0, or vice versa, optionally, one or more of said information recording units is absent, such absence being set as an additional information recording mode, corresponding to one of the ternary information bits, forming a ternary information bit recording mode with the Z-base containing and Z-base not containing.
3. The method of claim 2, wherein said sequencing is performed by a method selected from the group consisting of massively parallel sequencing, ion semiconductor sequencing, and nanopore sequencing.
4. Use of a DNA fragment in the manufacture of a DNA library for storing information, said information comprising multi-bit code bits, wherein said DNA fragment comprises one or more DNA fragments, wherein each of said one or more DNA fragments comprises one or more information recording units, each of said information recording units comprising an index portion and a data recording portion, wherein said data recording portion comprises or does not comprise a Z-base, and wherein a data recording portion comprising a Z-base is set to correspond to an information bit 1, and a DNA data recording portion not comprising a Z-base is set to correspond to an information bit 0, or vice versa, optionally, one or more of said information recording units is absent, such absence being set as an additional information recording mode, corresponding to one of the ternary information bits, forming a ternary information bit recording mode with the Z-base containing and Z-base not containing.
5. Use of a DNA library for storing information, said information comprising multi-bit code bits, wherein said DNA library comprises one or more DNA fragments, wherein each of the one or more DNA fragments comprises one or more information recording units, each of the information recording units comprising an index portion and a data recording portion, wherein the data recording portion contains a Z-base or does not contain a Z-base, and wherein a data recording portion containing a Z-base is set to correspond to an information bit 1, and a DNA data recording portion not containing a Z-base is set to correspond to an information bit 0, or vice versa, optionally, one or more of the information recording units is absent, such absence being set to correspond to one of the ternary information bits, forming a ternary information bit recording pattern with Z-base containing and Z-base not containing.
6. The method or use of any one of claims 1 to 5, wherein the information is privacy information or a key.
7. The method or use of any one of claims 1 to 5, wherein the index portion is a DNA fragment having a specific sequence.
8. The method or use of claim 7, wherein the index portion is a DNA fragment of at least 4 nucleotides in length.
9. The method or use of claim 7, wherein the index portion does not contain a Z-base.
10. The method or use of claim 7, wherein the index portion is located at the 5’ end, 3’ end or in the middle of the DNA.
11. The method or use of any one of claims 1 to 5, wherein the information comprises n bits of password, wherein n is greater than or equal to 2.
12. The method or use of claim 11, wherein n is greater than or equal to 16.
13. The method or use of any one of claims 1 to 5, wherein each of the one or more DNA fragments is 20 bp - 2000 bp in length.
14. The method or use of claim 13, wherein each of the one or more DNA fragments is 600 bp - 1000 bp in length.
15. The method or use of any one of claims 1 to 5, wherein one or more of the information recording units is absent, such absence being set to correspond to one of the ternary information bits, forming a ternary information bit recording pattern with Z-base containing and Z-base not containing.
16. An information storage system comprising a DNA library obtained by the method of any one of claims 1 and 6-15.
Citation Information
Patent Citations
DNA storage hierarchical representation and interleaved coding method capable of containing artificial bases
CN110569974A
Data information storage method based on recombinant plasmid DNA molecules
CN114758703A