An advanced data encryption protection technology for DNA information storage
By introducing random error and breaking steps in the DNA information storage, combined with fountain code encoding and block encryption, the data susceptibility and easy-to-break data caused by DNA replication errors in the prior art is solved, and efficient and secure data storage and protection are achieved.
Patent Information
- Application Number
- CN202311006001.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-10
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-08-10
AI Technical Summary
The existing DNA information storage encryption strategy fails to fully utilize the replication error characteristics of DNA, resulting in data being easily damaged during storage and reading, and there is a risk of being cracked, especially when large-scale data storage and reading, where replication errors accumulate, resulting in information loss or misunderstanding.
By introducing random error and random DNA fragment breaking steps, combined with fountain coding and block encryption design, the complex multimolecular copy properties of DNA are used to increase the difficulty of decryption and improve security.
It improves the security and long-term stability of data, reduces the complexity and calculation amount of cracking, enhances the protection of encrypted information, and ensures the integrity and reliability of data.
Smart Images

Figure CN119479838B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of bioinformatics and synthetic biology and is an advanced data encryption protection technology for DNA information storage. Background Art
[0002] As the demand for data storage continues to grow, traditional electronic storage methods face challenges with capacity limitations and security risks. Therefore, finding new data storage and encryption protection technologies has become particularly important. DNA information storage, as an emerging storage technology, has become a research hotspot with its extremely high density and long-term data preservation capabilities. Furthermore, the characteristics of DNA storage give it unique advantages in terms of security and privacy protection. Data encryption technology for DNA storage can be implemented in a variety of ways. The following are existing strategies and methods:
[0003] Hybrid encryption: This method mixes the encrypted message with a large number of random DNA sequences. It's like hiding the real message within a large number of random DNA sequences. First, a sufficient number of random DNA sequences are generated. Then, the encrypted message is encoded into a DNA sequence and mixed with the randomly generated sequence. During decryption, only someone with the correct code (which could be a marker or identification sequence) can decipher the real message from the sequence.
[0004] Hidden information: This method involves hiding information in a specific region of DNA, known to the decryptor but unknown to everyone else. First, the information is encoded as a DNA sequence and then inserted into a specific location within a larger DNA sequence. During decryption, only someone who knows this location can extract the encrypted information.
[0005] Molecular encryption: This involves encrypting information by performing specific biochemical reactions on DNA (e.g., by DNA cutting or gene editing). First, the information is encoded into a DNA sequence. Then, this sequence is subjected to a specific biochemical reaction. This reaction alters part of the DNA sequence while leaving the rest intact. During decryption, only someone who knows the specific reaction can decrypt the encrypted information.
[0006] Specific coding strategies: This involves encoding information into DNA using a specific coding method. This may involve specific sequence design and encoding rules. During decryption, only those who know the coding strategy can decipher the information.
[0007] Sequence permutation: This method encrypts information by altering the order of DNA sequences. First, the information is encoded into a DNA sequence, and then this sequence is rearranged. During decryption, only someone who knows the original sequence's ordering can rearrange the sequence and decrypt the information.
[0008] Point mutations: Encrypting information by introducing a point mutation (changing a single base) at a specific location in the DNA sequence. First, the information is encoded into a DNA sequence, and then the point mutation is introduced at a specific location within that sequence. During decryption, only someone who knows the specific mutations can decrypt the encrypted information.
[0009] Multi-strand encryption: Information is encrypted by distributing it across multiple DNA strands. First, the information is divided into multiple components, each encoded as a separate DNA sequence. These sequences can be mixed together or stored separately. To decrypt the information, all the sequences must be read and reassembled in the correct order to decrypt the complete message.
[0010] Paired strand encryption: This technology leverages the complementary pairing principle of DNA to encrypt information by designing specific complementary strands. First, the information is encoded into a DNA sequence, and then a complementary strand is designed based on this sequence. This complementary strand serves as both an encryption code and a part of the information. During decryption, only those who know the complementary strand can decrypt the information.
[0011] Shape-based encryption (DNA origami): This method uses DNA origami technology to encrypt information by modifying the three-dimensional structure of DNA. First, the information is encoded into a DNA sequence, and then a specific three-dimensional DNA structure is designed based on this sequence. This structure can serve as both an encryption code and a part of the information. During decryption, only those who know the three-dimensional structure can decipher the information.
[0012] Location-based encryption (DNA chip encryption): This method involves placing encrypted information at a specific location on a DNA chip. First, the information is encoded as a DNA sequence and then placed at a specific location on the DNA chip. This location serves as both an encryption key and a part of the information. During decryption, only those who know the location can decrypt the information.
[0013] Sequence-based encryption (DNA hybridization): This method encrypts information by designing a specific complementary sequence. First, the information is encoded as a DNA sequence, and then a complementary sequence is designed based on this sequence. This complementary sequence serves as both an encryption code and a part of the information. During decryption, only those who know the complementary sequence can decrypt the information.
[0014] While existing technologies and strategies each possess unique characteristics and advantages, they also suffer from significant shortcomings and limitations. A major issue is that these strategies fail to fully exploit the inherent properties of DNA as an information storage medium, particularly its unique multi-copy nature, which can contain random errors. Throughout the DNA information storage process, for each information unit, multiple copies exist, often containing random base errors. Existing encryption strategies are not designed to address or mitigate this characteristic. Some schemes attempt to create interference by using different copies of the DNA molecule, increasing the difficulty of decryption. However, these schemes typically introduce only a small number of base differences, which attackers can easily identify and target, eliminating them, and conducting cracking attacks. Furthermore, these strategies lack effective means to address random base errors or overlook the potential of these copies to enhance information security and reliability. For example, during the storage and retrieval of large-scale data, replication errors can accumulate, leading to data loss. For example, strategies such as hybrid encryption, hidden information, specialized encoding strategies, and sequence permutation can corrupt encrypted information if random errors generated during DNA replication are not properly addressed. Similarly, in shape-based encryption (DNA origami encryption) and position-based encryption (DNA chip encryption), any replication error may cause the encrypted information to become unrecognizable or misinterpreted. For strategies such as point mutation introduction and multi-chain encryption, although they can deal with the problem of replication errors to some extent, other problems still exist. For example, during the storage and reading of large-scale data, replication errors may accumulate, resulting in a large amount of data loss. In addition, for multi-chain encryption strategies, although errors can be detected and corrected by comparing the sequences of different chains, this requires a lot of computing resources and is not guaranteed to be successful in all cases.
[0015] To overcome the various limitations of existing DNA-based encryption methods, the present invention leverages the complex multi-molecular copy nature of DNA information storage to significantly enhance the protection of encrypted information. By introducing random errors and random DNA fragment fragmentation steps, the present invention effectively increases the interference intensity of random combination interference noise. Furthermore, through its innovative encryption and decoding system, it effectively mitigates the impact of large amounts of random combination interference noise on encrypted information. Furthermore, through its block-by-block encryption design and fountain code encoding scheme, the failure to decrypt any small block of data will also cause the decryption of the entire message to fail, further enhancing the security of the encrypted information. Users with the key can quickly and accurately assemble the correct DNA fragment with the aid of the key information, thereby quickly and accurately decrypting the data. However, those attempting to illegally steal information, lacking the correct key, must randomly brute-force crack each similar copy of the message. This wastes a significant amount of computing power on erroneous copies of the information. Furthermore, any small block of information that is incorrectly decrypted due to the use of an erroneous copy will result in a failure to decode the entire message, thus providing strong protection for the encrypted information. Furthermore, by incorporating a DNA assembly algorithm, the present invention effectively addresses DNA degradation, ensuring the long-term reliability of encrypted stored information.
[0016] Overall, the advantages of this invention lie in improving data security, significantly increasing the complexity of cracking, and providing long-term data stability. Furthermore, compared to some DNA encryption methods that rely on complex calculations, this solution requires less computation, improving encryption and decryption efficiency. These characteristics make this invention more capable of meeting the demand for efficient and secure DNA information encryption and protection. Summary of the Invention
[0017] The present invention aims to provide an advanced data encryption protection technology based on DNA information storage. Through the ingeniously designed DNA information storage encryption and decoding system and the introduced interference information enhancement means, the technology of the present invention fully exploits the complex multi-molecule copy characteristics of DNA information storage to greatly enhance the protection of encrypted information. The present invention effectively improves the interference intensity of random combination interference noise by introducing random errors and random DNA fragment breakage steps. And through the innovatively designed encryption and decoding system, it efficiently copes with the impact of a large amount of random combination interference noise on real information. In addition, through the design of block-by-block encryption and the fountain code encoding scheme, the failure of decryption of any small block of data will cause the decryption failure of the entire information, further enhancing the security of the encrypted information.
[0018] For users who possess the key, they can quickly and accurately assemble the correct DNA fragment with the help of the key information, thereby quickly and accurately decrypting the data. However, for those who attempt to steal the information illegally, lacking the correct key, they must randomly brute force each similar copy of the information. This wastes a large amount of cracking computing power on incorrect copies of the information, forming a strong protection barrier against encrypted information copies.
[0019] The present invention specifically provides:
[0020] A first aspect of the present invention provides a method for encrypting DNA stored information, comprising the following steps:
[0021] Step 1: Encode information blocks: Encode the data to be encrypted into a certain number of information blocks through erasure coding. Each data fragment includes an index and a data carrying area.
[0022] Step 2, encryption: perform encryption operations on the data of each small block of information;
[0023] Step 3: Add ECC checksum and scramble the byte sequence: Add ECC code to each encrypted information block and perform pseudo-random scrambling on the byte sequence with the ECC code.
[0024] Step 4: transcode the byte sequence of the encrypted and scrambled information block into a DNA sequence;
[0025] Step 5, synthesizing the DNA sequence obtained in step 4 by high-throughput DNA synthesis technology;
[0026] Step 6: random interference information is introduced and combined interference is enhanced.
[0027] In a specific embodiment of the present invention, the erasure coding in step 1 is fountain code, Reed-Solomon (RS) code or LDPC code.
[0028] In the preferred embodiment of the present invention, fountain codes are used. If a small number of data fragments are erroneous, the entire information will be lost. The fountain code scheme is more conducive to enhancing the encryption protection strength of the information.
[0029] In a specific embodiment of the present invention, DES, AES (Advanced Encryption Standard), and 3DES (Triple Data Encryption Standard) are used as optional encryption schemes in step 2; preferably, DES is used.
[0030] In a specific embodiment of the present invention, CRC32 is used as the ECC encoding scheme for the data segment in step 3.
[0031] In a specific embodiment of the present invention, the index value of the information block and the key of the encryption scheme in step 2 are used as the seed value of the pseudo-random function in step 3, thereby preventing information thieves without the key from successfully performing ECC verification.
[0032] In the specific embodiment of the present invention, the base-bit correspondence relationship used is:
[0033]
[0034]
[0035] Or any other combination among the 24 base bit combination correspondences.
[0036] In a specific embodiment of the present invention, the DNA sequence synthesized in step 5 has a linker.
[0037] In a specific embodiment of the present invention, in step 6, the introduction of interference information and the enhancement of combined interference are achieved through error-prone PCR, or high-temperature degradation treatment, or ultraviolet treatment, or DNAase I enzymatic treatment, or a combination of any of the above technologies.
[0038] The second aspect of the present invention provides a DNA library carrying encrypted information obtained by the method of the first aspect of the present invention.
[0039] The third aspect of the present invention provides a method for decrypting the DNA library carrying encrypted information in the second aspect:
[0040] Step 1, high-throughput sequencing: Perform high-throughput sequencing on the encrypted and stored DNA samples;
[0041] Step 2, DNA fragment assembly and reconstruction;
[0042] Step 3: Reconstruct the information block; reconstruct the scrambled and encrypted information block through the base-bit correspondence;
[0043] Step 4: reorder and perform data ECC verification on the reconstructed information blocks; the ordering method and verification method are based on the ECC check code and pseudo-random scrambling algorithm used in the encryption writing process;
[0044] Step 5: Decode the reordered information blocks using the key;
[0045] Step 6: Decode the original information based on the erasure coding scheme.
[0046] In a specific embodiment of the present invention, step 1 uses an Illumina sequencer for high-throughput sequencing, with a read length of 150 bp*2 for paired-end sequencing, or nanopore sequencing is used.
[0047] In a specific embodiment of the present invention, the base-bit correspondence relationship used in step 3 is:
[0048]
[0049] Or any other combination among the 24 base bit combination correspondences.
[0050] In a specific embodiment of the present invention, steps 2) to 4) are an assembly method based on the De Blain diagram, specifically comprising:
[0051] Step S1, calculating the initial k-mer of the current DNA sequence to be assembled according to the index information;
[0052] Step S2, searching for a k-mer connected to the k-mer in the De Braying graph (DBG) graph; if not found, the assembly of the current DNA fragment fails;
[0053] Step S3: If connected k-mers are found in the DBG graph, these k-mers are placed at the end of the search path; each k-mer found forms an independent terminal k-mer node;
[0054] Step S4: If the path length does not reach the length of the DNA fragment, repeat steps S2-S4 until the path length meets the length of the DNA fragment.
[0055] Step S5: When the path length reaches the length of the DNA sequence, the path is reordered according to the index and key information. This step is the reverse of the encoding process, which randomly scrambles the path according to the index and key information. After this step, the encrypted information bytes and ECC bytes are returned to their original positions.
[0056] Step S6, reading the ECC value of the path-related DNA sequence and performing ECC verification;
[0057] Step S7, judging the node, judging the number of paths that meet the ECC check; if the number of paths is zero or greater than 1, the DNA sequence assembly process in this direction fails; if the number of paths is 1, outputting the relevant DNA sequence;
[0058] Step S8: If the number of paths that meet the ECC check is 1, the DNA sequence contained in the path is output.
[0059] After the above steps are completed, the corresponding indexed DNA fragments and encrypted information blocks can be obtained.
[0060] A fourth aspect of the present invention provides a computer module that runs the method of the first aspect, or the method of the third aspect, or a combination of the above methods.
[0061] Beneficial technical effects
[0062] 1) High Data Security: By utilizing a combination of base-difference DNA molecule copies and random interruptions, the present invention increases interference during the decryption process, making it difficult for attackers to accurately obtain a correct copy of the encrypted information. This significantly increases the difficulty of decryption and improves data security.
[0063] 2) Efficient encryption and decryption: Compared to some DNA encryption methods that rely on complex calculations, the present invention requires relatively low computational effort. The random interruption and combined interference operations can be completed in a shorter time, improving encryption and decryption efficiency.
[0064] 3) High capacity and long-term storage capability: DNA information storage has extremely high density and long-term data storage capability, which can cope with large-scale data storage needs and maintain data integrity and reliability.
[0065] 4) Broad Application Prospects: The DNA information storage data encryption protection technology of this invention has broad application prospects in the field of data storage and protection. It can be applied to cloud storage, data backup, sensitive information storage, and other fields, providing reliable protection for data security.
[0066] In summary, this invention provides a data encryption and protection technology for DNA information storage. Through random interruption and combined interference steps, copies of DNA molecules with base differences interfere with the decryption process. Compared to existing technologies, this solution offers advantages such as concealment, difficulty in attack, and lower computational requirements. This technology has broad application prospects in the fields of data storage and protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 : Flowchart of the technical solution for encrypted storage and decrypted reading of the present invention (A: encryption process; B: decryption process);
[0068] Figure 2 : Schematic diagram of the principle of randomly interrupting DNA sequences to increase the combinatorial interference space;
[0069] Figure 3 :DNA fragment assembly and reconstruction flow chart;
[0070] Figure 4 : Electrophoresis analysis results of samples treated at 70 degrees for 30 days;
[0071] Figure 5 : Digital images verifying the method of this application. DETAILED DESCRIPTION
[0072] the term
[0073] DNA Fountain Codes and Erasure Codes
[0074] DNA The fountain code was first As reported by Erlich et al., Science 355, 950–954 (2017), this method The data file is cut into several segments by Fountain Code Algorithm Obtain the coding sequence and then convert it into The DNA sequence can be converted into a DNA sequence and can be screened by conditions such as GC content or homopolymers to eliminate unwanted sequences.
[0075] Fountain codes are a type of erasure code. Erasure codes are a type of coding error-tolerant technology first used in the communications industry to address the problem of data loss during transmission. Their basic principle is to segment the transmitted signal, add certain checks, and then establish a certain connection between the segments. Even if part of the signal is lost during transmission, the receiving end can still calculate the complete information through algorithms.
[0076] In the present invention, "data droplet", "information block", "byte sequence" and the like have the same meaning, and all refer to data fragments obtained after processing a data file through erasure coding.
[0077] DBG method and DBG diagram
[0078] The de Bruijn Graph (DBG) is an anti-intuition algorithm for assembling sequencing data. It is primarily used for assembling short, high-abundance fragments, particularly second-generation sequencing data. The DBG algorithm transforms the assembly process into the problem of finding an Eulerian path in the DBG graph (starting from a certain point and traversing all edges exactly once). An Eulerian path is a P-type problem, meaning that there are reliable necessary and sufficient conditions to prove the existence of an Eulerian path and that it can be solved within a certain timeframe. The method is as follows:
[0079] ① Split the reads into shorter k-mers of uniform length (reads shorter than k will be discarded);
[0080] ② Find the overlapping relationship between k-mers and build a DBG graph. That is, for any two k-mers u and w, if the last k-1 base sequences of u are the same as the first k-1 base sequences of w, then build a directed edge from u to w.
[0081] ③Find the Euler path in the DBG graph to obtain the assembly sequence.
[0082] Example
[0083] Example 1 Encrypted Writing Scheme
[0084] like Figure 1 As shown in (A), the information encryption and writing of the DNA information storage encryption and decryption technology of the present invention includes the following six main steps:
[0085] Step 1: Erasure coding: Encode the encrypted data into a certain number of small data droplets using fountain codes. Each data droplet includes an index and a data carrier area. This technical solution uses fountain codes as an erasure scheme. Other erasure codes, such as Reed-Solomon (RS) codes and LDPC codes, are also available as alternatives. Fountain codes can cause the entire information to be lost if a small number of data droplets are erroneous. Fountain codes are more conducive to enhancing the encryption protection strength of information.
[0086] Step 2: Encryption and scrambling. Encryption is performed on the data in each information droplet. This technical solution uses DES as the encryption algorithm. Other encryption algorithms such as AES (Advanced Encryption Standard) and 3DES (Triple Data Encryption Standard) are also available as optional encryption schemes.
[0087] Step 3, add ECC checksum and byte sequence scrambling. Add ECC code to each data droplet; this technical solution uses CRC32 as the ECC coding scheme for data droplets. Other ECC codes, such as RS code, can be used as optional coding schemes; in order to prevent information thieves from successfully performing ECC verification, it is necessary to pseudo-randomly scramble the sequence with the added ECC code. The index value and key of the data droplet serve as the seed value of the pseudo-random function, thereby preventing information thieves without the key from successfully carrying out ECC verification. The pseudo-random sequence function for random byte scrambling and the corresponding three Python functions for scrambling and resetting bytes in sequence used in this technical solution are as follows:
[0088]
[0089] Step 4: Transcode the binary sequence of the data droplet into a DNA sequence. This technical solution uses the transcoding rules shown in Table 1 to convert binary sequences into DNA sequences. Any other transcoding rules can be used as alternatives and can achieve similar transcoding effects.
[0090] Table 1 Base bit correspondence table used in technical solution 1
[0091]
[0092] Note: Any other base-bit correspondence can achieve similar technical effects.
[0093] Step 5: synthesize the ssDNA sequence obtained in step 4 using high-throughput DNA synthesis technology.
[0094] Step 6: Random interference information introduction and combined interference enhancement: The introduction of interference information and combined interference enhancement are achieved through error-prone PCR, high temperature degradation treatment, ultraviolet treatment, DNAase I enzymatic treatment, or a combination of any of the above techniques. Figure 2 This study demonstrates the principle of increasing the combinatorial interference space by randomly interrupting the DNA sequence. The processed DNA sample can then be used as an encrypted information medium, which can be protected through proper storage.
[0095] Example 2 Decryption Reading
[0096] like Figure 1 As shown in (B), the information decryption and reading of the DNA information storage encryption technology of the present invention includes the following six main steps:
[0097] Step 1: High-throughput sequencing: Perform high-throughput sequencing on the encrypted and stored DNA sample. This technical solution uses the Illumina NovaSeq sequencer for high-throughput sequencing, with a paired-end sequencing read length of 150bp*2. Other sequencing technologies, such as nanopore sequencing, are also available as options.
[0098] Step 2: DNA fragment assembly and reconstruction; the assembly and reconstruction process is as follows Figure 3 As shown, the detailed steps are described in Example 3;
[0099] Step 3: Reconstruct the data droplet; reconstruct the encrypted data droplet using the bit-DNA sequence transcoding rules shown in Table 1;
[0100] Step 4: reorder the reconstructed data droplets and perform data ECC verification; the ordering method and verification method are based on the ECC check code and byte scrambling algorithm used in the encryption writing process;
[0101] Step 5, using the key to decode the data;
[0102] Step 6: Use fountain code to decode the original information.
[0103] Example 3 Decryption and assembly of DNA fragments
[0104] The DNA fragments (small pieces of encrypted information) corresponding to each index value are assembled through the following sequence assembly process based on the Debreu graph. This sequence assembly process consists of 8 core steps:
[0105] Step S1, calculating the initial k-mer of the current DNA sequence to be assembled according to the index information;
[0106] Step S2, searching for a k-mer connected to the k-mer in the De Braying graph (DBG) graph; if not found, the assembly of the current DNA fragment fails;
[0107] Step S3: If connected k-mers are found in the DBG graph, these k-mers are placed at the end of the search path; each k-mer found forms an independent terminal k-mer node;
[0108] Step S4: If the path length does not reach the length of the DNA fragment, repeat steps S2-S4 until the path length meets the length of the DNA fragment.
[0109] Step S5: When the path length reaches the length of the DNA sequence, the path is reordered according to the index and key information. This step is the reverse operation of the encoding process, which randomly scrambles the path according to the index and key information. After this step, the encrypted information bytes and CRC bytes are returned to their original positions.
[0110] Step S6, reading the CRC value of the path-related DNA sequence and performing CRC verification;
[0111] Step S7, determine the node and the number of paths that meet the CRC check; if the number of paths is zero or greater than 1, the DNA sequence assembly process in this direction fails; if the number of paths is 1, the relevant DNA sequence is output;
[0112] Step S8: If the number of paths that meet the CRC check is 1, the DNA sequence contained in the path is output.
[0113] After the above steps are completed, the byte strings of the corresponding indexed DNA fragments and the encrypted information blocks can be obtained.
[0114] Example 4: Experiment on random errors introduced by error-prone PCR:
[0115] All error-prone polymerase chain reactions (PCR) were performed using the "Controlled Error Polymerase Chain Reaction Kit" (CAT#: 160903-100) produced by Beijing Tianenze Biotechnology Co., Ltd. The temperature cycling parameters were set as shown in Table 2: heating at 94°C for 3 minutes; 94°C for 30 seconds, 55°C for 1 minute, and 72°C for 30 seconds, repeated 30 times; and finally stored at 4°C.
[0116] Table 2 PCR reaction conditions
[0117]
[0118] Example 5: Random fragment fragmentation experiment (DNAse I method):
[0119] Use a commercial DNAse I product, such as NEB DNAse I (M0303L), and incubate at 37°C for 1 minute using the reaction system shown in Table 3.
[0120] Table 3 DNAse I random disruption conditions
[0121]
[0122] Example 6: Random fragment interruption experiment (high temperature degradation method):
[0123] The samples were placed at high temperatures of 70 degrees or above to accelerate the degradation of the fragments. Figure 4 The results of agarose electrophoresis analysis of samples treated at 70 degrees for 30 days are shown.
[0124] Example 7: Example of encrypted storage and decrypted reading of images:
[0125] This embodiment takes the encrypted storage and reading of a digital picture as an example to illustrate the entire encryption and decryption process of the present invention. Figure 5 As shown. The image size is about 308KB, MD5 (f84a4b1775e80d8f838f50f2ff0b5a67). Using the 32-bit password: 0b10101010101111011001010100110110 and the scheme of Example 1, the image is encoded into 12,000 196bp DNA sequences (Table 4 shows some of the sequence information). It contains two primer sequences p1 = 'CCTGCAGAGTAGCATGTC' and p2 = 'CTGACACTGATGCATCCG'. Other parameters: Fountain code random seed value is 2, delta = 0.01, c_value = 0.01, data chunk number 10517, each DNA sequence encodes 30 bytes of information, initial index value 211011011.
[0126] The encoded sequence was synthesized into an ssDNA library by a DNA synthesis company. Using the ssDNA synthesized by the company as a template, the library was amplified based on error-prone PCR according to the method of Example 4. The amplified sequence was processed by the method of Example 6 and sent to a sequencing company for sequencing. After sequencing, the sequencing reads shown in Table 5 were generated. The sequencing reads were decrypted using the encryption key 0b10101010101111011001010100110110 by the method of Example 2, and the encrypted data ( Figure 5 ), the MD5 value of the image is consistent with the original data. According to statistics, each fragment of this embodiment introduces an average of 10 4 The above interference paths greatly improve the security of encrypted data.
[0127] Table 4 Partial sequences of 12,000 DNA sequences generated by coding
[0128]
[0129]
[0130]
[0131]
[0132]
[0133] Table 5 Statistics of sequencing data
[0134]
[0135] Q20 (Q30) refers to the percentage of bases in a sequence with a quality score greater than or equal to 20 (30). It is used to assess the accuracy of sequencing. Q20 (Q30) indicates a 1% (0.1%) probability of a base being misread and a 99% (99.9%) accuracy. Generally speaking, the number of bases with a Q30 accuracy of at least 85% must be high.
Claims
1. A method for encrypting DNA stored information, comprising the following steps: Step 1: Encode information blocks: Encode the data to be encrypted into a certain number of information blocks through erasure coding. Each information block includes an index and a data carrying area. Step 2, encryption: perform encryption operations on the data of each small block of information; Step 3: Add ECC checksum and scramble the byte sequence: Add ECC code to each encrypted information block and perform pseudo-random scrambling on the byte sequence with the ECC code. Step 4: transcode the byte sequence of the encrypted and scrambled information block into a DNA sequence; Step 5, synthesizing the DNA sequence obtained in step 4 by high-throughput DNA synthesis technology; Step 6: random interference information is introduced and combined interference is enhanced.
2. The DNA storage information encryption method according to claim 1, wherein the erasure coding in step 1 is fountain code, Reed-Solomon code or LDPC code.
3. The DNA storage information encryption method according to claim 2, wherein the erasure coding uses fountain code.
4. The DNA storage information encryption method as claimed in claim 1, wherein DES, AES, or 3DES is used as the encryption scheme in step 2.
5. The DNA storage information encryption method as claimed in claim 1, wherein in step 3, CRC32 is used as the ECC encoding scheme for the information block.
6. In the DNA storage information encryption method as described in claim 5, in step 3, the index value of the information block and the key of the encryption scheme in step 2 are used as the seed value of the pseudo-random function, thereby preventing information thieves without the key from successfully performing ECC verification.
7. The DNA storage information encryption method according to claim 1, wherein the following base-bit correspondence is used in step 4: A: 00, C: 01, G: 10, and T: 11; or, T: 00, G: 01, C: 10, and A: 11; or, Any other combination among the 24 base-bit combination correspondences.
8. The DNA storage information encryption method according to claim 1, wherein in step 6, the interference information is introduced and the combined interference enhancement is achieved by error-prone PCR, or high temperature degradation treatment, or ultraviolet treatment, or DNAase I enzymatic treatment, or a combination of these means.
9. A DNA library carrying encrypted information obtained by the DNA storage information encryption method according to any one of claims 1 to 8.
10. A method for decrypting the DNA library carrying encrypted information according to claim 9: Step D1, high-throughput sequencing: performing high-throughput sequencing on the encrypted and stored DNA sample; Step D2, DNA fragment assembly and reconstruction; Step D3, reconstructing the information block: reconstructing the scrambled and encrypted information block through the base-bit correspondence; Step D4, reordering and performing data ECC verification on the reconstructed information blocks; the ordering method and verification method are based on the ECC check code and pseudo-random scrambling algorithm used in the encryption writing process; Step D5, decoding the reordered information blocks using the key; Step D6: Decode the original information based on the erasure coding scheme.
11. The decryption method according to claim 10, wherein step D1 uses an Illumina sequencer for high-throughput sequencing with a read length of 150 bp*2 for paired-end sequencing, or uses nanopore sequencing.
12. The decryption method according to claim 10, wherein step D3 adopts the following base-bit correspondence relationship: A: 00, C: 01, G: 10, and T: 11; or, T: 00, G: 01, C: 10, and A: 11; or, Any other combination among the 24 base-bit combination correspondences.
13. The decryption method according to claim 10, wherein steps D2 to D4 are an assembly method based on a De Bray graph, specifically comprising: Step S1: Calculate the initial sequence of the DNA sequence to be assembled based on the index information. k -mer; Step S2, find the k -mer connected k -mer; if not found, the assembly of the current DNA fragment fails; Step S3, if a connected k -mer, these k -mer is put at the end of the search path; each found k -mer forms an independent end k -mer node; Step S4: If the path length does not reach the length of the DNA fragment, repeat steps S2-S4 until the path length meets the length of the DNA fragment. Step S5: When the path length reaches the length of the DNA sequence, the paths are reordered according to the index and key information; This step is the reverse operation of the encoding process, which randomly scrambles the path based on the index and key information. After this step, the encrypted information bytes and ECC bytes will be returned to their original positions. Step S6, reading the ECC value of the path-related DNA sequence and performing ECC verification; Step S7, determining the node and the number of paths that meet the ECC check; If the path number is zero or greater than 1, the DNA sequence assembly process of the path fails; If the number of paths is 1, the relevant DNA sequence is output; Step S8: If the number of paths that meet the ECC check is 1, then output the DNA sequence contained in the path; After the above steps are completed, the corresponding indexed DNA fragments and encrypted information blocks can be obtained.
14. A computer device running the encryption method according to any one of claims 1 to 8, or the decryption method according to any one of claims 10 to 13, or a combination of the above methods.
Citation Information
Patent Citations
Polymer molecule information storage error correction coding and decoding system
CN110190858A
Encoding and decoding method and device for DNA information storage
CN113687976A