DNA storage encryption method, device and equipment with dual-noise collaborative enhancement and medium

By adopting a dual-noise collaborative encryption method in DNA storage, the data is modulated and encoded using normal and misleading key sequences, and structural errors are injected into the misleading sequence, the sequence alignment distortion problem caused by noise interference in DNA storage is solved, and efficient and secure data recovery and decryption are achieved.

CN120165860APending Publication Date: 2025-06-17GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455202.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Technical problems in DNA storage with inefficient sequence alignment distortion and consensus generation caused by multi-source noise interference such as synthesis and sequencing.

Method used

The DNA storage encryption method with dual noise synergistic enhancement is adopted to generate the normal sequence and misleading sequence of DNA by converting the plaintext information to be encrypted into an indexed binary sequence and modulating it with normal key sequence and misleading key sequence. Inject structural errors into misleading sequences and mix storage with normal sequences, thereby enabling efficient data filtering and decryption using a dual-key matching algorithm during the decryption phase.

Benefits of technology

In a high-noise environment, it significantly improves the reliable recovery and security of DNA storage data, ensures the integrity and uniqueness of data, and enhances the concealment of information storage and anti-lateral channel attack capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165860A_ABST
    Figure CN120165860A_ABST
Patent Text Reader

Abstract

The invention provides a DNA storage encryption method, device and equipment with double-noise collaborative enhancement and a medium, and the method comprises the steps: converting plaintext information to be encrypted into a binary sequence with an index; randomly generating a normal key sequence and a misleading key sequence; modulating and coding the binary sequence by using the normal key sequence and the misleading key sequence to generate a normal sequence and a misleading sequence of the DNA; injecting a structural error into the misleading sequence, mixing the misleading sequence with the normal sequence, and storing; obtaining a mixed sequencing set from a DNA sequence stored in a mixed manner, and screening and decrypting the mixed sequencing set by using the normal key sequence and the misleading key sequence to obtain plaintext information, the problems of sequence alignment distortion and low consensus generation efficiency caused by multi-source noise interference such as synthesis and sequencing in data recovery in DNA storage are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of DNA storage technology, and in particular to a dual-noise collaboratively enhanced DNA storage encryption method, device, equipment and medium. Background Art

[0002] With the exponential growth of digital data worldwide, traditional storage technologies such as hard disks and magnetic tapes are facing increasingly severe challenges in terms of power consumption, physical volume, reliability, and long-term durability. As a natural carrier of genetic information, DNA is considered to be a strong candidate medium for large-scale, long-term information storage in the future due to its extremely high storage density, ultra-long life, low maintenance cost, and high parallelism. As DNA storage systems gradually move towards the actual deployment stage, how to ensure the confidentiality and integrity of stored information has become an important issue that needs to be solved urgently. Since DNA storage involves complex and easily disturbed biological processes, encryption technology not only needs to have sufficient security, but also needs to be compatible with the physical and chemical constraints of biochemical systems.

[0003] A common feature of DNA storage channels is that the sequencing sequences used for information reading are inevitably interfered by various noises. These noises mainly come from technical errors in biological processes such as synthesis, storage, PCR (Polymerase Chain Reaction) amplification and sequencing, which manifest as insertion, deletion and substitution (IDS) of bases. Such errors are usually randomly distributed in different positions and sequences, seriously interfering with sequence alignment and consensus generation, and becoming a key problem that needs to be solved in the process of reliable data recovery. Therefore, any encryption method applied to the DNA storage environment must have fault tolerance under real channel noise.

[0004] At present, some studies have attempted to use this "natural error" for encryption purposes. Yao et al. proposed an image encryption method based on forward error correction DNA coding, which can achieve decryption robustness under up to 10% IDS errors. In addition, this method also introduces the concept of gene mutation, mapping the image pixel value to one of the eight mutation forms of the coding gene to achieve pixel-level diffusion encryption. However, its upper limit of tolerance to substitution errors is only 1%, which still has limitations in actual high-noise environments.

[0005] Therefore, there is an urgent need to propose a dual-noise synergistically enhanced DNA storage encryption method to solve the technical problems of sequence alignment distortion and low consensus generation efficiency caused by multi-source noise interference such as synthesis and sequencing in data recovery in DNA storage. Summary of the invention

[0006] To overcome the problems existing in the related technologies, the present disclosure provides a DNA storage encryption method, apparatus, device and medium with dual-noise collaborative enhancement, so as to solve the technical problems of sequence alignment distortion and low consensus generation efficiency caused by multi-source noise interference in DNA storage during data recovery in the related technologies.

[0007] One or more embodiments of this specification provide a DNA storage encryption method with dual-noise collaborative enhancement, including the following steps:

[0008] Convert the plaintext information to be encrypted into an indexed binary sequence;

[0009] Randomly generate a normal key sequence and a misleading key sequence;

[0010] Use the normal key sequence and the misleading key sequence to modulate and encode the binary sequence to generate a normal sequence and a misleading sequence of DNA;

[0011] Inject structural errors into the misleading sequence, and store it after mixing with the normal sequence;

[0012] Obtain a mixed sequencing set from the mixed-stored DNA sequences, and use the normal key sequence and the misleading key sequence to screen and decrypt the mixed sequencing set to obtain the plaintext information.

[0013] Preferably, the injecting structural errors into the misleading sequence specifically includes the following steps:

[0014] Adjust the synthesis conditions of the DNA strand at specific positions to increase the probability of errors occurring in the misleading sequence at the specific positions.

[0015] Preferably, it further includes the following steps:

[0016] Use the misleading key sequence to modulate and decode the misleading sequence, extract sequence fragments with residual information characteristics, and assist in voting for error correction;

[0017] Or, use the normal key sequence to decode the mixed-stored DNA sequences, and obtain sequence fragments from all the sequences;

[0018] Or, use the misleading key sequence to decode the mixed-stored DNA sequences, and capture sequence fragments that are misjudged due to noise perturbation but still contain information residues.

[0019] Preferably, the using the normal key sequence and the misleading key sequence to modulate and encode the binary sequence to generate a normal sequence and a misleading sequence of DNA specifically includes the following steps:

[0020] Modulate and encode the indexed binary sequence using the normal key sequence and the misleading key sequence respectively according to the modulation and encoding rules to generate a normal sequence and a misleading sequence;

[0021] After copying the misleading sequence a preset number of times, mix it with the normal sequence and store it.

[0022] Preferably, screening and decrypting the mixed sequencing set using the normal key sequence and the misleading key sequence to obtain the plaintext information, specifically including the following steps:

[0023] Screen the mixed stored DNA sequence using the normal key sequence and the misleading key sequence to obtain a set of normal sequences;

[0024] Based on the set of normal sequences, demodulate and restore each sequence to binary fragments using the modulation and decoding rules;

[0025] Cluster and sort the binary fragments using the index to obtain the plaintext information.

[0026] One or more embodiments of this specification provide a DNA storage encryption device with dual-noise collaborative enhancement, including a conversion module, a key generation module, an encoding module, an error injection module, and a decryption module;

[0027] The conversion module is used to convert the plaintext information to be encrypted into an indexed binary sequence;

[0028] The key generation module is used to randomly generate a normal key sequence and a misleading key sequence;

[0029] The encoding module is used to modulate and encode the binary sequence using the normal key sequence and the misleading key sequence to generate a normal sequence and a misleading sequence of DNA;

[0030] The error injection module is used to inject structural errors into the misleading sequence and store it after mixing with the normal sequence;

[0031] The decryption module is used to obtain a mixed sequencing set from the mixed stored DNA sequence, and screen and decrypt the mixed sequencing set using the normal key sequence and the misleading key sequence to obtain the plaintext information.

[0032] Preferably, it further includes an auxiliary module configured to demodulate and decode the misleading sequence using the misleading key sequence, extract sequence fragments with residual information characteristics, and assist in voting and error correction;

[0033] Or, decode the mixed stored DNA sequence using the normal key sequence, and obtain sequence fragments from all sequences;

[0034] Alternatively, use the misleading key sequence to decode the hybrid-stored DNA sequence, and capture the sequence fragments that are misjudged due to noise perturbation but still contain information fragments.

[0035] Preferably, the encoding module includes an encoding unit and a mixing unit;

[0036] The encoding unit is configured to perform modulation encoding on the indexed binary sequence using the normal key sequence and the misleading key sequence respectively according to the rules of modulation encoding, to generate a normal sequence and a misleading sequence;

[0037] The mixing unit is configured to copy the misleading sequence a preset number of times, and then mix and store it with the normal sequence.

[0038] One or more embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the double-noise collaborative enhanced DNA storage encryption method as described above is implemented.

[0039] One or more embodiments of this specification provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the double-noise collaborative enhanced DNA storage encryption method as described above are implemented.

[0040] A DNA storage encryption method, device, equipment and medium with dual-noise collaborative enhancement provided by the present disclosure has the advantages that by converting the plaintext information to be encrypted into an indexed binary sequence, data standardization processing is realized, format differences are eliminated through binary coding, and the indexing mechanism provides accurate positioning ability for subsequent encryption operations, establishing a structured data foundation; a normal key sequence and a misleading key sequence are randomly generated, the normal key ensures the encryption effectiveness, the misleading key constructs a security confusion layer, and the key generation based on a random algorithm makes the key space reach the theoretical security level to resist brute-force cracking; the normal key sequence and the misleading key sequence are used to modulate and encode the binary sequence to generate a normal sequence and a misleading sequence of DNA, the data density is increased by using the storage characteristics of biomolecules, the biological characteristics of the DNA sequence enhance the encryption concealment, and the multi-sequence confusion strategy forms an information fog, making it difficult for attackers to locate valid data; structural errors are injected into the misleading sequence and stored after being mixed with the normal sequence, controllable errors are implanted in the misleading sequence to construct an information interference layer, and the distributed storage strategy reduces the risk of data concentration, and the error checking mechanism is cooperated to ensure the decryption integrity; a mixed sequencing set is obtained from the DNA sequences stored in a mixed manner, the normal key sequence and the misleading key sequence are used to screen and decrypt the mixed sequencing set to obtain the plaintext information, efficient data screening is realized through a dual-key matching algorithm to improve the accuracy rate, and the decryption process based on the key ensures the uniqueness of information restoration, and the data integrity can still be maintained in a complex storage environment, realizing the complete process of encrypted storage and secure acquisition. It improves the analyzability, comparability and controllability of the DNA encryption scheme, and lays a foundation for constructing a systematic DNA encryption complexity framework. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic flowchart of a DNA storage encryption method with dual-noise collaborative enhancement provided for one or more embodiments of this specification;

[0043] Figure 2 It is a general schematic diagram of an encryption framework combining injected noise-channel noise provided for one or more embodiments of this specification, Figure 2 (a) is a schematic diagram of the encryption, storage and decryption processes, Figure 2 (b) is the modulation and encoding process; Figure 2 (c) is the error injection process;Figure 2 (d) is the modulation and decoding process;

[0044] Figure 3 It is a heat map of the optimized decryption process provided by one or more embodiments of this specification;

[0045] Figure 4 It is a schematic structural diagram of a DNA storage encryption device with dual-noise collaborative enhancement provided by one or more embodiments of this specification;

[0046] Figure 5 It is a schematic structural diagram of a computer device provided by one or more embodiments of this specification. Specific implementation manners

[0047] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this invention document.

[0048] The following will make a detailed description of the present invention with reference to the specific implementation manners and the accompanying drawings of the specification.

[0049] Method embodiments

[0050] According to an embodiment of the present invention, a DNA storage encryption method with dual-noise collaborative enhancement based on "injected noise + channel noise" is provided. As Figure 1 shown, it is a flowchart of the DNA storage encryption method with dual-noise collaborative enhancement provided by this embodiment. The DNA storage encryption method with dual-noise collaborative enhancement according to an embodiment of the present invention includes the following steps:

[0051] S110. Convert the plaintext information to be encrypted into an indexed binary sequence. Specifically, divide the plaintext information to be encrypted into several information blocks (for example, N = 100 blocks, each block is 160 bits), add a unique index (for example, 8 bits) before each information block, and generate N indexed binary message sequences. This index is used for subsequent information sorting and recombination.

[0052] S120. Randomly generate a normal key sequence \(k_n\) and a misleading key sequence \(k_m\). The lengths of the two modulation key sequences are \(l\) (e.g., \(l = 168\)), and there is a fixed edit distance between them (e.g., modifying 5% of the bits), so as to have a high similarity in decoding representation. The keys must meet common biochemical constraints such as: the GC content is controlled at 45% - 55%; the continuous length of the same base does not exceed 3.

[0053] S130. Use the normal key sequence \(k_n\) and the misleading key sequence \(k_m\) to modulate and encode the binary sequence to generate a normal DNA sequence \(R_n\) and a misleading sequence \(R_m\) of DNA.

[0054] Specifically, first, through the modulation and encoding method, the same original information is encoded into two groups of DNA sequences respectively:

[0055] One group is the normal DNA sequence, encoded with the normal modulation key, which is the main data source for legitimate users to decode.

[0056] The other group is the misleading DNA sequence, encoded with the misleading modulation key, which is used to construct interference data with high confusion ability.

[0057] S140. When synthesizing the misleading DNA sequence, inject structural errors into the misleading sequence \(R_m\). It is not difficult to inject errors manually during the synthesis process, which can be achieved by basic operations such as adding different concentrations of bases according to the synthesis equipment. The steps are as follows: In each misleading sequence, randomly select several columns (positions); in the selected columns, randomly introduce a structural error (insertion, deletion, or substitution); all error injections are not correctable, aiming to destroy the consensus consistency and construct a "pseudo-consensus structure" to mislead decoding; no errors are injected into the normal sequence, and it is only affected by subsequent channel noise.

[0058] After injecting the structural errors, merge all the normal sequences (N pieces) and the misleading sequences (N×t pieces) to form a DNA sequence pool (a total of N×(t + 1) pieces), perform mixed storage of the DNA sequences, and uniformly put them into a biological environment for preservation (such as freeze-drying, refrigeration, etc., in vitro methods), and wait for subsequent reading.

[0059] Next is the decryption stage, which includes steps such as sequencing, screening, error correction, and decoding. Since there are naturally sequencing errors (insertions, deletions, substitutions), loss rates, and uneven coverage in the DNA channel, the reading process itself constitutes a preliminary confounding environment. What the receiver receives is a DNA mixed sequencing set R that is a mixture of misleading and normal sequences and is full of noise effects, including read segments (CE) of normal sequences interfered by channel noise, and interfering read segments (IE + CE) formed after structural error injection of misleading sequences. To ensure the usability of the encryption method, the decryption stage needs to ensure that authorized users can accurately restore the original information under realistic noise conditions, and at the same time allow the use of misleading sequences to assist in improving the decoding quality under low coverage.

[0060] S150. Obtain the mixed sequencing set R from the DNA sequences stored in a mixed manner, and use the normal key sequence and the misleading key sequence to screen and decrypt the mixed sequencing set to obtain the plaintext information.

[0061] As Figure 2 shown, it is the overall schematic diagram of the encryption framework combining injection noise and channel noise provided by this embodiment. Figure 2 (a) is the schematic diagram of the encryption, storage, and decryption processes. Figure 2 (b) is the modulation coding process. Figure 2 (c) is the error injection process. Figure 2 (d) is the modulation decoding process.

[0062] The method provided in this embodiment realizes data standardization processing by converting the plaintext information to be encrypted into an indexed binary sequence, eliminates format differences through binary coding, and the indexing mechanism provides accurate positioning ability for subsequent encryption operations, establishing a structured data foundation; randomly generates a normal key sequence and a misleading key sequence, the normal key ensures encryption effectiveness, the misleading key constructs a security confusion layer, and the key generation based on a random algorithm makes the key space reach the theoretical security level to resist brute force cracking; uses the normal key sequence and the misleading key sequence to modulate and encode the binary sequence to generate a normal sequence and a misleading sequence of DNA, improves data density by utilizing the storage characteristics of biomolecules, the biological characteristics of the DNA sequence enhance encryption concealment, and the multi-sequence confusion strategy forms an information fog, making it difficult for attackers to locate valid data; after injecting structural errors into the misleading sequence R_m, performs hybrid storage of DNA sequences, implants controllable errors in the misleading sequence to construct an information interference layer, and the distributed storage strategy reduces the risk of data concentration, and cooperates with the error checking mechanism to ensure decryption integrity; obtains a hybrid sequencing set from the hybrid-stored DNA sequences, uses the normal key sequence k_n and the misleading key sequence k_m to screen and decrypt the hybrid sequencing set to obtain the plaintext information, realizes efficient data screening through a dual-key matching algorithm, improves the accuracy rate, and the decryption process based on the key ensures the uniqueness of information restoration, and can still maintain data integrity in a complex storage environment, realizing the complete process of encrypted storage and secure acquisition. It improves the analyzability, comparability and controllability of the DNA encryption scheme, and lays a foundation for constructing a systematic DNA encryption complexity framework.

[0063] When entering the DNA synthesis link, different process parameters are adopted to create differences:

[0064] 1. Synthesis of normal DNA sequence

[0065] The normal sequence adopts a standardized DNA synthesis process, and the conventional conditions include:

[0066] The concentration of the synthesis monomer (phosphoramidite) is sufficient;

[0067] The reaction time for each round of synthesis is sufficient;

[0068] The catalyst and the cleaning steps maintain normal efficiency. Such a process will inevitably introduce a certain amount of natural background errors (i.e., "channel noise"), and the error rate of the third-generation sequencing is generally between 5% and 15%, which belongs to the typical level of current commercial DNA synthesis.

[0069] 2. Synthesis of misleading DNA sequence

[0070] In one embodiment, structural errors are injected into the misleading sequence, which specifically includes the following steps:

[0071] Adjust the synthesis conditions of the DNA strand at specific sites to increase the probability of errors occurring in the misleading sequence at these specific sites. For example, for a sequence with a length of 100, randomly select 2 positions (50 and 100), and make it more likely for all synthesized DNA sequences to have errors at these two positions.

[0072] DNA synthesis is a process of base-by-base extension, and each round of synthesis corresponds to the insertion of a base at a specific position on the strand. If the monomer concentration of the target base in a certain round is reduced (such as one of the four monomers A, T, C, G), or the concentration of competitive incorrect bases is increased (such as increasing the proportion of "G" at the position where "A" should be added), the probability of incorrect insertion will be significantly increased.

[0073] In addition, by shortening the reaction time or changing the catalyst activity, the completion degree of the synthesis reaction in this round can be reduced, increasing the risk of deletion or mismatch at the target position.

[0074] For example: at the 50th base position, deliberately reduce the monomer concentration of "A" and increase the proportion of other bases, while shortening the reaction time; at the 100th base position, reduce the catalyst efficiency so that the target base cannot be successfully incorporated. After such design, the 50th and 100th positions will become "error-prone hotspots", and all sequences synthesized under this condition will be more likely to have errors at these two positions, artificially forming structured injection errors.

[0075] The method provided in this embodiment can precisely control the error sites, accurately create errors at the preset sites, and help construct a DNA model with specific structural defects. This method does not require special chemical modifications or additional complex processes, and can be achieved only by adjusting the existing synthesis process parameters, having good process operability and scalability for batch manufacturing.

[0076] In one embodiment, screening and decrypting the mixed sequencing set using the normal key sequence k_n and the misleading key sequence k_m to obtain the plaintext information specifically includes the following steps:

[0077] Use the normal key sequence k_n and the misleading key sequence k_m to screen the mixed stored DNA sequences, eliminate the sequences with a closer edit distance to k_m, and obtain the normal sequence set R_n.

[0078] Based on the normal sequence set R_n, use the modulation decoding rule to demodulate and restore each sequence into binary fragments.

[0079] Use the index to cluster and sort the binary fragments to obtain the plaintext information.

[0080] This decryption method can achieve an accuracy rate of ≥99% when the sequencing depth ≥20, and is applicable to the conventional biological noise environment (channel error rate pce is about 5% - 15%), which is the default recommended path for authorized users.

[0081] The method provided in this embodiment significantly improves the reliable recovery ability and security of DNA stored data in high-noise environments such as synthesis and sequencing through a three-level collaborative mechanism of "dual-key screening - modulation decoding - index clustering". The dual-key screening uses the normal key and the misleading key to collaboratively filter the mixed sequencing set, actively excluding noise interference and maliciously forged sequences, improving the purity of the normal sequence set R_n. At the same time, the misleading key is used to induce the attacker to decrypt invalid information, enhancing the anti-side-channel attack ability; the modulation decoding and binary reduction are based on the redundant mapping rule and index drive, tolerating base substitution errors and short fragment losses, achieving local error correction during the demodulation process, and reorganizing the disordered or fragmented binary fragments through index clustering to effectively repair the data logic disorder caused by insertion / deletion errors; the index clustering and sorting further combine the multi-copy consensus algorithm to eliminate the remaining error fragments, ensuring the integrity and accuracy of the plaintext information.

[0082] In extreme scenarios with low sequencing depth (such as depth = 10) or high channel noise (pce ≥ 15%), there may be problems with the failure to recover individual fragments in the standard method. Therefore, three auxiliary decryption strategies are further provided to enhance recoverability. In one embodiment, the following steps are also included:

[0083] Use the misleading key sequence k_m to perform modulation decoding on the misleading sequence R_m, and extract sequence fragments with residual information characteristics to assist in voting for error correction.

[0084] Or, use the normal key sequence k_n to decode the DNA sequence stored in a mixed manner, and obtain sequence fragments from all sequences.

[0085] Or, use the misleading key sequence k_m to decode the DNA sequence stored in a mixed manner, and capture sequence fragments that are misjudged due to noise perturbation but still contain information remnants.

[0086] Although the above three methods have a lower accuracy rate than the standard method, they can all provide auxiliary information at some sequence positions. In the specific implementation of the present invention, if the accuracy rate of the standard method is slightly lower (such as ≥92%), the above three decryption results can be introduced for cross-comparison and fusion with the majority voting method, so as to further eliminate local errors and improve the overall decryption integrity.

[0087] Such as Figure 3As shown, this is the heat map of the optimized decryption process provided by this embodiment. Generally, two keys are used to screen the sequencing set to obtain R_n, and then it is decrypted with K_n. Since sequence loss may be introduced during the sequencing process, the sequence copy number may be small during the reading of some sequences, making it difficult to perform multiple sequence alignment and error correction, resulting in a situation where the decryption accuracy is not high enough. Therefore, these three auxiliary decryption methods can be used to decrypt the plaintext information, and as a reference, correct the errors in normal decryption. As a fusion-type optimized decryption process, it is described as follows:

[0088] In the decryption system, the above-mentioned multiple methods can be integrated to form the following process: First step, use the normal key sequence k_n and the misleading key sequence k_m to simultaneously screen out the normal sequence set R_n and the misleading sequence R_m; Second step, use the normal key sequence k_n to decode the normal sequence set R_n as the main decoding path; Third step, use the misleading key sequence k_m to decode the misleading sequence R_m, the normal key sequence k_n to decode the mixed sequencing set R, and the misleading key sequence k_m to decode the mixed sequencing set R as the auxiliary path; Fourth step, in each information block, vote for error correction through the four groups of decoding results, and reorder and reorganize with reference to the index; This method shows stronger error self-repair ability in multiple experimental repetitions and is applicable to scenarios with a high sequence error rate or a small number of samples. Figure 3 The heat map shows the comparison of the decryption performance of four authorized decryption methods under the conditions of a sequencing depth of 10 and different combinations of injection error rate (pie) and channel error rate (pce). The text on the right side of the figure gives an example of the decryption results corresponding to the four methods under the condition of pie + pce = 0.2 + 0.2. By referring to the three decryption paths that retain the misleading sequence, the results of the standard decryption method (k_n + R_n) can be optimized and the errors can be corrected.

[0089] In one embodiment, the modulation and encoding of the binary sequence using the normal key sequence and the misleading key sequence to generate the normal sequence and the misleading sequence of DNA specifically includes the following steps:

[0090] Use the normal key sequence and the misleading key sequence to perform modulation and encoding on the indexed binary sequence respectively according to the rules of modulation and encoding to generate the normal sequence and the misleading sequence.

[0091] After copying the misleading sequence a preset number of times, mix it with the normal sequence and store it.

[0092] The method provided in this embodiment encrypts the original information by using the normal key sequence k_n and the misleading key sequence k_m to modulate and encode the indexed binary sequence according to the modulation and coding rules, converting the digital information into the form of a DNA sequence. The normal sequence carries the correct encrypted information, and the misleading sequence serves as an interference element to enhance the encryption complexity. The generated misleading sequence can effectively interfere with the analysis of potential attackers, increase the difficulty of cracking, and strengthen the security protection. The misleading sequence is copied a preset number of times and then mixed with the normal sequence to construct the misleading sequence, greatly increasing the degree of data confusion, making it difficult to identify and separate the normal sequence, enhancing the concealment of information storage, and reducing the risk of the normal sequence being discovered and attacked.

[0093] Device embodiment

[0094] According to an embodiment of the present invention, there is provided a DNA storage encryption device with dual-noise collaborative enhancement, as Figure 4 shown, which is a schematic structural diagram of the DNA storage encryption device with dual-noise collaborative enhancement provided in this embodiment. The DNA storage encryption device with dual-noise collaborative enhancement according to an embodiment of the present invention includes a conversion module 41, a key generation module 42, an encoding module 43, an error injection module 44, and a decryption module 45.

[0095] The conversion module 41 is configured to convert the plaintext information to be encrypted into an indexed binary sequence.

[0096] The key generation module 42 is configured to randomly generate a normal key sequence and a misleading key sequence.

[0097] The encoding module 43 is configured to use the normal key sequence and the misleading key sequence to modulate and encode the binary sequence to generate a normal sequence and a misleading sequence of DNA.

[0098] The error injection module 44 is configured to inject structural errors into the misleading sequence and store it after mixing with the normal sequence.

[0099] The decryption module 45 is configured to obtain a mixed sequencing set from the DNA sequences stored in a mixed manner, and use the normal key sequence and the misleading key sequence to screen and decrypt the mixed sequencing set to obtain the plaintext information.

[0100] The device provided in this embodiment converts the plaintext information to be encrypted into a binary sequence with an index through the conversion module 41, realizing data standardization processing. The format differences are eliminated through binary coding, and the indexing mechanism provides accurate positioning ability for subsequent encryption operations, establishing a structured data foundation. The key generation module 42 randomly generates a normal key sequence and a misleading key sequence. The normal key ensures the encryption effectiveness, and the misleading key constructs a security confusion layer. The key generation based on a random algorithm makes the key space reach the theoretical security level, resisting brute-force cracking. The encoding module 43 modulates and encodes the binary sequence using the normal key sequence and the misleading key sequence to generate the normal sequence and the misleading sequence R_m of DNA. The data density is increased by utilizing the biological molecule storage characteristics, the biological characteristics of the DNA sequence enhance the encryption concealment, and the multi-sequence confusion strategy forms an information fog, making it difficult for attackers to locate the valid data. After injecting structural errors into the misleading sequence R_m by the error injection module 44, the DNA sequences are mixed and stored. Controllable errors are implanted in the misleading sequence to construct an information interference layer. The distributed storage strategy reduces the risk of data concentration and, combined with the error checking mechanism, ensures the decryption integrity. The decryption module 45 obtains the mixed sequencing set from the mixed-stored DNA sequences, screens and decrypts the mixed sequencing set using the normal key sequence k_n and the misleading key sequence k_m to obtain the plaintext information. The efficient data screening is achieved through the dual-key matching algorithm, improving the accuracy. The decryption process based on the key ensures the uniqueness of information restoration and can still maintain data integrity in a complex storage environment, realizing the complete process of encrypted storage and secure acquisition. It improves the analyzability, comparability, and controllability of the DNA encryption scheme, laying a foundation for constructing a systematic DNA encryption complexity framework.

[0101] In one embodiment, an auxiliary module is further included, configured to modulate and decode the misleading sequence using the misleading key sequence, extract sequence fragments with residual information characteristics, and assist in voting for error correction.

[0102] Alternatively, the mixed-stored DNA sequences are decoded using the normal key sequence, and sequence fragments are obtained from all the sequences.

[0103] Alternatively, the mixed-stored DNA sequences are decoded using the misleading key sequence, and sequence fragments that are misjudged due to noise perturbation but still contain information remnants are captured.

[0104] The device provided in this embodiment has the following functions: First, it uses the misleading key sequence to modulate and decode the misleading sequence, extracts sequence fragments with residual information characteristics, provides strong assistance for vote error correction, and improves the accuracy of error correction. Second, it decodes the mixed-stored DNA sequences through the normal key sequence, can accurately obtain the required sequence fragments from all sequences, and helps locate key information. Third, by using the misleading key sequence to decode the mixed-stored DNA sequences, it can capture sequence fragments that, although misjudged due to noise disturbance, still retain information remnants, provide more possibilities for information recovery, enhance the ability to mine and utilize valid information in a complex storage environment, and ensure the integrity and reliability of information processing.

[0105] In one embodiment, the encoding module 43 includes an encoding unit and a mixing unit.

[0106] The encoding unit is used to perform modulation encoding on the indexed binary sequence respectively according to the modulation encoding rules using the normal key sequence and the misleading key sequence, and generate a normal sequence and a misleading sequence.

[0107] The mixing unit is used to copy the misleading sequence a preset number of times and then mix it with the normal sequence for storage.

[0108] The device provided in this embodiment uses the normal key sequence k_n and the misleading key sequence k_m to perform modulation encoding on the indexed binary sequence according to the modulation encoding rules, converts digital information into the form of DNA sequences, realizes the encryption of the original information. The normal sequence carries the correct encrypted information, and the misleading sequence is used as an interference element to enhance the encryption complexity. The generated misleading sequence can effectively interfere with the analysis of potential attackers, increase the difficulty of cracking, strengthen the security protection. Copy the misleading sequence a preset number of times and mix it with the normal sequence to construct a misleading sequence, which greatly increases the degree of data confusion, makes the normal sequence difficult to be recognized and separated, enhances the concealment of information storage, and reduces the risk of the normal sequence being discovered and attacked.

[0109] In one embodiment, the decryption module 45 includes a screening unit, a decoding unit, and a plaintext information obtaining unit.

[0110] The screening unit is used to screen the mixed-stored DNA sequences using the normal key sequence and the misleading key sequence to obtain a set of normal sequences.

[0111] The decoding unit is used to demodulate and restore each sequence to a binary fragment based on the set of normal sequences using the modulation decoding rules.

[0112] The plaintext information obtaining unit uses the index to cluster and sort the binary fragments to obtain the plaintext information.

[0113] The device provided in this embodiment significantly improves the reliable recovery ability and security of DNA stored data in high-noise environments such as synthesis and sequencing through a three-level collaborative mechanism of "dual-key screening - modulation decoding - index clustering". Dual-key screening uses a normal key and a misleading key to collaboratively filter the mixed sequencing set, actively excluding noise interference and maliciously forged sequences, enhancing the purity of the normal sequence set R_n. At the same time, the misleading key induces attackers to decrypt invalid information, enhancing the ability to resist side-channel attacks; modulation decoding and binary reduction are based on redundant mapping rules and index driving, tolerating base substitution errors and short fragment losses, achieving local error correction during the demodulation process, and reorganizing disordered or fragmented binary fragments through index clustering, effectively repairing data logic disorders caused by insertion / deletion errors; index clustering and sorting further combine the multi-copy consensus algorithm to eliminate residual error fragments, ensuring the integrity and accuracy of the plaintext information.

[0114] The embodiment of the present invention is a device embodiment corresponding to the above method embodiment. The specific operations of each module processing step can be understood with reference to the description of the method embodiment and will not be elaborated here.

[0115] As Figure 5 shown, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the DNA storage encryption method with dual-noise collaborative enhancement in the above embodiment, or when the computer program is executed by a processor, it implements the DNA storage encryption method with dual-noise collaborative enhancement in the above embodiment.

[0116] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0117] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and the content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.

Claims

1. A dual noise synergistically enhanced DNA storage encryption method, characterized in that: The following steps are involved: Convert the plaintext information to be encrypted into a binary sequence with an index; Randomly generate normal key sequences and misleading key sequences; Modulate and encode the binary sequence using the normal key sequence and the misleading key sequence to generate a normal sequence and a misleading sequence of DNA; Injecting structural errors into the misleading sequence and mixing it with the normal sequence before storing; A mixed sequencing set is obtained from the mixed stored DNA sequences, and the mixed sequencing set is screened and decrypted using the normal key sequence and the misleading key sequence to obtain plaintext information.

2. The dual noise synergistically enhanced DNA storage encryption method according to claim 1, characterized in that: The injecting of structural errors into the misleading sequence specifically includes the following steps: The synthesis conditions of the DNA chain at a specific site are adjusted to increase the probability of the misleading sequence making an error at the specific site.

3. The dual noise synergistically enhanced DNA storage encryption method according to claim 1, characterized in that: The following steps are also included: Using the misleading key sequence to modulate and decode the misleading sequence, extracting sequence segments with residual information characteristics, and assisting in voting error correction; or, using the normal key sequence to decode the mixed stored DNA sequence to obtain sequence fragments from the entire sequence; Or, the misleading key sequence is used to decode the mixed stored DNA sequence to capture sequence fragments that are misjudged due to noise disturbance but still contain information fragments.

4. The dual noise synergistically enhanced DNA storage encryption method according to claim 1, characterized in that: The method of using the normal key sequence and the misleading key sequence to modulate and encode the binary sequence to generate a normal sequence and a misleading sequence of DNA specifically includes the following steps: Using the normal key sequence and the misleading key sequence to modulate and encode the indexed binary sequence using a modulation and coding rule, respectively, to generate a normal sequence and a misleading sequence; The misleading sequence is copied for a preset number of times, mixed with the normal sequence and then stored.

5. The dual noise synergistically enhanced DNA storage encryption method according to claim 1, characterized in that: The using the normal key sequence and the misleading key sequence to screen and decrypt the mixed sequencing set to obtain plaintext information specifically includes the following steps: Using the normal key sequence and the misleading key sequence to screen the mixed sequencing set to obtain a normal sequence set; Based on the normal sequence set, each sequence is demodulated and restored to a binary fragment using a modulation decoding rule; The binary fragments are sorted using the index to obtain plaintext information.

6. A dual noise synergistically enhanced DNA storage encryption device, characterized in that: It includes a conversion module, a key generation module, an encoding module, an error injection module and a decryption module; The conversion module is used to convert the plaintext information to be encrypted into a binary sequence with an index; The key generation module is used to randomly generate a normal key sequence and a misleading key sequence; The encoding module is used to modulate and encode the binary sequence using the normal key sequence and the misleading key sequence to generate a normal sequence and a misleading sequence of DNA; The error injection module is used to inject structural errors into the misleading sequence and store them after mixing with the normal sequence; The decryption module is used to obtain a mixed sequencing set from the mixed stored DNA sequences, and use the normal key sequence and the misleading key sequence to screen and decrypt the mixed sequencing set to obtain plaintext information.

7. The dual noise synergistically enhanced DNA storage encryption device according to claim 6, characterized in that: It also includes an auxiliary module configured to use the misleading key sequence to modulate and decode the misleading sequence, extract sequence segments with residual information characteristics, and assist in voting error correction; or, using the normal key sequence to decode the mixed stored DNA sequence to obtain sequence fragments from the entire sequence; Or, the misleading key sequence is used to decode the mixed stored DNA sequence to capture sequence fragments that are misjudged due to noise disturbance but still contain information fragments.

8. The dual noise synergistically enhanced DNA storage encryption device according to claim 6, characterized in that: The encoding module includes an encoding unit and a mixing unit; The encoding unit is used to use the normal key sequence and the misleading key sequence to modulate and encode the indexed binary sequence using a modulation and coding rule to generate a normal sequence and a misleading sequence; The mixing unit is used to copy the misleading sequence a preset number of times, mix it with the normal sequence and store it.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the dual noise synergistically enhanced DNA storage encryption method as described in any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the dual-noise synergistically enhanced DNA storage encryption method as described in any one of claims 1 to 5 are implemented.