Method and apparatus for encryption and coding during DNA storage, and electronic device and storage medium
By using hyperchaotic pseudo-random sequence generator and DNA Raptor encoding in DNA storage technology, the encryption process is integrated into the DNA encoding process, which solves the problem of data security in DNA storage that is not considered, achieves the satisfaction of high information density and multiple constraints, and has good error correction capabilities.
Patent Information
- Application Number
- PCT/CN2023/139238
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-19
AI Technical Summary
There is a lack of effective data security guarantees in existing DNA storage technologies, and mainstream encoding methods cannot achieve high information density and data security at the same time.
The superchaotic pseudo-random sequence generator is used to generate random number sequences, combining DNA Raptor encoding and DNA base mapping rules, the encryption process is integrated into the DNA encoding process, and the target DNA sequence is generated.
It effectively improves the data security in DNA storage process, solves the problem that data security is not considered, and meets high information density and multiple constraints, and has good error correction capabilities.
Smart Images

Figure CN2023139238_19062025_PF_FP_ABST
Abstract
Description
Encryption coding method, device, electronic device and storage medium in DNA storage Technical Field
[0001] The present application relates to the field of DNA storage technology. Specifically, the present application relates to an encryption coding method, device, electronic device and storage medium in DNA storage. Background Art
[0002] DNA is a promising storage medium, offering high density, long durability, and low maintenance compared to traditional storage media. Theoretically, the entirety of human history could be stored in a space roughly the size of a double garage. These properties make DNA an ideal choice for information storage, promising large-scale practical applications in the future. Consequently, research into data storage using DNA has become a hot topic at the intersection of computing and biology. The general DNA storage process consists of six steps: encoding, synthesis, storage, retrieval, sequencing, and decoding. To meet the requirements of DNA storage, a wide range of encoding methods have been proposed, driven by considerations such as cost and biochemical technology. Mainstream encoding methods can be divided into two categories: encoding based on fixed rule mapping, such as DNA cryptography; and encoding based on filtering operations, such as DNA fountain codes.
[0003] However, both of these mainstream coding methods have significant limitations. For example, DNA fountain codes are suitable for DNA storage technology, but they fail to address data security concerns. The data is completely transparent during the encoding and decoding process, making it susceptible to privacy leaks. In theory, conventional encryption methods can only be used to encrypt data before encoding, making the encoding process more complex and completely dependent on the conventional encryption algorithm itself, making it easier to crack. DNA cryptography, on the other hand, focuses solely on encryption using DNA base calculation and mapping rules, and cannot directly generate DNA sequences that meet the high information density and specific constraints required for DNA storage.
[0004] It can be seen that how to ensure data security during DNA storage remains to be solved.
[0005] Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method, device, electronic device, and storage medium for DNA storage encryption. The technical solution is as follows:
[0007] According to one aspect of the present application, a method for encryption coding in DNA storage includes: obtaining data to be stored, and a first random number sequence and a second random number sequence generated by a hyperchaotic pseudo-random sequence generator; encrypting the data to be stored using the first random number sequence to obtain source data; performing DNA Raptor encoding on the source data, and introducing the second random number sequence into the DNA Raptor encoding process as a random number seed to obtain at least one encoded data; based on the DNA base mapping rule, converting each of the encoded data into a corresponding DNA sequence, and generating a target DNA sequence from the DNA sequence corresponding to each of the encoded data.
[0008] According to one aspect of the present application, an encryption coding device for DNA storage includes: a data acquisition module for acquiring data to be stored, and a first random number sequence and a second random number sequence generated by a hyperchaotic pseudo-random sequence generator; a data encryption module for encrypting the data to be stored using the first random number sequence to obtain source data; a DNA encoding module for performing DNA Raptor encoding on the source data and introducing the second random number sequence into the DNA Raptor encoding process as a random number seed to obtain at least one encoded data; a DNA mapping module for converting each of the encoded data into a corresponding DNA sequence based on a DNA base mapping rule, and generating a target DNA sequence from the DNA sequence corresponding to each of the encoded data
[0009] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the encryption encoding method in DNA storage as described above.
[0010] According to one aspect of the present application, a storage medium stores computer-readable instructions thereon, wherein the computer-readable instructions are executed by one or more processors to implement the encryption encoding method in DNA storage as described above.
[0011] According to one aspect of the present application, a computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the encryption encoding method in DNA storage as described above.
[0012] The beneficial effects of the technical solution provided by this application are:
[0013] In the above technical solution, first, a hyperchaotic pseudo-random sequence generator generates a first random number sequence and a second random number sequence. Then, on the one hand, the first random number sequence is used to encrypt the data to be stored to obtain source data. On the other hand, the source data is DNA Raptor encoded using DNA Raptor encoding that introduces the second random number sequence as a random number seed to obtain at least one encoded data. Finally, a target DNA sequence is generated by converting each encoded data according to the DNA base mapping rule. Therefore, based on the requirements of encoding and encryption in DNA storage, it is proposed to combine the hyperchaotic pseudo-random sequence generator in the chaotic system and the DNA Raptor encoding in the DNA fountain code in the DNA encryption coding scheme to integrate the data encryption process into the DNA encoding process, thereby fully ensuring the data security in the DNA storage process, thereby effectively solving the problem that the related technology does not consider data security in the DNA storage process. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.
[0015] FIG1 is a hardware structure diagram of an electronic device according to an exemplary embodiment;
[0016] FIG2 is a flow chart showing an encryption encoding method in DNA storage according to an exemplary embodiment;
[0017] FIG3 is a schematic diagram showing a pseudo-random sequence generation process according to an exemplary embodiment;
[0018] FIG4 is a schematic diagram showing encryption using a first random number sequence according to an exemplary embodiment;
[0019] FIG5a is a schematic diagram showing a constraint matrix according to an exemplary embodiment;
[0020] FIG5 b is a schematic diagram showing DNA Raptor encoding according to an exemplary embodiment;
[0021] FIG6 is a schematic diagram showing DNA mapping and screening according to an exemplary embodiment;
[0022] FIG7 is a schematic diagram of a specific implementation of an encryption encoding method in DNA storage in an application scenario;
[0023] FIG7a is a schematic diagram of data to be stored involved in the application scenario shown in FIG7;
[0024] FIG7 b is a schematic diagram of a target DNA sequence involved in the application scenario shown in FIG7 ;
[0025] FIG8 is a structural block diagram of an encryption encoding device in DNA storage according to an exemplary embodiment;
[0026] Fig. 9 is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0027] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0028] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0029] The following is an introduction and explanation of several terms involved in this application:
[0030] DNA storage: DNA storage is a data storage method that uses DNA (deoxyribonucleic acid) molecules as a medium to store digital data. DNA is the molecule that stores genetic information in organisms. Its excellent information density and long-term stability make it considered a potential high-efficiency, high-capacity, and long-lasting data storage medium. The basic concept of DNA storage is to encode digital data into DNA sequences, which are then synthesized through synthetic chemistry and stored in test tubes or other containers. When the data needs to be retrieved, the DNA sequence can be read out through DNA sequencing technology and decoded back into digital data.
[0031] DNA Cryptography: DNA cryptography is a field that exploits the properties and molecular structure of DNA to encrypt and store information. Researchers use biochemical techniques, DNA encoding / decoding, and base-counting rules to physically or logically encrypt data, ultimately synthesizing specific DNA sequences.
[0032] Chaotic systems: Chaotic systems are dynamical systems that are deterministic, non-periodic, and extremely sensitive to initial conditions. Such systems exhibit seemingly disordered, complex, and unpredictable behavior, even when their motion is described by simple nonlinear equations or rules. In practical applications, chaotic systems are also used in data encryption, random number generation, and information hiding. Chaotic cryptography, based on chaotic systems, has become a new branch of cryptography.
[0033] Fountain codes: Fountain codes are an error-correcting coding technique originating from communications research. They aim to generate an unlimited number of coded symbols at the transmitter to meet the receiver's required information, eliminating the need to know the length of the transmitted data in advance. They offer three key advantages: no fixed length, randomness, and efficient error correction.
[0034] Current research in the field of DNA storage coding focuses on encoding methods that achieve high information density (the number of binary data bits represented by each nucleotide) under specified constraints (based on biochemical technologies such as DNA synthesis, storage, and sequencing), while less attention is paid to the information security of the DNA storage process. At the same time, traditional DNA cryptography does not take measures to deal with sequence errors generated by the DNA storage process (i.e., errors in the stored data), which may result in incomplete or incorrect decrypted data.
[0035] To this end, the encryption coding method in DNA storage provided in this application can effectively improve the data security in DNA storage. Accordingly, the encryption coding method in DNA storage is suitable for an encryption coding device in DNA storage, and the encryption coding device in DNA storage can be deployed in an electronic device, which can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a laptop computer, a server, etc.
[0036] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0037] Fig. 1 shows a schematic structural diagram of an electronic device according to an exemplary embodiment.
[0038] It should be noted that the electronic device is only an example adapted for the present application and should not be considered to provide any limitation on the scope of use of the present application. The electronic device should not be interpreted as needing to rely on or necessarily having one or more components of the exemplary electronic device 200 shown in FIG1 .
[0039] The hardware structure of the electronic device 200 may vary greatly due to different configurations or performances. As shown in FIG1 , the electronic device 200 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .
[0040] Specifically, the power supply 210 is used to provide operating voltage for various hardware devices on the electronic device 200 .
[0041] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, as shown in FIG1 , and this is not a specific limitation.
[0042] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.
[0043] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 200 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0044] Application 253 is a computer-readable instruction that performs at least one specific task based on operating system 251. It may include at least one module (not shown in FIG1 ), each of which may contain computer-readable instructions for electronic device 200. For example, the encryption encoding device in DNA storage can be considered as application 253 deployed in electronic device 200.
[0045] The data 255 may be photos, pictures, etc. stored in a disk, or may be target files, DNA sequences, etc. stored in the memory 250 .
[0046] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer-readable instructions stored in the memory 250, thereby performing operations and processing on the massive amount of data 255 in the memory 250. For example, the encryption encoding method in DNA storage can be implemented by the central processing unit 270 reading a series of computer-readable instructions stored in the memory 250.
[0047] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.
[0048] Please refer to FIG2 . An embodiment of the present application provides an encryption coding method in DNA storage. The method is applicable to an electronic device, and the hardware structure of the electronic device is shown in FIG1 .
[0049] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.
[0050] As shown in FIG2 , the method may include the following steps:
[0051] Step 310: Acquire data to be stored, and a first random number sequence and a second random number sequence generated by a hyperchaotic pseudo-random sequence generator.
[0052] The data to be stored may be any type of data, such as text, pictures, audio, video, etc., which is not limited here.
[0053] The first random number sequence is used for encrypting the data to be stored, and the second random number sequence is used for encoding the data to be stored. In some embodiments, the first random number sequence and the second random number sequence are generated by a hyperchaotic pseudo-random sequence generator.
[0054] In some embodiments, a hyperchaotic pseudo-random sequence generator is obtained by using the calculation of a five-dimensional hyperchaotic system and accompanying additional operations such as perturbations, floating point rounding, large number processing, and variable exclusive OR. The calculation formula of the five-dimensional hyperchaotic system is as follows:
[0055] Among them, a, b, c, d, e, h, r, m, k, p, and q are constant parameters, while x, y, z, w, and u are state variables.
[0056] Figure 3 shows a schematic diagram of a hyperchaotic pseudo-random sequence generator constructed using a five-dimensional hyperchaotic system in one embodiment. In Figure 3, the input of the hyperchaotic pseudo-random sequence generator includes: state variables x, y, z, w, u, and the sequence length n of the first random number sequence / second random number sequence. The output includes: the first random number sequence / second random number sequence, and the first random number sequence / second random number sequence contains n keys.
[0057] As shown in Figure 3, the process of generating a first random number sequence / a second random number sequence by a hyperchaotic pseudo-random sequence generator may include the following steps: first, initializing the initial parameters a, b, c, d, e, h, r, m, k, p, and q configured for the hyperchaotic pseudo-random sequence generator; second, inputting the first and second keys into the hyperchaotic pseudo-random sequence generator to perform pseudo-random sequence generation operations, thereby obtaining the corresponding first and second random number sequences. The second step specifically includes: when the sequence lengths of the first and second random number sequences have not reached n, perturbing them at fixed intervals, calculating a specified number of times according to a calculation formula for the hyperchaotic system, selecting at least one set of a specified number of state variables to generate key values and adding them to the first and second random number sequences until the sequence lengths of the first and second random number sequences reach n. Furthermore, the process of generating a key value for each set of a specified number of state variables specifically includes: converting a real number into a fixed-length integer sequence, performing XOR operations on the five integer parts to obtain a key value, and adding the key value to the first and second random number sequences.
[0058] Step 330: Encrypt the data to be stored using the first random number sequence to obtain source data.
[0059] In some embodiments, the encryption process may include the following steps: obtaining a global hash value of the data to be stored, and using the global hash value to preprocess the data to be stored to obtain multiple encoding blocks; performing a logical operation on the first random data sequence and each encoding block to obtain the source data.
[0060] FIG4 shows a schematic diagram of encryption using a first random number sequence. In FIG4 , on the one hand, the SHA256 algorithm is used to perform a hash operation on the data to be stored (e.g., the target file) to obtain a global hash value, and the global hash value is stored in the first coding block, the identifier of which is 0, that is, the first coding block is coding block 0; on the other hand, the data to be stored is divided into blocks to obtain multiple coding blocks, each of which is identified by 1 to n, that is, these coding blocks are coding block 1, coding block 2, ..., coding block n. After obtaining the above n+1 coding blocks, these n+1 coding blocks are encrypted and traversed. The encryption traversal process includes: performing a logical operation on coding block 0 and coding block 1, and storing the logical operation result in coding block 1; performing a logical operation on coding block 1 and coding block 2, and storing the logical operation result in coding block 2; and so on, performing a logical operation on coding block n-1 and coding block n, and storing the logical operation result in coding block n; until the nth coding block (i.e., coding block n) is traversed, obtaining n+1 coding blocks. Based on these n+1 coding blocks and the first random number sequence, the source data can be obtained through logical operations to achieve encryption of the data to be stored.
[0061] It should be noted that the logical operation in this embodiment refers to XOR. Of course, in other embodiments, the logical operation may also include but is not limited to: AND, OR, NOT, XNOR, etc. This embodiment does not constitute a specific limitation to this.
[0062] Step 350 , performing DNA Raptor encoding on the source data, and introducing the second random number sequence into the DNA Raptor encoding process as reference table data to obtain at least one encoded data.
[0063] In some embodiments, DNA Raptor encoding includes pre-encoding based on a constraint matrix and LT encoding, wherein the pre-encoding uses the constraint matrix to encode source data into multiple intermediate symbols; and the LT encoding is used to encode multiple intermediate symbols into multiple encoded data.
[0064] Specifically, the DNA Raptor encoding process can include the following steps: expanding the source data and pre-encoding the expanded source data using a constraint matrix to obtain multiple intermediate symbols; performing LT encoding on each intermediate symbol to obtain multiple encoded data; and each encoded data corresponds one-to-one with each intermediate symbol. It should be noted that expansion here refers to dividing the source data into multiple source symbols, using K source symbols as a source block, adding (S+H) zero elements before the K source symbols in each source block to form an expanded source symbol, and ultimately obtaining the expanded source data from the multiple expanded source symbols.
[0065] FIG5a shows a schematic diagram of a constraint matrix in one embodiment. In FIG5a , the constraint matrix is composed of GLDPC , G Half , I S , I H , Z, G LT Among them, G LDPC is the S×K dimensional generator matrix of the LDPC symbol; G Half is the H×(K+S)-dimensional generator matrix of the Half symbol (Gray code); I S is the S×S dimensional identity matrix; I H is the H×H dimensional identity matrix; Z is the S×H dimensional zero matrix; G LT is a K×L dimensional generator matrix of LT encoding symbols. In some embodiments, the G in the constraint matrix LT The matrix, as well as the degree value and random number in the LT encoding process, are all generated by a random number generator using the second random number sequence as reference table data.
[0066] Figure 5b shows a schematic diagram of the DNA Raptor encoding process in one embodiment. As shown in Figure 5b, first, the expanded source data D is pre-encoded using the constraint matrix A. That is, the expanded source data D is multiplied by the inverse matrix of the constraint matrix A to obtain L intermediate symbols C. Then, the L intermediate symbols are LT-encoded to ultimately obtain n encoded data. Specifically, the process includes: first, using the degree distribution function as the degree value of the LT encoding, randomly selecting an integer value d from the range of 1 to n; second, generating d random numbers from the range of 1 to n, selecting the intermediate symbol corresponding to the random number from the n intermediate symbols, and performing an XOR operation on the selected intermediate symbol to generate a piece of encoded data; third, repeating the first and second steps until n pieces of encoded data are generated. In the above-mentioned DNA Raptor encoding process, a random generator (i.e., random function R(x)) is introduced, and the random base tables V0 and V1 of the random generator are composed of a second random number sequence generated by a hyperchaotic pseudo-random sequence generator inputted with the key Key1.
[0067] Step 370 : Based on the DNA base mapping rule, each encoding data is converted into a corresponding DNA sequence, and a target DNA sequence is generated from the DNA sequence corresponding to each encoding data.
[0068] After obtaining multiple coded data, each coded data can be mapped to a corresponding DNA sequence according to a given mapping rule.
[0069] In some embodiments, the given mapping rule includes a DNA base mapping rule, which substantially reflects the correspondence between different binary data and different DNA bases.
[0070] Specifically, as shown in FIG6 , the mapping process based on the DNA base mapping rule can include the following steps: for n pieces of coded data, determining the DNA base mapping rule corresponding to each piece of coded data, and generating a corresponding DNA sequence from each piece of coded data according to the determined DNA base mapping rule; using the identifier of the coded data as a random number seed, and obtaining the error correction code generated for the coded data; according to the determined DNA base mapping rule, converting the random number seed and the error correction code into corresponding DNA bases, and adding the converted DNA bases to the DNA sequence corresponding to the coded data. It is noted that the DNA base corresponding to the random number seed can be added before the DNA sequence, and the DNA base corresponding to the error correction code can be added after the DNA sequence, but this is not limited here.
[0071] In this way, the data error correction capability is improved by increasing the length of the DNA sequence, so that the DNA sequence itself has a certain redundancy and thus a certain error correction capability, and the stored information can be restored to the original file at a certain error rate.
[0072] After obtaining the DNA sequences corresponding to each coded data, these DNA sequences will be screened according to given constraints to obtain the final stored target DNA sequence.
[0073] In some embodiments, the given constraints include biological constraints, including but not limited to homopolymers, GC content, palindromes, etc.
[0074] Specifically, referring to FIG6 , the screening process based on biological constraints may include the following steps: performing biological constraint detection on the DNA sequence corresponding to each coding data; regenerating the coding data corresponding to the DNA sequence that does not meet the biological constraint through the LT code in the DNA Raptor code until the DNA sequence corresponding to the regenerated coding data meets the biological constraint; and generating a target DNA sequence from all n DNA sequences that meet the biological constraint.
[0075] Through the above process, considering the requirements of encoding and encryption in DNA storage, it is proposed to combine the hyperchaotic pseudo-random sequence generator in the chaotic system and the DNA Raptor code in the DNA fountain code in the DNA encryption coding scheme to integrate the data encryption process into the DNA coding process, thereby fully ensuring the data security in the DNA storage process, thereby effectively solving the problem that related technologies do not consider data security in the DNA storage process.
[0076] In addition, while ensuring high information density of DNA storage, it meets multiple custom constraints and solves the problem of a small number of sequence errors generated during the DNA storage process.
[0077] FIG7 is a schematic diagram showing a specific implementation of an encryption encoding method for DNA storage in an application scenario, including but not limited to semiconductor application scenarios and biological application scenarios.
[0078] In this application scenario, the encryption encoding process in DNA storage is divided into four parts: hyperchaotic pseudo-random sequence generation, file preprocessing, DNA Raptor encoding, and mapping screening. The specific steps are as follows:
[0079] The first step is to generate a hyperchaotic pseudo-random sequence. The steps are as follows:
[0080] (1) The initial parameters of the hyperchaotic system are determined based on the key Key1 (a set of state variables). The initial hyperchaotic pseudo-random sequence generation operation is performed to obtain 512 4-byte unsigned integers. These are used as the reference tables V0 and V1 (each containing 256 4-byte unsigned integers) for the random number generator Rand[X, i, m] in the Raptor code. Rand[X, i, m] is used to generate random numbers for the LT encoding process and is defined as follows:
[0081] Where X represents the input value, i is an index variable, and m is the modulus to be applied to the result. The values of V0 and V1 are combined using a bitwise XOR operation, and then the modulus operation is applied to get the final result, which is an integer between 0 and m-1.
[0082] (2) Generate a random number sequence for subsequent encryption using the same pseudo-random sequence generator based on the key Key2.
[0083] The second step is file preprocessing, the steps are as follows:
[0084] (1) Blocking. The target file is divided into blocks, each block is 30 bytes in size and recorded as a coding block. The blocks are numbered in units of blocks. Each block has a corresponding unique ID number, which is 4 bytes in size (the last 2 bytes record the number). The ID number is not involved in encryption. At the same time, the global hash value of the entire file is calculated using SHA256, and its 30 bytes are taken as the initial hash key and stored in the first coding block. The block generated by the target file starts from the second coding block. The maximum number of blocks in a single block is 20,000. The number of blocks is recorded in the first 2 bytes of the ID number. The number of blocks and the block number constitute the ID.
[0085] (2) Encryption. For each coded block, XOR it with the hash key and use the result as the new hash key. Generate 20 random numbers through the constructed hyperchaotic pseudo-random sequence generator. XOR the random numbers with the hashed coded block data to complete the encryption.
[0086] The third step is DNA Raptor encoding. The steps are as follows:
[0087] (1) Pre-coding generates intermediate symbols. The number of blocks K is counted, and all coding blocks constitute the source data matrix D. The intermediate symbols C are generated through the matrix A. The matrix A is shown in Figure 2. Its G LT The random number generator Rand[X, i, m] is used to generate the intermediate symbols. For example, for 10KB of data, K = 342. A table query yields S = 31, H = 10, and L = 383 (the parameters corresponding to different K values are recorded in a specific table). Formula (1.2) ultimately generates 383 intermediate symbols.
[0088] (2) LT encoding. Intermediate symbols are encoded using LT encoding to generate encoded data. The degree value and random number for LT encoding are generated by the random number generator Rand[X, i, m]. Each intermediate symbol corresponds to a line of encoded data. LT encoding requires generating more encoded symbols than the number of source symbols (i.e., the number of blocks K). Therefore, the lowest redundancy is selected, ultimately generating 350 lines of encoded data. By setting redundancy, a corresponding amount of additional encoded data can be generated.
[0089] The fourth step is mapping screening. The steps are as follows:
[0090] (1) Mapping. Based on the result of modulo 8 for each encoded data ID, the mapping rule from 1-byte binary data to DNA bases is determined, resulting in a 4-nt DNA sequence. The DNA mapping rules are shown in Table 1. Each row of encoded data generates a 120-nt DNA sequence. In addition, a 4-byte random number seed (i.e., block ID) is added before the sequence, and a 2-byte RS error correction code generated by the 30-byte encoded block is added after the sequence. These 6-byte data are then converted into DNA bases according to the mapping rules. Finally, a 144-nt DNA sequence is obtained.
[0091] Table 1 DNA mapping rules
[0092] (2) Screening. Each DNA sequence is tested for biological constraints, i.e., homopolymer and GC content. Based on the test results, DNA sequences that pass the constraints are retained, and DNA sequences that do not meet the constraints are removed. The above steps are restarted from the LT code until all the coding data are converted into qualified DNA sequences.
[0093] It is worth mentioning that DNA decryption and decoding refers to restoring the DNA sequence to the target file. Its overall process corresponds one-to-one to the encryption encoding process, that is, the DNA sequence is converted into the original target file through three steps: reverse mapping screening, DNA Raptor decoding, and file recovery. Each step is reversible, and the random numbers used are generated by the hyperchaotic pseudo-random sequence generator with the same key, so they are also equal, ensuring that the data obtained by decryption and decoding is correct and complete.
[0094] Based on the above process, the encryption coding method for DNA storage proposed in this application can encrypt and encode any type of data, and finally convert it into a DNA sequence for DNA storage. The redundancy of the DNA sequence (the number of additional DNA chains that meet the 100% decoding condition) and the constraint rules are customizable. The simulation experimental analysis is based on 20% redundancy and meeting two commonly used biological constraint conditions (GC content 40%-60%, homopolymer length less than 4).
[0095] As shown in Figure 7a, taking 10KB of text as the data to be stored as an example, the encryption encoding method of this application is used in DNA storage, and the target DNA sequence obtained after encryption encoding is shown in Figure 7b. The number of blocks corresponding to 10KB of data to be stored is 342, which theoretically corresponds to 342 DNA chains, each of which is 144nt in length. The characteristics of the Raptor code require slightly more than 342 DNA chains for successful decoding. Therefore, the preliminary encoding result is set to 350 DNA chains, that is, DNA chains are infinitely generated and selected until they meet the requirements for successful decoding and reach the number of DNA chains of 350. Taking these 350 DNA chains as the preliminary result, a redundancy of 0.26 and an information capacity of 1.59 bits / nt can be achieved (the theoretical upper limit is 2 bits / nt). In addition, the data error correction capability can be improved by increasing the number of DNA chains. Therefore, 20% redundancy is selected, and an additional 20% of DNA chains are generated, resulting in 411 DNA chains, achieving a redundancy of 0.512 and an information capacity of 1.32 bits / nt.
[0096] In this application scenario, in order to achieve the encryption effect, conventional encryption algorithm performance analyses such as key space analysis, key sensitivity analysis, correlation analysis, information entropy analysis, ciphertext change rate analysis, and randomness analysis were carried out to prove that the encryption effect of the present invention ensures data security; in terms of error correction performance, the decoding success rate of different error rates (based on sequence loss and base errors) under different redundancies was tested, reflecting good error correction performance.
[0097] The following is an embodiment of the device of the present application, which can be used to implement the encryption encoding method in DNA storage involved in this application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the encryption encoding method in DNA storage involved in this application.
[0098] Please refer to Figure 8. An embodiment of the present application provides an encryption encoding device 900 in DNA storage, including but not limited to: a data acquisition module 910, a data encryption module 930, a DNA encoding module 950 and a DNA mapping module 950.
[0099] The data acquisition module 910 is used to acquire the data to be stored, and the first random number sequence and the second random number sequence generated by the hyperchaotic pseudo-random sequence generator.
[0100] The data encryption module 930 is configured to encrypt the data to be stored using a first random number sequence to obtain source data.
[0101] The DNA encoding module 950 is configured to perform DNA Raptor encoding on the source data and introduce the second random number sequence into the DNA Raptor encoding process as reference table data to obtain at least one encoded data.
[0102] The DNA mapping module 970 is used to convert each encoding data into a corresponding DNA sequence based on the DNA base mapping rule, and generate a target DNA sequence from the DNA sequence corresponding to each encoding data.
[0103] In an exemplary embodiment, the apparatus 900 further includes: a random number sequence generation module.
[0104] Among them, the random number sequence generation module is used to determine the initial parameters configured for the hyperchaotic pseudo-random sequence generator; initialize the hyperchaotic pseudo-random sequence generator, and input the first key and the second key into the hyperchaotic pseudo-random sequence generator respectively to perform pseudo-random sequence generation operations to obtain the corresponding first random number sequence and second random number sequence.
[0105] In an exemplary embodiment, the data encryption module 930 is further used to obtain a global hash value of the data to be stored, and use the global hash value to preprocess the data to be stored to obtain multiple coding blocks; and perform logical operations on the first random data sequence and each coding block to obtain source data.
[0106] In an exemplary embodiment, the data encryption module 930 is further used to use the SHA256 algorithm to perform a hash operation on the data to be stored to obtain a global hash value, and store the global hash value in the first coding block; the identifier of the first coding block is 0; the data to be stored is divided into blocks to obtain multiple coding blocks; the identifiers of each coding block are 1 to n respectively; the n+1 coding blocks are encrypted and traversed, and the encryption traversal processing includes: performing a logical operation on the previous coding block and the current coding block, and storing the logical operation result in the current coding block; until the nth coding block is traversed, and n+1 coding blocks are obtained.
[0107] In an exemplary embodiment, the data encoding module 950 is further configured to expand the source data and pre-encode the expanded source data using a constraint matrix to obtain a plurality of intermediate symbols; perform LT encoding on each intermediate symbol to obtain a plurality of encoded data; each encoded data corresponds to each intermediate symbol one-to-one; wherein G in the constraint matrix LT The matrix, as well as the degree value and random number in the LT encoding process, are all generated by a random number generator using the second random number sequence as reference table data.
[0108] In an exemplary embodiment, the DNA encoding module 970 is further used to determine, for each piece of encoding data, a DNA base mapping rule corresponding to the encoding data, and generate a corresponding DNA sequence from the encoding data according to the determined DNA base mapping rule; use the identifier of the encoding data as a random number seed, and obtain an error correction code generated for the encoding data; convert the random number seed and the error correction code into corresponding DNA bases according to the determined DNA base mapping rule, and add the converted DNA bases to the DNA sequence corresponding to the encoding data.
[0109] In an exemplary embodiment, the DNA encoding module 970 is further used to perform biological constraint detection on the DNA sequence corresponding to each encoding data; regenerate the encoding data corresponding to the DNA sequence that does not meet the biological constraints through DNA Raptor encoding until the DNA sequence corresponding to the regenerated encoding data meets the biological constraints; and generate a target DNA sequence from all DNA sequences that meet the biological constraints.
[0110] It should be noted that the encryption coding device in DNA storage provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when performing encryption coding in DNA storage. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the encryption coding device in DNA storage will be divided into different functional modules to complete all or part of the functions described above.
[0111] In addition, the encryption coding device in DNA storage provided in the above embodiment and the encryption coding method in DNA storage belong to the same concept, and the specific way in which each module performs operations has been described in detail in the method embodiment and will not be repeated here.
[0112] Please refer to FIG9 . An electronic device 4000 is provided in an embodiment of the present application. The electronic device 4000 may include a desktop computer, a laptop computer, a server, etc.
[0113] In FIG. 9 , the electronic device 4000 includes at least one processor 4001 and at least one memory 4003 .
[0114] Data exchange between the processor 4001 and the memory 4003 can be achieved via at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG9 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0115] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0116] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0117] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions or codes in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.
[0118] Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .
[0119] The computer-readable instructions are executed by one or more processors 4001 to implement the encryption encoding method in DNA storage in the above-mentioned embodiments.
[0120] In addition, an embodiment of the present application provides a storage medium on which computer-readable instructions are stored. The computer-readable instructions are executed by one or more processors to implement the encryption encoding method in DNA storage as described above.
[0121] In an embodiment of the present application, a computer program product is provided, which includes computer-readable instructions. The computer-readable instructions are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the encryption encoding method in DNA storage as described above.
[0122] Compared to the existing fountain code encoding technology in DNA storage, this application integrates encryption operations into the encoding process, ensuring that the encoding results meet the requirements of DNA storage, providing data encryption effects, improving data security, and retaining its advantages in the field of DNA storage encoding. At the same time, a new hyperchaotic pseudo-random sequence generator is constructed. A new super five-dimensional chaotic system with good chaotic characteristics is constructed, and a series of data processing operations are added to it to form a corresponding hyperchaotic pseudo-random sequence generator to generate random numbers with strong randomness.
[0123] Furthermore, compared to encryption methods in DNA cryptography, this application is not limited to image encryption and does not require consideration of the type or size of the data to be stored. The length of the generated DNA sequence can be freely selected, the constraints can be freely set, and the amount of redundancy can be freely determined. This approach has a wider scope of application and can be adapted to any current DNA storage system, generating DNA sequences of arbitrary length and redundancy while satisfying any constraints, achieving similar encryption results. Compared to existing encryption methods, the DNA sequences generated by this application have a higher information density and can satisfy a greater number of constraints.
[0124] Finally, in response to the data errors caused by the current technology in the DNA storage process, this application has a certain error correction capability and can restore the stored information to the original file at a certain error rate.
[0125] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0126] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for encryption coding in DNA storage, characterized in that, The method includes: obtaining the data to be stored, as well as a first random number sequence and a second random number sequence generated by a hyperchaotic pseudorandom sequence generator; encrypting the data to be stored by using the first random number sequence to obtain source data; performing DNA Raptor encoding on the source data, and introducing the second random number sequence into the DNA Raptor encoding process as a random number seed to obtain at least one encoded data; converting each of the encoded data into a corresponding DNA sequence based on the DNA base mapping rule, and generating a target DNA sequence from the DNA sequences corresponding to each of the encoded data.
2. The method according to claim 1, characterized in that, Before obtaining the first random number sequence and the second random number sequence generated by the hyperchaotic pseudorandom sequence generator, the method further includes: initializing the initial parameters configured for the hyperchaotic pseudorandom sequence generator; inputting a first key and a second key into the hyperchaotic pseudorandom sequence generator respectively to perform pseudorandom sequence generation operations, and obtaining the corresponding first random number sequence and the second random number sequence.
3. The method according to claim 1, characterized in that, The encrypting the data to be stored by using the first random number sequence to obtain source data includes: obtaining the global hash value of the data to be stored, and preprocessing the data to be stored by using the global hash value to obtain a plurality of encoded blocks; performing a logical operation on the first random data sequence and each of the encoded blocks to obtain the source data.
4. The method according to claim 3, characterized in that, The obtaining the global hash value of the data to be stored, and preprocessing the data to be stored by using the global hash value to obtain a plurality of encoded blocks includes: performing a hash operation on the data to be stored by using the SHA256 algorithm to obtain the global hash value, and storing the global hash value into the first encoded block; the identifier of the first encoded block is 0; blocking the data to be stored to obtain a plurality of the encoded blocks; the identifiers of each of the encoded blocks are 1 to n respectively; performing an encrypted traversal process on the n + 1 encoded blocks, where the encrypted traversal process includes: taking the previous encoded block and performing a logical operation on the current encoded block, and storing the logical operation result into the current encoded block; until the nth encoded block finishes traversal, obtaining the n + 1 encoded blocks.
5. The method according to claim 1, characterized in that, The performing DNA Raptor encoding on the source data, and introducing the second random number sequence into the DNA Raptor encoding process as a random number seed to obtain at least one encoded data includes: extending the source data, and performing precoding on the extended source data by using a constraint matrix to obtain a plurality of intermediate symbols; performing LT encoding on each of the intermediate symbols to obtain a plurality of the encoded data; each of the encoded data corresponds to each of the intermediate symbols one by one; Among them, G in the constraint matrix LT matrix, as well as the degree value and random number in the LT coding process, are all generated by a random number generator using the second random number sequence as reference table data.
6. The method according to claim 1, characterized in that, The converting each of the encoded data into a corresponding DNA sequence based on the DNA base mapping rule includes: for each of the encoded data, determining the DNA base mapping rule corresponding to the encoded data, and generating a corresponding DNA sequence from the encoded data according to the determined DNA base mapping rule; Use the identifier of the encoded data as a random number seed, and obtain the error correction code generated for the encoded data; According to the determined DNA base mapping rule, convert the random number seed and the error correction code into corresponding DNA bases respectively, and add the converted DNA bases to the DNA sequence corresponding to the encoded data respectively.
7. The method according to any one of claims 1 to 6, characterized in that, The generating of the target DNA sequence from the DNA sequences corresponding to the respective encoded data includes: Perform a biological constraint detection on the DNA sequences corresponding to the respective encoded data; Regenerate the encoded data corresponding to the DNA sequences that do not meet the biological constraints through DNA Raptor encoding until the DNA sequences corresponding to the regenerated encoded data meet the biological constraints; Generate the target DNA sequence from all the DNA sequences that meet the biological constraints.
8. An encryption coding device in DNA storage, characterized in that, The apparatus includes: A data acquisition module, configured to acquire the data to be stored, as well as a first random number sequence and a second random number sequence generated by a hyperchaotic pseudo-random sequence generator ; A data encryption module, configured to encrypt the data to be stored by using the first random number sequence to obtain source data; A DNA encoding module, configured to perform DNA Raptor encoding on the source data, and introduce the second random number sequence into the DNA Raptor encoding process as reference table data to obtain at least one encoded data; A DNA mapping module, configured to convert each of the encoded data into a corresponding DNA sequence based on the DNA base mapping rule, and generate a target DNA sequence from the DNA sequences corresponding to the respective encoded data.
9. An electronic device, characterized in that, including: At least one processor and at least one memory, wherein, The memory stores computer-readable instructions; The computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the encryption and encoding method in DNA storage as described in any one of claims 1 to 7.
10. A storage medium, on which computer-readable instructions are stored, characterized in that, The computer-readable instructions are executed by one or more processors to implement the encryption and encoding method in DNA storage as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Color quantum image encryption and decryption method based on multi-chaos and DNA operation
CN113297606A
Extended coding method, system and related device of storage channel
CN114254748A
Image data DNA storage method and system, electronic equipment and storage medium
CN116258781A
Image encryption method and device, equipment and storage medium
CN116743935A
Color image encryption method and device based on two-dimensional hyperchaos and compressed sensing
CN116886269A
Cited By
Customer information data storage method and system
CN120750514A
Blood glucose data encryption transmission method and device, equipment and storage medium
CN121750345A