Data encryption method based on DNA encoding, data decryption method based on DNA encoding, intelligent terminal, and medium

By constructing a DNA encoding table based on four-base units and five-base units, encrypting and optimizing the encoded data, the problems of small amount of DNA information and poor encryption effects in the prior art are solved, and efficient data storage and encryption are achieved.

WO2025118369A1PCT designated stage expired Publication Date: 2025-06-12SHENZHEN INST OF ADVANCED TECH

Patent Information

Application Number
PCT/CN2023/141317
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2023-12-23
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the prior art, the amount of information stored based on DNA encoding is small, and the DNA information encryption effect is not ideal.

Method used

The DNA encoding table constructed based on four-base units and five-base units is used to encrypt the encoded data, and the encrypted DNA sequence is generated through DNA sequence structure optimization rules and chaotic mapping rearrangement.

Benefits of technology

It improves the storage density and encryption effect of DNA information, ensuring the stability and security of encrypted DNA sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023141317_12062025_PF_FP_ABST
    Figure CN2023141317_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention particularly relates to the technical field of data storage, and provides a data encryption method based on DNA encoding, a data decryption method based on DNA encoding, an intelligent terminal, and a medium. A solution comprises: encrypting acquired data to be encoded to obtain encrypted data to be encoded; on the basis of a pre-constructed DNA encoding table, performing DNA encoding on each byte of said encrypted data to obtain an initial base sequence, the DNA encoding table being constructed from four-base units and five-base units; and using a preset DNA sequence structure optimization rule to perform chaotic mapping adjustment and rearrangement on the initial base sequence to obtain an encrypted DNA sequence. According to the solution, the GC content in the encrypted DNA sequence can be kept in a balanced state, and the information density of the encrypted data is increased, such that the encryption effect is effectively improved, thereby reducing a storage space occupied by data to be encoded, improving the security and attack resistance of the data to be encoded, and thus guaranteeing the stability of the encrypted DNA sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Data encryption and decryption method, intelligent terminal and medium based on DNA coding Technical Field

[0001] The present invention relates to the technical field of data storage, and in particular to a data encryption and decryption method, an intelligent terminal and a medium based on DNA coding. Background Art

[0002] DNA information storage, as an emerging information storage technology, has attracted not only biologists but also experts in computer information and communications. DNA information storage has evolved from simple DNA encoding to more advanced areas such as DNA error correction, automated DNA information storage systems, and DNA encryption. As a promising future information storage method, DNA information storage must not only be low-cost and energy-efficient, but also possess strong security properties.

[0003] Although there are many studies on DNA information storage in the existing technology, there are fewer studies on the combination of DNA encryption and DNA coding. DNA encryption and DNA coding are usually integrated by establishing a simple one-to-one (adenine A-00, thymine T-01, cytosine C-10, guanine G-11) or two-to-one (A, T=0, C, G=1) mapping relationship. However, using only a simple mapping relationship cannot well meet the needs of storing large amounts of information based on DNA coding, resulting in the problem of unsatisfactory DNA information encryption effect.

[0004] Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a data encryption and decryption method, intelligent terminal and medium based on DNA coding, aiming to solve the problems in the prior art that the amount of information stored based on DNA coding is small and the DNA information encryption effect is not ideal.

[0006] In order to achieve the above objectives, the present invention provides a first aspect of a data encryption method based on DNA coding, comprising:

[0007] Get the data to be encoded;

[0008] Encrypting the data to be encoded to obtain encrypted data to be encoded;

[0009] Based on a pre-constructed DNA encoding table, DNA encoding is performed on each byte of the encrypted data to obtain an initial base sequence, wherein the DNA encoding table is constructed from four-base units and five-base units;

[0010] Using a preset DNA sequence structure optimization rule, the initial base sequence is adjusted to obtain an optimized base sequence;

[0011] The optimized base sequence is subjected to chaotic mapping rearrangement to obtain an encrypted DNA sequence.

[0012] Optionally, the DNA encoding of each byte of the encrypted data to be encoded based on a pre-constructed DNA encoding table to obtain an initial base sequence includes:

[0013] Obtaining a byte sequence based on the encrypted data to be encoded, where the byte sequence includes a plurality of bytes;

[0014] Based on the mapping rules between bytes and base units in the pre-constructed DNA encoding table, the base units corresponding to the bytes in the byte sequence are respectively obtained to obtain an initial base sequence.

[0015] Optionally, the process of constructing the DNA coding table includes:

[0016] According to the preset code table design principle, a hybrid code table is constructed using four-base units and five-base units;

[0017] Based on the mixed code table, a one-to-one correspondence between bytes and base units is established to obtain a mapping rule between bytes and base units.

[0018] Optionally, the hybrid code table is constructed using four-base units and five-base units according to a preset code table design principle, including:

[0019] Based on the four single bases, several four-base units and five-base units were constructed;

[0020] If the first base of the four-base unit is A or T, the fourth base is C or G;

[0021] If the first base of the four-base unit is C or G, the fourth base is A or T;

[0022] If the first base of the five-base unit is A, the fourth and fifth bases are both TC or TG;

[0023] If the first base of the five-base unit is T, the fourth and fifth bases are both AC or AG;

[0024] If the first base of the five-base unit is C, the fourth and fifth bases are both GA or GT;

[0025] If the first base of the five-base unit is G, the fourth base and the fifth base are both CA or CT.

[0026] Optionally, when the first base unit of two adjacent bases is a four-base unit, the initial base sequence is adjusted using a preset DNA sequence structure optimization rule to obtain an optimized base sequence, including:

[0027] If the last three bases of the previous base unit are the same as the first base of the adjacent next base unit, the last base of the previous base unit is replaced by the first base of the previous base unit;

[0028] If the last two bases of the previous base unit are the same as the first two bases of the adjacent next base unit, the last base of the previous base unit is replaced by the first base of the previous base unit;

[0029] If the last single base of the previous base unit is the same as the first three single bases of the adjacent next base unit, the last single base of the previous base unit is replaced by a preset complementary base and one complementary base is added.

[0030] Optionally, when the first base unit of two adjacent bases is a five-base unit, the initial base sequence is adjusted using a preset DNA sequence structure optimization rule to obtain an optimized base sequence, including:

[0031] If the last single base of the previous base unit is the same as the first three single bases of the adjacent next base unit, the last single base of the previous base unit is replaced by the first single base of the previous base unit.

[0032] A second aspect of the present invention provides a DNA-encoded data decryption method, which is applied to decrypt an encrypted DNA sequence obtained by any of the DNA-encoded data encryption methods described above, comprising:

[0033] The encrypted DNA sequence is rearranged based on the sequence information obtained by the initial setting of the chaotic map to restore the optimized base sequence;

[0034] Based on preset DNA sequence structure optimization rules and a pre-constructed DNA encoding table, DNA decoding is performed on the optimized base sequence to restore the decoded base sequence;

[0035] The decoded base sequence is decrypted to obtain restored data.

[0036] Optionally, performing DNA decoding on the optimized base sequence based on a preset DNA sequence structure optimization rule and a pre-constructed DNA encoding table to restore the decoded base sequence includes:

[0037] Based on a preset DNA sequence structure optimization rule, adjusting the base units in the optimized base sequence to obtain a quasi-base sequence;

[0038] Based on a pre-constructed DNA coding table, the quasi-base sequence is DNA decoded to restore the decoded base sequence.

[0039] The third aspect of the present invention provides an intelligent terminal, which includes a memory, a processor, and a DNA-encoding-based data encryption program stored in the memory and runnable on the processor. When the DNA-encoding-based data encryption program is executed by the processor, it implements any one of the steps of the above-mentioned DNA-encoding-based data encryption method.

[0040] A fourth aspect of the present invention provides a computer-readable storage medium, on which a DNA-encoding-based data encryption program is stored. When the DNA-encoding-based data encryption program is executed by a processor, the steps of any one of the above-mentioned DNA-encoding-based data encryption methods are implemented.

[0041] Compared with the existing technology, the beneficial effects of this solution are as follows:

[0042] The present invention encrypts the data to be encoded to obtain the encrypted data to be encoded, constructs a new DNA encoding table using four-base units and five-base units, forms a unique DNA encoding mapping rule, converts the encrypted data to be encoded into a unique initial base sequence, and on this basis, in order to avoid the generation of oligonucleotides and improve the information density, adopts the designed DNA sequence structure optimization rule to adjust the initial base sequence to obtain an optimized base sequence; and rearranges the optimized base sequence using chaotic mapping, so that the GC content in the generated encrypted DNA sequence is maintained in a balanced state and the information density of the encrypted data is improved, thereby effectively improving the encryption effect, thereby helping to reduce the storage space occupied by the data to be encoded, and at the same time helping to improve the security and anti-attack performance of the data to be encoded, thereby ensuring the stability of the encrypted DNA sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] FIG1 is a flow chart of a data encryption method based on DNA coding according to the present invention;

[0045] FIG2 is a schematic diagram of a data encryption method based on DNA coding according to the present invention;

[0046] FIG3 is a schematic diagram of the DNA sequence structure optimization rules of the present invention;

[0047] Figure 4 shows the Lena image before encryption;

[0048] FIG5 is a simulation diagram of GC content distribution of Lena image after AES encryption according to the present invention;

[0049] FIG6 is a simulation diagram of GC content distribution of Lena image after RSA encryption according to the present invention;

[0050] Figure 7 shows the Cameraman image before encryption;

[0051] FIG8 is a simulation diagram of the GC content distribution of the Cameraman image after AES encryption according to the present invention;

[0052] FIG9 is a simulation diagram of GC content distribution after RSA encryption of a Cameraman image according to the present invention;

[0053] FIG10 is a schematic structural diagram of the intelligent terminal of the present invention. DETAILED DESCRIPTION

[0054] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.

[0055] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0056] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0057] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0058] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0059] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0061] In response to the problems in the prior art that the amount of information stored based on DNA coding is small and the DNA information encryption effect is unsatisfactory, the present invention proposes a data encryption method based on DNA coding, which mainly includes four key steps: byte encryption, DNA coding, optimized coding and chaotic mapping. Specifically, in the byte encryption stage, symmetric encryption or asymmetric encryption is used for different data to obtain various types of data to be encoded in the form of byte sequences; in the DNA coding stage, a DNA coding table is constructed based on a mixed code table of four bases and five bases. The coding rules of the DNA coding table can effectively avoid the occurrence of extreme GC content and lay a good foundation for subsequent optimized coding; in the optimized coding stage, the main focus is on optimizing the generated initial DNA sequence structure, replacing and interrupting oligonucleotides; in the chaotic mapping stage, a chaotic function is used to generate an index sequence, and the DNA sequence after the adjusted structure is rearranged to make the distribution of each base in the sequence more uniform.

[0062] Exemplary Methods

[0063] The present invention provides a DNA-based data encryption method that can be deployed on computers, smartphones, servers, portable wearable devices, and other computing devices. It is applied to data that requires compression and / or encryption. The present invention does not limit the type of target data; it can be various types of data, including voice, images, videos, and text data. Specifically, as shown in Figures 1 and 2, the method of this embodiment includes the following steps:

[0064] Step S1000: obtaining data to be encoded;

[0065] Specifically, data to be encoded is obtained, such as various types of data such as voice, picture, video or text data.

[0066] Step S2000: encrypting the data to be encoded to obtain encrypted data to be encoded;

[0067] Specifically, symmetric encryption or asymmetric encryption is used to encrypt different data to be encoded in an appropriate encryption form. Symmetric encryption, such as the AES encryption algorithm, is used for large files and data that require identity authentication; asymmetric encryption, such as the RSA encryption algorithm, is used for small files and data used for signatures.

[0068] It should be noted that this embodiment uses the AES encryption algorithm and the RSA encryption algorithm as the symmetric encryption algorithm and the asymmetric encryption algorithm respectively. As other preferred implementation methods, other types of symmetric encryption algorithms and asymmetric encryption algorithms can be flexibly selected according to actual application conditions. The present invention does not impose specific restrictions.

[0069] Step S3000: Based on a pre-constructed DNA encoding table, DNA encoding is performed on each byte of the encrypted data to obtain an initial base sequence, wherein the DNA encoding table is constructed from four-base units and five-base units;

[0070] Specifically, since there are four nitrogenous bases that make up DNA, namely adenine (A), guanine (G), thymine (T), and cytosine (C), this embodiment constructs several types of tetrabase units and pentabase units based on these four types of bases, and constructs a DNA encoding table based on these tetrabase units and pentabase units to perform DNA encoding on the encrypted data to be encoded. Among them, the tetrabase unit and pentabase unit respectively represent base units composed of four and five single bases.

[0071] Step S4000: using a preset DNA sequence structure optimization rule to adjust the initial base sequence to obtain an optimized base sequence;

[0072] Specifically, in order to avoid problems such as the occurrence of oligonucleotides, this embodiment uses preset DNA sequence structure optimization rules to check each base unit (also called codeword) in the initial base sequence, and adjusts the type or number of bases in the base units that meet the base sequence rules, so that oligonucleotides are not formed within or between base units, thereby obtaining an optimized base sequence.

[0073] Step S5000: performing chaotic mapping rearrangement on the optimized base sequence to obtain an encrypted DNA sequence.

[0074] Specifically, chaotic mapping is used to reorder and randomize the optimized base sequence, making the distribution of various single bases in the base sequence more random, thereby achieving a balanced distribution of GC content.

[0075] For example, using the piecewise linear chaotic map (PWLCM) function, the function is as follows:

[0076] Where k represents the length of the generated DNA sequence, z k Represents the randomly generated iteration value, z k ∈(0,1), z k+1 Represents a randomly generated pseudo-random sequence, p represents a constant, p∈(0,1).

[0077] By randomly generating values ​​of constants p and z kThe initial value of z0 is denoted as z0. A set of numbers equal in length to the DNA sequence is generated, and these numbers are sorted in ascending order. The index value is taken as sequence S, and sequence S is used as the DNA sequence index to rearrange the DNA sequence to generate an encrypted DNA sequence, which serves as the final target DNA sequence. Because the numbers generated by the PWLCM function are ergodic and random, the resulting pseudo-random sequence is uniform, ensuring that the various bases in the rearranged base sequence of the optimized base sequence are evenly distributed.

[0078] It should be noted that this embodiment uses the PWLCM function for chaotic mapping. As other preferred implementation methods, other types of chaotic mapping algorithms or random number generation algorithms can be flexibly selected according to actual application conditions to rearrange the optimized base sequence. The present invention does not impose specific restrictions.

[0079] In this embodiment, a unique DNA coding mapping rule is formed by using a designed DNA coding table constructed from four-base units and five-base units, and a designed DNA sequence structure optimization rule is adopted, which can avoid the formation of oligonucleotides within and between base units and improve information density. Combined with encryption and chaotic mapping, the GC content in the generated encrypted DNA sequence is kept in a balanced state, and the encryption effect of the data to be encoded is improved, which is beneficial to reducing the storage space occupied by the data to be encoded, and at the same time is beneficial to improving the security and anti-attack ability of the data to be encoded, thereby ensuring the stability of the encrypted DNA sequence.

[0080] As shown in FIG2 , in one embodiment, the DNA encoding of each byte of the encrypted data to be encoded based on the pre-constructed DNA encoding table in step S3000 to obtain an initial base sequence includes:

[0081] Step S3100: obtaining a byte sequence based on the encrypted data to be encoded, where the byte sequence includes a plurality of bytes;

[0082] Specifically, using quaternary encoding to perform encoding conversion on the encrypted data to obtain quaternary bytes can improve the fault tolerance and accuracy of the obtained byte sequence.

[0083] For example, if the encrypted data to be encoded is an image file, after encoding the image file, the data to be processed is represented by a binary sequence, where each eight-bit binary unit represents a byte of the image data. Then, a byte sequence consisting of several bytes can be used to represent the original image data.

[0084] Step S3200: Based on the mapping rules between bytes and base units in the pre-constructed DNA encoding table, the base units corresponding to the bytes in the byte sequence are respectively obtained to obtain an initial base sequence.

[0085] Specifically, the mapping rules between bytes and base units in the pre-constructed DNA coding table can correspond bytes of different sizes to different base units one by one. In practical applications, the base unit corresponding to each byte can be found by looking up the DNA coding table to ensure the accuracy of the initial base sequence.

[0086] For example, through the preset mapping rules, each byte in the byte sequence used to represent the original image data is converted into a corresponding base unit. For example, three consecutive bytes are converted into 66, 77 and 88 respectively by quaternary encoding. In the pre-constructed DNA coding table, the byte data 66, 77 and 88 correspond to AAGG, ATCG and CGAT respectively. Then, the base sequence generated after these three consecutive bytes are encoded is AAGGATCGCGAT.

[0087] In this embodiment, by encoding the encrypted data to be encoded to generate a byte sequence, and using a unique DNA encoding table to convert the byte sequence into a base sequence according to a one-to-one mapping relationship, the accuracy of converting various types of encrypted data to be encoded into a unique base sequence can be effectively improved.

[0088] In one embodiment, the process of constructing the DNA coding table in step S3000 includes:

[0089] Step S3300: constructing a mixed code table using four-base units and five-base units according to the preset code table design principle;

[0090] Specifically, the code table design principles need to comply with the GC balanced distribution principle and the principle of avoiding oligonucleotide generation. By constraining the types of single bases at different positions in the base unit, a mixed code table consisting of several four-base units and five-base units is designed.

[0091] Step S3400: Based on the mixed code table, a one-to-one correspondence between bytes and base units is established to obtain a mapping rule between bytes and base units.

[0092] In this embodiment, the designed mapping rules can ensure that each byte of data corresponds to a unique base unit, and comply with the GC balanced distribution principle and the principle of avoiding oligonucleotide generation, thereby ensuring that the constructed DNA coding table is rigorous, scientific and accurate.

[0093] In one embodiment, the step S3300 constructs a hybrid code table using four-base units and five-base units according to a preset code table design principle, including:

[0094] Based on the four single bases, several four-base units and five-base units were constructed;

[0095] If the first base of the four-base unit is A or T, the fourth base is C or G;

[0096] If the first base of the four-base unit is C or G, the fourth base is A or T;

[0097] If the first base of the five-base unit is A, the fourth and fifth bases are both TC or TG;

[0098] If the first base of the five-base unit is T, the fourth and fifth bases are both AC or AG;

[0099] If the first base of the five-base unit is C, the fourth and fifth bases are both GA or GT;

[0100] If the first base of the five-base unit is G, the fourth base and the fifth base are both CA or CT.

[0101] Specifically, a hybrid code table was constructed using 128 four-base units and 128 five-base units, with each unit corresponding to a byte of data ranging from 0 to 255. The type of the first base in the four-base unit was used to constrain the type of the fourth base, while the type of the first base in the five-base unit was used to constrain the types of the fourth and fifth bases, thereby designing a DNA encoding table that adheres to the principle of balanced GC distribution and avoids the formation of oligonucleotides. The designed DNA encoding table, as shown in Table 1, increases the information density to 1.8 bits / nt, with a GC content distribution of approximately 50%.

[0102] Table 1 DNA coding table

[0103] In one embodiment, when the previous base unit of two adjacent bases is a four-base unit, the step S4000 of using a preset DNA sequence structure optimization rule to adjust the initial base sequence to obtain an optimized base sequence includes:

[0104] Step S4100: If the last three bases of the previous base unit are the same as the first base of the adjacent next base unit, the last base of the previous base unit is replaced by the first base of the previous base unit;

[0105] Step S4200: If the last two bases of the previous base unit are the same as the first two bases of the adjacent next base unit, the last base of the previous base unit is replaced by the first base of the previous base unit;

[0106] Step S4300: If the last base of the previous base unit is the same as the first three bases of the adjacent next base unit, the last base of the previous base unit is replaced with a preset complementary base and one complementary base is added.

[0107] Specifically, by constraining the type or quantity of single bases at different positions in the codewords in the DNA coding table, it is possible to effectively avoid the formation of oligonucleotides within the codewords (i.e., base units) and between two consecutive codewords, that is, to avoid the same single base from being repeated three or more times. The complementary bases preset in this embodiment refer to cytosine C and guanine G as complementary bases to each other, and adenine A and thymine T as complementary bases to each other. As other preferred embodiments, other similar complementary bases can also be set, which are not limited here.

[0108] For example, as shown in FIG3 , when the first base unit of two adjacent bases is a four-base unit, for case a, if the first base of the first base unit is adenine A, and the last three bases and the first base of the adjacent last base unit are all cytosine C, then an oligonucleotide in which four consecutive bases are all cytosine C is formed. At this time, the fourth base of the first base unit is replaced with adenine A, which is the same as the first base of the first base unit, and then the oligonucleotide is eliminated; similarly, for case b, if the first base of the first base unit is adenine A, and the last two bases and the first two bases of the adjacent last base unit are replaced with adenine A, then the oligonucleotide is eliminated. If the bases are all cytosine C, the fourth base of the previous base unit is replaced with adenine A, which is the same as the first base of the previous base unit; for case c, if the first base of the previous base unit is adenine A, and the last single base and the first three single bases of the adjacent subsequent base unit are all cytosine C, an oligonucleotide in which four consecutive single bases are all cytosine C is formed. At this time, in order to avoid the previous base unit becoming an oligonucleotide after the last base of the previous base unit is replaced with adenine A, the fourth base of the previous base unit is replaced with thymine T, the complementary base of adenine A, and a thymine T is added.

[0109] In this embodiment, for the case where the first base unit between two adjacent base units is a four-base unit, the DNA encoding table designed according to the above embodiment only allows for the aforementioned three possible oligonucleotide occurrences. This embodiment considers all possible oligonucleotide formation scenarios and designs corresponding substitution and complementation rules for each scenario, thereby ensuring that, in the case where the first base unit between two adjacent base units is a four-base unit, no oligonucleotide will appear in the generated corresponding DNA sequence fragment.

[0110] In one embodiment, when the previous base unit of two adjacent bases is a five-base unit, the step S4000 of using a preset DNA sequence structure optimization rule to adjust the initial base sequence to obtain an optimized base sequence includes:

[0111] Step S4400: If the last base of the previous base unit is the same as the first three bases of the adjacent next base unit, the last base of the previous base unit is replaced by the first base of the previous base unit.

[0112] For example, as shown in Figure 3, when the previous base unit of two adjacent bases is a five-base unit, for case d, if the first base of the previous base unit is adenine A, and the last single base and the first three single bases of the adjacent next base unit are all cytosine C, then an oligonucleotide in which four consecutive single bases are all cytosine C is formed, then the last single base cytosine C of the previous base unit is replaced with adenine A, which is the same as the first single base of the previous base unit, thereby eliminating the oligonucleotide.

[0113] In this embodiment, when the first base unit between two adjacent base units is a five-base unit, the DNA encoding table designed according to the above embodiment can only cause the above-mentioned oligonucleotide to appear. This embodiment takes into account the only possible oligonucleotide and designs corresponding substitution rules for this situation, thereby ensuring that when the first base unit between two adjacent base units is a five-base unit, the corresponding DNA sequence fragment generated will not contain an oligonucleotide.

[0114] It should be noted that the two adjacent base units referred to in the present invention only limit the first base unit to a four-base unit, and do not limit the type of the second base unit, that is, the type of the second base unit can be a four-base unit or a five-base unit.

[0115] Since the entire DNA sequence is composed of two types of base units, four-base units and five-base units, there is only a situation where the first base unit in adjacent base units is a four-base unit or a five-base unit. Therefore, by using the above-mentioned overall DNA sequence structure optimization rules to adjust the initial base sequence, an optimized base sequence without oligonucleotides can be obtained.

[0116] To verify the feasibility and effectiveness of the encryption algorithm described above, we used two classic images for experimental simulations: Figure 4 shows the Lena image before encryption, Figure 5 shows the GC content distribution of the Lena image after AES encryption, and Figure 6 shows the GC content distribution of the Lena image after RSA encryption. Figure 7 shows the Cameraman image before encryption, Figure 8 shows the GC content distribution of the Cameraman image after AES encryption, and Figure 9 shows the GC content distribution of the Cameraman image after RSA encryption. The horizontal axis in each simulation represents the base position in the encrypted DNA sequence, and the vertical axis represents the GC content percentage of each base position in the encrypted DNA sequence.

[0117] Comparative analysis of simulated GC content distributions in DNA sequences obtained after image data encryption using the AES and RSA algorithms reveals that the GC content distributions in both plots conform to the biological characteristics of DNA sequence synthesis and sequencing, thus validating the effectiveness of the proposed method. In practical applications, symmetric or asymmetric encryption can be flexibly selected to optimize file encryption based on factors such as the desired encryption level, encryption speed, and encrypted file size.

[0118] In one embodiment, the encrypted DNA sequence obtained by the DNA encoding data encryption method also provides a DNA encoding data decryption method, comprising:

[0119] The encrypted DNA sequence is rearranged based on the sequence information obtained by the initial setting of the chaotic map to restore the optimized base sequence;

[0120] Based on preset DNA sequence structure optimization rules and a pre-constructed DNA encoding table, DNA decoding is performed on the optimized base sequence to restore the decoded base sequence;

[0121] The decoded base sequence is decrypted to obtain restored data.

[0122] In this embodiment, the decryption process of the decryption method is the inverse process of the encryption process in the encryption method. By adopting the corresponding criteria adopted in the encryption stage in different decryption stages, after decrypting the encrypted DNA sequence, the restored data generated is exactly the same as the data information contained in the original data to be encoded.

[0123] Furthermore, the decryption method further includes performing DNA decoding on the optimized base sequence based on a preset DNA sequence structure optimization rule and a pre-constructed DNA encoding table to restore the decoded base sequence, including:

[0124] Based on a preset DNA sequence structure optimization rule, adjusting the base units in the optimized base sequence to obtain a quasi-base sequence;

[0125] Based on a pre-constructed DNA coding table, the quasi-base sequence is DNA decoded to restore the decoded base sequence.

[0126] In this embodiment, during the encryption stage, after the initial base sequence generated based on the DNA coding table is adjusted according to the preset DNA sequence structure optimization rules, abnormal base units that do not exist in the DNA coding table will be generated. Therefore, in order to avoid decryption errors, the abnormal base units are first restored to generate a quasi-base sequence, and then DNA decoding is performed on the quasi-base sequence through the DNA coding table to restore the decoded base sequence, thereby ensuring that the restored base sequence is completely correct, and further ensuring that the data information contained in the original data to be encoded can be accurately restored.

[0127] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown in Figure 10. The above intelligent terminal includes a processor, a memory, a network interface and a display screen connected via a system bus. The processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a data encryption program based on DNA coding. The internal memory provides an environment for the operation of the operating system and the data encryption program based on DNA coding in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal via a network connection. When the data encryption program based on DNA coding is executed by the processor, the steps of any one of the above-mentioned data encryption methods based on DNA coding are implemented. The display screen of the intelligent terminal can be a liquid crystal display or an electronic ink display.

[0128] Those skilled in the art will understand that the principle block diagram shown in Figure 10 is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the smart terminal to which the solution of the present invention is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0129] In one embodiment, a smart terminal is provided, which includes a memory, a processor, and a DNA-encoding-based data encryption program stored in the memory and executable on the processor. When the DNA-encoding-based data encryption program is executed by the processor, the steps of any one of the DNA-encoding-based data encryption methods provided in the embodiments of the present invention are implemented.

[0130] An embodiment of the present invention also provides a computer-readable storage medium, on which a data encryption program based on DNA coding is stored. When the data encryption program based on DNA coding is executed by a processor, the steps of any one of the data encryption methods based on DNA coding provided in an embodiment of the present invention are implemented.

[0131] It should be understood that the sequence numbers of the steps in the above embodiments do not imply a specific order of execution; the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0133] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0134] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0135] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units described above is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not implemented.

[0136] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. DNA-encoded data encryption method, characterized in that, it includes: Obtain the data to be encoded; Encrypt the data to be encoded to obtain encrypted data to be encoded; Based on a pre-constructed DNA coding table, perform DNA coding on each byte of the encrypted data to be encoded to obtain an initial base sequence, where the DNA coding table is constructed from four-base units and five-base units; Use a preset DNA sequence structure optimization rule to adjust the initial base sequence to obtain an optimized base sequence; Perform chaotic mapping rearrangement on the optimized base sequence to obtain an encrypted DNA sequence.

2. The DNA-encoded data encryption method according to claim 1, characterized in that, The step of performing DNA coding on each byte of the encrypted data to be encoded based on a pre-constructed DNA coding table to obtain an initial base sequence includes: Based on the encrypted data to be encoded, obtain a byte sequence, and the byte sequence includes a plurality of bytes; Based on the mapping rule between bytes and base units in the pre-constructed DNA coding table, respectively obtain the base units corresponding to each byte in the byte sequence to obtain an initial base sequence.

3. The DNA-encoded data encryption method according to claim 1, characterized in that, The construction process of the DNA coding table includes: Construct a mixed coding table using four-base units and five-base units according to a preset coding table design principle; Based on the mixed coding table, establish a one-to-one correspondence between bytes and base units to obtain a mapping rule between bytes and base units.

4. The DNA-encoded data encryption method according to claim 3, characterized in that, The step of constructing a mixed coding table using four-base units and five-base units according to a preset coding table design principle includes: Based on four single bases, construct several four-base units and five-base units; If the first base of the four-base unit is A or T, the fourth base is C or G; If the first base of the four-base unit is C or G, the fourth base is A or T; If the first base of the five-base unit is A, the fourth and fifth bases are both TC or TG; If the first base of the five-base unit is T, the fourth and fifth bases are both AC or AG; If the first base of the five-base unit is C, the fourth and fifth bases are both GA or GT; If the first base of the five-base unit is G, the fourth and fifth bases are both CA or CT.

5. The DNA-encoded data encryption method according to claim 3, characterized in that, When the previous base unit among two adjacent bases is a four-base unit, the step of using a preset DNA sequence structure optimization rule to adjust the initial base sequence to obtain an optimized base sequence includes: If the last three single bases of the previous base unit are the same as the first single base of the adjacent next base unit, replace the last single base of the previous base unit with the first single base of the previous base unit; If the last two single bases of the previous base unit are the same as the first two single bases of the adjacent next base unit, replace the last single base of the previous base unit with the first single base of the previous base unit; If the last single base of the previous base unit is the same as the first three single bases of the adjacent next base unit, replace the last single base of the previous base unit with a preset complementary base and add one such complementary base.

6. The DNA-encoding-based data encryption method according to claim 3, wherein, when the previous base unit among two adjacent bases is a five-base unit, the initial base sequence is adjusted by using the preset DNA sequence structure optimization rules to obtain an optimized base sequence, including: If the last single base of the previous base unit is the same as the first three single bases of the adjacent next base unit, replace the last single base of the previous base unit with the first single base of the previous base unit.

7. A DNA-encoding-based data decryption method, wherein, it is applied to decrypt the encrypted DNA sequence obtained by the DNA-encoding-based data encryption method according to any one of claims 1-6, including: Rearranging the order of the encrypted DNA sequence in reverse based on the initial setting of the chaotic mapping to obtain sequence information, and restoring the optimized base sequence; Based on the preset DNA sequence structure optimization rules and the pre-constructed DNA coding table, perform DNA decoding on the optimized base sequence to restore the decoded base sequence; Decrypt the decoded base sequence to obtain the restored data.

8. The DNA-encoding-based data decryption method according to claim 7, wherein, the performing DNA decoding on the optimized base sequence based on the preset DNA sequence structure optimization rules and the pre-constructed DNA coding table to restore the decoded base sequence includes: Based on the preset DNA sequence structure optimization rules, adjust the base units in the optimized base sequence to obtain a quasi-base sequence; Based on the pre-constructed DNA coding table, perform DNA decoding on the quasi-base sequence to restore the decoded base sequence.

9. An intelligent terminal, wherein, the intelligent terminal includes a memory, a processor, and a DNA-encoding-based data encryption program stored on the memory and executable on the processor. When the DNA-encoding-based data encryption program is executed by the processor, it implements the steps of the DNA-encoding-based data encryption method according to any one of claims 1-6.

10. A computer-readable storage medium, wherein, a DNA-encoding-based data encryption program is stored on the computer-readable storage medium. When the DNA-encoding-based data encryption program is executed by a processor, it implements the steps of the DNA-encoding-based data encryption method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image encryption algorithm based on quantum chaotic mapping and DNA encoding

    CN108665404A

  • Information coding method and device based on DNA storage, computer equipment and medium

    CN114974434A

  • GC content and homopolymer controllable DNA data storage method

    CN117095758A

  • Method for Information Encoding and Decoding, and Method for Information Storage and Interpretation

    US20230032409A1

  • Method and system for encrypting genetic data of a subject

    US20230317211A1

Cited By

  • Facial image encryption method, decryption method and image processing device

    CN120658372A

  • Face image encryption method, decryption method and image processing device

    CN120658372B