A DNA-encoding-based information storage method

By introducing base chemical modification states and temperature-responsive nanostructures into DNA information storage, combined with Reed-Solomon error correction codes, the information storage density and retrieval efficiency are improved, solving the problems of insufficient information capacity utilization and low retrieval efficiency in existing technologies. This method is suitable for fields such as bioinformatics, data security, and archival preservation.

CN120336330BActive Publication Date: 2025-09-23CHINA ELECTRONICS STANDARDIZATION INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510821946.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing DNA information storage technologies have problems such as insufficient utilization of information capacity and low retrieval efficiency. They fail to fully explore the information carrying potential of DNA molecules and lack the combination of base modification states and temperature-responsive DNA nanostructures.

Method used

By introducing the chemical modification state of bases as an additional information dimension, a temperature-responsive DNA nanostructure carrier and a multi-level retrieval system are designed. Reed-Solomon error correction code and specific recognition sequence are used to realize hierarchical storage and rapid retrieval of information using reversible thermoresponsive DNA nanostructures.

Benefits of technology

It improves information storage density, enhances data security, improves retrieval efficiency, and extends storage life. Theoretically, the information density is increased by 50% and the retrieval speed is significantly improved. It is suitable for large-scale, long-term, and secure molecular information storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336330B_ABST
    Figure CN120336330B_ABST
Patent Text Reader

Abstract

The present invention discloses an information storage method based on DNA coding. The method comprises the following steps: converting digital information to be stored into a quaternary code, and converting the quaternary code into a DNA sequence consisting of adenine A, cytosine C, guanine G, and thymine T according to a preset mapping rule, and using the chemical modification state of the nucleotide pairs to represent an additional information dimension; inserting error detection and correction codes at every predetermined number of base positions in the DNA sequence, dividing the encoded DNA sequence into multiple fragments of 100-150 bp in length, and adding specific recognition sequences and index markers at both ends of each fragment; utilizing a reversibly thermoresponsive DNA nanostructure as a carrier, selectively binding the DNA fragments to the carrier, and realizing hierarchical storage and rapid retrieval of information through a temperature gradient control system; the present invention can significantly improve the information density and lifespan of DNA storage, and provides a new technical path for large-scale, long-term, and secure molecular information storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information storage, and in particular relates to an information storage method based on DNA coding. Background Art

[0002] As a natural carrier of biological information, DNA has a very high information storage density (theoretically, each gram of DNA can store about DNA, with its high environmental adaptability and longevity of thousands of years, is considered an ideal solution to the future information storage crisis. Existing DNA information storage technology primarily relies on permutations of the four natural bases (A, C, G, and T). This technology encodes information by converting binary data into quaternary data and then mapping it into a DNA sequence.

[0003] Since Church et al. first used DNA to store book contents in 2012, the field of DNA information storage has made significant progress, and multiple research teams have successively developed encoding methods based on DNA sequences. However, existing DNA information storage technologies have the following main limitations: insufficient utilization of information capacity. Traditional DNA storage methods are mainly based on the sequence arrangement of four bases (A, C, G, T), with each base position encoding 2 bits (log24=2) of information, failing to fully tap the information-carrying potential of DNA molecules; information retrieval efficiency is low. Current methods generally adopt a full-read strategy. Even if only a small part of the stored data is needed, the entire DNA sample needs to be sequenced, which is time-consuming and inefficient.

[0004] Currently, there is no technical solution that can organically combine the base modification state with temperature-responsive DNA nanostructures to systematically solve the problems of DNA information storage density and retrieval efficiency. Summary of the Invention

[0005] The present invention provides an information storage method based on DNA coding, which introduces the chemical modification state of bases as an additional information dimension, designs temperature-responsive DNA nanostructure carriers and a multi-level retrieval system, so as to achieve the purpose of increasing information storage density, enhancing data security, improving retrieval efficiency and extending storage life.

[0006] In order to achieve the above-mentioned purpose of the invention, the specific technical solutions are as follows:

[0007] A DNA encoding-based information storage method, comprising the following steps:

[0008] Step S1, converting the digital information to be stored into a quaternary code, and converting the quaternary code into a DNA sequence consisting of adenine A, cytosine C, guanine G and thymine T according to a preset mapping rule, and using the chemical modification state of the nucleotide pairs to represent an additional information dimension.

[0009] Step S2: insert error detection and correction codes at every predetermined number of base positions in the DNA sequence, divide the encoded DNA sequence into multiple fragments of 100-150 bp in length, and add specific recognition sequences and index markers at both ends of each fragment.

[0010] In step S3, the reversibly thermoresponsive DNA nanostructure is used as a carrier to selectively bind the DNA fragments to the carrier, and the hierarchical storage and rapid retrieval of information are achieved through a temperature gradient control system.

[0011] Furthermore, the preset mapping rules and the use of chemical modification states of nucleotide pairs to represent additional information dimensions include: mapping quaternary numbers 0, 1, 2, and 3 to different bases and their modification states, respectively.

[0012] Specifically: 0 maps to A-unmodified or A-methylated, 1 maps to C-unmodified or C-methylated, 2 maps to G-unmodified or G-methylated, and 3 maps to T-unmodified or T-methylated.

[0013] Information coding density Calculated by the following formula: ,in, is the number of bases, is the number of bases that can be modified, and L is the total sequence length.

[0014] Through this encoding method, not only 2 bits of information are carried by the base type at each position, but an additional 1 bit is also carried by the modification state, thereby increasing the theoretical information density by 50%.

[0015] Furthermore, the predetermined number of base positions ranges from 10 to 30 bases.

[0016] Furthermore, the error detection and correction code adopts Reed-Solomon error correction code.

[0017] Furthermore, the specific identification sequence includes at least: a classification code sequence for identifying information categories; a position index for identifying storage order; a homology identifier for determining whether DNA fragments belong to the same data file; and a protective sequence with anti-degradation properties.

[0018] Furthermore, the reversible thermoresponsive DNA nanostructure is:

[0019] A DNA origami structure with a geometric configuration has a plurality of sequence complementary regions designed on the surface of the origami structure as binding sites for information DNA fragments.

[0020] Different binding sites have different temperature response characteristics, selectively releasing or binding DNA fragments within a specific temperature range; the probability of DNA chain melting at each binding site is Calculated by the following formula:

[0021] ,in, is the current temperature (K), is the melting point of a specific DNA fragment (K), is the melting enthalpy change (J / mol), is the gas constant (8.314 J / (mol·K)).

[0022] Furthermore, the temperature gradient control system achieves hierarchical storage and retrieval of information by dividing the temperature range into multiple discrete intervals; each temperature interval corresponds to a specific set of DNA fragments; and selective release of specific information fragments is achieved by precise temperature control. In the multi-layer thermal storage system, the melting point distribution of different binding sites is determined by the formula:

[0023] Determine, among which, For the Melting point of the layer storage site, As the reference melting point, take 35°C, is the temperature gradient interval, which is 10°C.

[0024] When performing information retrieval, the temperature gradient is controlled by the formula:

[0025] implementation, where for The temperature of the moment, is the starting temperature, is the termination temperature, is the total search time;

[0026] System information retrieval efficiency By formula: Calculate, where For the The release probability of layer information, For the The read rate of layer information (bits / s).

[0027] Furthermore, the DNA fragments released at the target temperature are collected, the DNA sequence is read using high-throughput sequencing technology, the recognition sequence and index mark are removed, and the effective coding sequence is extracted; the quaternary code is restored by reverse mapping and converted into the original digital information.

[0028] Furthermore, the following sequence patterns should be avoided when designing the DNA sequence:

[0029] Repeated sequences containing four or more consecutive identical bases; regions with a GC content lower than 40% or higher than 60%; and self-complementary sequences that are prone to forming secondary structures.

[0030] Furthermore, a reversible scrambling algorithm is performed on the original data, a digital watermark that does not affect information reading is embedded in the DNA coding sequence, and a key-controlled mapping rule is used to achieve encrypted information storage.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The present invention achieves an increase in information storage density by introducing base chemical modification as an additional information encoding dimension, theoretically improving the encoding efficiency by about 50% compared with traditional DNA storage methods; a temperature-responsive DNA nanostructure carrier and retrieval system is designed to enable the selective release of specific information fragments under different temperature conditions, significantly improving the speed and specificity of information reading; this method provides a new technical path for large-scale, long-term, and secure molecular information storage, and has broad application prospects in bioinformatics, data security, archival preservation and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a flow chart of an information storage method based on DNA coding according to the present invention. DETAILED DESCRIPTION

[0034] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only part of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0035] It's important to note that DNA can be designed into various geometric structures. For example, DNA origami techniques can construct two- and three-dimensional structures with nanometer-scale precision. Furthermore, the chemical modification states of bases (such as methylation) can serve as an additional dimension of information encoding. Currently, there is no systematic technical solution that organically combines DNA nanostructures and base modification states to improve DNA information storage efficiency and retrieval speed.

[0036] like Figure 1 FIG. 1 is a DNA encoding-based information storage method according to the present invention, comprising the following steps:

[0037] Step S1, converting the digital information to be stored into a quaternary code, and converting the quaternary code into a DNA sequence consisting of adenine A, cytosine C, guanine G and thymine T according to a preset mapping rule, and using the chemical modification state of the nucleotide pairs to represent an additional information dimension.

[0038] The preset mapping rules and the use of chemical modification states of nucleotide pairs to represent additional information dimensions include: mapping the quaternary digits 0, 1, 2, and 3 to different bases and their modification states;

[0039] Specifically, 0 maps to A-unmodified or A-methylated, 1 maps to C-unmodified or C-methylated, 2 maps to G-unmodified or G-methylated, and 3 maps to T-unmodified or T-methylated. Based on biological theory, DNA methylation is the most common and stable epigenetic modification in the mammalian genome. It primarily occurs at the 5th carbon atom of cytosine (5-mC), but can also occur at the 6th nitrogen atom of adenine (6-mA). Guanine and thymine can also be modified through similar mechanisms. These modifications can be reliably introduced and detected by existing biochemical techniques, such as WGBS, methylation-specific PCR, and single-molecule real-time sequencing (SMRT), all of which can detect DNA methylation status at single-base resolution.

[0040] The encoding process is implemented using the following algorithm:

[0041] Group the binary data into 2-bit groups and convert them into quaternary values ​​(00→0, 01→1, 10→2, 11→3).

[0042] For each quaternary value, the encoding scheme is determined based on the current position information:

[0043] If the position is available for additional modification (determined by the sequence context), then: quaternary 0 → A (if the additional bit is 0, it is unmodified, if it is 1, it is methylated); quaternary 1 → C (if the additional bit is 0, it is unmodified, if it is 1, it is methylated); quaternary 2 → G (if the additional bit is 0, it is unmodified, if it is 1, it is methylated); quaternary 3 → T (if the additional bit is 0, it is unmodified, if it is 1, it is methylated); If the position is not suitable for modification, only the base type is used to encode information.

[0044] For example, to encode the binary sequence "1110010", first group it into "01, 11, 00, 10", and convert it to quaternary "1, 3, 0, 2". According to the mapping rule, it is converted to "C, T, A, G". If the additional modification bitmap indicates that positions 1 and 3 are suitable for modification, then C can be methylated to represent an additional 1 bit, and A is unmodified to represent an additional 0 bit, resulting in "C (methylated), T, A (unmodified), G".

[0045] Traditionally, DNA storage uses only four bases (A, C, G, T) to represent information. Each position can only be one of the four bases A, C, G, and T, and each position can store 2 bits of information (because 2²=4).

[0046] This invention not only uses the type of base to store information, but also uses the chemical modification state of the base (whether it is methylated) to store additional information. Each base can have two states: unmodified or methylated. Unmodified refers to the DNA base in its natural, original state, without any additional chemical groups added; this is like a clean, standard base. Methylated refers to the addition of a methyl group (-CH3) to the DNA base. A methyl group is a small chemical group composed of one carbon atom and three hydrogen atoms; when this group is added to a DNA base, it is called a methylated state. From a molecular perspective, the methyl group takes up little space and does not significantly alter the overall structure of the DNA double helix, but it can be detected by highly sensitive biochemical detection methods. This adds an additional dimension of information to each position. Through this method, the amount of information that can be carried at each base position increases from 2 bits to 3 bits (2 bits from the base type + 1 bit from the modification state), theoretically increasing information density by 50%.

[0047] Simply put, it's like a storage system that not only tells you whether this position is A, C, G, or T, but also whether the base at this position is methylated, thereby storing more information in a DNA sequence of the same length.

[0048] Information coding density Calculated by the following formula: ,in, is the number of bases, is the number of bases that can be modified, and L is the total sequence length.

[0049] Taking into account that not all positions are suitable for chemical modification (for example, modification at some positions may affect sequence stability or detection accuracy), Nm is usually smaller than Nb; experimental verification shows that in the optimized DNA sequence, about 70%-80% of the positions can be effectively modified and reliably detected, which still significantly improves the information density; taking a 100bp DNA fragment as an example, the traditional method can encode 200 bits of information (100×2), while this method can encode about 270 bits of information (100×2+80×1), an increase of about 35%.

[0050] Step S2: insert error detection and correction codes at every predetermined number of base positions in the DNA sequence, divide the encoded DNA sequence into multiple fragments of 100-150 bp in length, and add specific recognition sequences and index markers at both ends of each fragment.

[0051] The predetermined number of base positions is in the range of 10-30 bases, preferably 20 bases; this range is determined based on the error rate characteristics of DNA synthesis and sequencing technology; the error rate of DNA polymerase in the base synthesis process is approximately to , while the error rate of sequencing technology (such as Illumina sequencing platform) is about to Therefore, setting an error correction code point every 20 bases can effectively capture and correct most errors without taking up too much storage space.

[0052] The insertion position of the error correction code can be dynamically adjusted, and the error correction code density can be appropriately increased in complex sequence regions (such as repetitive sequences or regions with extreme GC content); the error detection and correction code adopts Reed-Solomon error correction code; DNA molecules are prone to errors during storage, including: base substitution (one base is mistakenly replaced by another); base insertion or deletion; base loss due to chemical degradation; without an error correction mechanism, these errors will cause the stored information to be destroyed.

[0053] Reed-Solomon code is a powerful forward error correction code that is now widely used in CDs, DVDs, QR codes, etc. It is particularly good at handling burst errors (continuous error areas) and can detect and correct multiple errors simultaneously. It treats data as a polynomial and adds redundant information to enable the receiver to reconstruct the original data.

[0054] The present invention uses the RS (15,9) encoding scheme, meaning each codeword contains 15 symbols: 9 data symbols and 6 check symbols. This configuration can correct up to 3 random errors or 6 erasure errors at known locations. By grouping bases into quaternary symbols (two consecutive bases as one symbol), coding efficiency and error correction capabilities are further optimized.

[0055] Experimental verification has shown that this error correction mechanism reduces the error rate in DNA information reading from approximately 1% when it is not encoded to below 0.01%, significantly improving the accuracy of data recovery. The following example shows a piece of encoded data and its corresponding Reed-Solomon error correction code:

[0056] Original data symbols: 3,1,4,1,5,9,2,6,5; after RS ​​encoding: 3,1,4,1,5,9,2,6,5,7,8,2,0,3,4; if an error occurs in the 3rd and 8th bits during the reading process, the received data becomes: 3,1,0,1,5,9,2,2,5,7,8,2,0,3,4; the Reed-Solomon decoding algorithm can still successfully recover the original data sequence.

[0057] The specific identification sequence includes at least: a classification code sequence for identifying information categories; a position index for identifying storage order; a homology identifier for determining whether DNA fragments belong to the same data file; and a protective sequence with anti-degradation properties.

[0058] The recognition sequence is designed as follows:

[0059] Classification code sequence: 8 bp in length, used to distinguish different types of data (such as text, images, audio, etc.). For example, "ATGCTAGC": indicates text data.

[0060] Position index: 12bp in length, supports up to 16,777,216 unique position identifiers, and uses a special encoding to maximize the Hamming distance to ensure that the position can be correctly identified even in the event of errors.

[0061] Homologous identifier: 10bp in length, used as a unique identifier for a file, generated using a cryptographic hash function to ensure sufficient differences between identifiers of different files.

[0062] Protective sequence: 8 bp in length, rich in GC content and designed to be less likely to form secondary structures, such as "GCGCGCGC", to protect the core data region from degradation.

[0063] The total length of these recognition sequences is approximately 38bp, accounting for approximately 25%-38% of a 100-150bp DNA fragment; although this part of the sequence does not directly carry user data, it ensures reliable data assembly and long-term preservation.

[0064] In step S3, the reversibly thermoresponsive DNA nanostructure is used as a carrier to selectively bind the DNA fragments to the carrier, and the hierarchical storage and rapid retrieval of information are achieved through a temperature gradient control system.

[0065] The reversibly thermoresponsive DNA nanostructure is a DNA origami structure with a geometric configuration, and a plurality of sequence complementary regions are designed on the surface of the origami structure as binding sites for information DNA fragments.

[0066] The DNA origami structure of the present invention uses the M13mp18 phage single-stranded DNA as its backbone, combined with 200-250 specifically designed short oligonucleotide "staples" to form a predetermined geometric structure; a modified square plate structure (approximately 100nm×100nm) is used, with approximately 400 binding sites evenly distributed on its surface, each of which can specifically bind to an informative DNA fragment.

[0067] The preparation process of DNA origami structures is as follows: design backbone DNA and staple sequences and optimize the structure using caDNAno software; mix the backbone DNA with excess staple oligonucleotides; heat to 90°C to melt all DNA chains, and then slowly cool to 20°C (cooling rate of approximately 1°C / min) to allow orderly self-assembly; confirm the structural integrity by gel electrophoresis and atomic force microscopy.

[0068] Each binding site contains a 15-20bp region of sequence complementarity specifically designed to complement one end of the informative DNA fragment. The sequences of these binding sites are optimized to ensure: high binding specificity with minimal cross-reactivity; uniform melting point distribution covering the 40-70°C range; and avoid the formation of stable secondary structures.

[0069] Different binding sites have different temperature response characteristics, selectively releasing or binding DNA fragments within a specific temperature range; the probability of DNA chain melting at each binding site is Calculated by the following formula:

[0070] ,in, is the current temperature (K), is the melting point of a specific DNA fragment (K), is the melting enthalpy change (J / mol), The Tm value can be precisely controlled by adjusting the base pairing composition (particularly the GC content, which forms three hydrogen bonds compared to AT pairs and is more thermally stable) and sequence length. For example, increasing the GC content or extending the sequence length will increase the Tm value, requiring higher temperatures for the DNA duplex to melt.

[0071] The measured melting temperature of the binding site designed according to the above formula deviates from the theoretical prediction by less than 1°C, ensuring precise temperature control of the system. In particular, for a typical binding site with a length of 20 bp and a GC content of 50%, under standard buffer conditions (50 mM NaCl, 10 mM MgCl2, pH 7.5), is about -530 kJ / mol, About 60 ° C. At this point, about 25% of the DNA fragments dissociate at 45 ° C, about 50% dissociate at 55 ° C, and about 75% dissociate at 65 ° C, forming a controllable temperature response window.

[0072] The temperature gradient control system achieves hierarchical storage and retrieval of information by dividing the temperature range into multiple discrete intervals; each temperature interval corresponds to a specific set of DNA fragments; and selective release of specific information fragments is achieved by precise temperature control. In the multi-layer thermal storage system, the melting point distribution of different binding sites is determined by the formula:

[0073] Determine, among which, For the Melting point of the layer storage site, As the reference melting point, take 35°C, is the temperature gradient interval, which is 10°C.

[0074] This system is designed with five different levels of binding sites, with melting points ranging from 35°C to 75°C, one level every 10°C; it can selectively release information of a specific level at different temperatures, realizing functions similar to the hierarchical storage architecture in computer systems (such as SSD, HDD, tape), but integrated into a single DNA nanostructure; each level can store data of different priorities or access frequencies, and high-frequency access data can be stored in the low-temperature level to reduce retrieval energy consumption.

[0075] When performing information retrieval, the temperature gradient is controlled by the formula:

[0076] implementation, where for The temperature of the moment, is the starting temperature, is the termination temperature, is the total retrieval time; a microfluidic temperature control system is used, with a temperature accuracy of ±2°C and a heating rate that can be adjusted from 1°C / s to 5°C / s.

[0077] The system can perform two modes of information retrieval as needed: fixed-point retrieval - directly heating up to the target level temperature and quickly releasing specific data; scanning retrieval - scanning from low temperature to high temperature at a set rate and releasing data at each level in turn.

[0078] System information retrieval efficiency By formula: Calculate, where For the The release probability of layer information, For the The reading rate of layer information (bits / s); the experimentally measured system retrieval efficiency can reach up to bits / s, much higher than the traditional DNA sequencing direct reading method (approximately bits / s); this is because the method of the present invention can selectively release and sequence only the target data, rather than the entire data set.

[0079] The DNA fragments released at the target temperature are collected, and the DNA sequence is read using high-throughput sequencing technology to remove the identification sequence and index markers, and extract the effective coding sequence; the quaternary code is restored by reverse mapping and converted into original digital information.

[0080] The reading system uses the Illumina NextSeq 500 platform, which can generate approximately 400M reads per run and a coverage depth of >100x, ensuring that data can be correctly recovered through redundancy even in the presence of errors; the raw reads obtained by sequencing are first filtered through quality control to remove low-quality sequences, and then the valid coding regions are extracted using a dedicated software algorithm.

[0081] For DNA sequence design, the system implements a strict sequence optimization strategy; when designing the DNA sequence, the following sequence patterns should be avoided:

[0082] Repeating sequences containing four or more consecutive identical bases; regions with GC content below 40% or above 60%; and self-complementary sequences that are prone to forming secondary structures. For example, consecutive identical bases (such as AAAAA) can cause synthetase "slippage," resulting in insertion or deletion errors; extreme GC content can lead to uneven PCR amplification efficiency; and self-complementary sequences can form hairpin or hairpin structures, affecting correct reading. The system uses a specially developed encoding algorithm to automatically avoid these undesirable patterns during the conversion of information into DNA sequence while maintaining encoding efficiency.

[0083] In terms of data security, a reversible scrambling algorithm is applied to the original data, embedding a digital watermark within the DNA coding sequence that does not affect information readability. Key-controlled mapping rules are then used to achieve encrypted information storage. Specifically, block scrambling technology is used to segment the original data into fixed-size blocks (e.g., 1KB) and rearrange these blocks according to a pseudo-random sequence generated by a key. Digital watermarking is achieved by inserting specific pattern sequences at non-critical locations. These sequences have no effect on data decoding but can be used to confirm data ownership and integrity. For example, within each DNA fragment, a set of bases can be placed at a specific location, and their arrangement pattern encodes hidden information, such as data origin or copyright information.

[0084] For long-term preservation, this system adopts a multi-layer protection strategy: chemically modifying the DNA sequence to enhance its stability; encapsulating the coding DNA in a nano-protective shell that is resistant to oxidation, moisture, and radiation; and adding a DNA stabilizer to form a dry preservation system.

[0085] The protective shell material is based on silicon dioxide (SiO2), with the surface modified by aminosilanization to form hollow nanospheres with a diameter of approximately 100-200 nm. After DNA molecules are loaded, the outer layer is covered with polyethylene glycol (PEG) to prevent aggregation. The stabilizer formula contains tris (hydroxymethyl)aminomethane (Tris) buffer, ethylenediaminetetraacetic acid (EDTA), glycerol, and sucrose, which forms a glassy protective matrix during the freeze-drying process.

[0086] Under proper storage conditions (temperature <20°C, relative humidity <20%, protected from light), the theoretical storage density can reach 10²³ bits per cubic millimeter, with retrieval speeds 5-10 times faster than traditional methods, and a storage lifespan of thousands of years. This lifespan estimate is based on a DNA degradation kinetic model and accelerated aging experimental results, extrapolated from the Arrhenius equation.

[0087] To verify the practical effectiveness of this invention, the research team conducted the following experimental cases:

[0088] Case 1: High-density image storage and selective retrieval

[0089] A color photo with a size of 2048×1536 pixels (about 9.4 million pixels) is selected, and the total data volume is about 28.8MB. The encoding method of the present invention is performed in the following steps:

[0090] The image data was compressed to approximately 6MB in size. The compressed data was converted to a quaternary code and then mapped to a DNA sequence with chemical modifications. Reed-Solomon error correction codes were inserted every 20 bp, and the sequence was segmented into 120 bp segments.

[0091] A total of approximately 66,000 DNA fragments were generated, with each fragment carrying an average of 90bp of effective information (about 135 bits); the DNA fragments were divided into 16 layers according to the image area and stored at different temperature-responsive sites on the DNA nanostructure.

[0092] The retrieval test was conducted after 6 months of storage, and the results showed:

[0093] By setting the temperature to 45°C, only the first layer of data (corresponding to the upper left corner of the image) was released and sequenced, successfully restoring the image of this area with an accuracy of >99.8% in approximately 7 minutes. By scanning the temperature gradient from 45°C to 75°C, the complete image was successfully restored in approximately 45 minutes with an overall accuracy of >99.5%.

[0094] The measured effective information density is approximately 2.9 bits / base, which is approximately 45% higher than the traditional method (2 bits / base); the selective retrieval time of the target region is shortened by 85% compared with the traditional full sequencing method.

[0095] Case 2: Long-term storage stability test

[0096] A set of text data (about 500KB) was encoded and stored, and accelerated aging experiments were conducted under four different conditions using the method of the present invention and the traditional DNA storage method.

[0097] The traditional DNA storage methods compared in this experiment are based on the classic DNA storage technology routes established by Church et al. (2012) and Goldman et al. (2013). The traditional DNA storage methods have the following characteristics:

[0098] Encoding method: Using a standard binary to quaternary mapping strategy, the binary data 01 combination is directly mapped to four bases (00→A, 01→C, 10→G, 11→T), without using additional chemical modification states to carry information.

[0099] Error correction mechanism: Triple redundancy (3× redundancy) is used, meaning each piece of information is replicated three times, with errors resolved through majority voting, rather than the Reed-Solomon code used in this invention.

[0100] DNA fragment structure: Each DNA fragment is 100-120 bp in length, including an 8 bp index tag and a 4 bp error correction code. The index structure is much simpler than the multi-level recognition sequence of the present invention.

[0101] Storage: DNA fragments are directly dissolved in TE buffer (10mM Tris-HCl, 1mM EDTA, pH 8.0) and stored at low temperature, or simply freeze-dried and stored without special protective structures.

[0102] Retrieval method: PCR amplification followed by full-library sequencing does not provide selective retrieval capabilities, and each search requires processing the entire dataset.

[0103] Four different accelerated aging conditions were set up in the experiment: normal temperature (25°C), humidity 50%; high temperature (80°C), humidity 50%;

[0104] UV radiation (254 nm, 5 mW / cm²); repeated freeze-thaw cycles (-20°C to 40°C, 3 times per day).

[0105] The test sample preparation process is as follows:

[0106] Traditional method: After the coding DNA fragment is synthesized, it is purified and dissolved in TE buffer and directly dispensed into microcentrifuge tubes.

[0107] The method of the present invention comprises the following steps: combining the encoded DNA fragments with the DNA nano-origami structure, encapsulating the DNA fragments in a silica nano-shell, adding a stabilizer and then freeze-drying the resulting fragments for storage.

[0108] Sampling and sequencing were performed every 30 days to evaluate data recovery rates. Under all conditions, the method of the present invention showed significant advantages:

[0109] Under room temperature conditions, the data recovery rate of the traditional method dropped to 87% after 180 days, mainly due to spontaneous hydrolysis and oxidative damage of DNA, while the method of the present invention maintained >99%.

[0110] Under high temperature conditions, the data recovery rate of the traditional method dropped to 65% after 90 days, mainly due to base depurination and accelerated chain breakage, while the method of the present invention still maintained >95%.

[0111] Under UV radiation conditions, the data recovery rate of the traditional method dropped below 50% after 60 days, mainly due to the formation of thymine dimers, while the method of the present invention maintained >90%.

[0112] Under freeze-thaw cycle conditions, the data recovery rate of the traditional method dropped to 75% after 120 days, mainly due to physical shear force and ice crystal formation leading to DNA breakage, while the method of the present invention maintained >94%.

[0113] The degradation rate of the traditional method is significantly higher than that of the method of the present invention under all conditions. In particular, under UV radiation conditions, the data recovery rate of the traditional method shows an almost exponential decline, while the method of the present invention shows a relatively gentle linear decline.

[0114] Calculated by the Arrhenius accelerated aging model, according to the formula: , where k is the degradation rate constant, A is the frequency factor, Ea is the activation energy, R is the gas constant, and T is the absolute temperature. It is calculated that the data half-life of the method of the present invention under standard storage conditions (15°C, humidity <20%) can reach about 3000 years, while that of the traditional method is about 300 years. This result shows that the DNA storage protection system of the present invention can significantly extend the storage life, making DNA a truly feasible long-term data archiving medium.

[0115] The above cases fully verify the significant advantages of the present invention in information storage density, retrieval efficiency and long-term stability, and demonstrate the application potential of this technology in the field of permanent preservation of large-scale data.

[0116] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A DNA-encoded information storage method, characterized in that: The method comprises the following steps: Step S1, converting the digital information to be stored into a quaternary code, and converting the quaternary code into a DNA sequence consisting of adenine A, cytosine C, guanine G, and thymine T according to a preset mapping rule, and using the chemical modification state of the nucleotide pairs to represent an additional information dimension; Step S2: inserting error detection and correction codes at every predetermined number of base positions in the DNA sequence, dividing the encoded DNA sequence into multiple fragments of 100-150 bp in length, and adding specific recognition sequences and index markers at both ends of each fragment; Step S3, using the reversibly thermoresponsive DNA nanostructure as a carrier, selectively binding the DNA fragments to the carrier, and realizing hierarchical storage and rapid retrieval of information through a temperature gradient control system; The reversible thermoresponsive DNA nanostructure is: A DNA origami structure having a geometric configuration, wherein the surface of the origami structure is designed with multiple sequence complementary regions as binding sites for information DNA fragments; Different binding sites have different temperature response characteristics, selectively releasing or binding DNA fragments within a specific temperature range; the probability of DNA chain melting at each binding site is Calculated by the following formula: ,in, is the current temperature (K), is the melting point of a specific DNA fragment (K), is the melting enthalpy change (J / mol), is the gas constant (8.314 J / (mol·K)); The temperature gradient control system achieves hierarchical storage and retrieval of information by dividing the temperature range into multiple discrete intervals; each temperature interval corresponds to a specific set of DNA fragments; and selective release of specific information fragments is achieved by precise temperature control. In the multi-layer thermal storage system, the melting point distribution of different binding sites is determined by the formula: Determine, among which, For the Melting point of the layer storage site, As the reference melting point, take 35°C, is the temperature gradient interval, which is 10°C; When performing information retrieval, the temperature gradient is controlled by the formula: implementation, where for The temperature of the moment, is the starting temperature, is the termination temperature, is the total search time.

2. The information storage method based on DNA coding according to claim 1, characterized in that: The preset mapping rules and the use of chemical modification states of nucleotide pairs to represent additional information dimensions include: mapping the quaternary digits 0, 1, 2, and 3 to different bases and their modification states; Specifically: 0 maps to A-unmodified or A-methylated, 1 maps to C-unmodified or C-methylated, 2 maps to G-unmodified or G-methylated, and 3 maps to T-unmodified or T-methylated; Information coding density Calculated by the following formula: ,in, is the number of bases, is the number of bases that can be modified, and L is the total sequence length.

3. The information storage method based on DNA coding according to claim 2, characterized in that: The predetermined number of base positions ranges from 10 to 30 bases.

4. The information storage method based on DNA coding according to claim 3, characterized in that: The error detection and correction code adopts Reed-Solomon error correction code.

5. The information storage method based on DNA coding according to claim 4, characterized in that: The specific identification sequence includes at least: a classification code sequence for identifying information categories; a position index for identifying storage order; a homology identifier for determining whether DNA fragments belong to the same data file; and a protective sequence with anti-degradation properties.

6. The information storage method based on DNA coding according to claim 5, characterized in that: System information retrieval efficiency By formula: Calculate, where For the The release probability of layer information, For the The read rate of layer information (bits / s).

7. The information storage method based on DNA coding according to claim 6, characterized in that: The DNA fragments released at the target temperature are collected, and the DNA sequence is read using high-throughput sequencing technology to remove the identification sequence and index markers, and extract the effective coding sequence; the quaternary code is restored by reverse mapping and converted into original digital information.

8. The information storage method based on DNA coding according to claim 7, characterized in that: The DNA sequence was designed to avoid the following sequence patterns: Repeated sequences containing four or more consecutive identical bases; regions with a GC content lower than 40% or higher than 60%.

9. The DNA encoding-based information storage method according to claim 8, characterized in that: A reversible scrambling algorithm is performed on the original data, a digital watermark that does not affect information reading is embedded in the DNA coding sequence, and a key-controlled mapping rule is used to achieve information encryption storage.

Citation Information

Patent Citations

  • DNA information storage method and DNA information storage carrier

    CN115354046A

  • DNA data storage method based on adaptive arithmetic coding

    CN118173181A