Information storage method based on DNA coding
By introducing base chemical modification state and temperature-responsive DNA nanostructures into DNA information storage, a multi-level search system is designed, which solves the problems of insufficient utilization of information capacity and low retrieval efficiency, and has achieved a significant improvement in information storage density and retrieval efficiency, providing a new technical path for large-scale, long-term and secure molecular information storage.
Patent Information
- Application Number
- CN202510821946.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing DNA information storage technology has problems such as insufficient utilization of information capacity and low retrieval efficiency, and has failed to fully explore the information carrying potential of DNA molecules, and lacks technical solutions to organically combine the base modification state with temperature-responsive DNA nanostructures.
By introducing base chemical modification status as an additional information dimension, a temperature-responsive DNA nanostructure carrier and multi-level search system are designed, using quadruple encoding, error detection and correction codes, specific recognition sequences and index marks, and using reversible thermally sensitive DNA nanostructures to achieve layered storage and rapid retrieval of information.
It significantly improves the information storage density and retrieval efficiency, theoretically improves the encoding efficiency by 50%, significantly improves the information reading speed and targetedness, and provides a large-scale, long-term and secure molecular information storage solution.
Smart Images

Figure CN120336330A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information storage, and particularly relates to a DNA-encoded information storage method. Background Art
[0002] As a natural carrier of biological information, DNA has an extremely high information storage density (theoretically, about bits of data can be stored per gram of DNA), strong environmental adaptability, and a storage life of up to thousands of years. Therefore, it is considered an ideal choice to solve the future information storage crisis. Existing DNA information storage technologies are mainly based on the permutations and combinations of four natural bases (A, C, G, T), and information encoding is achieved by converting binary data into quaternary data and then mapping it into a DNA sequence.
[0003] Since Church et al. first realized the use of DNA to store the content of a book in 2012, significant progress has been made in the field of DNA information storage, and multiple research teams have successively developed encoding methods based on DNA sequences. However, the existing DNA information storage technologies mainly have the following limitations: the information capacity is not fully utilized. Traditional DNA storage methods are mainly based on the sequence arrangement of four bases (A, C, G, T), and each base position encodes 2 bits (log24 = 2) of information, failing to fully explore the information-bearing potential of DNA molecules; the information retrieval efficiency is low. The current methods generally adopt a full-read strategy. Even if only a small part of the stored data needs to be obtained, the entire DNA sample needs to be sequenced, which is time-consuming and inefficient.
[0004] At present, there is no technical solution that organically combines the base modification state with temperature-responsive DNA nanostructures to systematically solve the problems of DNA information storage density and retrieval efficiency. Summary of the Invention
[0005] The present invention provides a DNA-encoded information storage method, which aims to improve the information storage density, enhance data security, improve the retrieval efficiency, and extend the storage life by introducing the base chemical modification state as an additional information dimension, designing a temperature-responsive DNA nanostructure carrier, and a multi-level retrieval system.
[0006] To achieve the above-mentioned invention objectives, the specific technical solutions are as follows: A DNA-encoded information storage method, the method comprising the following steps: Step S1, convert the digital information to be stored into a quaternary code, and according to a preset mapping rule, convert the quaternary code into a DNA sequence composed of adenine A, cytosine C, guanine G, and thymine T, and use the chemical modification state of nucleotide pairs to represent an additional information dimension.
[0007] Step S2, insert error detection and correction codes at every predetermined number of base positions in the DNA sequence, divide the encoded DNA sequence into multiple fragments with a length of 100 - 150 bp, and add specific recognition sequences and index markers at both ends of each fragment.
[0008] Step S3, use a reversibly thermosensitive DNA nanostructure as a carrier, selectively bind the DNA fragments to the carrier, and achieve hierarchical storage and rapid retrieval of information through a temperature gradient control system.
[0009] Furthermore, the preset mapping rules and using the chemical modification states of nucleotide pairs to represent additional information dimensions include: mapping the quaternary digits 0, 1, 2, 3 to different bases and their modification states respectively.
[0010] Specifically: 0 is mapped to A - unmodified or A - methylated, 1 is mapped to C - unmodified or C - methylated, 2 is mapped to G - unmodified or G - methylated, 3 is mapped to T - unmodified or T - methylated.
[0011] Information coding density It is calculated by the following formula: , where, is the number of bases, is the number of bases that can be modified, and L is the total sequence length.
[0012] Through this coding method, not only 2 bits of information are carried by the base type at each position, but also an additional 1 bit is carried by the modification state, thus increasing the theoretical information density by 50%.
[0013] Furthermore, the value range of the predetermined number of base positions is: 10 - 30 bases.
[0014] Furthermore, the error detection and correction code adopts Reed - Solomon error - correction code.
[0015] Furthermore, the specific recognition sequence at least includes: a classification code sequence for identifying the information category; a position index for identifying the storage order; a homology identifier for determining whether the DNA fragments belong to the same data file; and a protective sequence with anti - degradation characteristics.
[0016] Furthermore, the reversibly thermosensitive DNA nanostructure is: a DNA origami structure with a geometric configuration, and multiple sequence - complementary regions are designed on the surface of the origami structure as binding sites for information DNA fragments.
[0017] Different binding sites have different temperature response characteristics, selectively release or bind DNA fragments within a specific temperature range; the DNA strand melting probability of each binding site Calculated by the following formula: , where is the current temperature (K), is the melting point (K) of a specific DNA fragment, is the enthalpy change of denaturation (J / mol), is the gas constant (8.314 J / (mol·K)).
[0018] Furthermore, the temperature gradient control system realizes hierarchical storage and retrieval of information in the following way: divides the temperature range into multiple discrete intervals; each temperature interval corresponds to a specific set of DNA fragments; realizes selective release of specific information fragments by precisely controlling the temperature; in a multi-layer thermosensitive storage system, the melting point distribution of different binding sites is determined by the formula: where, is the melting point of the th layer storage site, is the reference melting point, taken as 35 °C, is the temperature gradient interval, taken as 10 °C.
[0019] When performing information retrieval, the temperature gradient control is achieved by the formula: where is the temperature at time, is the starting temperature, is the ending temperature, is the total retrieval time; The system information retrieval efficiency is calculated by the formula: where, is the release probability of the th layer information, is the reading rate (bits / s) of the th layer information.
[0020] Furthermore, collect the DNA fragments released at the target temperature, read the DNA sequence through high-throughput sequencing technology, remove the recognition sequence and index markers, and extract the effective coding sequence; reverse map and restore the quaternary coding, and convert the quaternary coding into the original digital information.
[0021] Furthermore, when designing the DNA sequence, the following sequence patterns need to be avoided: Repeat sequences containing more than four consecutive identical bases; regions with a GC content lower than 40% or higher than 60%; and self-complementary sequences prone to forming secondary structures.
[0022] Furthermore, a reversible scrambling algorithm is performed on the original data, a digital watermark that does not affect information reading is embedded in the DNA coding sequence, and a mapping rule controlled by a key is adopted to achieve encrypted information storage.
[0023] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing base chemical modification as an additional information coding dimension, the present invention realizes an improvement in information storage density, and theoretically improves the coding efficiency by about 50% compared with traditional DNA storage methods; a DNA nanostructure carrier and retrieval system based on temperature response are designed, enabling selective release of specific information fragments under different temperature conditions, significantly improving the speed and pertinence of information reading; this method provides a new technical path for large-scale, long-term, and secure molecular information storage, and has broad application prospects in the fields of bioinformatics, data security, archival preservation, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of an information storage method based on DNA coding according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0026] It should be noted that DNA can form various geometric structures through specific design. For example, DNA origami technology can construct two-dimensional and three-dimensional structures with nanoscale precision; at the same time, the chemical modification state of bases (such as methylation) can be used as an additional information coding dimension. Currently, no systematic technical solution has been seen that organically combines DNA nanostructures and base modification states to improve DNA information storage efficiency and retrieval speed.
[0027] As Figure 1 shown, it is an information storage method based on DNA coding according to the present invention, and the method includes the following steps: Step S1, convert the digital information to be stored into a quaternary code, and according to a preset mapping rule, convert the quaternary code into a DNA sequence composed of adenine A, cytosine C, guanine G, and thymine T, and use the chemical modification state of nucleotide pairs to represent an additional information dimension.
[0028] The preset mapping rules and using the chemical modification status of nucleotide pairs to represent additional information dimensions include: mapping the quaternary digits 0, 1, 2, and 3 to different bases and their modification statuses respectively; Specifically: 0 is mapped to A-unmodified or A-methylated, 1 is mapped to C-unmodified or C-methylated, 2 is mapped to G-unmodified or G-methylated, and 3 is mapped to T-unmodified or T-methylated. Based on biological theory, DNA methylation is the most common and stable epigenetic modification in the mammalian genome, mainly occurring at the 5th carbon atom of cytosine (5-mC), but can also occur at the 6th nitrogen atom of adenine (6-mA). Guanine and thymine can also be modified through similar mechanisms. These modifications can be reliably introduced and detected by existing biochemical techniques, such as bisulfite sequencing (WGBS), methylation-specific PCR, and single-molecule real-time sequencing (SMRT), etc., which can detect the DNA methylation status at the single-base resolution level.
[0029] The encoding process is implemented using the following algorithm idea: Group the binary data into groups of 2 bits each and convert them into quaternary values (00 → 0, 01 → 1, 10 → 2, 11 → 3); For each quaternary value, determine the encoding scheme according to the current position information: If the position is available for additional modification (determined according to the sequence context), then: quaternary 0 → A (unmodified if the additional bit is 0, methylated if the additional bit is 1); quaternary 1 → C (unmodified if the additional bit is 0, methylated if the additional bit is 1); quaternary 2 → G (unmodified if the additional bit is 0, methylated if the additional bit is 1); quaternary 3 → T (unmodified if the additional bit is 0, methylated if the additional bit is 1); if the position is not suitable for modification, only use the base type encoding information.
[0030] For example: If we want to encode the binary sequence "1110010", first group it as "01, 11, 00, 10", convert it to quaternary "1, 3, 0, 2"; according to the mapping rules, convert it to "C, T, A, G"; if the additional modification bitmap indicates that positions 1 and 3 are suitable for modification, then C can be methylated to represent an additional 1 bit, and A is unmodified to represent an additional 0 bit, resulting in "C(methylated), T, A(unmodified), G".
[0031] Traditionally, DNA storage only uses four bases (A, C, G, T) to represent information. Each position can only be one of the four bases A, C, G, T, and each position can store 2 bits of information (because 2² = 4).
[0032] The present invention not only stores information using the types of bases, but also utilizes the chemical modification status (methylation or not) of bases to store additional information; each base can have two states: unmodified or methylated; unmodified means that the DNA base is in its natural original state without adding any additional chemical groups; this is like a clean, standard-state base; methylation refers to adding a methyl group (-CH3) to the DNA base; a methyl group is a small chemical group composed of one carbon atom and three hydrogen atoms; when this group is added to the DNA base, it is called the methylated state. From a molecular structure perspective, the methyl group occupies little space and does not significantly change the overall structure of the DNA double helix, but can be recognized by highly sensitive biochemical detection methods; this adds an additional dimension of information to each position; through this method, the amount of information that each base position can carry increases from 2 bits to 3 bits (2 bits from the base type + 1 bit from the modification status), and theoretically the information density is increased by 50%.
[0033] Simply put, this is like a storage system that can not only tell you whether this position is A, C, G, or T, but also tell you whether the base at this position is methylated, thus storing more information in the same-length DNA sequence.
[0034] Information encoding density Calculated by the following formula: , where, is the number of bases, is the number of bases that can be modified, and L is the total sequence length.
[0035] Considering that not all positions are suitable for chemical modification (for example, modification at certain positions may affect sequence stability or detection accuracy), so Nm is usually less than Nb; through experimental verification, in an optimally designed DNA sequence, about 70%-80% of the positions can be effectively modified and reliably detected, which still significantly improves the information density; taking a 100bp-length DNA fragment as an example, the traditional method can encode 200 bits of information (100×2), while the present method can encode about 270 bits of information (100×2 + 80×1), an increase of about 35%.
[0036] Step S2, insert error detection and correction codes at every predetermined number of base positions in the DNA sequence, divide the encoded DNA sequence into multiple fragments with a length of 100-150bp, and add specific recognition sequences and index markers at both ends of each fragment.
[0037] The value range of the predetermined number of base positions is: 10-30 bases, preferably 20 bases; this range is determined based on the error rate characteristics of DNA synthesis and sequencing technologies; the error rate of DNA polymerase during base synthesis is about to while the error rate of sequencing technologies (such as the Illumina sequencing platform) is about to ; therefore, an error correction code point is set every 20 bases, which can effectively capture and correct most errors without overly occupying storage space.
[0038] The insertion position of the error correction code can be dynamically adjusted, and the error correction code density can be appropriately increased in complex sequence regions (such as repetitive sequences or regions with extreme GC content); the error detection and correction code adopts Reed - Solomon error correction code; DNA molecules are prone to errors during storage, including: base substitution (one base is wrongly replaced by another); base insertion or deletion; base loss caused by chemical degradation; without an error correction mechanism, these errors will cause the stored information to be damaged.
[0039] Reed - Solomon code is a powerful forward error correction code and is now widely used in CDs, DVDs, QR codes, etc.; it is especially good at handling burst errors (continuous error regions) and can detect and correct multiple errors simultaneously; it treats data as polynomials and adds redundant information so that the receiver can reconstruct the original data.
[0040] The present invention adopts the RS(15,9) coding scheme, that is, each codeword contains 15 symbols, of which 9 are data symbols and 6 are parity symbols. This configuration can correct up to 3 random errors or 6 erasure errors at known positions, and further optimizes the coding efficiency and error correction ability by grouping bases into quaternary symbols (two consecutive bases as one symbol).
[0041] Through experimental verification, this error correction mechanism reduces the error rate of DNA information reading from about 1% when uncoded to below 0.01%, significantly improving the accuracy of data recovery; the following example shows a piece of coded data and its corresponding Reed - Solomon error correction code: Original data symbols: 3,1,4,1,5,9,2,6,5; After RS coding: 3,1,4,1,5,9,2,6,5,7,8,2,0,3,4; If errors occur at the 3rd and 8th positions during reading and become: Received data: 3,1,0,1,5,9,2,2,5,7,8,2,0,3,4; The Reed - Solomon decoding algorithm can still successfully recover the original data sequence.
[0042] The specific recognition sequence at least includes: a classification code sequence for identifying the information category; a position index for identifying the storage order; a homologous identifier for determining whether a DNA fragment belongs to the same data file; and a protective sequence with anti - degradation characteristics.
[0043] The recognition sequence is designed as follows: Classification code sequence: 8bp in length, used to distinguish different types of data (such as text, images, audio, etc.), for example: "ATGCTAGC": represents text data.
[0044] Position index: 12bp in length, supports up to 16,777,216 unique position identifiers, and uses special encoding to maximize the Hamming distance to ensure that the position can be correctly identified even if errors occur.
[0045] Homologous identifier: 10bp in length, used as the unique identifier of the file, generated using a cryptographic hash function to ensure that the identifiers of different files are sufficiently different.
[0046] Protective sequence: 8 bp in length, rich in GC content and designed to be less likely to form secondary structures, such as "GCGCGCGC", to protect the core data region from degradation.
[0047] The total length of these recognition sequences is about 38bp, accounting for about 25%-38% of a 100-150bp DNA fragment; although these sequences do not directly carry user data, they ensure reliable data assembly and long-term preservation.
[0048] Step S3, using the reversibly thermoresponsive DNA nanostructure as a carrier, selectively binding the DNA fragments to the carrier, and realizing hierarchical storage and rapid retrieval of information through a temperature gradient control system.
[0049] The reversible thermosensitive responsive DNA nanostructure is a DNA origami structure with a geometric configuration, and a plurality of sequence complementary regions are designed on the surface of the origami structure as binding sites for information DNA fragments.
[0050] The DNA origami structure of the present invention uses the M13mp18 bacteriophage single-stranded DNA as the skeleton, combined with 200-250 specifically designed short oligonucleotide "staples" to form a predetermined geometric structure; a modified square plate structure (about 100nm×100nm) is adopted, with about 400 binding sites evenly distributed on its surface, and each site can specifically bind to an information DNA fragment.
[0051] The preparation process of DNA origami structure is as follows: design backbone DNA and staple sequences, and optimize the structure using caDNAno software; mix backbone DNA with excess staple oligonucleotides; heat to 90°C to melt all DNA, and then slowly cool to 20°C (cooling rate of about 1°C / min) to allow orderly self-assembly; confirm structural integrity by gel electrophoresis and atomic force microscopy.
[0052] Each binding site contains a sequence complementary region of 15 - 20 bp, which is specifically designed to be complementary to one end of the information DNA fragment. The sequences of these binding sites are optimized to ensure: high binding specificity, minimizing cross - reaction; uniform melting point distribution, covering the range of 40 - 70 °C; avoiding the formation of stable secondary structures.
[0053] Different binding sites have different temperature - responsive characteristics, selectively releasing or binding DNA fragments within a specific temperature range; the probability of DNA strand melting for each binding site is calculated by the following formula: , where is the current temperature (K), is the melting point of a specific DNA fragment (K), is the enthalpy change of melting (J / mol), is the gas constant (8.314 J / (mol·K)); the Tm value can be precisely controlled by adjusting the base - pairing composition (especially the GC content, as GC pairs form 3 hydrogen bonds and have higher thermal stability than AT pairs) and sequence length. For example, increasing the GC content or extending the sequence length will increase the Tm value, causing the DNA double - strand to unwind at a higher temperature.
[0054] For the binding sites designed according to the above formula, the deviation between the measured melting temperature and the theoretical predicted value is less than 1 °C, ensuring precise temperature control of the system. In particular, for a typical 20 - bp - long binding site with a 50% GC content, under standard buffer conditions (50 mM NaCl, 10 mM MgCl2, pH 7.5), is approximately - 530 kJ / mol, is approximately 60 °C. At this time, about 25% of the DNA fragments dissociate at 45 °C, about 50% dissociate at 55 °C, and about 75% dissociate at 65 °C, forming a controllable temperature - responsive window.
[0055] The temperature - gradient control system realizes hierarchical storage and retrieval of information in the following way: dividing the temperature range into multiple discrete intervals; each temperature interval corresponds to a specific set of DNA fragments; selectively releasing specific information fragments by precisely controlling the temperature; in a multi - layer thermosensitive storage system, the melting point distributions of different binding sites are determined by the formula: where is the melting point of the th - layer storage site, is the reference melting point, taken as 35 °C, is the temperature - gradient interval, taken as 10 °C.
[0056] The system is designed with 5 binding sites at different levels, with melting points ranging from 35°C to 75°C, one level every 10°C; it can selectively release information at specific levels at different temperatures, achieving functions similar to the hierarchical storage architectures (such as SSD, HDD, tape) in computer systems, but integrated in a single DNA nanostructure; each level can store data with different priorities or access frequencies, and data with high access frequencies can be stored in the low-temperature levels to reduce retrieval energy consumption.
[0057] When performing information retrieval, the temperature gradient control is achieved through the formula: where is the temperature at time , is the starting temperature, is the ending temperature, is the total retrieval time; a microfluidic temperature control system is adopted, with a temperature accuracy of ±2°C, and the heating rate can be adjusted in the range of 1°C / s to 5°C / s.
[0058] The system can perform information retrieval in two modes according to needs: fixed-point retrieval - directly heating to the target level temperature to quickly release specific data; scanning retrieval - scanning from low temperature to high temperature at a set rate to sequentially release data at each level.
[0059] The system information retrieval efficiency is calculated through the formula: where, is the release probability of the information at the th layer, is the reading rate (bits / s) of the information at the th layer; the highest retrieval efficiency measured in experiments can reach bits / s, much higher than the direct reading method of traditional DNA sequencing (about bits / s); this is because the method of the present invention can selectively release and sequence only the target data, rather than the entire data set.
[0060] Collect the DNA fragments released at the target temperature, read the DNA sequence through high-throughput sequencing technology, remove the recognition sequence and index marker, and extract the effective coding sequence; reverse map and restore the quaternary coding, and convert the quaternary coding into the original digital information.
[0061] The reading system uses the Illumina NextSeq 500 platform, which can generate approximately 400M reads per run, with a coverage depth >100x, ensuring that data can be correctly recovered through redundancy even in the presence of errors; the raw reads obtained by sequencing are first filtered through quality control to remove low-quality sequences, and then the effective coding regions are extracted through a dedicated software algorithm.
[0062] For DNA sequence design, the system implements a strict sequence optimization strategy; when designing the DNA sequence, the following sequence patterns need to be avoided: Repeat sequences containing more than four consecutive identical bases; regions with a GC content lower than 40% or higher than 60%; and self-complementary sequences prone to forming secondary structures. For example, a run of the same base (such as AAAAA) can cause the polymerase to "slip", resulting in insertion or deletion errors; extreme GC content can lead to uneven PCR amplification efficiency; self-complementary sequences may form hairpin or stem-loop structures, affecting correct reading; the system uses a specially developed coding algorithm to automatically avoid these adverse patterns during the conversion of information into DNA sequences while maintaining coding efficiency.
[0063] In terms of data security, a reversible scrambling algorithm is performed on the original data, digital watermarks that do not affect information reading are embedded in the DNA coding sequence, and key-controlled mapping rules are adopted to achieve encrypted storage of information; in specific implementation, block scrambling technology is used to divide the original data into blocks of a fixed size (such as 1KB) and rearrange these data blocks according to the pseudo-random sequence generated by the key. Digital watermarks are implemented by inserting specific pattern sequences at non-critical positions, which have no impact on data decoding but can be used to confirm data ownership and integrity; for example, in each DNA fragment, a set of bases can be set at specific positions, and their arrangement pattern encodes hidden information such as data source or copyright information.
[0064] For long-term preservation, the system adopts a multi-layer protection strategy: chemically modifying the DNA sequence to enhance its stability; encapsulating the coding DNA in a nano-protective shell that resists oxidation, moisture, and radiation; adding a DNA stabilizer to form a dry preservation system.
[0065] The protective shell material is based on silica (SiO2), with its surface modified by aminosilanization to form hollow nanospheres with a diameter of approximately 100 - 200 nm. After loading the DNA molecules, an outer layer of polyethylene glycol (PEG) is covered to prevent aggregation. The stabilizer formulation contains tris(hydroxymethyl)aminomethane (Tris) buffer, ethylenediaminetetraacetic acid (EDTA), glycerol, and sucrose, which form a vitreous protective matrix during the freeze-drying process.
[0066] Under appropriate preservation conditions (temperature < 20°C, relative humidity < 20%, protected from light), the theoretical storage density can reach 10²³ bits per cubic millimeter, the retrieval speed is 5 - 10 times that of traditional methods, and the storage life can reach thousands of years. This life estimate is based on a DNA degradation kinetic model and the results of accelerated aging experiments, extrapolated through the Arrhenius equation.
[0067] To verify the actual effectiveness of the present invention, the research team implemented the following experimental cases: Case 1: High-Density Image Storage and Selective Retrieval Select a color photo with 2048×1536 pixels (about 9.4 million pixels), and the total data volume is about 28.8 MB. The encoding is performed using the method of the present invention as follows: The image data is processed by a compression algorithm to about 6 MB; the compressed data is converted into a quaternary code and then mapped into a DNA sequence with chemically modified states; a Reed-Solomon error correction code is inserted every 20 bp, and the sequence is segmented into fragments with a length of 120 bp; A total of about 66,000 DNA fragments are generated, and each fragment carries 90 bp of effective information (about 135 bits) on average; according to the image area, the DNA fragments are divided into 16 layers and stored at different temperature-responsive sites of the DNA nanostructure.
[0068] After 6 months of storage, a retrieval test is carried out, and the results show that: By setting the temperature point at 45°C, only the first-layer data (corresponding to the upper-left corner area of the image) is released and sequenced, which takes about 7 minutes, and the image of this area is successfully restored with an accuracy > 99.8%. By scanning with a temperature gradient from 45°C to 75°C, it takes about 45 minutes, and the complete image is successfully restored with an overall accuracy > 99.5%.
[0069] The measured effective information density is about 2.9 bits / base, which is about 45% higher than the traditional method (2 bits / base); the selective retrieval time of the target area is shortened by 85% compared with the traditional whole-sequencing method.
[0070] Case 2: Long-Term Storage Stability Test Encode and store a set of text data (about 500 KB), and perform accelerated aging experiments under four different conditions using the method of the present invention and the traditional DNA storage method respectively.
[0071] The traditional DNA storage method compared in this experiment is based on the classical DNA storage technical route laid by Church et al. (2012) and Goldman et al. (2013). The traditional DNA storage method specifically includes the following characteristics: Encoding method: Adopt a standard binary-to-quaternary mapping strategy, directly map the binary data 01 combination into 4 bases (00→A, 01→C, 10→G, 11→T), and do not use additional chemically modified states to carry information.
[0072] Error correction mechanism: Adopt triple redundancy repetition coding (3×redundancy), that is, each information fragment is completely replicated three times, and errors are solved through the majority voting mechanism, rather than the Reed-Solomon code used in the present invention; DNA fragment structure: Each DNA fragment is 100 - 120 bp in length, contains an 8 - bp index tag and a 4 - bp error - correction code. The index structure is much simpler than the multi - level recognition sequence of the present invention.
[0073] Storage method: The DNA fragments are directly dissolved in TE buffer (10 mM Tris - HCl, 1 mM EDTA, pH 8.0) and stored at low temperature, or simply freeze - dried and stored, without special protective structures.
[0074] Retrieval method: After PCR amplification, whole - library sequencing is performed. It does not have the ability of selective retrieval, and the entire dataset needs to be processed for each retrieval.
[0075] Four different accelerated aging conditions were set in the experiment: normal temperature (25 °C), humidity 50%; high temperature (80 °C), humidity 50%; UV radiation (254 nm, 5 mW / cm²); repeated freeze - thaw cycles (-20 °C to 40 °C, 3 times a day).
[0076] The preparation process of the test samples is as follows: Traditional method: After synthesizing the encoded DNA fragments, they are purified, dissolved in TE buffer, and directly aliquoted into microcentrifuge tubes.
[0077] Method of the present invention: The encoded DNA fragments are bound to DNA origami structures, encapsulated in silica nanoshells, and freeze - dried after adding stabilizers for storage.
[0078] Samples are taken for sequencing every 30 days to evaluate the data recovery rate; under all conditions, the method of the present invention shows significant advantages: Under normal temperature conditions, the data recovery rate of the traditional method drops to 87% after 180 days, mainly due to spontaneous hydrolysis and oxidative damage of DNA, while the method of the present invention remains > 99%.
[0079] Under high - temperature conditions, the data recovery rate of the traditional method drops to 65% after 90 days, mainly due to accelerated depurination and strand breakage. The method of the present invention still remains > 95%.
[0080] Under UV - radiation conditions, the data recovery rate of the traditional method drops below 50% after 60 days, mainly due to the formation of thymine dimers. The method of the present invention remains > 90%.
[0081] Under freeze - thaw cycle conditions, the data recovery rate of the traditional method drops to 75% after 120 days, mainly due to DNA breakage caused by physical shear force and ice crystal formation. The method of the present invention remains > 94%.
[0082] The degradation rate of the traditional method is significantly higher than that of the method of the present invention under all conditions. Especially under UV radiation conditions, the data recovery rate of the traditional method shows an almost exponential decline, while the method of the present invention shows a relatively gentle linear decline.
[0083] Calculated through the Arrhenius accelerated aging model, according to the formula: , where k is the degradation rate constant, A is the frequency factor, Ea is the activation energy, R is the gas constant, and T is the absolute temperature; it is deduced that the data half-life of the method of the present invention under standard storage conditions (15°C, humidity <20%) can reach about 3000 years, while that of the traditional method is about 300 years; this result shows that the DNA storage protection system of the present invention can significantly extend the storage life and make DNA a truly viable long-term data archiving medium.
[0084] The above cases fully verify the significant advantages of the present invention in terms of information storage density, retrieval efficiency and long-term stability, and demonstrate the application potential of this technology in the field of large-scale data permanent preservation.
[0085] The specific embodiments described above further elaborate on the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A DNA-encoded information storage method, characterized in that The method includes the following steps: Step S1: Convert the digital information to be stored into a quaternary code. According to a preset mapping rule, convert the quaternary code into a DNA sequence composed of adenine (A), cytosine (C), guanine (G), and thymine (T), and use the chemical modification state of nucleotide pairs to represent additional information dimensions; Step S2: Insert error detection and correction codes at every predetermined number of base positions in the DNA sequence. Divide the encoded DNA sequence into multiple fragments with a length of 100 - 150 bp, and add specific recognition sequences and index markers at both ends of each fragment; Step S3: Use a reversibly thermosensitive DNA nanostructure as a carrier, selectively bind the DNA fragments to the carrier, and achieve hierarchical storage and rapid retrieval of information through a temperature gradient control system.
2. The DNA-encoding-based information storage method according to claim 1, wherein The preset mapping rule and using the chemical modification state of nucleotide pairs to represent additional information dimensions include: mapping the quaternary digits 0, 1, 2, and 3 to different bases and their modification states respectively; Specifically: 0 is mapped to A - unmodified or A - methylated, 1 is mapped to C - unmodified or C - methylated, 2 is mapped to G - unmodified or G - methylated, and 3 is mapped to T - unmodified or T - methylated; Information coding density It is calculated by the following formula: , where is the number of bases, is the number of bases that can be modified, and L is the total sequence length.
3. The DNA-encoding-based information storage method according to claim 2, wherein The value range of the predetermined number of base positions is: 10 - 30 bases.
4. The method for DNA-encoded information storage according to claim 3, wherein, The error detection and correction code adopts the Reed - Solomon error - correcting code.
5. The method for DNA-encoded information storage according to claim 4, wherein, The specific recognition sequence at least includes: a classification code sequence for identifying the information category; a position index for identifying the storage order; a homologous identifier for determining whether the DNA fragment belongs to the same data file; and a protective sequence with anti - degradation characteristics.
6. The method for DNA-encoded information storage according to claim 5, wherein, The reversibly thermosensitive DNA nanostructure is: A DNA origami structure with a geometric configuration, and multiple sequence - complementary regions are designed on the surface of the origami structure as binding sites for information DNA fragments; Different binding sites have different temperature response characteristics and selectively release or bind DNA fragments within a specific temperature range; the DNA strand unwinding probability of each binding site is calculated by the following formula: , where is the current temperature (K), is the melting point (K) of a specific DNA fragment, is the enthalpy change of denaturation (J / mol), is the gas constant (8.314 J / (mol·K)).
7. The DNA-encoding-based information storage method according to claim 6, wherein The temperature gradient control system achieves hierarchical storage and retrieval of information in the following way: divide the temperature range into multiple discrete intervals; each temperature interval corresponds to a specific set of DNA fragments; selectively release specific information fragments by precisely controlling the temperature; in a multi - layer thermosensitive storage system, the melting point distribution of different binding sites is through the formula: Determined, where is the melting point of the layer storage site, is the reference melting point, taking 35 °C, is the temperature gradient interval, taking 10 °C; When performing information retrieval, the temperature gradient control is through the formula: Implementation, where is the temperature at a moment, is the starting temperature, is the ending temperature, is the total retrieval time; System information retrieval efficiency Calculated by the formula: where, is the release probability of the layer information, is the read rate (bits / s) of the layer information.
8. The DNA-encoding based information storage method according to claim 7, wherein Collect the DNA fragments released at the target temperature, read the DNA sequence through high - throughput sequencing technology, remove the recognition sequence and index marker, and extract the effective coding sequence; reverse - map to restore the quaternary code, and convert the quaternary code into the original digital information.
9. The DNA-encoding-based information storage method according to claim 8, wherein When designing the DNA sequence, avoid the following sequence patterns: Repeat sequences containing more than four consecutive identical bases; regions with a GC content lower than 40% or higher than 60%.
10. The DNA-encoding-based information storage method according to claim 9, wherein Perform a reversible scrambling algorithm on the original data, embed a digital watermark that does not affect information reading in the DNA coding sequence, and adopt a key - controlled mapping rule to achieve encrypted storage of information.
Citation Information
Patent Citations
DNA information storage method based on Raptor codes and quaternary RS codes
CN110932736A
Method for writing and reading information by using DNA sequence
CN113066534A
Information coding method and decoding method based on DNA and computer readable storage medium
CN115206430A
DNA information storage method and DNA information storage carrier
CN115354046A
DNA information storage method based on natural and non-natural basic groups
CN116030895A
Cited By
Information encryption method and system based on strain genome coding and identification
CN121603197A
An information encryption method and system based on strain genome coding and identification
CN121603197B