A selective access method for DNA data storage

By optimizing primer design and incorporating factors such as GC content, homopolymer, and free energy change ΔG, the problems of high reading cost, long reading time, and low accuracy in DNA data storage were solved, achieving efficient and accurate selective access.

CN116417071BActive Publication Date: 2026-04-03XIANGFU LAB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing DNA data storage technologies suffer from high reading costs, long reading times, and low selective reading accuracy, mainly because existing primer designs do not fully consider the impact of primer neck loop structure and mismatched single strands on DNA hybridization.

Method used

When designing primer sets, factors such as GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure are comprehensively considered. Selective amplification is achieved through multiplex PCR. Primer sequences are generated and optimized using Matlab or Python, and primers are modified to improve specificity.

Benefits of technology

It enables low-cost, high-accuracy DNA data reading, reduces non-specific amplification, and expands the capacity for DNA data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116417071B_ABST
    Figure CN116417071B_ABST
Patent Text Reader

Abstract

This invention provides a selective access method for DNA data storage, comprising the following steps: S1: Converting digital information into DNA sequences based on the mapping relationship between binary and DNA bases; S2: Designing primer sets based on comprehensive considerations of GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure; S3: Synthesizing DNA sequences; S4: Selecting the corresponding primer set for the file to be accessed and selectively amplifying the DNA sequence of the content to be accessed through multiplex PCR; S5: Reading the DNA sequence information through sequencing; S6: Decoding the read DNA sequence information to recover the file. This invention, based on a comprehensive consideration of GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure in designing PCR primers, ultimately achieves low-cost, high-accuracy, and rapid reading of target files from DNA databases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of DNA data storage technology, and more specifically to a selective access method for DNA data storage. Background Technology

[0002] With the development of the internet and computer technology, digital information is being generated at an ever-increasing pace. According to data from the International Data Corporation (IDC), global big data storage increased from 21.6 ZB in 2017 to 60 ZB in 2020, and is projected to reach 175 ZB by 2025, requiring at least 175 trillion 1GB portable hard drives for storage. Currently used storage media, such as hard drives and flash memory, are traditional magnetic or optical storage media with limited capacity and a lifespan of generally 5 to 10 years. Optical discs also have a lifespan of only 50 years. Therefore, current storage media are gradually failing to meet the global data storage needs. Thus, facing the ever-increasing volume of data storage and the current limitations of storage space and lifespan of storage media, there is an urgent need for a data storage material with high storage density and long-term durability.

[0003] DNA, or deoxyribonucleotide, is composed of four bases: adenine (A), thymine (T), cytosine (C), and guanine (G). DNA is a long-chain polymer that stores a vast amount of genetic information. The content of DNA data storage can be text, images, sound, and video. The DNA data storage process involves two parts: first, translating the 0s and 1s in the binary file into ATCG and converting them into a DNA sequence, then using high-throughput synthesis technology to synthesize the DNA sequence for information storage; second, sequencing and decoding the DNA sequence, which can be done using Sanger, Illumina, or nanopore sequencing technologies, and finally using different algorithms to decode the sequencing results into a binary file.

[0004] DNA data storage has the following advantages: First, it has a high storage density, approximately 10^6. 19 bit / cm 3 These are the data storage densities of hard drives and flash memory, respectively, at 10. 6 and 10 3First, DNA has several advantages: 1 gram can store up to 455 exabytes of data, and storing all the current global data would only require about 1 kilogram of DNA (Nucleic acids research, 2021, 49, 5451–5469.)(Nature, 2016, 537(7618), 22–24.). Second, it has good stability; under suitable conditions, some DNA fossils can exist stably for hundreds of thousands or even millions of years (Nature, 2019, 576, 262–265). Third, its storage and maintenance costs are low, requiring no large investment of manpower or financial resources; only the DNA data needs to be preserved in a low-temperature environment. Therefore, DNA molecules, with their high storage density, stability, and low maintenance costs, are expected to become a new generation of information storage medium.

[0005] The development of DNA data storage technology is still limited by the timeliness and economic cost of data writing and reading, but with the development of synthesis and sequencing technologies, the cost will gradually decrease. However, due to the complexity of DNA databases, which contain multiple sequences, accurately locating the desired sequence and avoiding the need for complete and expensive sequencing of the entire DNA database would be crucial for reducing the cost of DNA data reading. Therefore, establishing an addressing system for each DNA sequence to select the required file from the complex DNA library is essential.

[0006] Current PCR-based addressing systems utilize address-specific primers to achieve highly specific PCR amplification and selective enrichment of target sequences. However, while PCR primers can be designed using tools like NCBI and Primer Premier 5, these designs are based on Tm-based primer deduction and do not fully consider the impact of primer neck loop structure and mismatched single strands on DNA hybridization. Furthermore, primer sequences provided in literature or online reports may not be suitable for the intended research, and due to a lack of theoretical and algorithmic support, they are difficult to extend and improve to meet storage needs. Using patented primers requires obtaining the necessary authorization. Therefore, it is essential to develop a novel method for storing and accessing DNA data. Summary of the Invention

[0007] The purpose of this invention is to provide a selective access method for DNA data storage, thereby solving the problems of high reading costs, long reading times, and low selective reading accuracy in existing DNA data storage and access methods due to the complexity of DNA data libraries.

[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0009] A selective access method for DNA data storage is provided, comprising the following steps: S1: Based on the mapping relationship between binary and DNA bases, digital information is converted into DNA sequences, the DNA sequences are divided into several small information sequences, and an address sequence is added; S2: Primer sets that can be added to both ends of the sequences are designed, the primer sets are designed based on a comprehensive consideration of GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure; S3: DNA sequences are synthesized according to the designs in steps S1 and S2 and stored under suitable conditions, each synthesized DNA sequence includes a linked information sequence and an address sequence, and forward and reverse primer sequences added to both ends of it respectively; S4: The corresponding primer set of the file to be accessed is selected, and the DNA sequence of the content to be accessed is selectively amplified by multiplex PCR to enrich the target sequence and add sequencing adapters to both ends; S5: The DNA sequence information is read by sequencing; S6: The read DNA sequence information is decoded based on the mapping relationship between DNA bases and binary information to recover the file, thereby realizing selective access to DNA data storage.

[0010] According to a preferred embodiment of the present invention, step S2 includes: S21: Using the random function function of Matlab or Python, generate a pair of random primer sequences with a sequence length between 15 and 30 nt, verify whether their GC content is between 40% and 60% and whether the maximum homopolymer length does not exceed 4. If yes, the verification is passed; otherwise, the verification fails, and some bases are changed before re-verification; S22: Further calculate whether the primer binding free energy change ΔG is between -4 kcal / mol and -20 kcal / mol. If yes, the verification is passed; otherwise, the verification fails, and the primer is discarded, returning to step S21; S23: Further perform complementarity evaluation. At the 3' end position, the number of complementary bases with itself or another primer does not exceed 4, and at the middle position, it does not exceed 6. If yes, the verification is passed; otherwise, the primer is discarded, and the primer is discarded, returning to step S21; S24: Further examine the formation of its secondary structure. Primers with obvious secondary structures are discarded, and the primer is returned to step S21. Primers that pass the verification are added to the available primer library.

[0011] Preferably, in step S22, the primer binding free energy change ΔG is calculated based on the nearest neighbor model, using the corresponding thermodynamic parameters, to calculate the free energy ΔG=ΔH–TΔS, where the unit of T is Kelvin.

[0012] More preferably, in step S22, if the primer binding free energy change ΔG calculated under conventional PCR conditions is between -10.5 kcal / mol and -12.5 kcal / mol, then the verification is successful. As an example and not a limitation, conventional PCR conditions include Na...+ The ion concentration was 0.18 M and the temperature T was 60 °C.

[0013] According to a preferred embodiment of the present invention, step S2 further includes: verifying the amplification specificity of the primers using a sequence alignment tool.

[0014] According to a preferred embodiment of the present invention, step S2 further includes: the primer may be modified during design with any functional group selected from: RNA, LNA, PNA, XNA, dU, Spacer, PEG, fluorescent group, phosphorylated group, reverse dT, methylated base.

[0015] Preferably, step S4 includes: combining 1 to N pairs of primers as needed, mixing them in a certain proportion, selectively amplifying the DNA sequence and sequencing it under appropriate PCR conditions.

[0016] The purpose of this invention is to amplify specific target file sequences in a DNA database using designed primers via PCR technology, and to achieve selective reading of the target file sequences through primer sequences. This selective reading method avoids the waste of sequencing time and costs caused by sequencing the entire database to find a specific target file sequence.

[0017] In step S4, the PCR conditions include: 1) The PCR cycling steps can be performed according to the denaturation-annealing-extension method, or according to the denaturation-annealing + extension method; 2) The primer concentration range for PCR is between 1 nM and 100 μM; 3) The annealing temperature range for PCR is between 20℃ and 72℃, and the annealing temperature can also be changed with the increase of the number of cycles; 4) The polymerase used in PCR is a low-fidelity polymerase or a high-fidelity polymerase, where low-fidelity polymerases include: Taq, PowerUp, iTaq, Universal Blue, etc., and high-fidelity polymerases include: Kapa, ​​Phusion, Q5, etc.; 5) In the reaction solution used in PCR, Mg 2+ Concentrations between 1 mM and 100 mM, dNTP concentrations between 20 μM and 20 mM, DMSO concentrations between 1% and 30%, and Na... + The concentration is between 100mM and 10M.

[0018] According to the present invention, a selective access method for DNA data storage is provided, particularly a primer design method. This method designs suitable primer sets by comprehensively considering a series of factors, including GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure. This ultimately achieves low-cost, high-accuracy, and rapid reading of target files from DNA data. Existing methods do not consider non-specific amplification of the library due to primer structure, because the primer design methods used in the prior art are usually based on Tm prediction and do not take into account the influence of primer neck loop structure and unmatched single strands on DNA hybridization.

[0019] The key inventive point of this invention lies in designing primers that specifically bind to the DNA sequence of the target file based on the calculation of the primer binding free energy change ΔG. This allows for the detection of primer dimers and non-specific amplification, increasing the accuracy of selective reading. The designed primers are then used to selectively amplify the corresponding target file sequence in the entire DNA database using PCR technology. Finally, selective reading of the target file sequence is achieved based on the primer sequence.

[0020] In summary, the selective access method for DNA data storage provided by this invention, compared with existing technologies that design primers based on Tm values, utilizes a comprehensive consideration of GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure to design PCR primers. This method can check for primer dimerization and non-specific amplification, increasing the accuracy of selective reading. Furthermore, the primer design method of this invention has the potential to be extended to the design of multiple primer sets, enabling the creation of larger primer sets to increase the capacity of DNA data storage. Attached Figure Description

[0021] Figure 1 The entire process of DNA storage and retrieval is illustrated;

[0022] Figure 2 A schematic diagram illustrating the principle of selective access to target files in DNA storage based on PCR amplification technology is shown.

[0023] Figure 3 A flowchart of primer set design is shown in a selective access method for DNA data storage provided by the present invention;

[0024] Figure 4 The types of groups that can be modified on the primers are shown. Detailed Implementation

[0025] The present invention will be further described below with reference to specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] Combination Figure 1 , Figure 2 , Figure 3 The illustration shows a selective access method for DNA data storage according to a preferred embodiment of the present invention. The method mainly includes the following steps:

[0027] 1) DNA Sequence Encoding and Synthesis: Any file on a computer is essentially composed of binary information like 0101. The first step in DNA storage is to convert this binary information into a DNA sequence based on the mapping relationship between digital information and bases. Many different encoding methods have been developed for this purpose. For ease of understanding, we introduce two simple mapping methods: ① A, T, C, G correspond to 00, 01, 10, and 11 respectively; ② A / C corresponds to 1, and T / G corresponds to 0. Based on the chosen mapping method, the digital information of the file can be converted into a long DNA sequence. However, due to limitations in DNA synthesis technology, the sequence usually needs to be divided into many small fragments, and an address sequence needs to be added to determine the position of each fragment during restoration. Finally, annealing sites for forward primers (FP) and reverse primers (RP) are added to both ends of the sequence (annealing is the process of primers binding to the template), resulting in the final synthesized sequence. The designed DNA sequence is then synthesized using chemical synthesis or enzymatic methods and stored in a suitable environment.

[0028] 2) DNA Sequence Sequencing and Decoding: Reading can be divided into two sub-steps: library construction and sequencing. First, the corresponding primer set for the file to be accessed is selected, and the target sequence is enriched by polymerase chain reaction (PCR) amplification (to increase the copy number of nucleic acid molecules), with sequencing adapters added to both ends. Then, the constructed library is sequenced using sequencing instruments such as Illumina. Finally, the sequenced DNA sequence is converted into digital information according to the mapping table used during encoding, thereby recovering the file information stored in the DNA sequence.

[0029] Clearly, selective access is crucial in DNA storage. By designing PCR primers to select target files, it's unnecessary to read and decode all information from the DNA library, thus reducing the time and cost of decoding. First, each file contains a corresponding primer set for its encoded DNA sequence. Second, for the desired target file, PCR amplification using the primer set for that target file achieves specific amplification of the target file sequence. Third, following standard library construction procedures, sequencing adapters are added to both ends of the sequence, and sequencing is performed. Finally, the sequencing results are decoded and analyzed to obtain the corresponding file information.

[0030] Therefore, this invention provides a selective access method for DNA data storage, particularly a primer design method. By comprehensively considering factors such as GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure, a suitable primer set is designed, ultimately achieving selective access to stored DNA data. This method mainly includes the following steps:

[0031] S1: Based on the mapping relationship between binary and DNA bases, digital information is converted into a DNA sequence, which is then divided into several small information sequences and an address sequence is added.

[0032] S2: Design primer sets that can be added to both ends of a sequence, the primer sets being designed based on a comprehensive consideration of GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure;

[0033] S3: Synthesize DNA sequences according to the design in steps S1 and S2 and store them under appropriate conditions. Each synthesized DNA sequence includes an information sequence and an address sequence linked together, as well as forward primer sequences and reverse primer sequences added to both ends of it respectively.

[0034] S4: Select the corresponding primer set for the file to be accessed, selectively amplify the DNA sequence of the content to be accessed by multiplex PCR to enrich the target sequence and add sequencing adapters to both ends of it.

[0035] S5: Read DNA sequence information through sequencing;

[0036] S6: Based on the mapping relationship between DNA bases and binary information, the read DNA sequence information is decoded to recover the file, enabling selective access to DNA data storage.

[0037] According to a preferred embodiment of the present invention, in combination with Figure 3 As shown, the primer set design process involved in step S2 is described below, mainly including the following sub-steps:

[0038] S21: Use the random function function of Matlab or Python, such as random, to generate a pair of random primer sequences with a sequence length between 15 and 30 nt. Verify whether the GC content is between 40% and 60% and whether the maximum homopolymer length does not exceed 4. If yes, the verification is successful; otherwise, the verification fails. Change some bases and re-verify.

[0039] It should be understood that GC content is calculated by counting the number of A, T, C, and G using programming languages ​​such as Matlab or Python, and CG content = (C+G) / (A+T+C+G). Similarly, the maximum homopolymer length can be achieved using a programming language.

[0040] S22: Further calculate whether the primer binding free energy change ΔG under conventional PCR conditions is between -4kcal / mol and -20kcal / mol. If it is, the verification is successful. For primer sets with ΔG outside the range, discard them and return to step S21 to generate new sequences for optimization design.

[0041] It should be noted that the primer binding free energy change ΔG can be calculated based on the nearest neighbor model, using the corresponding thermodynamic parameters, such as ΔH and ΔS, to calculate the free energy ΔG = ΔH – TΔS, where T is in Kelvin.

[0042] More preferably, under conventional PCR conditions (including but not limited to Na) + (Ion concentration: 0.18 M, temperature: T: 60 °C) Primer binding free energy change ΔG ranges from -10.5 kcal / mol to -12.5 kcal / mol.

[0043] As an example: under normal PCR conditions, such as Na... + At a concentration of 0.18 M and a temperature T of 60 °C, primers GCTTTCC were designed with ΔG = -3.19 kcal / mol and a length of 8 nt. Synthesis at this point is difficult and costly. Excessive energy at this concentration may lead to DNA sequence amplification failure and subsequent selective reading. Conversely, if ΔG is too low, for example...

[0044] GCTCTTCCTCTCACATCTTTATTTAACCCATTAGAAAATCCTATCAGCTCTA GAC, ΔG=-26.39kcal / mol, length is 57nt. Primers that are too long are prone to self-complementation and secondary structures.

[0045] S23: Further complementarity assessment is performed. The number of complementary bases with itself or another primer should not exceed 4 at the 3' end and should not exceed 6 at the middle position. If so, the verification is successful. Otherwise, the verification fails and the sample is discarded. Return to step S21.

[0046] S24: Further examine the formation of its secondary structure. Primers with obvious secondary structures are discarded directly, and the process returns to step S21. Primers that pass the verification are added to the available primer library.

[0047] It should be understood that the verification of secondary structures can use nucleic acid structure and hybridization prediction software provided by some institutions, such as NUPACK and Mfold.

[0048] According to a preferred embodiment of the present invention, primers can be endowed with additional functions by modifying them with different functional groups, such as... Figure 4 The diagram illustrates the types of modifiable groups on primers. Modifying RNA or nucleic acid analogs such as LNA, PNA, and XNA can alter the primer's binding affinity to the template, thereby improving the primer's single-base resolution (preventing misclassification of individual bases). Modifying multiple dU bases in the middle of the primer, combined with uracil-DNA glycosylase (UDG enzyme, which specifically cleaves dU), can reduce primer dimer formation. Molecular beacons with fluorescent and quenching groups modified at the 5' end of the primer can enable real-time monitoring of the PCR amplification process.

[0049] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention. Various variations can be made to the above embodiments of the present invention. All simple and equivalent changes and modifications made in accordance with the claims and description of this application fall within the protection scope of the claims of this patent. All aspects not described in detail in this invention are conventional technical content.

Claims

1. A method for selective access to DNA data storage, characterized in that, Includes the following steps: S1: Based on the mapping relationship between binary and DNA bases, digital information is converted into a DNA sequence, which is then divided into several small information sequences and an address sequence is added. S2: Design primer sets that can be added to both ends of a sequence, the primer sets being designed based on a comprehensive consideration of GC content, homopolymer, primer binding free energy change ΔG, self-complementarity, and secondary structure; S3: Synthesize DNA sequences according to the design in steps S1 and S2 and store them under appropriate conditions. Each synthesized DNA sequence includes an information sequence and an address sequence linked together, as well as forward primer sequences and reverse primer sequences added to both ends of it respectively. S4: Select the corresponding primer set for the file to be accessed, selectively amplify the DNA sequence of the content to be accessed by multiplex PCR to enrich the target sequence and add sequencing adapters to both ends of it. S5: Read DNA sequence information through sequencing; S6: Based on the mapping relationship between DNA bases and binary information, the read DNA sequence information is decoded to recover the file, enabling selective access to DNA data storage; Step S2 includes the following sub-steps: S21: Use the random function function of Matlab or Python to generate a pair of random primer sequences with a sequence length between 15 and 30 nt. Evaluate whether the GC content is between 40% and 60% and whether the maximum homopolymer length does not exceed 4. If yes, the screening is passed; otherwise, the screening is failed. Change some bases and screen again. S22: Further calculate whether the primer binding free energy change ΔG is between -4 kcal / mol and -20 kcal / mol. If it is, the verification is successful; otherwise, the verification is unsuccessful. Discard the code and return to step S21. S23: Further evaluate complementarity. At the 3' end, the number of complementary bases with itself or another primer should not exceed 4, and at the middle position, it should not exceed 6. If so, the verification is successful; otherwise, the verification is unsuccessful, and the sample is discarded and returned to step S21. S24: Further examine the formation of its secondary structure. Primers with obvious secondary structures are discarded directly, and the process returns to step S21. Primers that pass the verification are added to the available primer library.

2. The selective access method according to claim 1, characterized in that, In step S22, the primer binding free energy change ΔG is calculated based on the nearest neighbor model, using the corresponding thermodynamic parameters, and the free energy is calculated as ΔG = ΔH – TΔS, where T is in Kelvin.

3. The selective access method according to claim 1, characterized in that, In step S22, if the primer binding free energy change ΔG under conventional PCR conditions is between -10.5 kcal / mol and -12.5 kcal / mol, then the verification is successful.

4. The selective access method according to claim 1, characterized in that, Step S2 also includes: verifying the amplification specificity of the primers using sequence alignment tools.

5. The selective access method according to claim 1, characterized in that, Step S2 further includes: the primers can be modified during design with any functional group selected from: RNA, LNA, PNA, XNA, dU, Spacer, PEG, fluorescent group, phosphorylated group, reverse dT, methylated base.

6. The selective access method according to claim 1, characterized in that, Step S4 includes: combining 1 to N pairs of primers as needed, mixing them in a certain proportion, selectively amplifying the DNA sequence and sequencing it under appropriate PCR conditions.

7. The selective access method according to claim 1, characterized in that, In step S4, the PCR conditions include: 1) PCR cycling steps can be performed in the order of denaturation-annealing-extension, or in the order of denaturation-annealing + extension. 2) The primer concentration range for PCR is between 1 nM and 100 μM; 3) The annealing temperature range for PCR is between 20℃ and 72℃, and the annealing temperature can also be changed as the number of cycles increases; 4) The polymerase used in PCR is either a low-fidelity polymerase or a high-fidelity polymerase. Low-fidelity polymerases include: Taq, PowerUp, iTaq, Universal Blue, and high-fidelity polymerases include: Kapa, ​​Phusion, Q5. 5) In the reaction solution used in PCR, Mg 2+ Concentrations ranged from 1 mM to 100 mM, dNTP concentrations from 20 μM to 20 mM, DMSO concentrations from 1% to 30%, and Na... + The concentration is between 100 mM and 10 M.

Citation Information

Patent Citations

  • Library construction method for improving data uniformity of amplicon library

    CN106283200A

  • Polynucleotide libraries having controlled stoichiometry and synthesis thereof

    CN110382752A