PLS DNA polymerase constructed by DNA polymerase-DNA binding protein and application thereof in DNA information storage
By constructing a PLS DNA polymerase, the problem of traditional storage devices being unable to meet ZB-level data storage requirements was solved, achieving efficient and low-cost DNA information storage and enhancing data fidelity and amplification capabilities.
Patent Information
- Application Number
- CN202510026580.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Traditional storage devices are unable to meet the ZB-level data storage requirements, and silicon-based data storage has a limited retention time, necessitating new storage media to bridge this gap.
PLS DNA polymerase was constructed by fusing 9°N DNA polymerase isolated from marine thermophilic archaea with Sso7d DNA-binding protein using protein engineering methods. This formed a stable fusion protein-primer-template complex, which enhanced the DNA polymerase's synthetic capacity and was used for DNA information storage.
It improves the fidelity and amplification efficiency of DNA information storage, reduces storage costs, enables stable amplification of longer DNA fragments in complex environments, and improves the accuracy and efficiency of data decoding.
Smart Images

Figure CN119751704B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of biochemistry and molecular biology, and particularly relates to a PLS DNA polymerase constructed by a DNA polymerase-DNA binding protein and application thereof in DNA information storage. BACKGROUND
[0002] In recent years, the shift of industrial network demand to the cloud and the proliferation of smart devices connected to the Internet have driven the rapid growth of global digital data output. With the explosive growth of data, traditional storage devices, such as magnetic, optical and solid-state devices, are approaching their physical limits and are difficult to meet the requirements of digital storage. If ZB-level data is to be stored, a large amount of physical space needs to be occupied. In addition, the preservation time of silicon-based data storage is limited, which is also a great challenge to existing storage devices. Therefore, there is an urgent need for alternative storage media to make up for this growing gap. The great demand for information storage in the field of big data has promoted the cross-fusion of biotechnology and information technology, giving rise to the DNA-based information storage paradigm. As a natural carrier of genetic information, DNA has the characteristics of environmental friendliness, long-term storage stability and information density scalability, and is considered a potential alternative to solve the problems of traditional information storage methods.
[0003] In the process of DNA-based data storage, binary 0 / 1 data is first converted into quaternary DNA sequences, and then the substrate containing a chemically active protecting group is activated step by step by phosphoramidite chemistry to achieve controllable synthesis of DNA. For data reading, a specific DNA sequence is selected and sequenced, and the measured DNA sequence information is converted back to the original data, thereby recovering the stored digital file. As a kind of tool enzyme widely used in DNA replication in molecular biology, DNA polymerase plays an important role in the whole process of DNA data storage: (1) through the design of primers, DNA polymerase can amplify specific DNA sequences from a large number of DNA sequences, thereby realizing the addressing and retrieval of specific files in mixed files; (2) as an important protein in nanopore sequencing, DNA polymerase controls the moving speed of DNA strand through the nanopore, thereby improving the accuracy of the current change signal generated when different bases pass through the nanopore, which helps to improve the accuracy of nanopore sequencing; (3) using DNA polymerase to extend the overlapping DNA fragments of hybridization, the assembly of long fragment DNA is realized, overcoming the length limitation of chemical synthesis of DNA from scratch.
[0004] Sso7d DNA binding protein is a small molecule DNA binding protein from extreme thermophilic archaea, which has the property of non-specific binding with double-stranded deoxyribonucleic acid (dsDNA) and plays a key role in the stability of the genome in thermophilic archaea. SUMMARY
[0005] The purpose of the present application is to provide a PLS DNA polymerase constructed by DNA polymerase-DNA binding protein and its application in DNA information storage.
[0006] By means of protein engineering, the present application fuses 9°N DNA polymerase isolated from marine thermophilic archaea with Sso7d DNA binding protein to construct 9°N-Sso7d fusion protein polymerase (named as PLS DNA polymerase). The PLS DNA polymerase can be fixed on dsDNA like a sliding clamp to form a more stable fusion protein-primer-template complex, thereby assisting DNA polymerase to quickly recognize and extend DNA, which helps to enhance the continuous synthesis ability of 9°N DNA polymerase. Therefore, the PLS DNA polymerase as a key DNA polymerase molecule in the process of DNA information storage can participate in the data writing and addressing process of DNA information storage, build a good platform for DNA information storage coding and decoding, and promote the application of DNA polymerase in the fields of in vivo information storage and in vitro information storage.
[0007] The PLS DNA polymerase is composed of 9°N DNA polymerase, flexible connection short peptide and Sso7d DNA binding protein. The N terminus of the flexible connection short peptide is connected with the C terminus of the 9°N DNA polymerase domain, and the C terminus of the flexible connection short peptide is connected with the N terminus of the Sso7d DNA binding protein domain; the flexible connection short peptide prevents the two-sided proteins from being incorrectly folded to cause activity impairment, and its amino acid sequence is shown in SEQ ID No. 2.
[0008] The PLS DNA polymerase gene is subjected to E. coli preferred codon optimization, the codon-optimized PLS DNA polymerase gene is constructed on a prokaryotic expression vector, and after the plasmid containing the fusion protein gene is synthesized by a chemical method, plasmid transformation and expression strain construction are performed, and high-purity PLS DNA polymerase is prepared by two-step purification of nickel column affinity chromatography and cation exchange chromatography; with DNA containing storage information as a template, after optimization of the amplification system, the addressing ability of the PLS DNA polymerase in DNA information storage and its effect of assembling short fragment DNA are evaluated.
[0009] The preparation method of the PLS DNA polymerase constructed by DNA polymerase-DNA binding protein according to the present application comprises the following steps:
[0010] (1) Construction of PLS DNA polymerase engineering bacteria
[0011] The gene sequence encoding the PLS DNA polymerase is subjected to codon optimization treatment of the engineering bacteria, and the optimized nucleotide sequence is shown as SEQ ID No. 1, which is used as the gene encoding the PLS DNA polymerase;
[0012] A 6×His tag is added to the 5' end of the gene encoding the PLS DNA polymerase, and the gene encoding the PLS DNA polymerase with the added 6×His tag is connected to the expression vector pET-30a through Nde I and Hind III enzyme cutting sites to construct a recombinant expression vector pET-30a / PLS;
[0013] The recombinant expression vector pET-30a / PLS is transformed into the competent cell Escherichia coli Trans5α, and then plated on a kanamycin-resistant plate for screening. After overnight culture at 37℃, a single colony is picked for enzyme cutting experiment verification. The recombinant expression vector pET-30a / PLS with correct verification results is extracted from the Escherichia coli Trans5α by a plasmid kit, and is transformed into the protein expression competent cell Escherichia coli BL21 (DE3) to construct the PLS DNA polymerase gene expression engineering bacteria;
[0014] (2) Acquisition of the PLS DNA polymerase
[0015] The PLS DNA polymerase gene expression engineering bacteria are inoculated in the LB culture medium at an inoculation amount of 1:80-120 (v / v), and are cultured at 37℃ until the optical density OD 600 of the culture reaches 0.6-0.8. The final concentration of 0.5 mM IPTG is added for induction at 37℃ for 4-8 h. The culture is centrifuged at 5000-8000 r / min for 20-30 min to collect the bacterial cells. The bacterial cells are added in 50 mM phosphate buffer at a ratio of 1:8-15 (w / v), mixed uniformly, and then are ultrasonically broken for 20-30 min. The obtained bacterial broken liquid is heated in a water bath at 80-90℃ for 20-30 min, cooled to room temperature, and then is centrifuged at 7000-8000 r / min for 15-20 min to collect the supernatant, thereby obtaining the crude enzyme liquid.
[0016] The crude enzyme liquid is incubated with the nickel column material at 4℃ for 1-3 h. The impurity proteins are eluted with 10 mM imidazole, and the target proteins are eluted with 75 mM imidazole. The eluted target protein sample is centrifuged at 5000-7000 r / min for 15-25 min through a 15 kDa ultrafiltration tube, and the centrifugation is repeated for 2-4 times to remove the imidazole in the eluted target proteins. A cation exchange chromatography column is selected, and the eluted sample is linearly eluted with buffer B solution containing 1 M sodium chloride. The eluted sample is centrifuged at 5000-7000 r / min for 15-25 min through a 15 kDa ultrafiltration tube, and the centrifugation is repeated for 2-4 times, thereby obtaining the PLS DNA polymerase, which is stored in a storage buffer.
[0017] The preparation system of LB medium (w / v) is as follows: 1-2% of yeast powder, 1-2% of peptone, 0.5-2% of sodium chloride, and the pH is 7.0; the above reagents are dissolved in 1000 mL of deionized water to obtain the LB medium;
[0018] The preparation system of buffer B solution is as follows: 50-60 mM of potassium phosphate buffer (pH is 7.4), 1-1.5 M of sodium chloride, and 5-10% of glycerol (v / v); the above reagents are dissolved in 1000 mL of deionized water to obtain the buffer B solution;
[0019] The preparation system of the storage buffer is as follows: 5-15 mM of tris-hydroxymethyl aminomethane hydrochloride (pH is 7.4), 50-120 mM of sodium chloride, 1-1.5 mM of dithiothreitol, 0.1-0.5 mM of ethylenediaminetetraacetic acid, 0.1-0.5% of 4-(1,1,3,3-tetramethylbutyl) phenyl-polyethylene glycol, and 20-50% of glycerol (v / v); the above reagents are dissolved in 1000 mL of deionized water to obtain the storage buffer.
[0020] The protein purification method used in the application can obtain high-purity PLS DNA polymerase, and the optimization of PCR reaction conditions can greatly improve the catalytic efficiency of PLS DNA polymerase, reduce the generation of non-specific bands, improve the quality of the amplified target product, and be beneficial to downstream purification and subsequent sequencing analysis. Compared with 9 °N DNA polymerase, the thermal stability, fidelity, salt ion tolerance and continuous synthesis ability of PLS DNA polymerase are improved. Good thermal stability is a necessary condition for PLS DNA polymerase to be applied to PCR cycle reaction, so that PLS DNA polymerase can tolerate high temperature in the cycle reaction condition and maintain high catalytic efficiency during the reaction, which is beneficial to efficient amplification of the target DNA. Fidelity endows PLS DNA polymerase with good correction ability, which can remove the incorporated error bases in time to repair the mismatched fragments, thereby improving the accuracy of synthesized target DNA. The improvement of salt ion tolerance broadens the resistance of PLS DNA polymerase to complex environment, and has a broader application prospect. The improvement of continuous synthesis ability makes PLS DNA polymerase amplify DNA faster, and can continuously and stably amplify longer DNA fragments.
[0021] The PLS DNA polymerase can improve the fidelity of DNA information storage. The oligonucleotide pool obtained by encoding and storing digital information by Reed-Solomon (Reed-Solomon) error correction code is used as a template, and after optimization of the PCR reaction amplification system, the obtained sample is purified by a column type DNA purification kit to obtain a high-purity DNA amplification product, and high-throughput sequencing analysis is performed. In addition, the PLS DNA polymerase can amplify the DNA sequence of a specified file code from a mixed DNA sample containing two different file codes through the design of orthogonal barcodes, realize in vitro random addressing of multi-file digital information storage, completely restore the stored digital information content, and effectively reduce the number of errors generated in the DNA replication process, especially the base substitution error type which mainly depends on the catalytic performance of the enzyme, reduce the error probability between purine and pyrimidine, and improve the accuracy of DNA information storage encoding and decoding.
[0022] The PLS DNA polymerase can reduce the cost of DNA information storage. Current DNA-based data storage usually involves de novo chemical synthesis of DNA. However, due to the limitation of synthesis technology, it is currently impossible to synthesize DNA fragments of more than 200 nt from scratch. Therefore, the original data must be divided into many blocks for separate synthesis to overcome the length limitation of de novo chemical synthesis. In contrast, enzyme-based methods allow DNA synthesis to have higher fidelity and longer length. By combining the advantages of high-throughput chemical synthesis and high-synthesis length of enzymatic synthesis, short DNA blocks are first chemically synthesized, and then a one-step splicing of multiple short DNA fragments is realized by using a DNA polymerase-based cyclic assembly reaction to form longer DNA fragments, thereby reducing the synthesis cost.
[0023] The gene encoding the PLS DNA polymerase has a nucleotide sequence as shown in SEQ ID No. 1. The encoded PLS DNA polymerase has good purity, yield and catalytic activity after purification, can reduce the cost of DNA information storage, and improve the quality and efficiency of information storage data decoding.
[0024] If not otherwise specified, the solution described in the present application is deionized water solution. The PCR reaction buffer system involved in the present application is 10 x ThermoPol Reaction Buffer, which can provide salt ions and metal ions for DNA polymerase catalytic reaction, and is purchased from NEB Company in the United States. The dNTP mixture involved in the present application can provide substrate for DNA polymerase replication reaction, wherein the concentration of the four dNTPs is 4 μM, and is purchased from NEB Company in the United States. The oligonucleotide pool involved in the present application is a mixture of oligonucleotides (DNA) containing a plurality of different sequences obtained by Reed-Solomon error correction code encoding of a digital file, used for storing data, synthesized by electrochemical method on a 12K chip, and purchased from Nanjing Kingsriver Company. The restriction endonuclease, DNA loading buffer and protein loading buffer involved in the present application are used for cutting DNA and terminating reaction to prepare electrophoresis samples, respectively, and are purchased from Beijing Baorijing Biotechnology Co., Ltd. The ssm13mp18 DNA and λ DNA involved in the present application are used as DNA templates for PCR reaction, and are purchased from NEB Company in the United States. The E. coli Trans5α and E. coli BL21 (DE3) involved in the present application are used as competent host bacteria carrying transformed plasmids, and are purchased from Beijing Quanshijin Biotechnology Co., Ltd. The primers involved in the present application are purchased from Nanjing Kingsriver Company. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 It is denatured polyacrylamide gel electrophoresis diagram of the recombinant expression vector obtained in Example 1 after enzyme digestion; from left to right lane respectively represent: M: nucleic acid molecular weight marker; 1: untreated plasmid; 2: plasmid obtained after Nde I digestion; 3: plasmid obtained after Hind III digestion; 4: plasmid obtained after Nde I and Hind III digestion.
[0026] Figure 2 It is denatured polyacrylamide gel electrophoresis diagram of IPTG induction expression time optimization analysis in Example 2; from left to right lane respectively represent: M: protein molecular weight marker; 1: uninduced; 2: induced for 1 h; 3: induced for 2 h; 4: induced for 4 h; 5: induced for 6 h; 6: induced for 8 h.
[0027] Figure 3 It is denatured polyacrylamide gel electrophoresis analysis of PLS DNA polymerase obtained by purification by the method of Example 3; from left to right lane respectively represent: M: protein molecular weight marker; 1: uninduced; 2: after induction; 3: nickel column affinity chromatography; 4: cation exchange column chromatography.
[0028] Figure 4The agarose gel electrophoresis analysis of the amplification results under different PLS DNA polymerase activities obtained by the method in Example 4 is shown. The lanes from left to right represent: M: nucleic acid molecular weight standard; 1: 0.08U; 2: 0.40U; 3: 0.80U; 4: 1.60U.
[0029] Figure 5 Agarose gel electrophoresis images of the 9°N DNA polymerase and the DNA polymerase obtained by the method in Example 5, evaluating their thermostability; Figure A shows the 9°N DNA polymerase, and Figure B shows the PLS DNA polymerase. Lanes from left to right represent: M: nucleic acid molecular weight standard; M: 0 min; 1: 5 min; 2: 10 min; 3: 20 min; 4: 30 min; 5: 40 min; 6: 60 min; 7: 90 min; 8: 120 min.
[0030] Figure 6 Agarose gel electrophoresis images show the salt tolerance evaluation of 9°N DNA polymerase and the DNA polymerase obtained by the method in Example 6; Figure A shows 9°N DNA polymerase, and Figure B shows PLS DNA polymerase. M: nucleic acid molecular weight standard; lanes from left to right represent KCl concentrations, 1: 0 mM; 2: 10 mM; 3: 20 mM; 4: 30 mM; 5: 40 mM; 6: 50 mM; 7: 60 mM.
[0031] Figure 7 The images are agarose gel electrophoresis diagrams of PCR efficiency analysis based on 9°N DNA polymerase and PLS DNA polymerase (with 20 mM KCl added) obtained by the method in Example 8; Figure A is 72°C for 30 s, Figure B is 72°C for 60 s; M: nucleic acid molecular weight standard; lanes from left to right represent the expected length of synthesized DNA, 1: 1 kb; 2: 2 kb; 3: 4 kb; 4: 6 kb; 5: 8 kb.
[0032] Figure 8 This is an overview diagram of the in vitro DNA information storage addressing and in vivo DNA information storage assembly of DNA fragments involved in Examples 9, 10 and 11.
[0033] Figure 9 A bar graph showing the error types and quantities amplified by the two DNA polymerases obtained through high-throughput sequencing analysis using the method in Example 10. The horizontal axis represents the DNA polymerase group, including 9°N DNA polymerase and PLS DNA polymerase, and the vertical axis represents the percentage of base types and quantities obtained after high-throughput sequencing alignment.
[0034] Figure 10This is a bar chart analyzing the base substitutions amplified by the two DNA polymerases obtained through the method in Example 10. The horizontal axis represents the desired base type, the vertical axis represents the substituted base type, and the values represent the probability of each base substitution occurring.
[0035] Figure 11 This is an agarose gel electrophoresis image of multiple DNA fragments assembled at different PLS DNA polymerase concentrations obtained by the method in Example 11. Lanes from left to right represent the concentrations of PLS DNA polymerase: M: Marker; 1: 1.0 pM; 2: 2.0 pM; 3: 3.0 pM; 4: 4.0 pM; 5: 5.0 pM; 6: 6.0 pM; 7: 7.0 pM; 8: 8.0 pM.
[0036] Figure 12 This is a bar chart showing the frequency of the three error types on different type blocks after the sequencing data obtained by the method in Example 11 were compared. Detailed Implementation
[0037] The embodiments given below are further illustrations of the present invention to enable those skilled in the art to more fully understand the invention. However, the given embodiments should not be construed as limiting the scope of protection of the present invention. Therefore, non-essential improvements and adjustments made by those skilled in the art based on the above-described invention should also fall within the scope of protection of the present invention.
[0038] Example 1:
[0039] The construction of the engineered strain for PLS DNA polymerase was carried out as follows: First, the gene sequence encoding PLS DNA polymerase was optimized using an online codon optimization tool (https: / / www.genscript.com.cn / gensmart-free-gene-codon-optimization.html). The optimized nucleotide sequence is shown in SEQ ID No. 1. The nucleotide sequence shown in SEQ ID No. 1 was then constructed into the expression vector pET-30a (a common E. coli expression vector) through NdeⅠ and HindⅢ restriction sites, resulting in the recombinant expression vector pET-30a / PLS, thereby enabling efficient expression of PLS DNA polymerase in E. coli.
[0040] The start codon for the recombinant expression vector pET-30a / PLS is ATG, and the stop codon is TAA. First, a 6×His tag was added to the 5' end of the gene encoding PLS DNA polymerase for subsequent purification. The gene encoding PLS DNA polymerase with the 6×His tag was then ligated into the expression vector pET-30a via NdeI and HindIII restriction sites to construct the recombinant expression vector pET-30a / PLS. The recombinant expression vector was synthesized by Nanjing Genscript Biotech Co., Ltd. according to the above nucleotide sequence. This recombinant expression vector was transformed into competent *E. coli* Trans5α cells, then plated on kanamycin-resistant plates for selection. After overnight incubation at 37°C, single clones were picked and subjected to restriction enzyme digestion experiments for verification. The results are as follows: Figure 1 As shown, groups 2 and 3 are the results obtained by digestion with NdeI or HindIII, respectively. Both are single bands, and the band size is smaller than that of the recombinant expression vector pET-30a / PLS in group 1. This proves that the NdeI and HindIII restriction site sequences on the recombinant expression vector are correctly constructed and can recognize the corresponding restriction sites on the recombinant expression vector and cut the recombinant expression vector. Group 4 is the result obtained by digestion with both NdeI and HindIII. The two bands are the expression vector pET-30a cut by the two restriction endonucleases and the gene encoding PLS DNA polymerase with a 6×His tag, respectively. This again proves that the NdeI and HindIII restriction site sequences on the recombinant expression vector are correctly constructed. At the same time, Changchun Kumei Company was commissioned to perform sequencing analysis on the sequence between the NdeI and HindIII restriction sites in the recombinant expression vector pET-30a / PLS. The sequencing results are completely consistent with the sequence described in SEQ ID No. 1, proving the successful construction of the recombinant expression vector pET-30a / PLS. The recombinant expression vector pET-30a / PLS was extracted from Escherichia coli Trans5α using a plasmid kit and transformed into protein-competent Escherichia coli BL21(DE3) cells to construct an engineered strain for PLS DNA polymerase gene expression.
[0041] Example 2:
[0042] The optimization method for PLS DNA polymerase gene expression engineered bacteria is as follows: The PLS DNA polymerase gene expression engineered bacteria constructed in Example 1 were inoculated into LB medium at a ratio of 1:100 (v / v) and cultured at 37°C until the optical density OD of the culture reached the specified level. 600The concentration was reached to 0.65. Then, isopropyl-β-D-thiogalactoside (IPTG) was added to a final concentration of 0.5 mM, and the culture was induced at 37°C for 0 h, 1 h, 2 h, 4 h, 6 h, and 8 h, respectively. The bacterial culture was centrifuged at 5000 r / min for 20 min to collect the bacterial cells, and commercial protein loading buffer was added. The mixture was heated in a constant temperature metal bath at 100°C for 10 min. After cooling, the sample was loaded and subjected to denaturing polyacrylamide gel electrophoresis. The sample was then stained with Coomassie brilliant blue staining solution, and the bands were observed after destaining with destaining solution. The results are as follows: Figure 2 As shown, the protein expression level increases with increasing induction time, and tends to reach equilibrium at 6 hours, which is the optimal time for protein induction expression.
[0043] The LB medium preparation system (w / v) is as follows: 1.5% yeast extract, 1.5% peptone, 1% sodium chloride, pH 7.0. Dissolve the above reagents thoroughly in 1000 mL of deionized water to obtain the LB medium.
[0044] Example 3:
[0045] In this embodiment, the PLS DNA polymerase described in this invention was extracted and purified from the PLS DNA polymerase gene-expressing engineered bacteria using a two-step protein purification method involving nickel column affinity chromatography and ion exchange chromatography. The specific steps are as follows: The PLS DNA polymerase gene-expressing engineered bacteria were inoculated into the LB medium of Example 2 at an inoculum of 1:100 (v / v) and cultured at 37°C until the optical density OD of the culture reached the specified level. 600 The concentration reached 0.65; 0.5 mM IPTG was added and induced at 37℃ for 6 h; the culture was centrifuged at 5000 r / min for 20 min to collect bacterial cells, and 50 mM phosphate buffer (pH 8.0) was added to the bacterial cells at a ratio of 1:10 (w / v), mixed thoroughly, and then sonicated for 20 min; the obtained bacterial lysate was heated in a water bath at 85℃ for 30 min, cooled to room temperature, centrifuged at 8000 r / min for 15 min, and the supernatant was collected to obtain crude enzyme solution; protein purification was performed by nickel column affinity chromatography, the crude enzyme solution and nickel column material were co-incubated at 4℃ for 1 h, 10 mM imidazole was used to elute impurities, and 75 mM imidazole was used to elute the target protein; the eluted target protein sample was centrifuged at 6000 r / min for 20 min through a 15 kDa ultrafiltration tube, and the centrifugation was repeated three times to remove the imidazole eluted from the target protein; then cation exchange chromatography was performed using a buffer containing 1 M sodium chloride. Linear elution with solution B was performed, and the eluted sample was collected. The eluted sample was then centrifuged at 6000 r / min for 20 min through a 15 kDa ultrafiltration tube. This centrifugation was repeated 3 times. The obtained PLS DNA polymerase was stored in storage buffer.
[0046] The buffer B solution was prepared as follows: 50 mM potassium phosphate buffer (pH 7.4) containing 1 M sodium chloride and 5% glycerol (v / v). These reagents were dissolved thoroughly in 1000 mL of deionized water to obtain buffer B solution. The storage buffer was prepared as follows: 10 mM tris(hydroxymethyl)aminomethane hydrochloride (pH 7.4) containing 100 mM sodium chloride, 1 mM dimercaptothreitol, 0.1 mM ethylenediaminetetraacetic acid, 0.1% 4-(1,1,3,3-tetramethylbutyl)phenyl-polyethylene glycol, and 25% glycerol (v / v). These reagents were dissolved thoroughly in 1000 mL of deionized water to obtain the storage buffer. The purified PLS DNA polymerase was stored in this storage buffer, which enabled it to maintain its catalytic activity and stability for a long period.
[0047] Bacterial precipitates without IPTG induction, bacterial precipitates induced with 0.5 mM IPTG for 6 h, protein samples eluted with 75 mM imidazole in nickel column affinity chromatography, and protein samples collected by ion exchange chromatography were taken separately. Commercial protein loading buffer was added, and the samples were heated in a metal bath at 100°C for 10 min. After cooling, the samples were loaded and subjected to denaturing polyacrylamide gel electrophoresis. The samples were then stained with Coomassie brilliant blue staining solution, destained with destaining solution, and the bands were observed. Figure 3 As shown, after two consecutive steps of protein purification by nickel column affinity chromatography and ion exchange chromatography, the ion exchange chromatography group showed only a single band, and the size of this band matched that of the target protein PLS DNA polymerase, indicating that high-purity PLS DNA polymerase was successfully prepared.
[0048] Example 4:
[0049] The method for optimizing the enzyme concentration of the PLS DNA polymerase amplification system is as follows: First, prepare a PCR amplification reaction buffer system, including 1 μM single primer F1 (5'-GACACTCGTATGCAGTAGCC-3'), 300 ng template M1 (5'-ACAACCATTTATGTAGCATTTATGAAATTTTT AAATCAATTTACTATTGGCTACTGCATACGAGTGTC-3'), 2.5 mM dNTP, and 10×ThermoPol Reaction Buffer (NEB). Add the PLS DNA polymerase obtained in Example 3 to the amplification reaction buffer system at concentration gradients of 0.08 U, 0.40 U, 0.80 U, and 1.60 U. Make up the reaction system to 100 μL with deionized water. Distribute the system evenly into PCR tubes, 25 μL per tube, and perform the amplification reaction under the preset PCR cycle program. The PCR cycling program was as follows: 95℃ for 5 min; 95℃ for 30 sec, 55℃ for 30 sec, 72℃ for 30 sec, for 25 cycles; 72℃ for 10 min. After the amplification reaction was complete, commercial DNA loading buffer was added to stop the reaction, and the mixture was thoroughly mixed by pipetting. The sample was then loaded onto a 2% agarose gel for electrophoresis at 120V for 25 min. The results were then observed and analyzed using a gel imaging system. Figure 4 As shown, when the activity of PLS DNA polymerase is 0.8 U, the PCR amplification reaction can obtain good template M1 gene amplification, with a single band and no non-specific band contamination.
[0050] Example 5:
[0051] The thermostability evaluation method for PLS DNA polymerase is as follows: Based on the optimized optimal PCR reaction system, 9°N DNA polymerase and PLS DNA polymerase were added to 1×ThermoPol Reaction Buffer diluted with deionized water, respectively. Samples were taken after incubation at 95°C for 0, 5, 10, 20, 30, 40, 60, 90, and 120 min, respectively. The samples were then incubated at 72°C for another 2 min to equilibrate the temperature. The reaction system was then completed as described in Example 4, and PCR was performed to evaluate the thermostability of PLS DNA polymerase. Figure 5 As shown, after incubation at 95℃ for 120 min, the activity of PLS DNA polymerase was significantly higher than that of 9°N DNA polymerase, indicating that the addition of DNA-binding protein gave the fusion protein PLS DNA polymerase better thermal stability and effectively improved the heat tolerance of DNA polymerase when applied to PCR reaction.
[0052] Example 6:
[0053] The method for evaluating the salt tolerance of PLS DNA polymerase is as follows: Based on the optimized optimal PCR reaction system, a PCR amplification reaction buffer system was first prepared, including 1 μM primer F1 (5'-TCATGCATTGCCTGCTCTGCC-3'), 1 μM primer R1 (5'-GCAATGGCGATGACGCATCCTCACG-3'), 300 ng λ DNA template, 2.5 mM dNTPs, 10×ThermoPolReaction Buffer, and 0.8 U of PLS DNA polymerase or 9°N DNA polymerase. Then, 0 mM, 10 mM, 20 mM, 30 mM, 40 mM, 50 mM, and 60 mM KCl were added to each PCR tube, respectively. PCR reactions were performed according to the complete reaction system described in Example 4 to evaluate the salt tolerance of PLS DNA polymerase. The results are as follows: Figure 6 As shown, with the increase of KCl salt ion concentration, the enzyme activity of 9°N DNA polymerase was inhibited and the amplification band darkened. However, PLS DNA polymerase still had good amplification ability, indicating that the salt ion tolerance of PLS DNA polymerase is higher than that of 9°N DNA polymerase. This means that the addition of DNA binding protein significantly improves the salt tolerance of the fusion protein PLS DNA polymerase, which is expected to be suitable for a wider range of applications.
[0054] Example 7:
[0055] The evaluation method for the sustained synthesis capacity of PLS DNA polymerase is as follows: Fluorescently labeled primers M13-F (5'-FAM-GTTTTCCCAGTCACGACGTTGTAAAACGACGGCC-3') and ssm13mp18 DNA were pre-annealed, followed by the sequential addition of 300 μM dNTPs and DNA polymerase. Different concentrations of DNA polymerase (DNA polymerase:template after primer annealing = 1:100-1:2000) were added to different reaction tubes to explore the conditions for different DNA polymerases to reach a sustained synthesis state (generally, enzymes with high sustained synthesis capacity require a higher DNA polymerase:template ratio to achieve sustained synthesis). The amplification reaction was carried out at 37°C in the dark. At the end of the reaction, 50 mM EDTA was added to each reaction tube to terminate the reaction. The reaction products were sent to Sangon Biotech (Shanghai) Co., Ltd. for AFLP sequencing to obtain the length data of amplified products of different lengths and the corresponding fluorescence signal intensity of each amplified product. The average product amplification length of DNA polymerase can be calculated using the following formula:
[0056] log(nI / nT)=(n-1)logPI+log(1-PI)
[0057] The length of each signal peak above the background is denoted as nI, and the sum of the signal intensities of all products is nT. A plot of log(nI / nT) against n-1 is generated, where n is the number of bases on the primer terminal polymerase. Based on the above formula, the average primer amplification length is calculated as 1 / (1-PI), where PI represents the probability that polymerization does not stop at I. The results are shown in Table 1. The average primer amplification length of PLS DNA polymerase is higher than that of 9°N DNA polymerase, indicating that the fusion of DNA-binding proteins enhances the continuous synthesis capacity of PLS DNA polymerase.
[0058] Table 1: Evaluation of DNA polymerase's sustained synthesis capacity
[0059] Microscopic parameters of continuous synthesis (PI) Average primer extension length (nt) [1 / (1±PI)] 9°N 0.927±0.01 13.7±2.1 PLS 0.962±0.003 26.8±1.7
[0060] Example 8:
[0061] The PCR efficiency evaluation method for PLS DNA polymerase is as follows: DNA fragments of different lengths were amplified from a DNA template at the same time, using the reaction system described in Example 6. In Example 6, we found that the addition of the DNA-binding protein domain improved the salt tolerance of PLS DNA polymerase, and compared with the 9°N polymerase, the optimal salt concentration for PLS DNA polymerase to achieve its catalytic efficiency was 20 mM. Therefore, we performed PCR efficiency analysis under conditions of increased salt ions to evaluate the amplification capacity of PLS DNA polymerase. The results are as follows: Figure 7 As shown, in the absence of salt ions, the 9°N DNA polymerase can amplify a product length of 4kb, while under 20mM KCl concentration, the PLS DNA polymerase can amplify a product length of up to 8kb, and its catalytic efficiency increases with increasing extension time. These results indicate that the addition of DNA-binding proteins can improve the catalytic efficiency of PLS DNA polymerase in PCR applications, extending the same length of target DNA in a shorter time or amplifying longer DNA fragments within the same extension time.
[0062] Example 9:
[0063] The information storage and encoding process is as follows: The addressing application part of in vitro DNA information storage is as follows... Figure 8 As shown, the images of the "Jilin University logo in JPG format" and the "Jilin University School of Life Sciences logo in JPG format" were encoded using Reed-Solomon erasure coding, respectively, and converted into nucleotide sequences (the Reed-Solomon erasure coding is open-source code, and its code and specific usage methods can be downloaded from the link uploaded by the developer on Github, the download URL is https: / / github.com / reinhardh / dna_rs_coding), resulting in an oligonucleotide pool containing encoding DNA sequence information.
[0064] An oligonucleotide pool is a general customized product that can be used in DNA information storage research. By synchronously synthesizing thousands of oligonucleotide sequences on a microarray chip at one time, and then cutting the synthesized oligonucleotides and dissolving them in a single small tube, an oligonucleotide pool is obtained. After the DNA polymerase amplifies the oligonucleotide pool, high-throughput sequencing analysis is performed on the amplified sample. Taking the homologous arm sequence as the clustering feature, the sequenced data is subjected to clustering analysis. The nucleotide sequence is converted back to the binary sequence through the decoding algorithm in the Reed-Solomon erasure code to restore the stored data information.
[0065] The encoding scheme for converting the above binary sequence into a DNA sequence uses the method of converting two bits into one base. The preset corresponding relationship for the conversion is: 00→A, 01→C, 10→G, 11→T. For example, the binary number represented by the Chinese character "Ji" is 11100101 10010000 10001001, and the converted nucleotide sequence is TGCC GCAAGAGC. The "Jilin University School Emblem Picture" and the "Jilin University College of Life Sciences School Emblem Picture" are respectively encoded to obtain two oligonucleotide pools containing the above files (the capacity of the encoded oligonucleotide pool is related to the size of the file to be encoded). They are physically mixed at the same concentration to obtain a mixed oligonucleotide pool containing both files. For the DNA sequences encoded for the two files, barcodes are designed at both ends respectively. The barcode sequences of the two files have orthogonality and file uniqueness to avoid sequence collisions between different files affecting subsequent addressing. The orthogonal barcode serves as the binding site for the subsequent PCR amplification primer and is used to address and amplify a specific file. The oligonucleotide pool added with the orthogonal barcode and the amplification primer sequence are both entrusted to Nanjing GenScript Corporation for chemical synthesis for subsequent experiments. The orthogonal barcode sequence and the primer sequences for amplifying the two files respectively are as follows.
[0066] Table 2: Orthogonal barcode sequence and amplification primer sequence
[0067]
[0068] Example 10:
[0069] The PLS DNA polymerase amplifies the mixed oligonucleotide pool to randomly address and decode the school emblem picture. The process is as follows: Prepare the PCR amplification system as shown in Table 3. The oligonucleotide pool is replicated and amplified by the PLS DNA polymerase. The PCR cycling program is: 95°C for 5 min; 95°C for 30 sec, 55°C for 30 sec, 72°C for 30 sec, cycle 25 times; 72°C for 10 min.
[0070] Table 3: PCR amplification system
[0071] 10 x ThermoPol Reaction Buffer 10 μL dNTP mix 4 μM 2 μL School logo F 100 μM 2 μL School logo R 100 μM 2 μL Oligo pool (10 ng / μL) 5 μL PLS DNA polymerase 0.8U Deionized water q.s. to 100 μL
[0072] The amplified products obtained from PCR cyclic replication were purified using a DNA purification kit to remove impurities such as salt ions and proteins until the optical density OD was reduced. 260 / OD 280 With a ratio of 1.8 to 2.0 and a concentration ≥25 ng / μL, high-throughput sequencing analysis was performed.
[0073] Using barcode sequences as clustering features, the sequencing results were clustered. The processed data was then decoded using a Reed-Solomon erasure coding algorithm (decoding steps are described in the user manual provided by the code developer, as in Example 9), which enabled the recovery of the "Jilin University emblem image" stored in the DNA. Additionally, SNP calling analysis was performed on the sequencing data, comparing the sequenced DNA sequences with the coding DNA sequences to analyze the types and number of errors. The comparison results are as follows: Figure 9 As shown, the correct base ratio after amplification with 9°N DNA polymerase was 98.63%, while the erroneous base ratio was 1.37%. The proportions of the three common error types—insertion, deletion, and substitution—were 0.70%, 0.51%, and 0.22%, respectively. After amplification with PLS DNA polymerase, the correct base ratio was 98.79%, with the proportions of the three common error types—insertion, deletion, and substitution—being 0.60%, 0.44%, and 0.15%, respectively. Insertion and deletion errors can be further reduced through improvements and refinements to error correction algorithms, but base substitution errors are mainly introduced by the PCR process in DNA information storage. Therefore, high-fidelity DNA polymerases are a key strategy for reducing base substitution errors in DNA information storage. Compared to 9°N DNA polymerase, PLS DNA polymerase effectively reduces base substitution errors, which are primarily dependent on the fidelity of the DNA polymerase. This demonstrates the advantages of PLS DNA polymerase in amplifying oligonucleotide pools and reducing errors during replication, which is crucial for further improving the robustness of information storage technologies.
[0074] Using the coding DNA sequence as the expected sequence and the sequenced DNA sequence as the actual sequence, the two sequences were compared one by one to further analyze base substitution errors. The horizontal axis represents the expected base type, and the vertical axis represents the substitution base type; the values represent the probability of each base substitution occurring. The results are as follows: Figure 10 As shown, compared with 9°N DNA polymerase, PLS DNA polymerase can better reduce the number of errors of various types of bases, especially the probability of errors between purines and pyrimidines, thereby improving the accuracy of information storage.
[0075] Example 11:
[0076] The application of assembled DNA fragments in in vivo DNA information storage, such as... Figure 8 As shown. PLS DNA polymerase assembles DNA type blocks to obtain DNA type units, and further assembly of DNA type units yields longer DNA fragments. The file to be stored uses the ancient poem "Viewing the Waterfall at Mount Lu" for type storage. Each type unit represents a complete character, including an index type block, a check type block, and a content type block. The entire poem contains 44 Chinese characters, punctuation marks, and line breaks, ultimately resulting in 44 type units. The index type block indicates the position of the type within the entire file; the check type block is used for error correction, tolerating a small number of errors and ensuring normal decoding of the type; the content type block is the DNA sequence obtained after encoding the type. Based on the principle of DNA polymerase cyclic assembly reaction, the concentration gradient of PLS DNA polymerase was optimized. Using three types of type blocks—index type, check type, and content type—as primers and templates, PLS DNA polymerase concentrations of 1.0 pM, 2.0 pM, 3.0 pM, 4.0 pM, 5.0 pM, 6.0 pM, 7.0 pM, and 8.0 pM were added, and PCR reactions were performed according to the reaction conditions of Example 10. The results are as follows... Figure 11 As shown, after optimization according to the above enzyme concentration gradient, an 8.0 pM PLS DNA polymerase concentration can achieve one-step assembly of index, check, and address characters, obtaining complete character units. These units are then preserved on plasmids via homologous recombination and stored within bacteria, achieving in vivo character array-based DNA information storage. Sequencing of the assembled fragments is shown in Table 4. After sequencing alignment, the correct base ratio reached 99.9079%, while incorrect bases accounted for only approximately 0.092%, including approximately 0.0039% insertion errors, 0.0764% deletion errors, and approximately 0.0117% substitution errors. The positions of the errors corresponding to the corresponding character block positions are shown in the table below. Figure 12 This mostly occurs in the positions of index and check characters, and has little impact on content characters.
[0077] Table 4: Sequencing Results of Type Assembly
[0078]
[0079] This embodiment can connect 44 movable type blocks to assemble a ~50bp DNA short fragment into a ~7kb DNA long fragment. Since chemically synthesized DNA fragments are less than 200bp in length and the cost increases significantly with length, the technology used in this patent uses short DNA fragments as raw materials and performs enzymatic assembly using PLS DNA polymerase, eliminating the need for de novo chemical synthesis of long DNA fragments. Therefore, it can effectively reduce the cost of DNA information storage.
[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A PLS DNA polymerase constructed from a DNA polymerase-DNA binding protein, characterized in that: The PLS DNA polymerase consists of a 9°N DNA polymerase, a flexible linker peptide, and an Sso7d DNA-binding protein. The N-terminus of the flexible linker peptide is linked to the C-terminus of the 9°N DNA polymerase domain, and the C-terminus of the flexible linker peptide is linked to the N-terminus of the Sso7d DNA-binding protein domain. The nucleotide sequence encoding the PLS DNA polymerase gene is shown in SEQ ID No. 1, and the amino acid sequence of the flexible linker peptide is shown in SEQ ID No.
2.
2. The method for preparing PLS DNA polymerase constructed from DNA polymerase-DNA binding protein as described in claim 1, comprising the following steps: (1) Construction of PLS DNA polymerase engineered bacteria The gene sequence encoding PLS DNA polymerase was optimized using engineered bacteria codons. The optimized nucleotide sequence is shown in SEQ ID No. 1 and is used as the gene encoding PLS DNA polymerase. A 6×His tag was added to the 5' end of the gene encoding PLS DNA polymerase. The gene encoding PLS DNA polymerase with the 6×His tag was then ligated into the expression vector pET-30a via the Nde I and Hind III restriction sites to construct the recombinant expression vector pET-30a / PLS. The recombinant expression vector pET-30a / PLS was transformed into competent Escherichia coli Trans5α cells, then plated on kanamycin-resistant plates for selection. After overnight incubation at 37°C, single clones were picked for enzyme digestion experiments for verification. The recombinant expression vector pET-30a / PLS, which had been verified to be correct, was extracted from Escherichia coli Trans5α cells using a plasmid kit and transformed into protein-expressing competent Escherichia coli BL21(DE3) cells to construct the PLS DNA polymerase gene expression engineered bacteria. (2) Obtaining PLS DNA polymerase The PLS DNA polymerase gene expression engineered bacteria were inoculated into LB medium at an inoculum ratio of 1:80~120 (v / v) and cultured at 37°C until the optical density OD of the culture reached the specified level. 600 The concentration of the enzyme was increased to 0.6-0.
8. IPTG was added to a final concentration of 0.5 mM and induced at 37°C for 4-8 h. The culture was centrifuged at 5000-8000 r / min for 20-30 min to collect the bacterial cells. The bacterial cells were then added to 50 mM phosphate buffer at a ratio of 1:8-15 (w / v), mixed thoroughly, and sonicated for 20-30 min. The resulting bacterial lysate was heated in a water bath at 80-90°C for 20-30 min, cooled to room temperature, centrifuged at 7000-8000 r / min for 15-20 min, and the supernatant was collected to obtain the crude enzyme solution. The crude enzyme solution and nickel column stock were co-incubated at 4°C for 1-3 hours. Impurities were eluted with 10 mM imidazole, and the target protein was eluted with 75 mM imidazole. The eluted target protein sample was centrifuged at 5000-7000 rpm for 15-25 minutes using a 15 kDa ultrafiltration tube. This centrifugation was repeated 2-4 times to remove the imidazole eluted during the target protein elution. Then, a cation exchange chromatography column was selected, and linear elution was performed using buffer B solution containing 1 M sodium chloride. The eluted sample was collected and centrifuged at 5000-7000 rpm for 15-25 minutes using a 15 kDa ultrafiltration tube. This centrifugation was repeated 2-4 times to obtain PLS DNA polymerase, which was then stored in storage buffer.
3. The method for preparing PLS DNA polymerase constructed from DNA polymerase-DNA binding protein as described in claim 2, characterized in that: The LB medium preparation system (w / v) consists of 1%–2% yeast extract, 1%–2% peptone, and 0.5%–2% sodium chloride, pH 7.
0. Dissolve these reagents thoroughly in 1000 mL of deionized water to obtain the LB medium. The buffer B solution preparation system consists of 50–60 mM potassium phosphate buffer (pH 7.4), containing 1–1.5 M sodium chloride and 5%–10% glycerol (v / v). Dissolve these reagents thoroughly in 1000 mL of deionized water to obtain the buffer solution. Solution B; The preparation system of the storage buffer is as follows: 5-15 mM tris(hydroxymethyl)aminomethane hydrochloride (pH 7.4), 50-120 mM sodium chloride, 1-1.5 mM dimercaptothreitol, 0.1-0.5 mM ethylenediaminetetraacetic acid, 0.1%-0.5% 4-(1,1,3,3-tetramethylbutyl)phenyl-polyethylene glycol, and 20-50% glycerol (v / v). Dissolve the above reagents thoroughly in 1000 mL of deionized water to obtain the storage buffer.
4. The application of the PLS DNA polymerase constructed from a DNA polymerase-DNA binding protein as described in claim 1 in in vitro DNA information storage and in vivo DNA information storage.
Citation Information
Patent Citations
Thermostable reverse transcriptase
CN112442493A
Recombinant DNA polymerase and application thereof
CN113774039A
Recombinant 9-degree N DNA polymerase and application thereof in DNA information storage
CN117535265A