DNA polymerase and use thereof

By modifying specific sites of DNA polymerase, especially amino acid mutations in the exonuclease active sequence region and motif A, the problem of limited types of existing DNA polymerases and limited modification effects has been solved, and efficient polymerization of nucleotides with 3' blocking modifications has been achieved, which is suitable for high-throughput gene sequencing.

WO2025208292A1PCT designated stage Publication Date: 2025-10-09SHENZHEN HUADA GENE INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/085347
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

There are relatively few types of DNA polymerases available, and the modified DNA polymerases have limited effect on improving the polymerization ability of nucleotides with 3' blocking modifications, making it difficult to meet the needs of high-throughput gene sequencing.

Method used

Provided is a DNA polymerase having an amino acid sequence with substitutions, deletions or additions at specific sites shown in SEQ ID NO: 1, having DNA polymerase activity, and being capable of effectively polymerizing deoxyribonucleotides with 3'-O-reversible blocking modifications, including amino acid mutations in exoactive sequence regions such as D180, E182, N252, F256, D257, Y354 or D358, and modifications to the motif A sites of L452, Y453 or P454.

Benefits of technology

It achieves efficient polymerization of nucleotides with 3' blocking modifications, is suitable for high-throughput gene sequencing technology, and improves enzyme activity and selectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024085347_09102025_PF_FP_ABST
    Figure CN2024085347_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a DNA polymerase and a use thereof. The DNA polymerase comprises: (a) a protein containing an amino acid sequence as shown in SEQ ID NO: 1; (b) a protein having DNA polymerase activity, wherein substitution, deletion, and / or addition of one or more amino occurs at at least one site in an exonuclease active sequence region and / or motif A of the amino acid sequence as shown in in SEQ ID NO: 1; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% or more homology to the protein in (a) or (b) and having DNA polymerase activity. The present invention can solve the problem of few DNA polymerase types in the prior art and is applicable to the field of DNA polymerases.
Need to check novelty before this filing date? Find Prior Art

Description

DNA polymerase and its application Technical Field

[0001] The present invention relates to the field of DNA polymerases, and in particular to a DNA polymerase and applications thereof. Background Art

[0002] DNA polymerase uses a single DNA strand as a template and four deoxyribonucleotides as substrates, replicating and synthesizing a new DNA chain complementary to the template strand starting from the 5' end. DNA polymerase can add free nucleotides to the 3' end of the newly formed chain, thereby extending the new chain in the 5'-3' direction.

[0003] In next-generation sequencing (NGS) technology, based on the sequencing-by-synthesis (SBS) principle, the substrates used are nucleotides modified with reversible blocking groups at their 3' ends. During the sequencing process, DNA polymerase, using one DNA strand as a template, adds this modified nucleotide to the 3' end of the other DNA primer strand. The presence of the 3'-O-blocking group prevents the incorporation of subsequent nucleotides, thus halting the reaction at this step. An optical imaging system then detects the unique fluorescent group signal attached to each nucleotide to determine the base type of the incorporated nucleotide. The AT or CG pairing is then used to determine the base type of the nucleotide at the corresponding position on the DNA template strand, thereby completing DNA sequencing. After reading a nucleotide sequence, a chemical cleavage reaction removes the reversible blocking group from the 3' end of the DNA primer strand, restoring the natural free 3' hydroxyl group and enabling polymerization of the next blocking nucleotide. This cycle of "polymerization-optical imaging-chemical cleavage" completes the sequencing process for the DNA template strand.

[0004] Natural DNA polymerases are incapable of polymerizing nucleotides with 3'-blocking modifications. The substrate specificity of natural DNA polymerases is often modified to allow them to polymerize nucleotides with 3'-blocking modifications. These modified DNA polymerases can be used in gene sequencing technologies. Currently, DNA polymerases used in NGS are typically modified from archaeal B-family DNA polymerases, including KOD (Thermococcus kodakarensis), 9°N (Thermococcus sp. 9oN-7), Tgo (Thermococcus gorgonarius), DTok (Desulfurococcus sp. Tok), Pab (Pyrococcus abyssi), and Deep Vent (Deep Vent DNA Polymerase Pyrococcus). These enzymes share over 80% sequence identity, but modifying these highly similar DNA polymerases offers limited options and limited enzyme activity enhancement, making it difficult to meet demand.

[0005] Summary of the Invention

[0006] The main purpose of the present invention is to provide a DNA polymerase and its application to solve the problem of limited types of DNA polymerases in the prior art.

[0007] To achieve the above-mentioned object, according to a first aspect of the present invention, a DNA polymerase is provided, which comprises: (a) a protein comprising the amino acid sequence shown in SEQ ID NO: 1; (b) a protein having DNA polymerase activity in which one or more amino acids are substituted, deleted and / or added at least one site in the exoactive sequence region and / or motif A of the amino acid sequence shown in SEQ ID NO: 1; or (c) a protein having DNA polymerase activity that has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology with the protein in (a) or (b).

[0008] Furthermore, in (b), the sites in the exoactive sequence region include any one or more of the following sites: D180, E182, N252, F256, D257, Y354 or D358; the sites where motif A undergoes substitution, deletion and / or addition of one or several amino acids include any one or more of the following sites: L452, Y453 or P454.

[0009] Furthermore, the types of amino acids substituted at each site in the exoactive sequence region are independently selected from the following: the types of amino acids substituted at D180 include D180A; the types of amino acids substituted at E182 include E182A; the types of amino acids substituted at N252 include N252A; the types of amino acids substituted at F256 include F256A; the types of amino acids substituted at D257 include D257A; the types of amino acids substituted at Y354 include Y354A; the types of amino acids substituted at D358 include D358A; preferably, the types of amino acids substituted at each site in motif A are independently selected from the following: the types of amino acids substituted at L452 include L452A, ... L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S, L452T, L452V, L452Y, L452F or L452W; the type of amino acid substituted by Y453 includes Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; the type of amino acid substituted by P454 includes P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S or P454V.

[0010] Furthermore, the types of amino acids substituted at each site include any of the following combinations: D180A+E182A, D180A+E182A+L452A+Y453A+P454G, D180A+E182A+L452A+Y453A, D180A+E182A+L452A+Y453A+P454T, D180A+E182A+L452A+Y453A+P454V, D180A+E182A+L452A+Y453D+P454G, D180A+E182A+L452A+Y453G+P454A, D180A+E182A+L452A+Y453A+P454I, D180A+E182A+L452A+Y453A+P454T. 0A+E182A+L452A+Y453G+P454G, D180A+E182A+L452A+Y453G, D180A+E182A +L452A+Y453G+P454Q、D180A+E182A+L452A+Y453G+P454T、D180A+E182A+L 452A+Y453M+P454K, D180A+E182A+L452A+Y453S+P454G, D180A+E182A+L45 2A+Y453T+P454A, D180A+E182A+L452A+Y453V+P454A, D180A+E182A+L452G+ Y453A+P454G, D180A+E182A+L452G+Y453A, D180A+E182A+L452G+Y453G+P4 54T, D180A+E182A+L452G+Y453G+P454V, D180A+E182A+L452H+Y453A+P454 A. D180A+E182A+L452H+Y453G+P454G, D180A+E182A+L452H+Y453M+P454G, D180A+E182A+L452H+Y453V+P454G, D180A+E182A+L452I+Y453A+P454A, D18 0A+E182A+L452I+Y453A+P454T, D180A+E182A+L452I+Y453A+P454V, D180A +E182A+L452I+Y453G+P454A, D180A+E182A+L452I+Y453G, D180A+E182A+L 452I+Y453S, D180A+E182A+L452I+Y453T+P454A, D180A+E182A+L452I+Y45 3T+P454G, D180A+E182A+L452I+Y453T, D180A+E182A+L452K+Y453A+P454G,D180A+E182A+L452K+Y453G+P454A、D180A+E182A+L452L+Y453A+P454A、D180A+E182A+L452L+Y453A+P454G、D180A+E182A+L452L+Y453A+P454L、D180A+E182A+L452L+Y453A+P454M、D180A+E182A+L452L+Y453A、D180A+E182A+L452L+Y453A+P454T、D180A+E182A+L452L+Y453A+P454V、D180A+E182A+L452L+Y453G+P454A、D180A+E182A+L452L+Y453G+ P454G、D180A+E182A+L452L+Y453G+P454I、D180A+E182A+L452L+Y453G+P454L、D180A+E182A+L452L+Y453S+P454A、D180A+E182A+L452L+Y453T+P454G、D180A+E182A+L452L+Y453V+P454G、D180A+E182A+L452M+Y453A、D180A+E182A+L452M+Y453A+P454T、D180A+E182A+L452M+Y453A+P454V、D180A+E182A+L452M+Y453G+P454A、D180A+E182A+L452M+Y453G+P454G、D180A+E182A+L452M+Y453T+P454G、D180A+E182A+L452Q+Y453A+P454AD180A+E182A+L452S+Y453G+P454E、D180A+E182A+L452T+Y453A+P454A、D180A+E182A+L452T+Y453A、D180A+E182A+L452T+Y453A+P454V、D180A+E182A+L452T+Y453G+P454A、D180A+E182A+L452T+Y453G+P454T、D180A+E182A+L452T+Y453P+P454G、D180A+E182A+L452T+Y453S+P454T、D180A+E182A+L452T+Y453S+P454V、D180A+E182A+L452T+Y453T+P454S、D180A+E182A+L452T+Y453V+P454G、D180A+E182A+L452V+Y453A+P454A, D180A+E182A+L452V+Y453A+P454G, D180A+E182 A+L452V+Y453A+P454T, D180A+E182A+L452V+Y453A+P454V, D180A+E182A+L452V+Y45 3G+P454A, D180A+E182A+L452V+Y453G+P454G, D180A+E182A+L452V+Y453G, D180A+E 182A+L452V+Y453G+P454V, D180A+E182A+L452V+Y453H+P454A, D180A+E182A+L452V+ Y453S+P454A, D180A+E182A+L452V+Y453T+P454A, D180A+E182A+L452V+Y453T+P454 G. D180A+E182A+L452V+Y453T+P454S, D180A+E182A+L452V+Y453T+P454T, D180A+E18 2A+L452V+Y453T+P454V, D180A+E182A+L452V+Y453V+P454A, D180A+E182A+L452V+Y453V+P454G, D180A+E182A+L452V+Y453V+P454V or D180A+E182A+L452Y+Y453P+P454G. In this application, “+” means “and”.

[0011] Further, the DNA polymerase defined in (b) has polymerization activity for deoxyribonucleotides having a blocking modification; preferably, the blocking modification comprises a 3'-O-reversible blocking modification; preferably, the 3'-O-reversible blocking modification comprises a 3'-O-azidomethyl modification.

[0012] In order to achieve the above-mentioned purpose, according to the second aspect of the present invention, a DNA molecule is provided, which encodes any one of the above-mentioned DNA polymerases; preferably, the DNA molecule includes a DNA molecule having a nucleotide sequence shown in SEQ ID NO: 2 or DEQ ID NO: 3, or a DNA molecule having more than 70% homology with the nucleotide sequence shown in SEQ ID NO: 2 or DEQ ID NO: 3.

[0013] In order to achieve the above object, according to the third aspect of the present invention, a recombinant plasmid is provided, wherein the recombinant plasmid is connected to the above DNA molecule.

[0014] In order to achieve the above object, according to a fourth aspect of the present invention, a host cell is provided, wherein the host cell contains the above DNA molecule or the above recombinant plasmid.

[0015] Furthermore, the host cell includes a prokaryotic cell or a eukaryotic cell; preferably, the prokaryotic cell includes Escherichia coli.

[0016] In order to achieve the above-mentioned purpose, according to the fifth aspect of the present invention, a PCR kit is provided, which comprises any one of the above-mentioned DNA polymerases and optionally any one or more of the following components: 1) a buffer for providing a PCR amplification environment; 2) PCR primers; 3) a reagent for extracting target DNA.

[0017] In order to achieve the above objectives, according to the sixth aspect of the present invention, there is provided a use of any of the above DNA polymerases or any of the above PCR kits in incorporating modified deoxyribonucleotides into polynucleotides, PCR amplification, library construction or sequencing.

[0018] In order to achieve the above-mentioned object, according to the seventh aspect of the present invention, a method for incorporating modified deoxyribonucleotides is provided, wherein any one of the above-mentioned DNA polymerases is used to incorporate modified deoxyribonucleotides into a polynucleotide; the modification includes a 3'-O-blocking modification; and the 3'-O-blocking modification includes a 3'-O-azidomethyl modification.

[0019] In order to achieve the above object, according to an eighth aspect of the present invention, a PCR amplification method is provided, which comprises: performing PCR amplification using any of the above-mentioned DNA polymerases or the above-mentioned PCR kits.

[0020] Furthermore, the PCR amplification method comprises: using DNA polymerase to incorporate modified deoxyribonucleotides during the amplification process; preferably, the modification comprises a 3'-O-blocking modification; preferably, the 3'-O-blocking modification comprises a 3'-O-azidomethyl modification.

[0021] To achieve the above objectives, according to the ninth aspect of the present invention, a library construction kit is provided, which includes any one of the above-mentioned DNA polymerases and optionally any one or more of the following components: 1) an adapter sequence for library construction; 2) a buffer for providing a library construction environment; 3) enzymes for library construction, including one or more of a DNA shearing enzyme, an end-repair enzyme, or a ligase; and 4) modified deoxyribonucleotides.

[0022] In order to achieve the above object, according to the tenth aspect of the present invention, a library construction method is provided, which comprises: using any one of the above-mentioned DNA polymerases or the above-mentioned library construction kit to construct the library.

[0023] Furthermore, the library construction method includes: using DNA polymerase to perform PCR amplification on the fragments to be sequenced connected with sequencing adapters to obtain a DNA library.

[0024] To achieve the above objectives, according to the eleventh aspect of the present invention, a sequencing kit is provided, which comprises any one of the above-mentioned DNA polymerases and optionally any one or more of the following components: 1) a primer for complementary pairing with an adapter sequence; 2) a dideoxynucleotide; 3) dNTPs; 4) a nucleic acid probe; 5) an enzyme for sequencing, including a DNA ligase and / or an endonuclease; 6) a buffer for eluting the nucleic acid probe and / or dNTPs; and 7) a deoxyribonucleotide containing modifications.

[0025] In order to achieve the above object, according to the twelfth aspect of the present invention, a sequencing method is provided, which comprises: performing sequencing using any one of the above-mentioned DNA polymerases or the above-mentioned sequencing kits.

[0026] Furthermore, the sequencing method includes: using sequencing technology to perform sequencing-by-synthesis on the DNA library to obtain sequencing results; preferably, using DNA polymerase to incorporate modified deoxyribonucleotides to perform sequencing-by-synthesis.

[0027] By applying the technical solution of the present invention, a DNA polymerase sequence having a sequence shown in SEQ ID NO: 1 has low homology with existing DNA polymerases, and based on it, one or more amino acids can be substituted, deleted, and / or added to at least one site in the exoactive sequence region and / or motif A of the amino acid sequence, thereby providing a variety of DNA polymerases with high polymerization ability for nucleotides with 3' blocking modifications, which can be applied to gene sequencing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0029] FIG1 shows a schematic diagram of the comparison results of 35°S DNA polymerase and other DNA polymerases according to Example 1 of the present invention.

[0030] FIG2 shows a schematic diagram of the protein structure model prediction of 35°S DNA polymerase according to Example 1 of the present invention.

[0031] FIG3 shows a schematic diagram of the purification results of 35°S DNA polymerase protein according to Example 3 of the present invention.

[0032] FIG4 shows a schematic diagram of the polymerization activity determination principle according to Example 4 of the present invention.

[0033] FIG5 is a schematic diagram showing the results of the exo-activation test of 35°S DNA polymerase and KOD DNA polymerase according to Example 6 of the present invention. DETAILED DESCRIPTION

[0034] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.

[0035] As mentioned in the background art, natural DNA polymerases have no polymerization ability for nucleotides with 3' blocking modifications. It is usually necessary to modify the substrate specificity of natural DNA polymerases so that they have the ability to polymerize 3' blocking modified nucleotides. Currently, DNA polymerases used for NGS are usually modified from B family DNA polymerases of archaea, but the sequences of several commonly used modified DNA polymerases have a high similarity. Among them, the amino acid sequence of KOD DNA polymerase was BLAST-aligned with the amino acid sequences of 9°N DNA polymerase, Tgo DNA polymerase, DTok DNA polymerase, and PabDeep Vent DNA polymerase, and the results are shown in Table 1. Coverage refers to the percentage of the length of the sequence region involved in the alignment in the amino acid sequence / nucleotide sequence to the total length of the sequence; for example, in Table 1, the coverage of Tgo DNA polymerase is 92.75%, which means that the length of the sequence region involved in the homology alignment in the amino acid sequence of the enzyme accounts for 92.75% of its total sequence length.

[0036] Table 1

[0037] Modification of these enzymes with high sequence similarity has limited effect on improving enzymatic activity, has few options, and is difficult to meet demand. Therefore, it is urgent to find a new DNA polymerase with significant sequence differences from the above-mentioned DNA polymerases and with modification potential for modification. In this application, the inventors attempt to develop a DNA polymerase and its application, and thus propose a series of protection schemes in this application.

[0038] In a first typical embodiment of the present application, a DNA polymerase is provided, which DNA polymerase mutant includes: (a) a protein having the amino acid sequence shown in SEQ ID NO: 1; (b) a protein having a substitution, deletion and / or addition of one or more amino acids in at least one site in the exoactive sequence region and / or motif A of the amino acid sequence shown in SEQ ID NO: 1; and having DNA polymerase activity; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homology with the protein in (a) or (b) and having DNA polymerase activity.

[0039] The DNA polymerase shown in the above SEQ ID NO: 1 is a new type of B family DNA polymerase derived from deep-sea sediments (34.9 degrees south latitude, 179.0 degrees east longitude, 1611 meters seabed depth), named 35°S. The sequence of the 35°S DNA polymerase in this application is novel. When compared with the amino acid sequences of the DNA polymerases commonly used for modification mentioned in the background art (KOD DNA polymerase, 9°N DNA polymerase, Tgo DNA polymerase, DTok DNA polymerase, Pab DNA polymerase and Deep Vent DNA polymerase), the sequence similarity is 41.2%-42.9%, and the homology is relatively low. Therefore, the 35°S DNA polymerase in this application has a large room for modification. In addition, it also has good thermal stability, 5'-3' polymerization activity and 3'-5' exo-cleavage activity.

[0040] In a preferred embodiment, in (b), the sites in the exoactive sequence region include any one or more of the following sites: D180, E182, N252, F256, D257, Y354 or D358; the sites where motif A undergoes substitution, deletion and / or addition of one or more amino acids include any one or more of the following sites: L452, Y453 or P454.

[0041] D180, E182, N252, F256, D257, Y354 and D358 are sites in the exo-active sequence region of the 35 ° S DNA polymerase in the present application, relevant to its exo-active. When any one or more amino acid mutations in these sites relevant to the exo-active 35 ° S DNA polymerase, can inactivate its exo-active, prevent this DNA polymerase from being polymerized to the non-natural nucleotides on the template and excising. Therefore, when transforming DNA polymerase to it and can be used for polymerization of non-natural nucleotides, usually can the site mutation relevant to the exo-active dna polymerase. In the present application, the amino acid mutation combination in the L452+Y453+P454 site in motif A is conducive to the polymerization of non-natural nucleotides, and the combination of the amino acid mutation that occurs in the site in the above-mentioned exo-active sequence region is then conducive to preventing the non-natural nucleotides that are polymerized to the template chain from being cut by the DNA polymerase. Mutating the above-mentioned sites located in the exoactive sequence region can change the exoactivity of 35°S DNA polymerase and prevent enzymatic cleavage of non-natural nucleotides during polymerization. In this application specification, only D180A+E182A is used as an example to embody this type of mutation combination that changes the exoactivity of DNA polymerase, but in fact, this type of mutation is not limited to the mutation combination of this application.

[0042] Moreover, the DNA polymerase mutant obtained by simple mutation modification of the above-mentioned 35°S DNA polymerase has polymerization activity for 3'-O-reversibly blocked modified deoxyribonucleotides and can be applied to high-throughput sequencing technology based on 3'-O-reversibly blocked modified deoxyribonucleotides.

[0043] In a preferred embodiment, D180 and E182 are each independently substituted by a non-negatively charged amino acid; preferably, the non-negatively charged amino acid includes alanine, glycine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, tyrosine, asparagine, cysteine, glutamine, serine, threonine, arginine, histidine or lysine; preferably, the type of amino acid substituted at each site in the exoactive sequence region is independently selected from the following: the type of amino acid substituted by D180 includes D180A; the type of amino acid substituted by E182 includes E182A; the type of amino acid substituted by N252 includes N252A; the type of amino acid substituted by F256 includes F256A; the type of amino acid substituted by D257 includes D257A; the type of amino acid substituted by Y354 includes Y354A; the type of amino acid substituted by D358 includes D358A; preferably, the type of amino acid substituted at each site in motif A is independently selected from the following: L45 The type of amino acid substituted by L452A includes L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S, L452T, L452V, L452Y, L452F or L452W; preferably, the type of amino acid substituted by L452 includes L452A, L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S, L452T, L452V, L452Y, L452F or L452W; , L452V and L452Y; the type of amino acid substituted by Y453 includes Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; the type of amino acid substituted by P454 includes P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S or P454V.

[0044] Mutating sites D180 and E182 within the exo-active sequence region of 35°S DNA polymerase to non-negatively charged amino acids can inactivate the exo-active activity of 35°S DNA polymerase. Non-negatively charged amino acids include alanine, glycine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, tyrosine, asparagine, cysteine, glutamine, serine, threonine, arginine, histidine, or lysine. This application note uses D180A+E182A as an example to illustrate this type of mutation combination that alters the exo-active activity of DNA polymerase, but this type of mutation is not limited to the mutation combinations of this application.

[0045] In a preferred embodiment, in (b), the amino acid types substituted at each site include any one of the following combinations: D180A+E182A, D180A+E182A+L452A+Y453A+P454G, D180A+E182A+L452A+Y453A, D180A+E182A+L452A+Y453A+P454T, D180A+E182A+L452A+Y453A+P454V, D180A+ E182A+L452A+Y453D+P454G, D180A+E182A+L452A+Y453G+P454A, D180A+E182A+L452A+Y453A+P454I, D1 80A+E182A+L452A+Y453G+P454G, D180A+E182A+L452A+Y453G, D180A+E182A+L452A+Y453G+P454Q, D180A +E182A+L452A+Y453G+P454T, D180A+E182A+L452A+Y453M+P454K, D180A+E182A+L452A+Y453S +P454G, D180A+E182A+L452A+Y453T+P454A, D180A+E182A+L452A+Y453V+P454A, D180A+E182A+ L452G+Y453A+P454G, D180A+E182A+L452G+Y453A, D180A+E182A+L452G+Y453G+P454T, D180A+E 182A+L452G+Y453G+P454V, D180A+E182A+L452H+Y453A+P454A, D180A+E182A+L452H+Y453G+P4 54G, D180A+E182A+L452H+Y453M+P454G, D180A+E182A+L452H+Y453V+P454G, D180A+E182A+L45 2I+Y453A+P454A, D180A+E182A+L452I+Y453A+P454T, D180A+E182A+L452I+Y453A+P454V, D180 A+E182A+L452I+Y453G+P454A, D180A+E182A+L452I+Y453G, D180A+E182A+L452I+Y453S, D180A +E182A+L452I+Y453T+P454A, D180A+E182A+L452I+Y453T+P454G, D180A+E182A+L452I+Y453T,D180A+E182A+L452K+Y453A+P454G、D180A+E182A+L452K+Y453G+P454A、D180A+E182A+L452L+Y453A+P454A、D180A+E182A+L452L+Y453A+P454G、D180A+E182A+L452L+Y453A+P454L、D180A+E182A+L452L+Y453A+P454M、D180A+E182A+L452L+Y453A、D180A+E182A+L452L+Y453A+P454T、D180A+E182A+L452L+Y453A+P454V、D180A+E182A+L452L+Y453G+P454A、D180A+E182A+L452L+Y453G+P454G、D180A+E182A+L452L+Y453G+P454I、D180A+E182A+L452L+Y453G+P454L、D180A+E182A+L452L+Y453S+P454A、D180A+E182A+L452L+Y453T+P454G、D180A+E182A+L452L+Y453V+P454G、D180A+E182A+L452M+Y453A、D180A+E182A+L452M+Y453A+P454T、D180A+E182A+L452M+Y453A+P454V、D180A+E182A+L452M+Y453G+P454A、D180A+E182A+L452M+Y453G+P454G、D180A+E182A+L452M+Y453T+P454G、D180A+E182A+L452Q+Y453A+P454AD180A+E182A+L452S+Y453G+P454E、D180A+E182A+L452T+Y453A+P454A、D180A+E182A+L452T+Y453A、D180A+E182A+L452T+Y453A+P454V、D180A+E182A+L452T+Y453G+P454A、D180A+E182A+L452T+Y453G+P454T、D180A+E182A+L452T+Y453P+P454G、D180A+E182A+L452T+Y453S+P454T、D180A+E182A+L452T+Y453S+P454V、D180A+E182A+L452T+Y453T+P454S、D180A+E182A+L452T+Y453V+P454G、D180A+E182A+L452V+Y453A+P454A, D180A+E182A+L452V+Y453A+P454G, D180A+E182 A+L452V+Y453A+P454T, D180A+E182A+L452V+Y453A+P454V, D180A+E182A+L452V+Y45 3G+P454A, D180A+E182A+L452V+Y453G+P454G, D180A+E182A+L452V+Y453G, D180A+E 182A+L452V+Y453G+P454V, D180A+E182A+L452V+Y453H+P454A, D180A+E182A+L452V+ Y453S+P454A, D180A+E182A+L452V+Y453T+P454A, D180A+E182A+L452V+Y453T+P454 G. D180A+E182A+L452V+Y453T+P454S, D180A+E182A+L452V+Y453T+P454T, D180A+E18 2A+L452V+Y453T+P454V, D180A+E182A+L452V+Y453V+P454A, D180A+E182A+L452V+Y453V+P454G, D180A+E182A+L452V+Y453V+P454V or D180A+E182A+L452Y+Y453P+P454G. In this application, "+" means "and", for example: "D180A+E182A+L452A+Y453A+P454G" means that the amino acids at the five sites D180A, E182A, L452A, Y453A and P454G have all been mutated.

[0046] In a preferred embodiment, the DNA polymerase defined in (b) has polymerization activity on deoxyribonucleotides having a blocking modification; preferably, the blocking modification includes but is not limited to a 3'-O-reversible blocking modification; preferably, the 3'-O-reversible blocking modification includes but is not limited to a 3'-O-azidomethyl modification.

[0047] The DNA polymerase mutant exhibits polymerization activity towards 3'-O-reversibly blocked deoxyribonucleotides. This DNA polymerase mutant can also be used to amplify PCR reactions that are difficult for conventional DNA polymerases, thus enabling applications in fields such as high-throughput sequencing.

[0048] The above amino acid mutations were all experimentally investigated in the examples of this application, and all had DNA polymerase activity compared to the parent having the amino acid sequence shown in SEQ ID NO: 1. The above mutation sites are all mutations made around the amino acid active site, which can improve the activity of the protein. For mutations away from the active site, the effect on enhancing the activity of the protein is small, and thus proteins having 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% or more homology with the above amino acid sequence and having the same ability to improve enzyme activity can be obtained.

[0049] The term "identity" used herein refers to the "homology" between amino acid sequences, that is, the total ratio of identical amino acid residues in an amino acid sequence. The homology of amino acid sequences can be determined using alignment programs such as BLAST (Basic Local Alignment Search Tool) and FASTA.

[0050] The above proteins have 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (such as 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even 99.9% or more) homology with the protein provided by sequence (a) and have the same function, and their active sites, active pockets, active mechanisms, protein structures, etc. are most likely the same as those of the protein provided by sequence (a), and are homologous proteins obtained by amino acid mutations.

[0051] Sequences with the aforementioned homology can be obtained by amino acid substitution or replacement. Substitution or replacement rules generally apply, and amino acids with similar properties generally have similar effects when substituted with each other. For ease of description, the amino acid residue abbreviations are listed below: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine ​​(Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).

[0052] Amino acid substitutions or replacements, for example, in the above-mentioned homologous proteins, conservative amino acid substitutions may occur. "Conservative amino acid substitutions" include but are not limited to:

[0053] Hydrophobic amino acids (Ala, Cys, Gly, Pro, Met, Val, Ile, Leu) are replaced by other hydrophobic amino acids;

[0054] Substitution of bulky hydrophobic amino acids and aromatic amino acids (Phe, Tyr, Trp) with other bulky hydrophobic amino acids;

[0055] Amino acids with positively charged side chains (Arg, His, Lys) are replaced by other amino acids with positively charged side chains;

[0056] Amino acids with polar and uncharged side chains (Ser, Thr, Asn, Gln) are replaced by other amino acids with polar and uncharged side chains.

[0057] Those skilled in the art may also perform conservative substitutions on amino acids according to amino acid substitution rules well known to those skilled in the art, such as the "blosum62 scoring matrix" in the prior art.

[0058] In a second typical embodiment of the present application, a DNA molecule is provided, which encodes any one of the above-mentioned DNA polymerases; preferably, the DNA molecule includes a DNA molecule having a nucleotide sequence shown in SEQ ID NO: 2 or DEQ ID NO: 3, or a DNA molecule having more than 70% homology with the nucleotide sequence shown in SEQ ID NO: 2 or DEQ ID NO: 3, including 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (such as 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even 99.9% or more) homology and having the same function.

[0059] The above SEQ ID NO: 2 is the nucleotide sequence of 35°S DNA polymerase, and SEQ ID NO: 3 is the nucleotide sequence of a 35°S DNA polymerase mutant (D180A+E182A).

[0060] It should be noted that due to the principle of codon degeneracy, the nucleotide sequence of the translated amino acid sequence is not the only constant sequence. Any nucleotide sequence that can encode the amino acid sequence of the above-mentioned DNA polymerase is a nucleic acid sequence within the scope of this patent.

[0061] Furthermore, those skilled in the art can flexibly add nucleic acid sequences for expressing commonly used elements in the prior art near the above-mentioned DNA molecules, including but not limited to promoters for regulating transcription and translation levels, molecular tags for purifying proteins, signal peptides for localizing proteins, and other known sequences in the prior art.

[0062] In the third typical embodiment of the present application, a recombinant plasmid is provided, which is connected to the above-mentioned DNA molecule. The above-mentioned DNA can encode the above-mentioned DNA polymerase and can be connected to the recombinant plasmid to form circular DNA. The above-mentioned DNA and recombinant plasmid can both be transcribed and translated under the action of RNA polymerase, ribosomes, tRNA, etc. to obtain the above-mentioned DNA polymerase mutant. For different host species of DNA molecules or recombinant plasmids, the nucleotide sequence can be flexibly codon-optimized using existing technology to obtain a nucleotide sequence with higher transcription and translation efficiency.

[0063] In a fourth typical embodiment of the present application, a host cell is provided, wherein the host cell contains the above-mentioned DNA molecule or the above-mentioned recombinant plasmid.

[0064] The host cells described above can replicate recombinant plasmids within the host cells, and can also transcribe and translate DNA molecules carried on the recombinant plasmids, thereby obtaining a large amount of DNA polymerase. DNA polymerase can be obtained by disrupting the host cells using existing technologies, purifying proteins after disruption, or other methods. The host cells are not of plant origin.

[0065] In a preferred embodiment, the host cell comprises a prokaryotic cell or a eukaryotic cell; preferably, the prokaryotic cell comprises Escherichia coli; preferably, the host cell is not a plant cell.

[0066] In a fifth typical embodiment of the present application, a PCR kit is provided, which includes any of the above-mentioned DNA polymerases and optionally any one or more of the following components: 1) a buffer for providing a PCR amplification environment; 2) PCR primers; 3) a reagent for extracting target DNA.

[0067] The above-mentioned PCR primers include but are not limited to universal primers for amplifying specific target DNA, including but not limited to 16S universal amplification primers for prokaryotic bacteria, 18S universal amplification primers for eukaryotic bacteria, ITS universal amplification primers for fungi, or other PCR primers designed for specific organisms or specific target fragments.

[0068] The aforementioned buffer is a common PCR buffer in the prior art and is used to provide reaction conditions suitable for the aforementioned DNA polymerase mutant. The DNA polymerase mutant of the present application is similar to wild-type Taq DNA polymerase and can perform PCR reactions in PCR buffers in the prior art. Those skilled in the art can also flexibly optimize the components of the aforementioned PCR buffer to obtain a buffer more suitable for such DNA polymerase.

[0069] The above reagents for extracting target DNA can extract DNA from the target sample and provide an amplification template for the subsequent PCR reaction.

[0070] In a sixth typical embodiment of the present application, a method is provided for using any of the above-mentioned DNA polymerases or any of the above-mentioned PCR kits in PCR amplification, library construction or sequencing.

[0071] The aforementioned DNA polymerase mutants can be used to achieve PCR amplification, library construction, or high-throughput sequencing of target samples or sequences. In particular, the aforementioned DNA polymerase mutants that possess polymerization activity toward 3'-O-reversibly blocked deoxyribonucleotides can be used to polymerize DNA using nucleotides containing blocking modifications during PCR amplification or library construction, thereby labeling the target sequence. Alternatively, during sequencing, the aforementioned DNA polymerase mutants can be used to polymerize DNA containing blocking modified nucleotides, thereby completing DNA sequencing. Alternatively, these mutants can be used in the synthesis of other non-natural nucleotide (XNA, xeno-nucleic acid) polymers, such as the biosynthesis of non-natural nucleotide aptamers.

[0072] In a seventh typical embodiment of the present application, a method for incorporating modified deoxyribonucleotides is provided, the method comprising incorporating modified deoxyribonucleotides into a polynucleotide using any of the above-mentioned DNA polymerases; preferably, the modification comprises a 3'-O-blocking modification; preferably, the 3'-O-blocking modification comprises a 3'-O-azidomethyl modification.

[0073] Preferably, the DNA polymerase incorporating modified deoxyribonucleotides comprises a DNA polymerase having an amino acid sequence as shown in SEQ NO ID: 1, and a combination of amino acid substitutions occurring at L452, Y453 or P454 based on the D180A+E182A mutations; preferably, the types of amino acids substituted at L452, Y453 or P454 are independently selected from the following: the types of amino acids substituted at L452 include L452A, L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S, L452T, L452V or L452Y; the types of amino acids substituted at Y453 include Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; the types of amino acids substituted with P454 include P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S or P454V.

[0074] The method comprises: using any of the above DNA polymerases to add a modified nucleotide, such as a 3'-O-blocking group, to the 3' end of a polynucleotide to obtain a deoxyribonucleotide containing the modification.

[0075] In an eighth typical embodiment of the present application, a PCR amplification method is provided, which comprises: performing PCR amplification using any of the above-mentioned DNA polymerases or any of the above-mentioned PCR kits.

[0076] In a preferred embodiment, the PCR amplification method comprises: using a DNA polymerase to incorporate a modified deoxyribonucleotide during the amplification process; preferably, the modification comprises a 3'-O-blocking modification; preferably, the 3'-O-blocking modification comprises a 3'-O-azidomethyl modification. Preferably, the DNA polymerase incorporating the modified deoxyribonucleotide comprises a DNA polymerase having an amino acid sequence as shown in SEQ NOID: 1, and a combination of amino acid substitutions occurring at L452, Y453 or P454 based on the D180A+E182A mutation; preferably, the amino acid types substituted at L452, Y453 or P454 are independently selected from the following: the types of amino acids substituted at L452 include L452A, L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S ... 52T, L452V or L452Y; the type of amino acid substituted by Y453 includes Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; the type of amino acid substituted by P454 includes P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S or P454V.

[0077] The method comprises: using any of the above DNA polymerases to add a modified nucleotide, such as a 3'-O-blocking group, to the 3' end of a polynucleotide during PCR amplification to obtain a deoxyribonucleotide containing the modification.

[0078] In a ninth typical embodiment of the present application, a library construction kit is provided, wherein the library construction kit comprises any one of the above-mentioned DNA polymerases.

[0079] In a preferred embodiment, the library construction kit further includes any one or more of the following components: 1) an adapter sequence for library construction; 2) a buffer for providing a library construction environment; 3) enzymes for library construction, including one or more of DNA shearing enzymes, end repair enzymes, or ligases; 4) modified deoxyribonucleotides.

[0080] Preferably, the above modifications include but are not limited to 3'-O-blocking modifications; preferably, the 3'-O-blocking modifications include but are not limited to 3'-O-azidomethyl modifications.

[0081] In a tenth typical embodiment of the present application, a library construction method is provided, which comprises: constructing a library using any of the above-mentioned DNA polymerases or any of the above-mentioned library construction kits.

[0082] In a preferred embodiment, the library construction method comprises: using a DNA polymerase to perform PCR amplification on the fragments to be sequenced, which are connected to sequencing adapters, to obtain a DNA library; preferably, the library construction method further comprises: using the DNA polymerase to incorporate modified deoxyribonucleotides during the DNA amplification process. Preferably, the above-mentioned modifications include, but are not limited to, 3'-O-blocking modifications; preferably, 3'-O-blocking modifications include, but are not limited to, 3'-O-azidomethyl modifications.

[0083] Preferably, the DNA polymerase incorporating modified deoxyribonucleotides comprises a DNA polymerase having an amino acid sequence as shown in SEQ NO ID: 1, based on the amino acid substitutions D180A+E182A, and a combination of amino acid substitutions at L452, Y453 or P454; preferably, the amino acid substitutions at L452, Y453 or P454 are independently selected from the following: the amino acid substitutions at L452 include L452A, L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S, L452T, L452V or L452Y; the type of amino acid substituted by Y453 includes Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; the type of amino acid substituted by P454 includes P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S or P454V.

[0084] Using the above-mentioned library construction kit and / or library construction method, the above-mentioned DNA polymerase can be used to PCR amplify and enrich the fragments to be sequenced connected with sequencing adapters (including but not limited to full-length adapters or truncated adapters), providing sufficient samples for subsequent sequencing.

[0085] In an eleventh typical embodiment of the present application, a sequencing kit is provided, comprising any one of the above-mentioned DNA polymerases and optionally any one or more of the following components: 1) a primer for complementary pairing with an adapter sequence; 2) a dideoxynucleotide; 3) dNTPs; 4) a nucleic acid probe; 5) an enzyme for sequencing, including a DNA ligase and / or an endonuclease; 6) a buffer for eluting the nucleic acid probe and / or dNTPs; and 7) a solution containing modified deoxyribonucleotides.

[0086] Preferably, the above modifications include but are not limited to 3'-O-blocking modifications; preferably, the 3'-O-blocking modifications include but are not limited to 3'-O-azidomethyl modifications.

[0087] In a twelfth typical embodiment of the present application, a sequencing method is provided, which comprises: performing sequencing using any one of the above-mentioned DNA polymerases or any one of the above-mentioned sequencing kits.

[0088] In a preferred embodiment, the sequencing method comprises: performing sequencing-by-synthesis on a DNA library using a sequencing technology to obtain sequencing results; preferably, performing sequencing-by-synthesis using a DNA polymerase to incorporate modified deoxyribonucleotides.

[0089] Preferably, the above modifications include but are not limited to 3'-O-blocking modifications; preferably, the 3'-O-blocking modifications include but are not limited to 3'-O-azidomethyl modifications.

[0090] In the above sequencing method, a DNA polymerase is used to perform PCR amplification on the sample to be sequenced to obtain a DNA library; the DNA library is sequenced using a sequencing technology to obtain a sequencing result; preferably, the sequencing method further comprises: using the DNA polymerase to incorporate a modified deoxyribonucleotide during the PCR amplification process. Preferably, the above DNA polymerase incorporating the modified deoxyribonucleotide comprises a DNA polymerase having an amino acid sequence as shown in SEQ NO ID: 1, and a combination of amino acid substitutions occurring at L452, Y453 or P454 based on the D180A+E182A mutation; preferably, the types of amino acids substituted at L452, Y453 or P454 are independently selected from the following: the types of amino acids substituted at L452 include L452A, L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S ... 2T, L452V or L452Y; the type of amino acid substituted by Y453 includes Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; the type of amino acid substituted by P454 includes P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S or P454V.

[0091] Utilizing the above-mentioned sequencing kit and / or sequencing method, sequencing of the sample to be sequenced can be achieved by methods such as Sanger sequencing, second-generation sequencing or third-generation sequencing in the prior art. In sequencing, using the above-mentioned DNA polymerase of the present application, common PCR reactions in existing sequencing technologies such as emulsion PCR, bridge PCR, and rolling circle amplification can be achieved. The above-mentioned emulsion PCR can refer to the emulsion PCR operation in the Soild sequencing method of ABI, or the emulsion PCR operation in the Roche 454 sequencing method. The above-mentioned bridge PCR can refer to the bridge PCR operation in the Solexa sequencing method of illumina. The above-mentioned rolling circle amplification technology can refer to the method for preparing DNA nanoballs (DNBs) by amplifying the rolling circle amplification technology (RCA) in the MGI sequencing method of BGI.

[0092] The aforementioned DNA polymerases can be used to achieve PCR amplification, library construction, or high-throughput sequencing of target samples or sequences. In particular, using the aforementioned DNA polymerases (DNA polymerase mutants) that have polymerization activity toward modified single bases (including but not limited to 3'-O-blocking modified deoxyribonucleotides), nucleotides containing blocking modifications can be used to polymerize DNA during PCR amplification or library construction, thereby labeling the target sequence. Alternatively, during sequencing, the aforementioned DNA polymerase mutants can be used to amplify DNA containing blocking modified nucleotides, thereby completing DNA sequencing.

[0093] The following will further explain the beneficial effects of the present application in detail with reference to specific examples. In the examples of the present application, the following exploration is conducted:

[0094] (1). The sequence of this novel B family DNA polymerase 35°S DNA polymerase was derived from the metagenomic sequencing data analysis of deep-sea sediment samples.

[0095] (2) After gene synthesis, vector construction, heterologous expression and protein purification, the polymerization activity, exolytic activity and thermal stability of the sequence were tested.

[0096] (3). Based on the novel B family DNA polymerase 35°S, its amino acid sequence was analyzed and predicted. First, two conserved amino acids in the exo-active sequence region were mutated, D180A+E182A.

[0097] (4). Subsequently, the three sites L452, Y453, and P454 of the polymerase motif A were subjected to three-site combined mutations and the mutant protein was induced to express and purified.

[0098] (5) The polymerization activity of 35°S polymerase mutant was tested on a microplate reader using 3'-O-reversibly blocked deoxyribonucleotides.

[0099] Example 1

[0100] 1. Sequence alignment

[0101] Using the Clustal Omega online sequence alignment website, 35°S DNA polymerase was aligned with several representative B-family DNA polymerases, including KOD DNA polymerase, 9°N DNA polymerase, Tgo DNA polymerase, DTok DNA polymerase, Pab DNA polymerase, and Deep Vent DNA polymerase (sequences shown in SEQ ID NOs: 4-9, respectively). The alignment results are shown in Figure 1, where "cov" refers to coverage and "pid" refers to percent identity. The sequence similarity between 35°S DNA polymerase and these enzymes ranged from 40.5% to 42.9%. The specific alignment results are shown in Table 2.

[0102] Table 2

[0103] 2. Structural model prediction

[0104] The protein structure model of the 35°S DNA polymerase sequence was predicted based on AlphaFold, and the predicted structure is shown in Figure 2.

[0105] Example 2 Construction of 35°S DNA polymerase mutants

[0106] The expression plasmid pET-28a(+) / 35°S containing the 35°S DNA polymerase gene sequence (SEQ ID NO: 2) and the expression plasmid pET-28a(+) / 35°SM containing the 35°S DNA polymerase mutant (D180A+E182A) gene sequence (SEQ ID NO: 3) were commissioned to Changzhou Xinyisheng Life Science Technology Co., Ltd. for synthesis and construction. The pET-28a(+) / 35°SM recombinant plasmid was transformed into Escherichia coli BL21 (DE3) competent cells (Cat. No. CB105-02) purchased from Tiangen Biochemical Technology (Beijing) Co., Ltd. for subsequent expression and purification. The transformation steps are as follows: add 0.5 μL of 50 ng / μL pET-28a(+) / 35°SM plasmid to 100 μL of BL21(DE3) competent cells, gently flick to mix, and place on ice for 5 minutes; heat shock at 42°C for 1 minute and place on ice again for 5 minutes; add 400 μL of resistance-free LB medium and shake at 220 rpm at 37°C for 1 hour; spread 100 μL of the culture on a kanamycin-resistant LB plate and incubate at 37°C for 12-16 hours.

[0107] Example 3 Expression and purification of 35°S DNA polymerase protein

[0108] The affinity chromatography column and anion exchange column used for protein purification were HisTrap FF 5 mL (brand: Cytiva, product number: 17528601) and HiTrap Q FF 5 mL (brand: Cytiva, product number: 17515601), respectively. The specific purification steps were as follows:

[0109] (1) Pick 3-6 well-growing single colonies from the plate and inoculate them into a 50 / 250mL LB liquid conical flask. Incubate at 37℃ for 5-7h until the OD600 reaches 0.6-4.0. Then inoculate the above bacterial solution into 2L / 5L LB medium at a 1% inoculum volume and incubate at 37℃ for 2-4h until the OD600 reaches 0.8-1.0. Pre-cool the original shaker to 16℃. Add IPTG to the culture medium to a final concentration of 0.5mM. Induce expression for 12-16h in a shaker at 16℃, 220rpm.

[0110] (2) The cells were collected by centrifugation at 8000 g for 30 min, and then resuspended in Ni column-A buffer at a ratio of 1:20. The cells were then broken by ultrasound in an ice bath.

[0111] (3) Heating treatment: Preheat the water bath to 75-80℃, place the broken bacteria into the water bath, shake and mix, use a clean thermometer to check the internal temperature of the bacteria solution. When it reaches 75℃, start timing for 30 minutes. During this period, shake and mix 3 times every 10 minutes to ensure uniform heating.

[0112] (4) The heat-treated ultrasonic disrupted liquid was centrifuged at 12000 rpm for 60 min at 4°C, and the supernatant was filtered with a 0.22 μM filter membrane as the sample for the purification column.

[0113] (5) Load the sample onto the pretreated HisTrap FF column at a rate of 3 mL / min. After loading, rinse with Ni-column-A Buffer for 20 column volumes. Then, perform a linear elution with Ni-column-B Buffer (0-70%) for 10.5 CV. Collect the eluted protein when the UV absorption peak reaches 100 mAu and stop collecting when the UV absorption peak drops to 200 mAu.

[0114] (6) The collected eluate was diluted 6-fold with 35°S diluent and loaded onto a pre-treated Q column (HiTrap Q FF 5 ml). After loading, the column was rinsed with Q column-A buffer for 30 CV until the baseline was stable at a flow rate of 5 mL / min.

[0115] (7) Use Q column B buffer to perform gradient elution (0-100% Q column-B buffer, 10CV) of the target protein at a flow rate of 5mL / min. The collected samples were analyzed by SDS-PAGE to determine the protein purity.

[0116] (8) After the purified protein sample is dialyzed and its concentration is determined, it is stored in storage solution for subsequent functional activity analysis.

[0117] The specific components of the buffer used in the purification process are as follows:

[0118] Ni column-A Buffer: 20 mM Tris-HCl, 300 mM NaCl, 20 mM imidazole, 5% glycerol, pH 8.5.

[0119] Ni column-B Buffer: 20 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole, 5% glycerol, pH 8.5.

[0120] Q column-A Buffer: 20 mM Tris-HCl, 100 mM NaCl, 5% glycerol, pH 8.5.

[0121] Q column-B Buffer: 20 mM Tris-HCl, 500 mM NaCl, 5% glycerol, pH 8.5.

[0122] Diluent: 20 mM Tris-HCl, 5% glycerol, pH 8.5.

[0123] 2× Dialysate: 20 mM Tris-HCl, 200 mM KCl, 2 mM DTT, 0.2 mM EDTA, 5% glycerol, pH 8.0 @ 25°C.

[0124] Storage solution: 10 mM Tris-HCl, 100 mM KCl, 1 mM DTT, 0.1 mM EDTA, 50% glycerol, pH 8.0 @ 25°C.

[0125] The purification results of 35°S DNA polymerase protein are shown in FIG3 .

[0126] Example 4 Determination of the thermal stability of 35°S DNA polymerase

[0127] Protein Thermal Shift purchased from Thermo Fisher Scientific TMThe protein stability of 35°S DNA polymerase was determined using a dye kit (Cat. No. 4461146). KOD DNA polymerase (SEQ ID NO: 4) and Pfu DNA polymerase were used as controls. The test reaction system consisted of 5 μL Protein Thermal Shift Buffer, 2 μL test protein, 2.5 μL 8× Protein Thermal Shift Dye, and 10.5 μL nuclease-free water. The reaction system was mixed and placed in a StepOne TM A temperature experiment of 25-99°C was performed on a real-time fluorescence quantitative PCR system (Applied Biosystems™) to monitor the changes in ROX fluorescence signal. The results are shown in Table 3.

[0128] Table 3

[0129] The 35°S DNA polymerase in the present application has a Tm value of over 85°C and good thermal stability.

[0130] Example 5 35°S DNA polymerase polymerization activity assay

[0131] Polymerization activity was measured using a primed M13 ssDNA substrate. The schematic diagram of the process is shown in Figure 4. In the presence of polymerization activity, the primers on the primed M13 ssDNA extend along the ssDNA in a 5' to 3' direction to generate dsDNA. The generated dsDNA can be quantified using the Qubit™ dsDNA HS Quantitation Kit (Thermo Fisher Scientific, Cat. No. Q32854).

[0132] The specific reaction system and components are shown in Table 4.

[0133] Table 4

[0134] The reaction mixture was incubated at 72°C for 5 minutes, then terminated with 2 μL of 0.5 M EDTA. The dsDNA concentration was determined using the Qubit™ dsDNA HS Quantitation Kit. A negative control was performed using an equal volume of enzyme stock solution instead of 35°S DNA polymerase. The ΔQubit value was calculated by subtracting the Qubit value of the negative control from the Qubit value of the experimental group. The results, shown in Table 5, indicate that 35°S DNA polymerase exhibited polymerization activity at 72°C.

[0135] Table 5

[0136] Example 6 35°S DNA polymerase exo-activity assay

[0137] The exo-cleavage activity of 35°S DNA polymerase was qualitatively tested using a terminal mismatch fluorescent probe. The probe sequence was:

[0138] SEQ ID NO: 10: ATCAGCAGGCCACACGTTAAACTGT-BHQ2;

[0139] SEQ ID NO: 11: FAM-TGTCTTTAACGTGTGGCCTCTGAT;

[0140] The two were mixed in equimolar amounts (final concentration 10 μM) and annealed to serve as fluorescent probe substrates for exo-activity assays. The reaction system (total volume 25 μL) used in the exo-activity assay was as follows: 2.5 μL Reaction Buffer (NEB, Catalog No. B9004S), 0.25 μL of 10 μM fluorescent probe substrate, 2 μL of enzyme solution (enzyme concentration 0.18 mg / mL), and 20.25 μL of nuclease-free water were used. The positive control group consisted of KOD DNA polymerase. For the negative control group, the enzyme solution components were replaced with an equal volume of enzyme stock solution. The reaction system was set up on ice. After preparation, the plate was transferred to a 384-well plate and placed in a microplate reader (BioTek Synergy H1, Agilent) for fluorescence signal detection. The excitation and emission lights were set at 492 nm and 518 nm, respectively, at a reaction temperature of 37°C. Fluorescence signals were collected every 30 seconds for a total reaction duration of 1 hour. The results of the 35°S DNA polymerase exolytic activity test are shown in Figure 5 (fluorescence values ​​are subtracted from the corresponding blank control values). The exolytic activity of 35°S DNA polymerase is comparable to that of KOD DNA polymerase.

[0141] Example 7 Construction of 35°S DNA polymerase mutant vector

[0142] Using the 35°S DNA polymerase mutant expression plasmid pET-28a(+) / 35°SM as a template, different mutations were simultaneously introduced into the L452, Y453, and P454 sites of 35°SM (SEQ ID NO: 12) by rapid PCR amplification.

[0143] The rapid PCR reaction system consisted of 5 μL of Pfu DNA polymerase 10× buffer with MgSO₄, 1 μL of dNTP Mix (10 mM each), 1.5 μL of forward primer (10 μM), 1.5 μL of reverse primer (10 μM), 50 ng of template DNA, 0.5 μL of Pfu DNA polymerase (3 U / μL, Promega, Cat. No. M7741), and nuclease-free water to a total volume of 50 μL. The rapid PCR reaction program was as follows: initial denaturation at 95°C for 2 min; 16 cycles of 95°C for 30 s, 58°C for 30 s, and 72°C for 8 min; a final extension at 72°C for 5 min, and storage at 4°C.

[0144] Add 1 μL of DpnI (20 U / μL, NEB, Catalog No. R0176) to the PCR product and digest the template in a 37°C metal bath for 2 hours. The reaction product was then transformed into E. coli DH5α competent cells (Catalog No. CB101-02) purchased from Tiangen. The transformation steps were as follows: add 10 μL of the reaction product to 100 μL of DH5α competent cells, gently flick to mix, and incubate on ice for 30 minutes. Heat shock at 42°C for 1 minute, then incubate on ice again for 10 minutes. Add 400 μL of antibiotic-free LB medium and shake at 37°C at 220 rpm / min for 1 hour. Centrifuge the culture at 3000 g for 3 minutes, discard a portion of the supernatant, and retain approximately 100 μL. Resuspend and mix thoroughly, then plate on a kanamycin-resistant LB plate and incubate at 37°C for 12-16 hours. Pick several individual colonies and culture them overnight at 37°C in a shaker for small-scale plasmid DNA extraction. The plasmid was sent to Liuhe BGI Genomics Co., Ltd. for Sanger sequencing to verify that the mutation site was correctly introduced. The mutation site information is shown in Table 6.

[0145] Table 6

[0146] The primers and their sequences corresponding to the different mutants in this example are shown in Table 7, including primer 1 (SEQ ID NOs: 13-98) and primer 2 (SEQ ID NOs: 99-184).

[0147] Table 7

[0148] Example 8 Testing of the ability of mutants to polymerize 3'-O-blocking nucleotides

[0149] dATP labeled with Cy5 fluorescent dye and modified with a 3'-O-azidomethyl reversible blocking group was used as a substrate, and double-stranded DNA labeled with Cy3 fluorescent dye was used as a template / primer (PT-1 (SEQ ID NO: 185): Cy3-CGTGTATGC GTAATAGGATCCCGACTCACTATGGACG; PT-2 (SEQ ID NO: 186): Cy3-CGTGTATCGT CCATAGTGAGTCGGGATCCTATTACGC; PT-1 and PT-2, each modified with Cy3 at the 5' end, were mixed in equal proportions, incubated at 80°C for 10 minutes, and then cooled to room temperature to obtain the template / primer). This simulated the incorporation of reversibly blocked modified nucleotides during high-throughput sequencing. The polymerization activity of 35°S wild-type (35°S-WT), 35°SM DNA polymerase and its mutants towards the incorporation of modified single bases was detected by detecting the FRET-Cy5 fluorescence signal generated by Cy3 and Cy5 being within the interaction distance when the substrate was polymerized onto the template chain on a microplate reader.

[0150] The detection method is as follows: 10× reaction buffer is 200mM Tris-HCl, 100mM (NH4)2SO4, 100mM KCl, 20mM MgSO4, pH 8.5@25℃.

[0151] A 50 μL reaction system was prepared as follows: 5 μL 10× reaction buffer, 40 μM dT / G / CTP, 0.1 mg / mL BSA, 1 μM Cy3-DNA double-stranded template, 4 μM Cy5-modified 3'-O-AzidoMethyl-dATP, 0.5 μg 35°SM DNA polymerase or its mutants, and nuclease-free water was used to make up the reaction system to 50 μL.

[0152] The reaction system was set up on ice, transferred to a 384-well plate, and placed in a microplate reader (BioTek Synergy H1, Agilent) for fluorescence signal detection. The reaction temperature was 40°C, and fluorescence signals at 530 / 568 nm and 630 / 676 nm (FRET Cy5) were collected every 60 seconds. The entire reaction lasted 2 hours.

[0153] For the negative control, an equal volume of enzyme stock solution (i.e., the stock solution described in Example 3) was used to replace 35° SM DNA polymerase. After the reaction, raw data was exported and the maximum slope, i.e., the ΔFRET-Cy5 fluorescence value per unit time (ΔRFU), was calculated to characterize the polymerase's polymerization activity for the incorporation of modified single bases.

[0154] The experimental results are shown in Table 8. No FRET Cy5 fluorescence signal was detected in the negative control group. However, an increase in fluorescence signal was detected in the experimental group of 86 DNA polymerase mutants, including 35°S-M1-M86. This demonstrates that the combined mutagenesis of 35°S DNA polymerase has polymerization activity towards 3'-blocking modified nucleotides. Effective mutagenesis of 35°S DNA polymerase provides mutants with great potential for high-throughput sequencing.

[0155] Table 8

[0156] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0157] Compared to existing DNA polymerases or other DNA polymerases, the 35°S DNA polymerase in this application is a novel DNA polymerase. The wild-type version of this polymerase already possesses excellent thermal stability and PCR amplification capabilities, with good original enzymatic performance and excellent PCR amplification performance. Furthermore, the enzyme's sequence is novel and can be further optimized and modified based on the application scenario, offering significant room for improvement and potential for application. Furthermore, by performing simple mutational modification on the wild-type 35°S DNA polymerase, DNA polymerase mutants with polymerization activity towards deoxyribonucleotides with blocking modifications, such as 3'-O-reversible blocking modifications, can be obtained. This has application value in high-throughput sequencing technologies based on deoxyribonucleotides with blocking modifications, such as 3'-O-reversible blocking modifications.

[0158] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A DNA polymerase, characterized in that The DNA polymerase comprises: (a) a protein comprising the amino acid sequence shown in SEQ ID NO: 1; (b) a protein having DNA polymerase activity, wherein one or more amino acids are substituted, deleted, and / or added at least one site in the exo-active sequence region and / or motif A of the amino acid sequence shown in SEQ ID NO: 1; or (c) a protein that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% homologous to the protein in (a) or (b) and has DNA polymerase activity.

2. The DNA polymerase according to claim 1, wherein In (b), the sites of the exo-active sequence region include any one or more of the following sites: D180, E182, N252, F256, D257, Y354 or D358; The sites where the motif A undergoes substitution, deletion and / or addition of one or more amino acids include any one or more of the following sites: L452, Y453 or P454.

3. The DNA polymerase according to claim 2, characterized in that D180 and E182 were each independently substituted with a non-negatively charged amino acid; Preferably, the non-negatively charged amino acids include alanine, glycine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, tyrosine, asparagine, cysteine, glutamine, serine, threonine, arginine, histidine or lysine; Preferably, the types of amino acids substituted at each site in the exoactive sequence region are independently selected from the following: the types of amino acids substituted by D180 include D180A; Types of amino acids substituted by E182 include E182A; Types of amino acids substituted at N252 include N252A; Types of amino acids substituted by F256 include F256A; Types of amino acids substituted with D257 include D257A; Types of amino acids substituted with Y354 include Y354A; The types of amino acids substituted by D358 include D358A; preferably, the types of amino acids substituted at each position in the motif A are independently selected from the following: The types of amino acids substituted with L452 include L452A, L452G, L452H, L452I, L452K, L452L, L452M, L452Q, L452S, L452T, L452V, L452Y, L452F, or L452W; The types of amino acids substituted with Y453 include Y453A, Y453D, Y453G, Y453I, Y453M, Y453S, Y453T, Y453V, Y453H or Y453P; The types of amino acids substituted by P454 include P454G, P454T, P454V, P454A, P454I, P454E, P454I, P454K, P454L, P454M, P454S, or P454V.

4. The DNA polymerase according to claim 3, characterized in that The types of amino acids substituted at each site include any of the following combinations: D180A+E182A、D180A+E182A+L452A+Y453A、D180A+E182A+L452A+Y453G、D180A+E182A+L452G+Y453A、D180A+E182A+L452I+Y453G、D180A+E182A+L452I+Y453S、D180A+E182A+L452I+Y453T、D180A+E182A+L452L+Y453A、D180A+E182A+L452M+Y453A、D180A+E182A+L452T+Y453A、D180A+E182A+L452V+Y453G、D180A+E182A+L452A+Y453A+P454G、D180A+E182A+L452A+Y453A+P454T、D180A+E182A+L452A+Y453A+P454V、D180A+E182A+L452A+Y453D+P454G、D180A+E182A+L452A+Y453G+P454A、D180A+E182A+L452A+Y453A+P454I、D180A+E182A+L452A+Y453G+P454G、D180A+E182A+L452A+Y453G+P454Q、D180A+E182A+L452A+Y453G+P454T、D180A+E182A+L452A+Y453M+P454K、D180A+E182A+L452A+Y453S+P454G、D180A+E182A+L452A+Y453T+P454A、 D180A+E182A+L452A+Y453V+P454A、D180A+E182A+L452G+Y453A+P454G、D180A+E182A+L452G+Y453G+P454T、D180A+E182A+L452G+Y453G+P454V、D180A+E182A+L452H+Y453A+P454A、D180A+E182A+L452H+Y453G+P454G、D180A+E182A+L452H+Y453M+P454G、D180A+E182A+L452H+Y453V+P454G、D180A+E182A+L452I+Y453A+P454A、D180A+E182A+L452I+Y453A+P454T、D180A+E182A+L452I+Y453A+P454V、D180A+E182A+L452I+Y453G+P454A、D180A+E182A+L452I+Y453T+P454A、D180A+E182A+L452I+Y453T+P454G、D180A+E182A+L452K+Y453A+P454G、D180A+E182A+L452K+Y453G+P454A、D180A+E182A+L452L+Y453A+P454A、D180A+E182A+L452L+Y453A+P454G、D180A+E182A+L452L+Y453A+P454L、D180A+E182A+L452L+Y453A+P454M、D180A+E182A+L452L+Y453A+P454T、D180A+E182A+L452L+Y453A+P454V、D180A+E182A+L452L+Y453G+P454A、D180A+E182A+L452L+Y453G+P454G、D180A+E182A+L452L+Y453G+P454I、D180A+E182A+L452L+Y453G+P454L、 D180A+E182A+L452L+Y453S+P454A、D180A+E182A+L452L+Y453T+P454G、D180A+E182A+L452L+Y453V+P454G、D180A+E182A+L452M+Y453A+P454T、D180A+E182A+L452M+Y453A+P454V、D180A+E182A+L452M+Y453G+P454A、D180A+E182A+L452M+Y453G+P454G、D180A+E182A+L452M+Y453T+P454G、D180A+E182A+L452Q+Y453A+P454A、D180A+E182A+L452S+Y453G+P454E、D180A+E182A+L452T+Y453A+P454A、D180A+E182A+L452T+Y453A+P454V、D180A+E182A+L452T+Y453G+P454A、D180A+E182A+L452T+Y453G+P454T、D180A+E182A+L452T+Y453P+P454G、D180A+E182A+L452T+Y453S+P454T、D180A+E182A+L452T+Y453S+P454V、D180A+E182A+L452T+Y453T+P454S、D180A+E182A+L452T+Y453V+P454G、D180A+E182A+L452V+Y453A+P454A、D180A+E182A+L452V+Y453A+P454G、D180A+E182A+L452V+Y453A+P454T、D180A+E182A+L452V+Y453A+P454V、D180A+E182A+L452V+Y453G+P454A、D180A+E182A+L452V+Y453G+P454G、D180A+E182A+L452V+Y453G+P454V、 D180A+E182A+L452V+Y453H+P454A, D180A+E182A+L452V+Y453S+P454A, D180A+E182A+L452V+Y453 T+P454A, D180A+E182A+L452V+Y453T+P454G, D180A+E182A+L452V+Y453T+P454S, D180A+E182A+L45 2V+Y453T+P454T, D180A+E182A+L452V+Y453T+P454V, D180A+E182A+L452V+Y453V+P454A, D180A+E182A+L452V+Y453V+P454G, D180A+E182A+L452V+Y453V+P454V and D180A+E182A+L452Y+Y453P+P454G.

5. The DNA polymerase according to claim 4, characterized in that The DNA polymerase defined in (b) has polymerization activity on deoxyribonucleotides having blocking modifications; Preferably, the blocking modification comprises a 3'-O-reversible blocking modification; Preferably, the 3'-O-reversible blocking modification comprises a 3'-O-azidomethyl modification.

6. A DNA molecule, characterized in that the DNA molecule encodes the DNA polymerase according to any one of claims 1 to 5; The DNA molecule includes a DNA molecule having a nucleotide sequence shown in SEQ ID NO: 2 or DEQ ID NO: 3, or a DNA molecule having a homology of more than 70% with the nucleotide sequence shown in SEQ ID NO: 2 or DEQ ID NO:

3.

7. A recombinant plasmid, characterized in that The recombinant plasmid is connected to the DNA molecule according to claim 6.

8. A host cell, characterized in that The host cell contains the DNA molecule according to claim 6 or the recombinant plasmid according to claim 7.

9. The host cell according to claim 8, characterized in that The host cell includes a prokaryotic cell or a eukaryotic cell; Preferably, the prokaryotic cell comprises Escherichia coli; Preferably, the host cell is not a plant cell.

10. A PCR kit, characterized in that: The PCR kit comprises the DNA polymerase according to any one of claims 1 to 5 and optionally any one or more of the following components: 1) Buffer for providing a PCR amplification environment; 2) PCR primers; 3) Reagents for extracting target DNA.

11. Use of the DNA polymerase according to any one of claims 1 to 5 or the PCR kit according to claim 10 in incorporating modified deoxyribonucleotides into a polynucleotide, PCR amplification, library construction or sequencing.

12. A method for incorporating modified deoxyribonucleotides, characterized in that: Incorporating modified deoxyribonucleotides into a polynucleotide using the DNA polymerase according to any one of claims 1 to 5; The modifications include 3'-O-blocking modifications; The 3'-O-blocking modification includes a 3'-O-azidomethyl modification.

13. A PCR amplification method, characterized in that: The PCR method comprises: performing PCR amplification using the DNA polymerase according to any one of claims 1 to 5 or the PCR kit according to claim 10.

14. The PCR amplification method according to claim 13, characterized in that The PCR amplification method comprises: using the DNA polymerase to incorporate modified deoxyribonucleotides during the amplification process; Preferably, the modification comprises a 3'-O-blocking modification; Preferably, the 3'-O-blocking modification comprises a 3'-O-azidomethyl modification.

15. A library construction kit, characterized in that The library construction kit comprises the DNA polymerase according to any one of claims 1 to 5 and optionally any one or more of the following components: 1) Adapter sequence for library construction; 2) Buffer for providing library construction environment; 3) Enzymes used to construct the library, including one or more of DNA shearing enzymes, end-repair enzymes, or ligases; 4) Contains modified deoxyribonucleotides.

16. A method for constructing a library, characterized in that: The library construction method comprises: constructing a library using the DNA polymerase according to any one of claims 1 to 5 or the library construction kit according to claim 15.

17. The library construction method according to claim 16, characterized in that The library construction method comprises: The DNA polymerase is used to perform PCR amplification on the fragment to be sequenced connected with the sequencing adapter to obtain a DNA library.

18. A sequencing kit, characterized in that: The sequencing kit comprises the DNA polymerase according to any one of claims 1 to 5 and optionally any one or more of the following components: 1) a primer for complementary pairing with the adapter sequence; 2) dideoxynucleotides; 3) dNTPs; 4) Nucleic acid probes; 5) Enzymes for sequencing, including DNA ligase and / or endonuclease; 6) a buffer for eluting the nucleic acid probe and / or dNTPs; 7) Contains modified deoxyribonucleotides.

19. A sequencing method, characterized in that: The sequencing method comprises: performing sequencing using the DNA polymerase according to any one of claims 1 to 5 or the sequencing kit according to claim 18.

20. The sequencing method according to claim 19, characterized in that The sequencing method comprises: Sequencing the DNA library by synthesis and sequencing technology to obtain sequencing results; Preferably, the DNA polymerase is used to incorporate modified deoxyribonucleotides to perform sequencing by synthesis.

Citation Information

Patent Citations

  • 9-degree N DNA polymerase mutant

    CN111349623A

  • Recombinant KOD polymerase

    CN112639089A

  • Heat-resistant B family DNA polymerase mutant and application thereof

    CN117396600A

  • Modified polymerases for improved incorporation of nucleotide analogues

    US20160362664A1

  • Thermophilic DNA polymerase mutants

    US20210147817A1