Mutant DNA polymerases and methods of use thereof
By introducing specific amino acid mutations into DNA polymerase, the problems of low efficiency and poor resistance of existing DNA polymerases have been solved, resulting in more efficient DNA synthesis and sequencing performance, especially improvements under high-salt environments and complex templates.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LIFE TECHNOLOGIES CORP
- Filing Date
- 2024-07-29
- Publication Date
- 2026-05-26
AI Technical Summary
Existing DNA polymerases are inefficient in DNA synthesis, unable to efficiently read all regions of the template, and their efficiency decreases at high salt concentrations. They also have 5'-3' nuclease activity, which hinders the incorporation of fluorescently labeled nucleotides, and are sensitive to inhibitors, making it difficult to effectively sequence complex templates.
By introducing specific amino acid mutations, such as F667Y and G46D, into DNA polymerase, as well as combinations of other options such as E681I, D732N, A743H, E507N, E742H, M747K, and S543N, the performance of the enzyme is enhanced, its affinity for DNA substrates is increased, its 5'-3' nuclease activity is reduced, its incorporation of fluorescently labeled nucleotides is enhanced, its template readability is improved, and its high salt resistance and sequencing read length are increased.
This resulted in increased DNA polymerase speed, improved affinity for DNA substrates, reduced 5'-3' nuclease activity, enhanced incorporation of fluorescently labeled nucleotides, improved template readability, improved performance at high salt concentrations, and increased sequencing read length.
Smart Images

Figure CN122095103A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 516,396, filed July 28, 2023, and U.S. Provisional Application No. 63 / 516,412, filed July 28, 2023, and U.S. Provisional Application No. 63 / 516,435, filed July 28, 2023. Technical Field
[0003] This invention relates to mutant DNA polymerases. Background Technology
[0004] DNA polymerase is an enzyme that synthesizes DNA molecules using a template DNA strand and complementary synthetic primers annealed to a portion of the template. For a detailed description of DNA polymerases and their enzymatic characterization, see Kornberg (1989).
[0005] The amino acid sequences of many DNA polymerases have been determined, and sequence comparisons of different DNA polymerases have identified many homologous regions between the enzymes. Studies on the tertiary structure and amino acid sequence comparisons of DNA polymerases have revealed many structural similarities between different DNA polymerases. Generally, DNA polymerases have a large cleft, which is thought to be used to accommodate the binding of double-stranded DNA. This cleft is formed by two sets of helices, the first set of which is called the "finger" region and the second set of which is called the "thumb" region. The bottom of the cleft is formed by antiparallel β plates and is called the "palm" region. A review of DNA polymerase structures can be found in Joyce and Steitz (1994). Computer-readable data files describing the three-dimensional structures of some DNA polymerases have been publicly distributed.
[0006] DNA polymerases have a variety of uses in molecular biology techniques applicable to both research and clinical applications. Among these techniques, the most important are DNA sequencing and polynucleotide amplification techniques (such as polymerase chain reaction (PCR)).
[0007] However, despite their wide application, existing DNA polymerases may exhibit several properties that could reduce the efficiency of the enzyme used to synthesize DNA, including: the polymerase may not be able to efficiently read all regions of the template; the polymerase may have reduced efficiency at higher salt concentrations; the polymerase may exhibit 5'-3' nuclease activity; and / or the polymerase may hinder the efficient incorporation of fluorescently labeled nucleotides into the resulting DNA strand.
[0008] Therefore, there is a need for DNA polymerases that have improved efficiency in synthesizing DNA molecules from, for example, fluorescently labeled nucleotides. Summary of the Invention
[0009] This article provides mutant polymerases that can be used, for example, for sequencing DNA. In some embodiments, the mutant polymerase is mutated in the following ways: (1) increasing the polymerase speed relative to the wild-type DNA polymerase; (2) enhancing the affinity for DNA substrates; (3) increasing resistance to DNA polymerase inhibitors; (4) reducing 5'-3' nuclease activity; (5) enabling more efficient incorporation of fluorescently labeled nucleotides into the resulting DNA strand; (6) improving the polymerase's ability to read templates, for example, those with secondary structures; (7) increasing resistance to higher salt concentrations; and / or (8) increasing the sequencing read length of the DNA template.
[0010] Therefore, certain embodiments of the present invention provide a mutant DNA polymerase comprising F667Y and G46D mutations and one or more substitutions selected from the group consisting of E681I, D732N, A743H, E507N, E742H, and M747K. In a further embodiment, the mutant DNA polymerase comprises an E189K substitution. In a further embodiment, the mutant DNA polymerase comprises an S543N substitution.
[0011] Other embodiments of the present invention provide a mutant DNA polymerase comprising E189K, F667Y and G46D substitutions and one or more substitutions selected from the group consisting of D732N, E507N, S543N, E742H and M747K.
[0012] In some embodiments, the present invention also provides polynucleotides encoding the polymerases of the present invention, including expression cassettes and vectors of such polynucleotides, and cells containing such polymerases and polynucleotides.
[0013] A method for synthesizing polynucleotides in a reaction is also provided, comprising contacting at least one polymerase of the present invention with an initiating template and a nucleotide (e.g., a fluorescently labeled nucleotide) under conditions conducive to polynucleotide synthesis. In some embodiments, the present invention also provides a kit containing packaging material and at least one polymerase of the present invention.
[0014] Methods for sequencing polynucleotides (e.g., sequencing DNA sequences) using the polymerase of the present invention are also provided. Attached Figure Description
[0015] Figure 1 shows the protein yields obtained by the various mutant DNA polymerases of the present invention. All 20 mutant DNA polymerases and the control enzyme successfully achieved protein expression, with an average yield of 2.6 mg of active DNA polymerase activity per 1 L of bacterial culture (specific activity of 86,000 U / mg).
[0016] Figure 2 illustrates the ability of the mutant DNA polymerase of the present invention to shorten the time of thermal cycling.
[0017] Figure 3 depicts the results obtained from evaluating the effect of extension time on the success rate of sequencing reactions performed using the mutant DNA polymerase of the present invention according to the BigDyeTerminator (“BDT”) v3.1 thermal cycling protocol. Extension time in the cycle sequencing was varied (5 to 60 seconds), with pGEM used as a control plasmid in the DNA replication BDT v3.1 sequencing reaction, covering double strands (forward (F) and reverse (R)). Performance was measured in sequential read length QV20 (labeled QV20 in the figure). Success rate was determined against the quality standard limit (LSL = 850 nt) of the reference run module (50 cm, POP-7 StdSeq). The data show that extending the extension time resulted in an increased mean CRL, reduced read length variability, and improved success rate (relative to the quality standard limit).
[0018] Figure 4 depicts the situation in GeneAmp. TM Results obtained by evaluating the polymerase rate of the mutant DNA polymerase of this invention using multiplex PCR in PCR buffer II. The maximum fragment length was recorded for each enzyme and each extension time using multiplex PCR agarose gel electrophoresis. Three mutations (E507K, E742H, and E189K) caused GeneAmp TM The polymerase rate in PCR Buffer II increased to approximately 3-fold. Two additional mutations (E681I and A743H) resulted in a milder approximately 1.5-fold increase in polymerase rate (dashed arrows).
[0019] Figure 5 illustrates the results obtained when identifying the mutant DNA polymerases of the present invention using high-speed polymerase activity. All mutant enzymes were sequenced in both directions of plasmid DNA (pGEM) using very short (5 s and 10 s) cycle sequencing extension times in ReadyReaction sequencing (RR) buffer. Performance of all reactions was measured in sequential read lengths of QV20. Compared to the control DNA polymerases with G46D and F667Y mutations, the E189K mutation alone, and its combination with additional mutations (E507K, E742H, S543N), resulted in a significant increase in polymerase speed during dye-terminated sequencing. Polymerase speed assessment by sequencing confirmed the results of polymerase speed assessment by PCR.
[0020] Figure 6 shows the sequencing read lengths of various mutant DNA polymerases of the present invention under conditions of using a 5-second extension time. The five fastest mutant DNA polymerases that produced the longest read lengths using a 5-second thermal cycling extension time all carried the E189K mutation.
[0021] Figure 7 illustrates the results obtained in assessing the impact of E189K, G46D, and F667Y mutations in DNA polymerases on sequencing speed. Of the six mutant DNA polymerases tested, five showed unexpectedly longer sequencing read lengths due to the presence of the E189K mutation at short thermal cycling extension times of 5 seconds and 10 seconds.
[0022] Figure 8 illustrates the effect of the E189K, G46D, and F667Y mutations in DNA polymerases on peak quality. In parallel comparisons, for five of the six mutant DNA polymerases of the present invention tested, the addition of the E189K mutation produced better peak quality (as measured by Trace Score) at shorter thermal cycling extension times of 5 seconds and 10 seconds.
[0023] Figure 9 illustrates that the mutant DNA polymerase of the present invention can generate high-quality sequencing data using control plasmid template DNA. The sequencing quality of the mutant DNA polymerase of the present invention is compared with that of BDT using pGEM as control plasmid DNA and 5x ReadyReaction buffer 3.1. TM Compared to v3.1 (control), all 21 mutant DNA polymerases tested generated high-quality sequencing data. This figure depicts the sequencing results of the E507K, G46D, and F667Y mutant DNA polymerases.
[0024] Figure 10 shows the results observed when the mutant DNA polymerase of the present invention is used with a difficult target template, compared to BDT v3.1 DNA polymerase. Plasmids pGEM (control) and p4009-1 (difficult template) were sequenced in repeat reactions using BDT v3.1, covering double strands (F and R). Performance was measured in sequential read lengths of QV20. Success rate was determined as the percentage of reactions that met the quality criteria (LSL = 850 nt) for the run module (50 cm POP-5 StdSeq). Difficult-to-sequence DNA templates resulted in a reduced average CRL, increased read length variability, and a lower sequencing reaction success rate (relative to the quality criterion limit) compared to conventional DNA templates.
[0025] Figure 11 illustrates the read length (CRL QV-20) analysis of BDT v3.1 DNA polymerase relative to the mutant DNA polymerases of this invention. Two conventional DNA templates (pGEM, AAV400-polyA-v01) and seven difficult DNA templates, covering double strands (F and R), were sequenced in repeat reactions using BDT v3.1 and six mutant DNA polymerases of this invention carrying the E189K mutation. Performance was measured as the average base quality TraceScore. For all nine DNA templates tested, the DNA polymerases carrying the E189K mutation generally provided read lengths greater than or equal to those of the BDTv3.1 DNA polymerase.
[0026] Figure 12 illustrates the base quality (TraceScore) analysis of the BDT v3.1 DNA polymerase relative to the mutant DNA polymerases of the present invention. Using BDT v3.1 and six mutant DNA polymerases of the present invention carrying the E189K mutation, sequencing was performed on two conventional DNA templates (pGEM, AAV400-polyA-v01) and seven difficult DNA templates, covering double strands (F and R), in repeat reactions. Performance was measured as the average base quality TraceScore. For all nine DNA templates tested, the DNA polymerases carrying the E189K mutation provided an average base quality score higher than or equal to that of BDT v3.1.
[0027] Figure 13 shows the read lengths (CRL QV-20) observed by the two mutant DNA polymerases of the present invention compared to BDT v3.1 when using two conventional DNA templates and seven difficult DNA templates. For all nine DNA templates tested, the mutant DNA polymerases of the present invention performed equal to or better than BDT v3.1.
[0028] Figure 14 illustrates the effect of additional mutations (D732N, E507K, E742H, M747K, and S543N) in the E189K, G46D, and F667Y mutant DNA polymerases on the sequencing read length of two different plasmid DNA templates. Among all six mutant DNA polymerases tested, the presence of the E189K mutation resulted in longer sequencing read lengths for both plasmid DNA templates at the standard (240 sec) thermal cycling extension time.
[0029] Figure 15 illustrates the effect of additional mutations (D732N, E507K, E742H, M747K, and S543N) in the E189K, G46D, and F667Y mutant DNA polymerases on the sequencing read length of pGEM-3Zf(+) sequencing standard plasma DNA templates. Of the six mutant DNA polymerases tested in this invention, the presence of the E189K mutation resulted in longer sequencing read lengths in five of them under standard (240 sec) thermal cycling extension times.
[0030] Figure 16 illustrates the effect of additional mutations (D732N, E507K, E742H, M747K, and S543N) in the E189K, G46D, and F667Y mutant DNA polymerases on the sequencing read length of the p4009-1 plasmid DNA template (“difficult-to-sequence” template). Of the six mutant DNA polymerases tested in this invention, the presence of the E189K mutation resulted in longer sequencing read lengths in five of them at a standard (240 sec) thermal cycling extension time.
[0031] Figure 17 illustrates the multiplex amplification of 15 DNA targets of different lengths using the DNA mutant polymerase of the present invention.
[0032] Figure 18 shows a 1.3 KB amplification of λDNA using the DNA mutant polymerase of the present invention. Detailed Implementation
[0033] This document describes combinatorial mutations to produce polymerases that can be used, for example, for sequencing DNA. In some embodiments, these mutations: (1) increase polymerase speed relative to wild-type DNA polymerase; (2) enhance affinity for DNA substrates; (3) increase resistance to DNA polymerase inhibitors; (4) reduce 5'-3' nuclease activity; (5) enable more efficient incorporation of fluorescently labeled nucleotides into the resulting DNA strand; (6) improve the polymerase's ability to read templates, for example, those with secondary structures; (7) increase resistance to higher salt concentrations; and / or (8) increase the sequencing read length of the DNA template.
[0034] Therefore, certain embodiments of the present invention provide a mutant DNA polymerase comprising an amino acid substitution at amino acid position 189 and one or more substitutions selected from the group consisting of G46D, S543N, F667Y, D732N, E742H, and M747K. In some embodiments, the amino acid substitution at the amino acid position is a Lys residue. In some embodiments, the mutant DNA polymerase comprises E189K, S543N, F667Y, and E742H substitutions. In a further embodiment, the mutant DNA polymerase comprises E507K substitution.
[0035] Some embodiments of the present invention provide a mutant DNA polymerase comprising F667Y and G46D mutations and one or more substitutions selected from the group consisting of E681I, D732N, A743H, E507N, E742H and M747K.
[0036] Other embodiments of the present invention provide a mutant DNA polymerase comprising E189K, F667Y and G46D substitutions and one or more substitutions selected from the group consisting of D732N, E507N, S543N, E742H and M747K.
[0037] The DNA polymerase may be a thermostable Taq DNA polymerase. In some embodiments, the DNA polymerase may comprise SEQ ID NO: 3 to 8.
[0038] The present invention also provides a polynucleotide encoding the polymerase of the present invention, as well as a cassette and vector comprising such a polynucleotide. The polynucleotide can be operatively linked to a promoter. Cells containing the polymerase, polynucleotide, cassette, and / or vector of the present invention are also provided.
[0039] A wild-type polymerase from *Thermophyton aquatilis* is SEQ ID NO: 1. A nucleotide sequence encoding this wild-type polymerase is SEQ ID NO: 2. (See accession number 104636)
[0040] (SEQ ID NO: 1)
[0041] MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVD LLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLL EEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANL WGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSG DENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0042] In the following sequence, the start codon (atg) at position 121 is underlined. In some embodiments of the invention, mutations may occur to produce the codon for the polymerase of the invention, which is also underlined.
[0043] (SEQ ID NO: 2)
[0044] 1 aagctcagat ctacctgcct gagggcgtcc ggttccagct ggcccttccc gagggggaga
[0045] 61 gggaggcgtt tctaaaagcc cttcaggacg ctacccgggg gcgggtggtg gaagggtaac
[0046] 121 atg aggggga tgctgcccct ctttgagccc aagggccggg tcctcctggt ggacggccac
[0047] 181 cacctggcct accgcacctt ccacgccctg aagggcctca ccaccagccg gggggagccg
[0048] 241 gtgcaggcgg tctac ggc tt cgccaagagc ctcctcaagg ccctcaagga ggacggggac
[0049] 301 gcggtgatcg tggtctttga cgccaaggcc ccctccttcc gccacgaggc ctacgggggg
[0050] 361 tacaaggcgg gccgggcccc cacgccggag gactttcccc ggcaactcgc cctcatcaag
[0051] 421 gagctggtgg acctcctggg gctggcgcgc ctcgaggtcc cgggctacga ggcggacgac
[0052] 481 gtcctggcca gcctggccaa gaaggcggaa aaggagggct acgaggtccg catcctcacc
[0053] 541 gccgacaaag acctttacca gctcctttcc gaccgcatcc acgtcctcca ccccgagggg
[0054] 601 tacctcatca ccccggcctg gctttgggaa aagtacggcc tgaggcccga ccagtgggcc
[0055] 661 gactaccggg ccctgaccgg ggacgagtcc gacaaccttc ccggggtcaa gggcatcggg
[0056] 721 gagaagacgg cgaggaagct tctggaggag tgggggagcc tggaagccct cctcaagaac
[0057] 781 ctggaccggc tgaagcccgc catccgggag aagatcctgg cccacatgga cgatctgaag
[0058] 841 ctctcctggg acctggccaa ggtgcgcacc gacctgcccc tggaggtgga cttcgccaaa
[0059] 901 aggcgggagc ccgaccggga gaggcttagg gcctttctgg agaggcttga gtttggcagc
[0060] 961 ctcctccacg agttcggcct tctggaaagc cccaaggccc tggaggaggc cccctggccc
[0061] 1021 ccgccggaag gggccttcgt gggctttgtg ctttcccgca aggagcccatgtgggccgat
[0062] 1081 cttctggccc tggccgccgc cagggggggc cgggtccacc gggcccccgagccttataaa
[0063] 1141 gccctcaggg acctgaagga ggcgcggggg cttctcgcca aagacctgagcgttctggcc
[0064] 1201 ctgagggaag gccttggcct cccgcccggc gacgacccca tgctcctcgcctacctcctg
[0065] 1261 gacccttcca acaccacccc cgagggggtg gcccggcgct acggcggggagtggacggag
[0066] 1321 gaggcggggg agcgggccgc cctttccgag aggctcttcg ccaacctgtgggggaggctt
[0067] 1381 gagggggagg agaggctcct ttggctttac cgggaggtgg agaggcccctttccgctgtc
[0068] 1441 ctggcccaca tggaggccac gggggtgcgc ctggacgtgg cctatctcagggccttgtcc
[0069] 1501 ctggaggtgg ccgaggagat cgcccgcctc gaggccgagg tcttccgcctggccggccac
[0070] 1561 cccttcaacc tcaactcccg ggaccagctg gaaagggtcc tctttgacgagctagggctt
[0071] 1621 cccgccatcg gcaagacgga gaagaccggc aagcgctcca ccagcgccgccgtcctggag
[0072] 1681 gccctccgcg aggcccaccc catcgtggag aagatcctgc agtaccgggagctcaccaag
[0073] 1741 ctgaag agc a cctacattga ccccttgccg gacctcatcc accccaggacgggccgcctc
[0074] 1801 cacacccgct tcaaccagac ggccacggcc acgggcaggc tagtagctccgatcccaac
[0075] 1861 ctccagaaca tccccgtccg caccccgctt gggcagagga tccgccgggccttcatcgcc
[0076] 1921 gaggaggggt ggctattggt ggccctggac tatagccaga tagagctcagggtgctggcc
[0077] 1981
[0078] 2041 gagaccgcca gctggatgtt cggcgtcccc cgggaggccg tggaccccctgatgcgccgg
[0079] 2101 gcggccaaga ccatcaac ttc ggggtcctc tacggcatgt cggcccaccgcctctcccag
[0080] 2161 gagctagcca tcccttacga ggaggcccag gccttcattg agcgctactttcagagcttc
[0081] 2221 cccaaggtgc gggcctggat tgagaagacc ctggagagg gcaggaggcgggggtacgtg
[0082] 2281 gagaccctct tcggccgccg ccgctacgtg ccagacctag aggcccgggtgaagagcgtg
[0083] 2341 cgggaggcgg ccgagcgcat ggccttcaac atgcccgtcc agggcaccgccgccgacctc
[0084] 2401 atgaagctgg ctatggtgaa gctcttcccc aggctggagg aaatgggggccaggatgctc
[0085] 2461 cttcaggtcc acgacgagct ggtcctcgag gccccaaaag agagggcggaggccgtggcc
[0086] 2521 cggctggcca aggaggtcat ggagggggtg tatcccctgg ccgtgcccctggaggtggag
[0087] 2581 gggatag gggaggactg gctctccgcc aaggagtgat accacc
[0088] The mutant DNA polymerases of the present invention (E189K, G46D, F667Y; SEQ ID NO: 3) are described below. The mutant amino acids are underlined in the following text.
[0089] (SEQ ID NO: 3) MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVY D FAKSLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGD KSDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN Y GVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0090] The mutant DNA polymerase of the present invention (E189K, G46D, E507K, F667Y; SEQ ID NO:4) is provided below. The mutated amino acids are underlined below.
[0091] (SEQ ID NO: 4) MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVY D FAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDK SDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKT K KTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN Y GVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0092] The mutant DNA polymerase of the present invention (E189K, G46D, S543N, F667Y; SEQ ID NO:5) is provided below. The mutated amino acids are underlined below.
[0093] (SEQ ID NO: 5) MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVY DFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGD K SDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLK N TYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN Y GVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0094] The mutant DNA polymerase of the present invention (E189K, G46D, F667Y, D732N; SEQ ID NO:6) is provided below. The mutated amino acids are underlined below.
[0095] (SEQ ID NO: 6) MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVY D FAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGD K SDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN Y GVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVP N LEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0096] The mutant DNA polymerase of the present invention (E189K, G46D, F667Y, E742H; SEQ ID NO:7) is provided below. The mutated amino acids are underlined below.
[0097] (SEQ ID NO: 7)
[0098] MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVY D FAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGD K SDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN Y GVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVR H AAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0099] The mutant DNA polymerase of the present invention (E189K, G46D, F667Y, M747K; SEQ ID NO:8) is provided below. The mutated amino acids are underlined below.
[0100] (SEQ ID NO: 8)
[0101] MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVY D FAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGD K SDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN Y GVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAER KAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKE
[0102] Some embodiments of the invention also provide methods for synthesizing polynucleotides in a reaction, these methods comprising contacting at least one DNA polymerase of the invention with an initiated template and nucleotides. The reaction may be, for example, a chain-termination sequencing reaction or a polymerase chain reaction. The nucleotides may include labeled nucleotides, for example, fluorescently labeled nucleotides.
[0103] This type of polymerase chain reaction can be used in many laboratory and clinical techniques, including DNA fingerprinting, bacterial or viral detection, and genetic disease diagnosis.
[0104] Some embodiments of the present invention also provide a kit comprising packaging material and the DNA polymerase of the present invention. The kit may contain nucleotides, such as labeled nucleotides, for example, fluorescently labeled nucleotides. The kit may also contain unlabeled nucleotides. The kit may also include at least one primer.
[0105] Therefore, a novel polymerase has been developed, which incorporates various mutations to produce an enhanced polymerase suitable for applications such as DNA sequencing. These mutations may include: G46D, which reduces (e.g., eliminates) 5'-3' nuclease activity; F667Y, which enables more efficient dideoxynucleotide incorporation; and S543N, which enhances the polymerase's sustained synthetic capacity. S543N also improves the polymerase's ability to read regions of the template with secondary structure that would normally impair its sequencing capabilities. Furthermore, the S543N mutation enhances the polymerase's salt tolerance.
[0106] Therefore, methods utilizing certain polymerases of the present invention will exhibit reduced sequencing failures due to template secondary structure. Some polymerases also possess improved salt tolerance, reducing their sensitivity to salts (e.g., salts from template preparation or residual salts from PCR reactions). The use of certain polymerases also reduces the number of false terminations in dye-priming reactions. Mutations in certain polymerases also improve the ability of the polymerases of the present invention to tolerate dITPs and dUTPs in the extended strand.
[0107] The polymerase of the present invention can be used to prepare, for example, dye-terminated sequencing kits or dye-labeled primer kits. The polymerase of the present invention can also be used, for example, in direct PCR sequencing chemistry processes, such as in combination with polymerases that do not contain the F667Y mutation. In some embodiments of the present invention, the polymerase can be used in conjunction with, for example, dye-labeled primers and / or dye-labeled terminators to perform, for example, simultaneous amplification and sequencing.
[0108] Therefore, embodiments of the present invention include mutant polymerases and polynucleotide sequences encoding the mutant polymerases. The polynucleotide sequences encoding the mutant polymerases of the present invention can be used for the recombinant production of the mutant polymerases. The polynucleotide sequences encoding the mutant polymerases can be produced by various methods. One method for producing a polynucleotide sequence encoding a mutant polymerase is to introduce a desired mutation into the polynucleotide sequence encoding a parental wild-type polymerase using site-directed mutagenesis.
[0109] The polynucleotide encoding the mutant polymerase of the present invention can be used for recombinant expression of the mutant polymerase. Generally, recombinant expression of mutant polymerases is achieved by introducing the polynucleotide encoding the mutant polymerase into an expression vector regulated to suit a specific type of host cell. Therefore, another aspect of the present invention provides a vector comprising the polynucleotide encoding the mutant polymerase of the present invention, such that the polynucleotide encoding the mutant polymerase is functionally inserted into the vector. The present invention also provides a host cell comprising the vector of the present invention. The host cell for recombinant expression can be a prokaryotic cell or a eukaryotic cell. Examples of host cells include bacterial cells, yeast cells, cultured insect cell lines, and cultured mammalian cell lines. Various vectors (e.g., expression vectors) are well known in the art, and the expression of polymerases in recombinant cell systems is a well-established technique.
[0110] This invention also provides kits for synthesizing polynucleotides (e.g., fluorescently labeled polynucleotides). These kits are adaptable to perform specific polynucleotide synthesis procedures, such as DNA sequencing or PCR. Kits in certain embodiments of the invention include a mutant DNA polymerase of the invention. The kits preferably include instructions on how to perform the procedures for adapting the kit. Optionally, the kit may further include at least one other reagent for performing the methods adapted to be performed by the kit. Examples of such additional reagents include labeled nucleotides, unlabeled nucleotides, buffers, cloning vectors, restriction endonucleases, sequencing primers, and amplification primers. Reagents in the kits of the present invention can be supplied in pre-measured units to provide greater precision and accuracy.
[0111] The following terms are used to describe sequence relationships between two or more polynucleotides or polypeptides: (a) “reference sequence”, (b) “comparison window”, (c) “sequence identity”, (d) “percentage of sequence identity”, and (e) “substantial identity”.
[0112] As used in this article, a "reference sequence" refers to a sequence defined as the basis for sequence comparison. A reference sequence can be a fragment or the entirety of a specified sequence.
[0113] As used herein, a “comparison window” refers to a contiguous and specified segment of a polynucleotide or polypeptide sequence, wherein the polynucleotide or polypeptide sequence in the comparison window may include additions or deletions (i.e., vacancies) compared to a reference sequence (which does not include additions or deletions) to achieve optimal sequence alignment. Generally, the length of the comparison window is at least 5, 10, or 20 consecutive nucleotides or polypeptides, and optionally 30, 40, 50, 100, or longer. Those skilled in the art will understand that, to avoid high similarity to a reference sequence due to the presence of vacancies in the polynucleotide or polypeptide sequence, a vacancy penalty mechanism may be introduced and deducted from the number of matches.
[0114] Sequence alignment methods used for comparison are well known in the art. Therefore, the determination of the percentage of identity between any two sequences can be accomplished using mathematical algorithms. Preferred, non-limiting examples of such mathematical algorithms include: the algorithm of Myers and Miller, CABIOS, 4:11 (1988); the local homology algorithm of Smith et al., Adv. Appl. Math., 2:482 (1981); the homology alignment algorithm of Needleman and Wunsch, JMB, 48:443 (1970); the similarity search method of Pearson and Lipman, PNAS, 85:2444 (1988); and the algorithm of Karlin and Altschul, PNAS, 87:2264 (1990), modified as in Karlin and Altschul, PNAS, 90:5873 (1993).
[0115] Computer implementations of these mathematical algorithms can be used for sequence comparisons to determine sequence identity. Such implementations include, but are not limited to: CLUSTAL in the PC / Gene program (available from Intelligenetics (Mountain View, Calif.)); the ALIGN program; and GAP, BESTFIT, BLAST, PASTA, and TFASTA in the Wisconsin Genetics software package. These programs can be used with default parameters for alignment. The CLUSTAL program is described in detail in the following literature: Higgins et al., Gene, 73:237 (1988); Higgins et al., CABIOS, 5:151 (1989); Carpet et al., Nucl. Acids Res., 16:10881 (1988); Huang et al., CABIOS, 8:155 (1992); and Pearson et al., Meth. Mal. Biol., 24:307 (1994). The ALIGN program is based on the algorithms of Myers and Miller (ibid.). The BLAST procedure of Altschul et al., JMB, 215:403 (1990); Nucl. Acids Res., 25:3389 (1990) is based on the algorithm of Karlin and Altschul (ibid.).
[0116] The software used for BLAST analysis is publicly available through the National Center for Biotechnology Information (NCBI). This algorithm generally involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence. When compared with words of the same length in a database sequence, these HSPs match or satisfy a positive threshold score T. T is called the neighborhood word score threshold. These initial neighborhood word hits act as seeds to begin the search for longer HSPs containing them. Subsequently, word hits extend along each sequence in both directions as long as the cumulative alignment score can increase. For nucleotide sequences, the cumulative score is calculated using parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatched residues; always < 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of word hits in each direction is interrupted when the cumulative alignment score decreases by an amount X from its maximum value; when the cumulative score becomes zero or below zero due to the accumulation of one or more negative residue alignments; or when the end of any sequence is reached.
[0117] In addition to calculating the percentage of sequence identity, the BLAST algorithm also performs statistical analysis on the similarity between two sequences. One similarity metric provided by the BLAST algorithm is the minimum total probability (P(N)), which provides an indication of the probability that a match will occur accidentally between two nucleotide or amino acid sequences. For example, if the minimum total probability of the tested polynucleotide sequence compared to the reference sequence is less than about 0.1, more preferably less than about 0.01, and most preferably less than about 0.001, the polynucleotide sequence is considered similar to the reference sequence.
[0118] To obtain vacancy alignments for comparative purposes, Gapped BLAST can be used, as described by Altschul et al., Nucleic Acids Res. 25:3389 (1997). Alternatively, PSI-BLAST can be used for iterative searches to detect distant relationships between molecules. See Altschul et al., ibid. When using BLAST, Gapped BLAST, or PSI-BLAST, the default parameters of the respective programs can be used (e.g., BLASTN for nucleotide sequences, BLASTX for proteins). The BLASTN program (for nucleotide sequences) uses a word length (W) of 11, an expected value (E) of 10, a cutoff value of 100, M=5, N=-4, and a comparison of two strands as default values. For amino acid sequences, the BLASTP program defaults to a word length of 3, an expected value (E) of 10, and a BLOSUM62 scoring matrix. See http: / / www.ncbi.nlm.nih.gov. Manual alignment can also be performed by inspection.
[0119] For the purposes of this invention, to determine the percentage sequence identity between a sequence and the sequences disclosed herein, it is preferred to use the BlastN program (version 1.4.7 or later) with its default parameters or any equivalent program for sequence comparison. An "equivalent program" refers to any of the following sequence comparison programs that, for any two sequences in question, generate an alignment with the same nucleotide or amino acid residue match or the same percentage sequence identity as the corresponding alignment generated by the preferred program.
[0120] As used herein, in the context of two polynucleotide or polypeptide sequences, “sequence identity” or “identity” refers to a specified percentage of identical residues in two sequences when the maximum correspondence is matched within a specified comparison window by sequence comparison algorithms or by visual inspection. When using a sequence identity percentage for a reference protein, it should be recognized that differences in the positions of dissimilar residues are often due to conserved amino acid substitutions, where an amino acid residue is replaced by another amino acid residue with similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule. When sequences differ in terms of conserved substitutions, the sequence identity percentage can be increased to correct for the conservatism of the substitutions. The difference lies in the fact that sequences with such conserved substitutions are referred to as having “sequence similarity” or “identity.” The means for making such adjustments are well known to those skilled in the art. Typically, the adjustment involves scoring the conserved substitutions as partial mismatches rather than complete mismatches, thereby increasing the sequence identity percentage. Thus, for example, when identical amino acids are assigned a score of 1 and non-conservative substitutions are assigned a score of 0, the conservative substitutions are assigned a score between 0 and 1. Calculate the fraction of conservative substitutions, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, Calif.).
[0121] As used herein, "sequence identity percentage" refers to the value determined by comparing two best-aligned sequences within a comparison window. This comparison window may include additions or deletions (i.e., vacancies) compared to a reference sequence (excluding additions or deletions) to achieve optimal alignment. The percentage is calculated as follows: the number of positions in both sequences containing the same polynucleotide base or amino acid residue is determined to generate the number of matching positions. This number of matching positions is then divided by the total number of positions in the comparison window, and the result is multiplied by 100 to obtain the sequence identity percentage.
[0122] The term "substantially identical" in sequence means that a sequence having at least about 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68% or 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78% or 79%, preferably at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88% or 89%, more preferably at least 90%, 91%, 92%, 93% or 94%, and most preferably at least 95%, 96%, 97%, 98% or 99% sequence identity with respect to a reference sequence using one of the alignment procedures and standard parameters.
[0123] Another indicator of substantially identical sequences is whether the two molecules hybridize with each other under stringent conditions (see below). Generally, stringent conditions are chosen to be about 5°C below the thermodynamic melting point (Tm) of a particular sequence at a specified ionic strength and pH. However, stringent conditions cover temperatures ranging from about 1°C to 20°C, depending on the desired stringency (as detailed elsewhere in this document).
[0124] In sequence comparison, typically, one sequence serves as a reference sequence to be compared with the test sequence. When using a sequence comparison algorithm, the test and reference sequences are input into the computer, with subsequence coordinates specified if necessary, and the sequence algorithm program parameters are also specified. The sequence comparison algorithm then calculates the percentage of sequence identity between the test sequence and the reference sequence based on the specified program parameters.
[0125] As mentioned above, another indication that two sequences are substantially identical is that the two molecules hybridize with each other under stringent conditions. The phrase "specific hybridization to" refers to the phenomenon that, under stringent conditions, when a specific nucleotide sequence is present in a complex mixture of DNA or RNA (e.g., total cells), the molecule binds, doubles, or hybridizes only with that sequence. "Substantially binding" refers to complementary hybridization between the probe polynucleotide and the target polynucleotide, and includes slight mismatches that can be accommodated by reducing the stringency of the hybridization medium, thereby enabling the desired detection of the target polynucleotide sequence.
[0126] In the context of multinucleotide hybridization experiments such as Southern and Northern hybridization, "strict hybridization conditions" and "strict hybridization washing conditions" are sequence-dependent and differ under different environmental parameters. Longer sequences undergo specific hybridization at higher temperatures. mThe temperature at which 50% of the target sequence hybridizes to a perfectly matched probe (at a specified ionic strength and pH). Specificity typically varies with post-hybridization washing, with the key factors being the ionic strength and temperature of the final wash solution. For DNA-DNA hybrids, Tm can be estimated according to the formula Meinkoth and Wahl, Anal. Biochem., 138:267 (1984); Tm 81.5°C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% form) - 500 / L; where M is the molar concentration of the monovalent cation, % GC is the percentage of guanine to cytosine nucleotides in the DNA, % form is the percentage of formamide in the hybridization solution, and L is the length of the hybrid (in base pairs). Each 1% mismatch reduces Tm by approximately 1°C; therefore, Tm, hybridization conditions, and / or washing conditions can be adjusted to hybridize to sequences with desired identity. For example, if a sequence with >90% identity is sought, Tm can be lowered by 10°C. Generally, stringent conditions are chosen to be higher than the pyrolysis temperature (Tm) of a specific sequence and its complementary sequence at a given ionic strength and pH. m Approximately 5°C lower. However, under extremely stringent conditions, hybridization and / or washing temperatures 1°C, 2°C, 3°C, or 4°C lower than the specific thermal decomposition temperature (Tm) can be utilized; under moderately stringent conditions, hybridization and / or washing temperatures 6°C, 7°C, 8°C, 9°C, or 10°C lower than the specific thermal decomposition temperature (Tm) can be utilized; under low stringent conditions, specific thermal decomposition temperatures (Tm) can be utilized. mHybridization and / or washing temperatures 11°C, 12°C, 13°C, 14°C, 15°C, or 20°C lower. Using this equation, the hybridization and washing compositions, and the desired T, those skilled in the art will understand that a variation in the stringency of the hybridization and / or washing solutions is essentially described. If the desired mismatch rate results in T less than 45°C (aqueous solution) or 32°C (formamide solution), it is preferable to increase the SSC concentration so that a higher temperature can be used. A detailed guide to hybridization of polynucleotides can be found in Tijssen's "Laboratory Techniques in Biochemistry and Molecular Biology Hybridization with Nucleic Acid Probes," Volume I, Chapter 2, "Overview of principles of hybridization and the strategy of polynucleotide probe assays" (Elsevier, New York) (1993). Generally, highly stringent hybridization and washing conditions are chosen to be higher than the pyrolysis temperature (T) of a particular sequence at a specified ionic strength and pH. m The temperature is about 5°C lower.
[0127] An example of highly stringent wash conditions is washing at 72°C with 0.15 M NaCl for approximately 15 minutes. An example of stringent wash conditions is washing at 65°C with 0.2x SSC for 15 minutes (see Sambrook (below) for a description of the SSC buffer). Typically, a low-stringent wash is performed before a highly stringent wash to remove background probe signal. An example of a moderately stringent wash for duplexes with, for example, more than 100 nucleotides is washing at 45°C with 1x SSC for 15 minutes. An example of a low-stringent wash for duplexes with, for example, more than 100 nucleotides is washing at 40°C with 4-6x SSC for 15 minutes. For short probes (e.g., about 10 to 50 nucleotides), stringent conditions typically involve a Na+ ion concentration (or other salt) of less than about 1.5 M (more preferably about 0.01 to 1.0 M) at pH 7.0 to 8.3, and a temperature typically of at least about 30°C, and for long probes (e.g., > 50 nucleotides) at at least about 60°C. Stringent conditions can also be achieved by adding a destabilizing agent (such as formamide). Generally, a signal-to-noise ratio of twice that of an irrelevant probe is observed in a specific hybridization assay, indicating that specific hybridization has been detected. Polynucleotides that do not hybridize to each other under stringent conditions remain substantially identical if they encode substantially the same protein. This occurs, for example, when copies of polynucleotides are created using the maximum codon degeneracy allowed by the genetic code.
[0128] The very strict condition is chosen to be equal to T of a specific probe. mIn Southern or Northern blotting, an example of stringent hybridization conditions for a filter membrane containing complementary nucleic acid sequences with more than 100 complementary residues is: hybridization at 37°C with 50% formamide, for example, in 50% formamide, 1 M NaCl, and 1% SDS, followed by washing at 60°C to 65°C in 0.1x SSC solution. An exemplary low-stringency condition includes hybridization at 37°C with a buffer of 30% to 35% formamide, 1M NaCl, and 1% SDS (sodium dodecyl sulfate), followed by washing at 50°C to 55°C with 1x to 2x SSC (20×SSC = 3.0 M NaCl / 0.3 M trisodium citrate). Exemplary moderately stringent conditions include hybridization at 37°C in 40% to 45% formamide, 1.0 M NaCl, and 1% SDS, followed by washing at 55°C to 60°C with 0.5x to 1x SSC.
[0129] Therefore, certain embodiments of the present invention involve nucleotide sequences that specifically hybridize to or are substantially identical to the polypeptide sequence of the polymerase of the present invention, and nucleotide sequences encoding such polypeptide sequences. The activity of such polymerases can be determined using assays known to those skilled in the art.
[0130] Polymerases in certain embodiments of the present invention include polymerases having substitutions of at least one amino acid residue in the polypeptide. In some embodiments of the present invention, amino acid substitutions falling within the scope of the present invention include those that do not significantly differ in their effect on maintaining: (a) the structure of the peptide backbone within the substituted region, (b) the charge or hydrophobicity of the molecule at the target site, and (c) the volume of the side chain. Naturally occurring residues are classified into the following categories based on common side chain characteristics:
[0131] (1) Hydrophobicity: Leucine, Met, Ala, Val, Leu, Il;
[0132] (2) Neutral hydrophilicity: cys, ser, thr;
[0133] (3) Acidic: asp, glu;
[0134] (4) Alkaline: asn, gin, his, lys, arg;
[0135] (5) Residues affecting chain orientation: gly, pro; and
[0136] (6) Aromatics: trp, tyr, phe.
[0137] Hydrophilicity can also be used to substitute amino acids of the same class. As detailed in US Patent No. 4,554,101, the hydrophilicity values assigned to the amino acid residues are as follows: arginine (+3.0); lysine (+3.0); aspartic acid (+3.0 ± 1); glutamic acid (+3.0 ± 1); serine (+0.3); asparagine (+0.2); glutamine (+0.2); glycine (0); proline (-0.5 ± 1); threonine (-0.4); alanine (-0.5); histidine (-0.5); cysteine (-1.0); methionine (-1.3); valine (-1.5); leucine (-1.8); isoleucine (-1.8); tyrosine (-2.3); phenylalanine (-2.5); tryptophan (-3.4). In such variations, the substitution of amino acids can result in hydrophilicity values ranging from ±2, ±1, or ±0.5.
[0138] In one embodiment of the invention, the polymerase has conserved amino acid substitutions, such as aspartic acid-glutamic acid as acidic amino acids; lysine / arginine / histidine as basic amino acids; leucine / isoleucine, methionine / valine, and alanine / valine as hydrophobic amino acids; and serine / glycine / alanine / threonine as hydrophilic amino acids. The conserved amino acid substitutions also include groupings based on side chains. For example, the group of amino acids with aliphatic side chains is glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic hydroxyl side chains is serine and threonine; the group of amino acids with amide-containing side chains is asparagine and glutamine; the group of amino acids with aromatic side chains is phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains is lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains is cysteine and methionine.
[0139] Exemplary replacements include those in Table 1.
[0140] Table 1
[0141]
[0142] After the introduction of the substitution, those skilled in the art can use their known assays to screen the activity of the resulting polymerase.
[0143] The positions of amino acid residues in DNA polymerases are indicated by numbers or a combination of numbers and letters. The numbering begins at the amino-terminal residue. The letters represent the single-letter amino acid codes for the amino acid residues at the indicated positions in naturally occurring polymerases derived from mutants. Unless otherwise explicitly stated, amino acid residue position names should be interpreted as referring to similar positions in all DNA polymerases, but single-letter amino acid codes specifically refer to the amino acid residues at the indicated positions in Taq DNA polymerases.
[0144] Single substitution mutations are indicated by a combination of letters / numbers / letters. Letters are single-letter codes for amino acid residues. Numbers indicate the amino acid residue sequence position at the mutation site. The numbering system begins at the amino-terminal residue. Residue numbering in Taq DNA polymerase is as described in U.S. Patent No. 5,079,352. The amino acid sequence identity between different DNA polymerases allows for the assignment of corresponding positions to amino acid residues in DNA polymerases other than Taq. Unless otherwise stated, a given number refers to a position in Taq DNA polymerase. The first letter (i.e., the letter to the left of the number) indicates the amino acid residue at the indicated position in the non-mutant polymerase. The second letter indicates the amino acid residue at the same position in the mutant polymerase. For example, the term "R660D" indicates that arginine at position 660 has been replaced by an aspartic acid residue.
[0145] The gene encoding the DNA polymerase has been isolated and sequenced. This sequence information is available through publicly accessible DNA sequence databases such as GENBANK. Compilations of amino acid sequences of DNA polymerases from various organisms can be found in Braithwaite and Ito (1993). This information can be used to design various embodiments of the polymerases of the present invention, as well as the polynucleotides encoding these polymerases. The publicly available sequence information can also be used to clone the gene encoding the DNA polymerase using techniques such as gene library screening with hybridization probes.
[0146] Example 1
[0147] Multiple amplification
[0148] Amplification of 15 targets (99, 131, 160, 199, 251, 300, 345, 400, 516, 613, 735, 908, 1,005, 1,190, and 1,606 bp) was performed using 200 ng of human genomic DNA in 50 μL of a reaction containing 100 nM primers, 1x Platinum™ II PCR buffer, and 4 U polymerase. The cycling protocol was: one cycle at 94°C for 2 min; followed by 35 cycles at 94°C for 15 s, 60°C for 30 s, and 68°C for 96 s. Different extension step times were tested to identify faster polymerase variants: 1. 68°C for 96 s; 2. 68°C for 60 s; 3. 68°C for 30 s. 4.68°C for 10 seconds. Primer sequences are provided in Table 2. The “Mut 4” DNA polymerase carries the mutations G46D, F667Y, and E507K. The “Mut 9” DNA polymerase carries the mutations G46D, F667Y, and D732N. The “Mut 10” DNA polymerase carries the mutations G46D, Y667Y, and E742H. The “Mut 12” DNA polymerase carries the mutations G46D, F667Y, and M747K. The “Mut 15” DNA polymerase carries the mutations G46D, F667Y, and E189K. The “Mut 16” DNA polymerase carries the mutations G46D, F667Y, E189K, and E507K. The “Mut 17” DNA polymerase carries the mutations G46D, F667Y, E189K, and S542N. The “Mut 18” DNA polymerase carries the mutations G46D, F667Y, E189K, and D732N. The “Mut19” DNA polymerase carries the mutations G46D, F667Y, E189K, and E742H. The “Mut21” DNA polymerase carries the mutations G46D and F667Y.
[0149] Table 2
[0150]
[0151] Example 2
[0152] 1.3 KB amplification of λDNA
[0153] A 1.3 kb fragment was amplified from 10 ng λ DNA in 50 μL of a reaction containing 400 nM primers, 1x Platinum™ II PCR buffer, and 4 U polymerase. The cycling protocol was: one cycle at 94°C for 2 min; followed by 25 cycles at 94°C for 15 s, 60°C for 15 s, and 68°C for 60 s. Different extension step times were tested to identify faster polymerase variants: 1. 68°C for 60 s; 2. 68°C for 30 s; 3. 68°C for 15 s; 4. 68°C for 0 s. Primer sequences are provided in Table 3. The “Mut 4” DNA polymerase carries the G46D, F667Y, and E507K mutations. The “Mut 9” DNA polymerase carries the mutations G46D, F667Y, and D732N. The “Mut 10” DNA polymerase carries the mutations G46D, Y667Y, and E742H. The “Mut 12” DNA polymerase carries the mutations G46D, F667Y, and M747K. The “Mut 15” DNA polymerase carries the mutations G46D, F667Y, and E189K. The “Mut 16” DNA polymerase carries the mutations G46D, F667Y, E189K, and E507K. The “Mut 17” DNA polymerase carries the mutations G46D, F667Y, E189K, and S542N. The “Mut 18” DNA polymerase carries the mutations G46D, F667Y, E189K, and D732N. The “Mut19” DNA polymerase carries the mutations G46D, F667Y, E189K, and E742H. The “Mut 21” DNA polymerase carries the mutations G46D and F667Y.
[0154] Table 3.
[0155]
Claims
1. A kind Thermomyces (Taq) DNA polymerase, wherein the Taq DNA polymerase comprises F667Y and G46D mutations and at least one substitution selected from the group consisting of E681I, D732N, A743H, E507N, E742H and M747K, and wherein the Taq DNA polymerase does not retain 5' to 3' exonuclease activity.
2. A Taq DNA polymerase, wherein the Taq DNA polymerase comprises E189K, F667Y and G46D mutations and at least one substitution selected from the group consisting of D732N, E507N, S543N, E742H and M747K, and wherein the Taq DNA polymerase does not retain 5' to 3' exonuclease activity.
3. The DNA polymerase according to any one of claims 1 to 2, wherein it has 5 or fewer of the said substitutions.
4. The DNA polymerase according to any one of claims 1 to 3, wherein it has four or fewer of the said substitutions.
5. The DNA polymerase according to any one of claims 1 to 4, wherein it has three or fewer of the said substitutions.
6. The DNA polymerase according to any one of claims 1 to 5, wherein it has two of the said substitutions.
7. The DNA polymerase according to any one of claims 1 to 6, wherein the DNA polymerase further comprises an E189K substitution.
8. The DNA polymerase according to any one of claims 1 to 7, wherein the DNA polymerase further comprises an S543 substitution.
9. A polynucleotide comprising a sequence encoding a DNA polymerase according to any one of claims 1 to 7.
10. A vector comprising the polynucleotide according to claim 9.
11. The vector of claim 9, further comprising a promoter operatively linked to the polynucleotide.
12. A cell comprising the DNA polymerase according to any one of claims 1 to 8.
13. A cell comprising the polynucleotide according to claim 9.
14. A cell comprising the carrier according to claim 10.
15. A method for synthesizing a polynucleotide in a reaction, comprising contacting a DNA polymerase according to any one of claims 1 to 8 with an initiating template and nucleotides.
16. The method of claim 15, wherein the reaction is a chain termination sequencing reaction.
17. The method of claim 15, wherein the reaction is a polymerase chain reaction.
18. The method of claim 15, wherein the nucleotide comprises a labeled nucleotide.
19. The method of claim 18, wherein the labeled nucleotide is a fluorescently labeled nucleotide.
20. A kit for nucleic acid amplification, comprising a DNA polymerase according to any one of claims 1 to 8, and one or more reagents for DNA synthesis reactions under specific conditions to synthesize novel DNA complementary to target DNA.
21. The kit according to claim 20, further comprising labeled nucleotides.
22. The kit of claim 21, wherein the labeled nucleotide is a fluorescently labeled nucleotide.
23. The kit according to claim 22, further comprising unlabeled nucleotides.
24. The kit according to claim 20, further comprising at least one primer.
25. A method for determining the nucleic acid sequence of a nucleic acid molecule, comprising: (a) Contacting the nucleic acid molecule with a primer, ddNTP, and Taq DNA polymerase according to any one of claims 1 to 8, capable of hybridizing with the nucleic acid molecule. (b) Hybridize the primers into the nucleic acid molecule; (c) Incorporate ddNTPs into the 3' end of the primers to form extended primer products; as well as (d) The nucleic acid sequence of the nucleic acid molecule is determined based on the ddNTP incorporated at the 3' end of the extended primer product.
26. The method of claim 25, wherein the ddNTP is ddATP, ddTTP, ddCTP, ddGTP, ddUTP, derivatives thereof, or a combination thereof.
27. The method of claim 25, wherein the ddNTP is fluorescently labeled.
28. The method of claim 25, wherein the method further comprises a combination of dNTPs, wherein the combination of dNTPs is selected from one or more of dATP, dGTP, dCTP, dTTP, dUTP, dTTP or derivatives thereof.
29. The method of claim 25, wherein the determination comprises separating the extended primer product based on molecular weight and / or capillary electrophoresis.
30. The method of claim 25, wherein the nucleic acid sequence of the nucleic acid molecule is determined by Sanger sequencing.
31. The method of claim 25, wherein the nucleic acid sequence of the nucleic acid molecule is determined by PCR.