Modified polymerases for improved incorporation of nucleotide analogs
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2015-06-25
- Publication Date
- 2026-08-11
Smart Images

Figure CN106536728B_ABST
Abstract
Description
Technical Field
[0001] This application relates to polymerases with modified incorporation for the improvement of nucleotide analogs.
[0002] background
[0003] All organisms rely on DNA polymerases to replicate and maintain their genomes. DNA polymerases enable high-fidelity DNA replication by detecting base complementarity and recognizing additional structural features of the bases. There is a need for polymerases with modified incorporation of nucleotide analogs, particularly nucleotides modified at the 3' sugar hydroxyl group.
[0004] sequence list
[0005] This application is submitted together with an electronic sequence listing. The sequence listing is provided as a file entitled IP1152.TXT, created on May 28, 2014, and is 186 Kb in size. Information from the electronic sequence listing is incorporated herein by reference in its entirety.
[0006] Brief
[0007] This article presents polymerases for improved incorporation of nucleotide analogs, particularly those modified at the 3' sugar hydroxyl group such that the substituent is larger in size than that of naturally occurring 3' hydroxyl nucleotides. The inventors have unexpectedly identified certain modified polymerases that exhibit the desired improved incorporation of analogs and possess many other related advantages.
[0008] In some embodiments, the modified polymerase contains at least one amino acid substitution mutation at the Lys477 position in the amino acid sequence of the 9°N DNA polymerase, which is functionally equivalent to that of the wild-type 9°N DNA polymerase. The amino acid sequence of the wild-type 9°N DNA polymerase is as listed in SEQ ID NO:5. In some embodiments, the substitution mutation includes a mutation to a residue having a smaller side chain. In some embodiments, the substitution mutation includes a mutation to a residue having a hydrophobic side chain. In some embodiments, the substitution mutation includes a mutation to a methionine residue.
[0009] This article also presents a recombinant DNA polymerase comprising at least 60%, 70%, 80%, 90%, 95%, and 99% of the same amino acid sequence as SEQ ID NO:29, wherein the recombinant DNA polymerase is functionally equivalent to the 9°N DNA polymerase containing at position Lys477 in the amino acid sequence.
[0010] This article also presents a recombinant DNA polymerase comprising at least 60%, 70%, 80%, 90%, 95%, and 99% of the same amino acid sequence as SEQ ID NO:31, wherein the recombinant DNA polymerase is functionally equivalent to the 9°N DNA polymerase containing at position Lys477 in the amino acid sequence.
[0011] This article also presents a recombinant DNA polymerase comprising at least 60%, 70%, 80%, 90%, 95%, and 99% of the same amino acid sequence as SEQ ID NO:5, wherein the recombinant DNA polymerase is functionally equivalent to the 9°N DNA polymerase containing at position Lys477 in the amino acid sequence.
[0012] In some embodiments, the polymerase is a DNA polymerase. A modified polymerase, functionally equivalent to a 9°N DNA polymerase, contains at least one amino acid substitution mutation at the Lys477 position in its amino acid sequence, wherein the DNA polymerase is a B-family DNA polymerase. The polymerase may be, for example, a B-family archaea DNA polymerase, human DNA polymerase-α, T4, RB69, and phi29 phage DNA polymerase. In some embodiments, the B-family archaea DNA polymerase is selected from a genus selected from the group consisting of *Thermococcus*, *Pyrococcus*, and *Methanococcus*. For example, the polymerase may be selected from the group consisting of Vent polymerase, Deep Vent polymerase, 9°N polymerase, and Pfu polymerase. In some embodiments, the B-family archaea DNA polymerase is a 9°N polymerase.
[0013] In some embodiments, in addition to the mutations described above, the modified polymerase may also contain substitution mutations at positions functionally equivalent to Leu408 and / or Tyr409 and / or Pro410 in the 9°N DNA polymerase amino acid sequence. For example, substitution mutations may include substitution mutations homologous to Leu408Ala and / or Tyr409Ala and / or Pro410Ile in the 9°N DNA polymerase amino acid sequence.
[0014] In some embodiments, the modified polymerase contains reduced exonuclease activity compared to the wild-type polymerase. For example, in some embodiments, the modified polymerase contains substitution mutations at the positions of Asp141 and / or Glu143 in the amino acid sequence of a 9°N DNA polymerase that are functionally equivalent.
[0015] In some embodiments, the modified polymerase also contains a substitution mutation at the position of Ala485, which is functionally equivalent to the amino acid sequence of the 9°N DNA polymerase. For example, in some embodiments, the polymerase contains a substitution mutation of Ala485Leu or Ala485Val, which is functionally equivalent to the amino acid sequence of the 9°N polymerase.
[0016] In some embodiments, the modified polymerase also contains a substitution mutation at the position of Cys223 in the amino acid sequence of the 9°N DNA polymerase, which is functionally equivalent to this substitution mutation. For example, in some embodiments, the modified polymerase contains a substitution mutation of Cys223Ser in the amino acid sequence of the 9°N polymerase.
[0017] In some embodiments, at least one substitution mutation includes a mutation at a position equivalent to Thr514 and / or Ile521. For example, in some embodiments, the modified polymerase contains substitution mutations of Thr514Ala, Thr514Ser, and / or Ile521Leu that are functionally equivalent to the amino acid sequence of the 9°N polymerase.
[0018] In some embodiments, the modified polymerase may contain additional substitution mutations to remove internal methionine. For example, in some embodiments, the modified polymerase contains a substitution mutation at the position of Met129 in the 9°N DNA polymerase amino acid sequence, which is functionally equivalent to this substitution mutation. In some embodiments, the modified polymerase contains a substitution mutation of Met129Ala in the 9°N polymerase amino acid sequence, which is functionally equivalent to this substitution mutation.
[0019] This document also presents modified polymerases comprising substitution mutations of a semi-conserved domain comprising the amino acid sequence of any one of SEQ ID NO: 1-4, wherein the substitution mutation comprises a mutation at position 3 to any residue other than Lys, Ile, or Gln. In some embodiments, the modified polymerase comprises a mutation to Met at position 3 of any one of SEQ ID NO: 1-4.
[0020] In some embodiments, in addition to the mutations described above, the modified polymerase may also contain substitution mutations at positions functionally equivalent to Leu408 and / or Tyr409 and / or Pro410 in the 9°N DNA polymerase amino acid sequence. For example, substitution mutations may include substitution mutations homologous to Leu408Ala and / or Tyr409Ala and / or Pro410Ile in the 9°N DNA polymerase amino acid sequence.
[0021] In some embodiments, the modified polymerase contains reduced exonuclease activity compared to the wild-type polymerase. For example, in some embodiments, the modified polymerase contains substitution mutations at the positions of Asp141 and / or Glu143 in the amino acid sequence of a 9°N DNA polymerase that are functionally equivalent.
[0022] In some embodiments, the modified polymerase also contains a substitution mutation at the position of Ala485, which is functionally equivalent to the amino acid sequence of the 9°N DNA polymerase. For example, in some embodiments, the polymerase contains a substitution mutation of Ala485Leu or Ala485Val, which is functionally equivalent to the amino acid sequence of the 9°N polymerase.
[0023] In some embodiments, the modified polymerase also includes mutations at positions equivalent to Thr514 and / or Ile521 in the 9°N DNA polymerase amino acid sequence. For example, in some embodiments, the modified polymerase includes substitution mutations functionally equivalent to Thr514Ala, Thr514Ser, and / or Ile521Leu in the 9°N polymerase amino acid sequence.
[0024] In some embodiments, the modified polymerase also contains a substitution mutation at the position of Cys223 in the amino acid sequence of the 9°N DNA polymerase, which is functionally equivalent to this substitution mutation. For example, in some embodiments, the modified polymerase contains a substitution mutation of Cys223Ser in the amino acid sequence of the 9°N polymerase.
[0025] In some embodiments, the modified polymerase may contain additional substitution mutations to remove internal methionine. For example, in some embodiments, the modified polymerase contains a substitution mutation at the position of Met129 in the 9°N DNA polymerase amino acid sequence, which is functionally equivalent to this substitution mutation. In some embodiments, the modified polymerase contains a substitution mutation of Met129Ala in the 9°N polymerase amino acid sequence, which is functionally equivalent to this substitution mutation.
[0026] This article also presents modified polymerases containing any one of the amino acid sequences of SEQ ID NO: 6-8, 10-12, 14-16, 18-20, 22-24, 26-28, 30 and 32.
[0027] This article also presents nucleic acid molecules encoding modified polymerases as defined in any of the more than one implementation schemes. This article also presents expression vectors comprising the nucleic acid molecules described above. This article also presents host cells comprising the vectors described above.
[0028] This document also presents a method for incorporating modified nucleotides into DNA, the method comprising allowing the following components to interact: (i) a modified polymerase according to any of the above embodiments, (ii) a DNA template; and (iii) a nucleotide solution. In some embodiments, the DNA template comprises a clustered array.
[0029] This document also presents a kit for performing a nucleotide incorporation reaction, the kit comprising: a polymerase as defined in any of the embodiments above and a nucleotide solution. In some embodiments, the nucleotide solution comprises a labeled nucleotide. In some embodiments, the nucleotide comprises a synthetic nucleotide. In some embodiments, the nucleotide comprises a modified nucleotide. In some embodiments, the modified nucleotide has been modified at the 3' hydroxyl group such that the substituent is larger in size than the naturally occurring 3' hydroxyl group. In some embodiments, the modified nucleotide comprises a modified nucleotide or nucleoside molecule comprising a purine or pyrimidine base and a ribose or deoxyribose moiety having a removable 3'-OH blocking group covalently attached thereto such that the 3' carbon atom has been attached to the following structure.
[0030] -OZ
[0031] Where Z is any one of -C(R')2-OR”, -C(R')2-N(R”)2, -C(R')2-N(H)R”, -C(R')2-SR”, and -C(R')2-F.
[0032] Each R” is a removable protecting group or part of a removable protecting group;
[0033] Each R' is independently a hydrogen atom, alkyl, substituted alkyl, arylalkyl, alkenyl, alkynyl, aryl, heteroaryl, heterocyclic group, acyl, cyano, alkoxy, aryloxy, heteroaryloxy, or amido, or a detectable marker attached by a linking group; or (R')2 represents an alkylene group of the formula =C(R”')2, wherein each R”' can be the same or different and is selected from the group consisting of hydrogen atoms, halogen atoms, and alkyl groups; and
[0034] The molecules can react to produce an intermediate in which each R' is exchanged for H, or when Z is -C(R')2-F, F is exchanged for OH, SH, or NH2, preferably OH, and the intermediate dissociates under aqueous conditions to provide molecules with free 3'OH;
[0035] The condition is that when Z is -C(R')2-SR”, neither of the two R' groups is H.
[0036] In some embodiments, the R' of the modified nucleotide or nucleoside is an alkyl or substituted alkyl group. In some embodiments, the -Z of the modified nucleotide or nucleoside is of the formula -C(R')2-N3. In some embodiments, Z is an azide methyl group.
[0037] In some embodiments, the modified nucleotide is fluorescently labeled to allow its detection. In some embodiments, the modified nucleotide comprises a nucleotide or nucleoside having a base attached to a detectable marker via a cleavable linker. In some embodiments, the detectable marker includes a fluorescent marker. In some embodiments, the kit also comprises one or more DNA template molecules and / or primers.
[0038] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objectives, and advantages will be apparent from the specification and drawings, as well as from the claims. Brief description of the attached diagram
[0040] Figure 1 This is a schematic diagram showing the alignment of the amino acid sequences of polymerases derived from *Thermococcus* sp. 9°N-7 (9°N), the 9°N polymerase T514S / I521L mutant (Pol957), *Thermococcus gorgonarius* (TGO), *Thermococcus kodakaraensis* (KOD1), *Pyrococcus furiosus* (Pfu), *Methanococcus maripaludis* (MMS2), and RB69 phage DNA polymerase. The numbers shown represent the amino acid residues in the 9°N polymerase.
[0041] Figure 2 It is a display Figure 1 A schematic diagram of the two protruding parts shown in the comparison.
[0042] Detailed description
[0043] This article presents polymerases for improved incorporation of nucleotide analogs, particularly those modified at the 3' sugar hydroxyl group such that the substituent is larger in size than that of naturally occurring 3' hydroxyl nucleotides. The inventors have unexpectedly identified certain modified polymerases that exhibit the desired improved incorporation of analogs and possess many other related advantages.
[0044] As described in more detail below, the inventors have unexpectedly discovered that mutations in one or more residues of the polymerase lead to a significant increase in turnover rate and a reduction in pyrophosphate hydrolysis. These modified polymerases exhibit improved performance in DNA synthesis sequencing (SBS) and result in reduced phasing and / or pre-phasing errors.
[0045] As used herein, the term "phasing" refers to the phenomenon in SBS caused by the incomplete incorporation of nucleotides into portions of the DNA strand within a cluster by a polymerase during a given sequencing cycle. The term "pre-phasing" refers to the phenomenon in SBS caused by the incorporation of nucleotides without a valid 3' terminator, resulting in the incorporation event being advanced by one cycle. The intensity of phasing and pre-phasing for a particular cycle consists of the signal from the current cycle and noise from preceding and following cycles. As the cycle number increases, the proportion of sequences / clusters affected by phasing increases, hindering the recognition of correct bases. Phasing can be caused, for example, by polymerases performing the reverse reaction of nucleotide incorporation, which is known to occur under conditions favorable to pyrophosphate digestion. Therefore, the discovery of modified polymerases that reduce the incidence of phasing and / or pre-phasing was unexpected and has provided significant advantages in SBS applications. For example, the modified polymerase provides faster SBS cycle times, lower phasing and pre-phasing values, and longer sequencing read lengths. The characteristics of the modified polymerases described herein are listed in the following Examples section.
[0046] In some embodiments, substitution mutations include mutations to residues having smaller side chains. The relative sizes of amino acid side chains are well known in the art and can be compared using any known metric, including steric effects and / or electron density. Thus, an example of amino acids listed in ascending order of size would be G, A, S, C, V, T, P, I, L, D, N, E, Q, M, K, H, F, Y, R, W. In some embodiments, substitution mutations include mutations to residues having hydrophobic side chains, such as A, I, L, V, F, W, Y.
[0047] This article also presents polymerases modified with substitution mutations in the semi-conserved domains of the polymerase. As used herein, the term "semi-conserved domain" refers to a polymerase portion that is completely or at least partially conserved across multiple species. It has been unexpectedly found that mutations in one or more residues of the semi-conserved domain, in the presence of 3'-blocked nucleotides, affect polymerase activity, leading to a significant increase in turnover and a reduction in pyrophosphate hydrolysis. These modified polymerases exhibit improved performance in DNA synthesis sequencing and result in reduced phasing errors, as described in the Examples section below.
[0048] In some embodiments, the semi-conserved domain comprises amino acids having the sequence listed in any one of SEQ ID NO: 1-4. SEQ ID NO: 1-4 correspond to residues in the semi-conserved domain in multiple species. SEQ ID NO: 4 corresponds to residues 475-492 of the 9°N DNA polymerase amino acid sequence listed herein as SEQ ID NO: 5. Figure 1 and 2 The comparisons showing the conservation of semi-conserved domains in multiple species are listed. Figure 1 and 2 The polymerase sequences shown were obtained from Genbank database accession numbers Q56366 (9°N DNA polymerase), NP_577941 (Pfu), YP_182414 (KOD1), NP_987500 (MMS2), AAP75958 (RB69), and P56689 (TGo).
[0049] Unexpectedly, mutations to one or more residues in a semi-conserved domain have been found to increase turnover and reduce pyrophosphate hydrolysis, resulting in reduced phasing errors. For example, in some embodiments of the modified polymerases presented herein, substitution mutations include mutations at position 3 of any of SEQ ID NO: 1-4 to any residue other than Lys, Ile, or Gln. In some embodiments, the modified polymerases contain a mutation to Met at position 3 of any of SEQ ID NO: 1-4.
[0050] In some embodiments, the polymerase is a DNA polymerase. In some embodiments, the DNA polymerase is a B-family DNA polymerase. The polymerase can be, for example, a B-family archaea DNA polymerase, human DNA polymerase-α, and a bacteriophage polymerase. Any bacteriophage polymerase can be used in the embodiments presented herein, including, for example, bacteriophage polymerases such as T4, RB69, and phi29 bacteriophage DNA polymerases.
[0051] Family B archaeal DNA polymerases are well known in the art, as exemplified by the disclosure of U.S. Patent No. 8,283,149, which is incorporated herein by reference in its entirety. In some embodiments, the archaeal DNA polymerase is derived from hyperthermophilic archaea, meaning that the polymerase is generally thermostable. Thus, in another preferred embodiment, the polymerase is selected from Vent polymerase, Deep Vent polymerase, 9°N polymerase, and Pfu polymerase. Vent and Deep Vent are commercial names for family B DNA polymerases isolated from the hyperthermophilic archaea *Thermococcus litoralis*. A 9°N polymerase has also been identified from the genus *Thermococcus* sp. Pfu polymerase was isolated from *Thermococcus viridans*.
[0052] In some embodiments, the B-family archaea DNA polymerase comes from genera such as *Thermococcus*, *Pyrococcus*, and *Methanococcus*. Members of the genus *Thermococcus* are well known in the art and include, but are not limited to, *Thermococcus 4557*, *Thermococcus barophilus*, *Thermococcus gammatolerans*, *Thermococcus onnurineus*, *Thermococcus sibiricus*, *Thermococcus kodakarensis*, and *Thermococcus gorgonarius*. Members of the genus *Pyrococcus* are well known in the art, and include, but are not limited to, *Pyrococcus NA2*, *Pyrococcus abyssi*, *Pyrococcus horikoshii*, *Pyrococcus yayanosii*, *Pyrococcus endeavori*, *Pyrococcus glycovorans*, and *Pyrococcus woesei*. Members of the genus *Methanococcus* are well known in the art, and include, but are not limited to, *M. aeolicus*, *M. maripaludis*, *M. vannielii*, *M. voltae*, *M. thermolithotrophicus*, and *M. jannaschii*.
[0053] For example, the polymerase can be selected from the group consisting of Vent polymerase, Deep Vent polymerase, 9°N polymerase, and Pfu polymerase. In some embodiments, the B-family archaea DNA polymerase is a 9°N polymerase.
[0054] Sequence comparison, identity and homology
[0055] In the context of two or more nucleic acid or polypeptide sequences, the term “identical” or “identity percentage” refers to two or more sequences or subsequences that are identical or have a specified percentage of the same amino acid residues or nucleotides when compared and aligned for maximum correspondence, such as using one of the sequence comparison algorithms described below (or other algorithms available to a technician) or measured by visual inspection.
[0056] In the context of two nucleic acids or polypeptides (e.g., DNA encoding a polymerase or the amino acid sequence of a polymerase), the phrase "substantially identical" refers to two or more sequences or subsequences that, when compared and aligned for maximum correspondence, have at least about 60%, about 80%, about 90-95%, about 98%, about 99%, or greater nucleotide or amino acid residue identity, as measured by sequence comparison algorithms or by visual inspection. Such "substantially identical" sequences are generally considered "homologous" rather than referring to an actual ancestor. Preferably, "substantially identical" is present in regions of sequences of at least about 50 residues in length, more preferably in regions of at least about 100 residues, and most preferably, the sequences are substantially identical over at least about 150 residues or over the full length of the two sequences to be compared.
[0057] When a protein and / or protein sequence is naturally or artificially derived from a common ancestral protein or protein sequence, the protein and / or protein sequence is "homologous." Similarly, when a nucleic acid and / or nucleic acid sequence is naturally or artificially derived from a common ancestral nucleic acid or nucleic acid sequence, the nucleic acid and / or nucleic acid sequence is "homologous." Homology is typically inferred from the sequence similarity between two or more nucleic acids or proteins (or their sequences). The precise percentage of similarity between sequences used to establish homology varies depending on the nucleic acid and protein in question, but only 25% sequence similarity for 50, 100, 150, or more residues is routinely used to establish homology. Higher levels of sequence similarity, such as 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% or higher, can also be used. Methods for determining the percentage of sequence similarity (e.g., BLASTP and BLASTN using default parameters) are described herein and are generally available.
[0058] For sequence comparison and homology determination, a reference sequence is typically used, and the test sequence is compared to it. When using a sequence comparison algorithm, the test and reference sequences are input into the computer, and the coordinates of the subsequences are specified if necessary, along with the sequence algorithm program parameters. The sequence comparison algorithm then calculates the percentage of sequence identity between the test sequence and the reference sequence based on the specified program parameters.
[0059] The optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics software package, Genetics Computer Group, 575 Science Dr., Madison, Wis.) or by visual inspection (see broadly Current Protocols in Molecular Biology, Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., 2004 Supplement).
[0060] One example of an algorithm suitable for determining sequence identity percentages and sequence similarity is the BLAST algorithm, described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software used for BLAST analysis is publicly available from the National Center for Biotechnology Information. The algorithm involves first identifying high-scoring sequence pairs (HSPs) of length W in the query sequence, where the short words of length W match or satisfy a positive threshold score T when aligned with words of the same length in a database sequence. T is referred to as the adjacent word score threshold (Altschul et al., ibid.). These initial adjacent word hits act as seeds to initiate a search for longer HSPs containing them. Word hits then extend along each sequence in both directions as far as the cumulative alignment score can no longer be increased. For nucleotide sequences, the cumulative score is calculated using parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatched residues; always <0). For amino acid sequences, a score matrix is used to calculate the cumulative score. The cumulative alignment score decreases by an amount X from its maximum value due to the accumulation of one or more negatively scored residue alignments; when the cumulative score reaches zero or below zero, word hits stop extending in each direction. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the algorithm. The BLASTN program (for nucleotide sequences) uses word length (W) 11, expectation (E) 10, truncation 100, M=5, N=-4, and comparison of the two strands as defaults. For amino acid sequences, the BLAST program uses word length (W) 3, expectation (E) 10, and a BLOSUM62 score matrix (see Henikoff & Henikoff (1989) Proc. Natl. Acad. Sci. USA 89:10915) as defaults.
[0061] In addition to calculating the percentage of sequence identity, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, for example, Karlin & Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993)). One similarity measure provided by the BLAST algorithm is the minimum total probability (P(N)), which provides an indication of the probability that a match between two nucleotide or amino acid sequences will occur by chance. For example, if the minimum total probability in a comparison of the test nucleic acid with a reference nucleic acid is less than about 0.1, more preferably less than about 0.01, and most preferably less than about 0.001, the nucleic acid is considered similar to the reference sequence.
[0062] In studies using entirely different polymerases, "functionally equivalent" means that the control polymerase will contain amino acid substitutions at positions where they are believed to occur in other polymerases and have the same function. For example, a mutation from tyrosine to valine at position 412 (Y412V) in Vent DNA polymerase is functionally equivalent to a mutation from tyrosine to valine at position 409 (Y409V) in 9°N polymerase.
[0063] Typically, functionally equivalent substitution mutations in two or more different polymerases occur at homologous amino acid positions in the polymerase's amino acid sequence. Therefore, the use of the term "functionally equivalent" herein also includes mutations that are "positionally equivalent" or "homologous" to a given mutation, regardless of whether the specific function of the mutated amino acid is known. It is possible to identify positionally equivalent or homologous amino acid residues in the amino acid sequences of two or more different polymerases based on sequence alignment and / or molecular modeling. Examples of sequence alignment used to identify positionally equivalent and / or functionally equivalent residues are listed below. Figure 1 and 2 In the middle. Therefore, for example, such as Figure 2 As shown, residues in the semi-conserved domain are identified as positions 475-492 of the 9°N DNA polymerase amino acid sequence. Corresponding residues in the TGO, KOD1, Pfu, MmS2, and RB69 polymerases are identified as vertically aligned in the figure and are considered to be positionally and functionally equivalent to their corresponding residues in the 9°N DNA polymerase amino acid sequence.
[0064] The modified polymerase described above may include additional substitution mutations known to enhance polymerase activity in the presence of a 3'-blocked nucleotide and / or in DNA sequencing applications. For example, in some embodiments, in addition to any of the above mutations, the modified polymerase may also include substitution mutations at positions Leu408 and / or Tyr409 and / or Pro410 in the amino acid sequence of a 9°N DNA polymerase. Any of a plurality of substitution mutations at positions 408-410 in the amino acid sequence of a 9°N DNA polymerase, each of which results in the incorporation of an increased blocking nucleotide, as known in the art and exemplified by the disclosures of US 2006 / 0240439 and US 2006 / 0281109, each of which is incorporated herein by reference in its entirety. For example, substitution mutations may include substitution mutations homologous to Leu408Ala and / or Tyr409Ala and / or Pro410Ile in the 9°N DNA polymerase amino acid sequence. In some embodiments, in addition to any of the above mutations, the modified polymerase also contains a substitution mutation at a position functionally equivalent to Ala485 in the 9°N DNA polymerase amino acid sequence. For example, in some embodiments, the polymerase contains a substitution mutation functionally equivalent to Ala485Leu or Ala485Val in the 9°N polymerase amino acid sequence.
[0065] In some embodiments, in addition to any of the mutations described above, the modified polymerase may contain reduced exonuclease activity compared to the wild-type polymerase. Any of a plurality of substitution mutations at one or more positions known to result in reduced exonuclease activity may be performed, as known in the art and exemplified by the incorporated materials of US 2006 / 0240439 and US 2006 / 0281109. For example, in some embodiments, in addition to the mutations described above, the modified polymerase may also contain substitution mutations at the positions of Asp141 and / or Glu143 in the amino acid sequence of a 9°N DNA polymerase that are functionally equivalent.
[0066] In some embodiments, in addition to any of the mutations described above, the modified polymerase also contains a substitution mutation at the position of Cys223 in the amino acid sequence of the 9°N DNA polymerase, which is functionally equivalent to that of the 9°N DNA polymerase, as known in the art and exemplified by the incorporated material of US 2006 / 0281109. For example, in some embodiments, the modified polymerase contains a substitution mutation of Cys223Ser in the amino acid sequence of the 9°N polymerase.
[0067] In some embodiments, in addition to any of the mutations described above, the modified polymerase may contain one or more mutations at positions equivalent to Thr514 and / or Ile521 in the 9°N DNA polymerase amino acid sequence, as known in the art and exemplified by the disclosure of PCT / US2013 / 031694, which is incorporated herein by reference in its entirety. For example, in some embodiments, the modified polymerase contains substitution mutations functionally equivalent to Thr514Ala, Thr514Ser, and / or Ile521Leu in the 9°N polymerase amino acid sequence.
[0068] In some embodiments, in addition to any of the mutations described above, the modified polymerase may contain one or more mutations at the position of Arg713 in the amino acid sequence of the 9°N DNA polymerase, as known in the art and exemplified by the disclosure of U.S. Patent No. 8,623,628, which is incorporated herein by reference in its entirety. For example, in some embodiments, the modified polymerase contains substitution mutations that are functionally equivalent to Arg713Gly, Arg713Met, or Arg713Ala in the amino acid sequence of the 9°N polymerase.
[0069] In some embodiments, in addition to any of the mutations described above, the modified polymerase may contain one or more mutations at positions equivalent to Arg743 and / or Lys705 in the 9°N DNA polymerase amino acid sequence, as known in the art and exemplified by the disclosure of U.S. Patent No. 8,623,628, which is incorporated herein by reference in its entirety. For example, in some embodiments, the modified polymerase contains substitution mutations that are functionally equivalent to Arg743Ala and / or Lys705Ala in the 9°N polymerase amino acid sequence.
[0070] In some embodiments, in addition to any of the mutations described above, the modified polymerase may contain one or more additional substitution mutations to remove the internal methionine. For example, in some embodiments, the modified polymerase contains a substitution mutation to a different amino acid at the position of Met129 in the amino acid sequence of the 9°N DNA polymerase, which is functionally equivalent to this. In some embodiments, the modified polymerase contains a substitution mutation of Met129Ala in the amino acid sequence of the 9°N polymerase.
[0071] Mutant polymerase
[0072] For example, based on polymerase models and model predictions as discussed above, or using random or semi-random mutagenesis methods, various types of mutagenesis are optionally used in this disclosure, for example, to modify polymerases to produce variants. Generally, any available mutagenesis procedure can be used to prepare polymerase mutants. Such mutagenesis procedures optionally include selecting mutated nucleic acids and peptides for one or more activities of interest (e.g., for a given nucleotide analogue, such as reduced pyrophosphate hydrolysis, increased turnover). Procedures that can be used include, but are not limited to: site-directed mutagenesis, random site-directed mutagenesis, in vitro or in vivo homologous recombination (overlap PCR of DNA rearrangement and combination), mutagenesis using a template containing uracil, oligonucleotide-directed mutagenesis, phosphate-thiolated DNA mutagenesis, mutagenesis using nicked dimer DNA, site-directed mismatch repair, mutagenesis using repair-deficient host strains, restriction selection and restriction purification, deletion mutagenesis, mutagenesis by whole-genome synthesis, degenerate PCR, double-strand break repair, and many other procedures known to those skilled in the art. The initiating polymerase used for mutation can be any of the polymerases mentioned herein, including available polymerase mutants such as those identified, for example, in US2006 / 0240439 and US 2006 / 0281109, each of which is incorporated herein by reference in its entirety.
[0073] Optionally, the mutagenesis may be guided by information from naturally occurring polymerase molecules or known modified or mutated polymerases (e.g., using existing mutant polymerases as mentioned in the preceding references), such as the sequences discussed above, sequence comparisons, physical properties, crystal structures, etc. However, in another type of implementation, the modification may be substantially random (e.g., as in classical or “family” DNA rearrangements, see, for example, Crameri et al. (1998) "DNA shuffling of a family of genes from diverse species accelerates directed evolution" Nature 391:288-291).
[0074] Further information on the mutant forms can be found in: Sambrook et al., Molecular Cloning—A Laboratory Manual (3rd Edition), Volumes 1–3, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, 2000 (“Sambrook”); Current Protocols in Molecular Biology, eds., Ausubel et al., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc. (2011 Supplement) (“Ausubel”); and PCR Protocols: A Guide to Methods and Applications (eds., Innis et al.), Academic Press Inc., San Diego, Calif. (1990) (“Innis”). Further details of the mutation forms are provided in the following publications and references cited: Arnold, Protein engineering for unusual environments, Current Opinion in Biotechnology 4:450-455 (1993); Bass et al., Mutant Trp repressors with new DNA-binding specificities, Science 242:240-245 (1988); Bordo and Argos (1991) Suggestions for "Safe" Residue Substitutions in Site-directed Mutagenesis 217:721-729; Botstein & Shortle, Strategies and applications of in vitro mutagenesis, Science 229:1193-1201 (1985); Carter et al., Improved oligonucleotide site-directed mutagenesis using M13 vectors, Nucl. Acids Res.13:4431-4443(1985); Carter, Site-directed mutagenesis, Biochem.J.237:1-7(1986); Carter, Improved oligonucleotide-directed mutagenesis using M13 vectors, Methods in Enzymol. 154:382-403(1987); Dale et al., Oligonucleotide-directed random mutagenesis using the phosphorothioate method, Methods Mol. Biol. 57:369-374(1996); Eghtedarzadeh & Henikoff, Use of oligonucleotides to generate large deletions, Nucl. Acids Res. 14:5115(1986); Fritz et al., Oligonucleotide-directed construction of mutations: a gapped duplex DNA procedure without enzymatic reactions in vitro, Nucl. Acids Res. 16:6987-6999(1988); Grundstrom et al., Oligonucleotide-directed mutagenesis by microscale `shot-gun` gene synthesis, Nucl. Acids Res. 13:3305-3316(1985); Hayes(2002) Combining Computational and Experimental Screening for rapid Optimization of Protein Properties PNAS99(25)15926-15931; Kunkel, The efficiency of oligonucleotide directed mutagenesis, in Nucleic Acids & Molecular Biology (Eckstein, F. and Lilley, D.M.J. eds, Springer Verlag, Berlin))(1987); Kunkel, Rapid and efficient site-specific mutagenesis without phenotypic selection, Proc. Natl. Acad.Sci.USA 82:488-492(1985); Kunkel et al., Rapid and efficient site-specific mutagenesis without phenotypic selection, Methods in Enzymol. 154, 367-382(1987); Kramer et al., The gapped duplex DNA approach to oligonucleotide-directed mutation construction, Nucl. Acids Res. 12:9441-9456(1984); Kramer & Fritz Oligonucleotide-directed construction of mutations via gapped duplex DNA, Methods in Enzymol. 154:35-367(1987); Kramer et al., Point Mismatch Repair, Cell 38:879-887(1984); Kramer et al., Improved enzymatic in vitro reactions in the gapped duplex DNA approach to oligonucleotide-directed construction of mutations, Nucl. Acids Res. 16:7207(1988); Ling et al., Approaches to DNA mutagenesis: an overview, Anal Biochem. 254(2):157-178(1997); Lorimer and Pastan Nucleic Acids Res. 23, 3067-8(1995); Mandecki, Oligonucleotide-directed double-strand break repair in plasmids of Escherichia coli: a method for site-specific mutagenesis, Proc. Natl. Acad. Sci.USA, 83:7177 - 7181(1986); Nakamaye & Eckstein, Inhibition of restriction endonuclease Nci I cleavage by phosphorothioate groups and its application to oligonucleotide - directed mutagenesis, Nucl. Acids Res. 14:9679 - 9698(1986); Nambiar et al., Total synthesis and cloning of a gene coding for the ribonuclease S protein, Science 223:1299 - 1301(1984); Sakamar and Khorana, Total synthesis and expression of a gene for the a - subunit of bovine rod outer segment guanine nucleotide - binding protein (transducin), Nucl. Acids Res. 14:6361 - 6372(1988); Sayers et al., Y - T Exonucleases in phosphorothioate - based oligonucleotide - directed mutagenesis, Nucl. Acids Res. 16:791 - 802(1988); Sayers et al., Strand specific cleavage of phosphorothioate - containing DNA by reaction with restriction endonucleases in the presence of ethidium bromide, (1988) Nucl. Acids Res. 16:803 - 814; Sieber, et al., Nature Biotechnology, 19:456 - 460(2001); Smith, In vitro mutagenesis, Ann. Rev. Genet. 19:423 - 462(1985); Methods in Enzymol. 100:468 - 500(1983); Methods in Enzymol.154:329 - 350(1987); Stemmer, Nature 370, 389 - 91(1994); Taylor et al., The use of phosphorothioate - modified DNA in restriction enzyme reactions to prepare nicked DNA, Nucl. Acids Res. 13:8749 - 8764(1985); Taylor et al., The rapid generation of oligonucleotide - directed mutations at high frequency using phosphorothioate - modified DNA, Nucl. Acids Res. 13:8765 - 8787(1985); Wells et al., Importance of hydrogen - bond formation in stabilizing the transition state of subtilisin, Phil. Trans. R. Soc. Lond. A 317:415 - 423(1986); Wells et al., Cassette mutagenesis: an efficient method for generation of multiple mutations at defined sites, Gene 34:315 - 323(1985); Zoller & Smith, Oligonucleotide - directed mutagenesis using M 13 - derived vectors: an efficient and general procedure for the production of point mutations in any DNA fragment, Nucleic Acids Res. 10:6487 - 6500(1982); Zoller & Smith, Oligonucleotide - directed mutagenesis of DNA fragments cloned into M13 vectors, Methods in Enzymol.100:468-500(1983); Zoller&Smith, Oligonucleotide-directed mutagenesis: asimple method using two oligonucleotide primers and a single-stranded DNAtemplate, Methods in Enzymol.154:329-350(1987); Clackson et al. (1991) "Makingantibody fragments using phage display libraries" Nature 352:624-628; Gibbs et al. (2001) "Degenerate oligonucleotide gene shuffling (DOGS): a method for enhancing the frequency of recombination with family shuffling" Gene 271:13-20; and Hiraga and Arnold (2003) "General method for sequence-independent Site-directed chimeragenesis: J. Mol. Biol. 330:287-296 Available at Methods in. Further detailed descriptions of many of the above methods can be found in Volume 154 of Enzymology, which also describes useful comparisons for solving fault problems using various mutation methods.
[0075] Preparation and isolation of recombinant polymerase
[0076] Typically, nucleic acids encoding polymerases as presented herein can be prepared by cloning, recombination, in vitro synthesis, in vitro amplification, and / or other available methods. Various recombination methods can be used to express expression vectors encoding polymerases as presented herein. Methods for preparing recombinant nucleic acids, expressing them, and isolating the expressed products are well known in the art and are well described. Numerous exemplary mutations and combinations of mutations, as well as strategies for designing desired mutations, are described herein. Methods for generating and selecting mutations at the active site of the polymerase, including methods for modifying stereofeatures in or near the active site to allow improved entry of nucleotide analogs, are found above and, for example, WO 2007 / 076057 and PCT / US2007 / 022459, which are incorporated herein by reference in their entirety.
[0077] Other useful references for mutation, recombination, and in vitro nucleic acid manipulation methods (including cloning, expression, PCR, etc.) include: Berger and Kimmel, Guide to Molecular Cloning Techniques, Methods in Enzymology, Vol. 152, Academic Press, Inc., San Diego, Calif. (Berger); Kaufman et al. (2003), Handbook of Molecular and Cellular Methods in Biology and Medicine, 2nd Edition, edited by Ceske, CRC Press (Kaufman); and The Nucleic Acid Protocols Handbook, edited by Ralph Rapley (2000), Cold Spring Harbor, Humana Press Inc. (Rapley); Chen et al. (edited), PCR Cloning Protocols, 2nd Edition (Methods in Molecular Biology, Vol. 192), Humana Press; and Viljoen et al. (2005), Molecular Diagnostic PCR Handbook. Springer, ISBN 1402034032.
[0078] In addition, many kits are commercially available for purifying plasmids or other related nucleic acids from cells (see, for example, EasyPrep.TM. and FlexiPrep.TM. from Pharmacia Biotech; StrataClean.TM. from Stratagene; and QIAprep.TM. from Qiagen). Any isolated and / or purified nucleic acids can be further manipulated to produce other nucleic acids, for transfecting cells, incorporated into relevant vectors to infect organisms for expression, etc. Typical cloning vectors contain transcription and translation terminators, transcription and translation initiation sequences, and promoters for regulating the expression of specific target nucleic acids. Vectors optionally contain a universal expression cassette containing at least one independent terminator sequence, a sequence that allows the cassette to replicate in eukaryotes or prokaryotes or both (e.g., shuttle vectors), and selection markers for both prokaryotic and eukaryotic systems. Vectors are adapted for replication and integration in prokaryotes, eukaryotes, or both.
[0079] Other useful references for cell isolation and culture (e.g., for subsequent nucleic acid isolation) include Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, 3rd ed., Wiley-Liss, New York, and references cited therein; Payne et al. (1992) Plant Cell and Tissue Culture in Liquid Systems, John Wiley & Sons, Inc., New York, NY; Gamborg and Phillips (eds.) (1995) Plant Cell, Tissue and Organ Culture; Fundamental Methods Springer Lab Manual, Springer-Verlag (Berlin Heidelberg, New York); and Atlas and Parks (eds.) The Handbook of Microbiological Media (1993) CRC Press, Boca Raton, Fla.
[0080] The nucleic acid encoding the recombinant polymerase disclosed herein is also a feature of the embodiments presented herein. Specific amino acids can be encoded by multiple codons, and certain translation systems (e.g., prokaryotic or eukaryotic cells) tend to exhibit codon bias; for example, different organisms often prefer one of several synonymous codons encoding the same amino acid. Therefore, the nucleic acids presented herein are optionally "codon-optimized," meaning that the nucleic acid is synthesized to contain codons preferred by the specific translation system used to express the polymerase. For example, when it is desired to express a polymerase in a bacterial cell (or even a specific strain of bacteria), the nucleic acid can be synthesized to contain the codons most frequently found in the genome of that bacterial cell for efficient polymerase expression. A similar strategy can be used when it is desired to express a polymerase in a eukaryotic cell; for example, the nucleic acid can contain codons preferred by that eukaryotic cell.
[0081] Various protein isolation and detection methods are known and can be used, for example, to isolate polymerases from recombinant cultures of cells expressing the recombinant polymerases presented herein. Various protein isolation and detection methods are well known in the art, including, for example, those listed in the following references: R. Scopes, Protein Purification, Springer-Verlag, NY (1982); Deutscher, Methods in Enzymology, Vol. 182: Guide to Protein Purification, Academic Press, Inc., NY (1990); Sandana (1997) Bioseparation of Proteins, Academic Press, Inc.; Bollag et al. (1996) Protein Methods, Second Supplement, Wiley-Liss, NY; Walker (1996) The Protein Protocols Handbook, Humana Press, NJ; Harris and Angal (1990) Protein Purification Applications: A Practical Approach, IRL Press at Oxford, Oxford, England; Harris and Angal Protein Purification Methods: A Practical Approach, IRL Press at Oxford, Oxford, England; Scopes (1993) Protein Purification: Principles and Practice, 3rd Supplement, Springer Verlag, NY; Janson and Ryden (1998) Protein Purification: Principles, High Resolution Methods and Applications, 2nd Edition, Wiley-VCH, NY; and Walker (1998) Protein Protocols on CD-ROM, Humana Press, NJ; and the references cited therein. Further detailed information on protein purification and detection methods can be found in Satinder Ahuja, ed., Handbook of Bioseparations, Academic Press (2000).
[0082] How to use
[0083] The modified polymerase presented herein can be used in sequencing procedures such as sequencing-by-synthesis (SBS) technology. In short, SBS can be initiated by contacting a target nucleic acid with one or more labeled nucleotides, a DNA polymerase, etc. When the target nucleic acid is used as a template to extend primers, those features are incorporated into the detectable labeled nucleotides. Optionally, the labeled nucleotides may also include a reversible termination feature that terminates further primer extension after the nucleotide has been added to the primer. For example, nucleotide analogs with a reversible termination motif can be added to the primer so that subsequent extension cannot occur until a desealing agent is delivered to remove the motif. Thus, for embodiments using reversible termination, a desealing agent can be delivered to a flow cell (before or after detection). Washing can be performed between multiple delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS programs, fluid systems, and detection platforms that can be readily adapted for use with arrays generated by the methods of this disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Patent Nos. 7,057,026, 7,329,492, 7,211,414, 7,315,019, or 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082A1, each of which is incorporated herein by reference.
[0084] Other sequencing procedures that utilize cyclic reactions, such as pyrosequencing, can be used. Pyrosequencing detects the release of inorganic pyrosequencing (PPi) as specific nucleotides are incorporated into the nascent nucleic acid chain (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al. Science). U.S. Patent Nos. 281 (5375), 363 (1998); 6,210,891, 6,258,568, and 6,274,320, each of which is incorporated herein by reference. In pyrosequencing, the released PPi can be detected by conversion to adenosine triphosphate (ATP) by ATP sulfatase, and the resulting ATP can be detected by photons generated by luciferase. Thus, the sequencing reaction can be monitored via a luminescent detection system. The excitation radiation source for the fluorescence-based detection system is not required for the pyrosequencing procedure. Useful fluidic systems, detectors, and procedures that can be used for array applications of pyrosequencing to the present disclosure are described, for example, in WIPO Patent Application Serial No. PCT / US11 / 57111, U.S. Patent Application Publication No. 2005 / 0191698 A1, U.S. Patent Nos. 7,595,883, and 7,244,559, each of which is incorporated herein by reference.
[0085] Some implementations may utilize methods including real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between a fluorophore-loaded polymerase and a γ-phosphate-labeled nucleotide, or using a zero-mode waveguide. Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al., Science 299, 682–686 (2003); Lundquist et al., Opt. Lett. 33, 1026–1028 (2008); Korlach et al., Proc. Natl. Acad. Sci. USA 105, 1176–1181 (2008), the contents of which are incorporated herein by reference.
[0086] Some SBS implementations include detecting protons released after nucleotide incorporation into the extension product. For example, sequencing based on the detection of released protons can use electronic detectors and related technologies commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary), or sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082A1, 2009 / 0127589A1, 2010 / 0137143A1, or 2010 / 0282617A1, each of which is incorporated herein by reference.
[0087] Accordingly, this document presents a method for incorporating nucleotide analogs into DNA, the method comprising allowing the following components to interact: (i) a modified polymerase according to any of the above embodiments, (ii) a DNA template; and (iii) a nucleotide solution. In some embodiments, the DNA template comprises a cluster array. In some embodiments, the nucleotide is modified at the 3' sugar hydroxyl group, and the 3' sugar hydroxyl group contains a modification such that the substituent is larger in size than the naturally occurring 3' hydroxyl group.
[0088] Nucleic acids encoding modified nucleotides
[0089] This article also presents the nucleic acid molecule encoding the modified polymerase presented herein. For any given modified polymerase, where the amino acid sequence, and preferably the wild-type nucleotide sequence encoding the polymerase, is known, and which is a mutant form, it is possible to obtain the nucleotide sequence encoding the mutant based on fundamental principles of molecular biology. For example, given that the wild-type nucleotide sequence encoding the 9°N polymerase is known, it is possible to deduce the nucleotide sequence encoding any given mutant form of 9°N with one or more amino acid substitutions using the standard genetic code. Similarly, for other polymerases such as, for example, Vent... TM Mutations of Pfu, Tsp, JDF-3, Taq, etc., can be used to readily obtain nucleotide sequences. Then, nucleic acid molecules with the desired nucleotide sequences can be constructed using standard molecular biology techniques known in the art.
[0090] According to the embodiments presented herein, the defined nucleic acid includes not only the same nucleic acid but also any small base variation, which specifically includes substitutions in conserved amino acid substitutions where degenerate codons result in synonymous codons (different codons specifying the same amino acid residue). The term "nucleic acid sequence" also includes the complementary sequence of any single-stranded sequence given with respect to base variation.
[0091] The nucleic acid molecules described herein can also be advantageously contained in suitable expression vectors to express polymerase proteins encoded by said expression vectors in a suitable host. Incorporation of cloned DNA into suitable expression vectors for subsequent transformation of said cells and selection of subsequently transformed cells is well known to those skilled in the art, as provided in Sambrook et al. (1989), Molecular cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, which is incorporated herein by reference in its entirety.
[0092] Such expression vectors include vectors having nucleic acids according to embodiments presented herein, said nucleic acids being operatively linked to regulatory sequences, such as promoter regions, capable of influencing the expression of said DNA fragments. The term "operatively linked" refers to a juxtaposition in which the components described herein are in a relationship that allows them to function in their intended manner. Such vectors can be transformed into suitable host cells to provide expression of proteins according to embodiments presented herein.
[0093] Nucleic acid molecules can encode mature proteins or proteins having a presequence, said presequence including a presequence encoding a leader sequence on a precursor protein, said leader sequence being cleaved by a host cell to form the mature protein. The vector can be, for example, a plasmid vector, a viral vector, or a phage vector having an origin of replication and, optionally, a promoter for the expression of said nucleotides and, optionally, a regulator of said promoter. The vector can contain one or more optional markers, such as, for example, an antibiotic resistance gene.
[0094] The regulatory elements required for expression include: a promoter sequence that binds to RNA polymerase and directs the appropriate level of transcription initiation; and a translation initiation sequence for ribosome binding. For example, bacterial expression vectors may include: a promoter, such as the lac promoter; a Shine-Dalgarno sequence for translation initiation; and a start codon AUG. Similarly, eukaryotic expression vectors may include a heterologous or homologous promoter for RNA polymerase II, a downstream polyadenylation signal, a start codon AUG, and a stop codon for ribosome dissociation. Such vectors are commercially available or can be assembled from the described sequences using methods well known in the art.
[0095] Transcription of DNA encoding polymerases by higher eukaryotes can be optimized by including enhancer sequences in vectors. Enhancers are cis-acting DNA elements that function on promoters to increase the level of transcription. In addition to optional markers, vectors typically also contain origins of replication.
[0096] Example 1
[0097] General determination methods and conditions
[0098] The following paragraphs describe the general measurement conditions used in the examples presented below.
[0099] 1. Gel-based determination
[0100] This section describes the gel-based assay used in the examples below for monitoring the pyrophosphate hydrolysis activity of polymerases.
[0101] In summary, the pyrophosphate hydrolytic activity of an exemplary modified polymerase was measured by mixing 300 nM of enzyme and 100 nM of double-stranded primer-template DNA with different concentrations (0 mM, 0.125 mM, 0.25 mM, 0.5 mM, 1 mM, 2 mM, and 4 mM) of sodium pyrophosphate in a reaction buffer containing: 50 mM Tris-HCl (pH 9.0), 50 mM NaCl, 1 mM EDTA, 6 mM MgSO4, and 0.05% (v / v) Tween-20. The reaction was carried out at 55 °C for 1 minute and terminated by adding an equal volume of a 2X quenching solution containing 0.025% bromophenol blue, 30 mM EDTA, and 95% deionized formamide.
[0102] The reaction products were denatured at 95°C for 5 min and resolved by 15% 7M urea-polyacrylamide gel electrophoresis (urea-PAGE). The results were visualized by scanning with a GE Healthcare Typhoon 8000 PhosphorImager.
[0103] Double-stranded primer-template duplexes are formed by annealing the following oligonucleotides:
[0104] Primer: 5′- GCTTGCACAGGTGCGTTCGT* -3′
[0105] Template: 5′-CGTTAGTCCA CGAACGCACCTGTGCAAGC -3′
[0106] The primer contains 6-carboxytetramethylrhodamine (TAMRA) fluorescent dye linked to the 5' end of the oligonucleotide. The final "T*" contains a 3'-O-azidomethyl blocking moiety on the nucleotide. Lanes showing degradation of the labeled primer indicate enhanced pyrophosphate digestion (the reverse reaction of DNA polymerization) compared to the control.
[0107] 2. Cloning and expression of polymerases
[0108] This section describes methods for cloning and expressing various polymerase mutants used in the examples below.
[0109] Mutagenesis was performed on the gene encoding the backbone sequence of the polymerase using standard site-directed mutagenesis methods. For each mutation performed, the correct sequence of the mutated gene was confirmed by sequencing the cloned gene.
[0110] The polymerase gene was subcloned into the pET11a vector and transformed into BL21Star(DE3) expression cells from Invitrogen. Transformed cells were cultured at 37°C in 2.8 L Fernbock flasks until an OD600 of 0.8 was achieved. Protein expression was then induced by the addition of 1 mM IPTG, followed by an additional 3 hours of growth. The culture was then centrifuged at 7000 rpm for 20 minutes. Cell clumps were stored at -20°C until purification.
[0111] Bacterial cell lysis was performed by resuspending frozen cultures in 10x w / v lysis buffer (Tris pH 7.5, 500 mM NaCl, 1 mM EDTA, 1 mM DTT). An EDTA-free protease inhibitor (Roche) was added to the resuspended cell clumps. All lysis and purification steps were performed at 4°C. The resuspended cultures were passed through a microfluidizer four times to complete cell lysis. The lysates were then centrifuged at 20,000 rpm for 20 minutes to remove cell debris. Polyethyleneimine (final concentration 0.5%) was slowly added to the supernatant, and the mixture was stirred for 45 minutes to precipitate bacterial nucleic acids. The lysates were then centrifuged at 20,000 rpm for 20 minutes; the clumps were discarded. The lysates were then precipitated with ammonium sulfate using two volumes of cold-saturated (NH4)2SO4 in sterile dH2O. The precipitated proteins were centrifuged at 20,000 rpm for 20 minutes. The protein clumps were resuspended in 250 mL of buffer A (50 mM Tris pH 7.5, 50 mM KCl, 0.1 mM EDTA, 1 mM DTT). The resuspended lysates were then purified using a 5 mL SP FastFlow (GE) column pre-equilibrated in buffer A. A 50 mL gradient elution column was used, from 0.1 M to 1 M KCl. Peak fractions were pooled and diluted with buffer C (Tris pH 7.5, 0.1 mM EDTA, 1 mM DTT) until the conductivity was equal to that of buffer D (Tris pH 7.5, 50 mM KCl, 0.1 mM EDTA, 1 mM DTT). The pooled fractions were then loaded onto a 5 mL HiTrap Heparin Fastflow column. Polymerase was then eluted using a 100 mL gradient from 50 mM to 1 M KCl. The peak fractions were collected, dialyzed into storage buffer (20 mM Tris pH 7.5, 300 mM KCl, 0.1 mM EDTA, 50% glycerol) and frozen at -80°C.
[0112] 3. Phasing / Pre-phasing analysis
[0113] This section describes methods for analyzing the performance of polymerase mutants used in the following examples in synthetic sequencing assays.
[0114] Short 12-cycle sequencing experiments were used to generate phasing and pre-phasing values. Experiments were performed on an Illumina Genome Analyzer system (Illumina, Inc., San Diego, CA) converted to run MiSeq Fast chemistry, following the manufacturer's instructions. For example, for each polymerase, a separate incorporation mixture (IMX) was generated, and 4 x 12 cycles were run for each IMX using a different location. Standard MiSeq reagent formulations were used, with the polymerase being tested substituted for the standard polymerase. DNA libraries were prepared from PhiX genomic DNA (provided with Illumina reagents as a control) according to the standard TruSeq HT protocol. Phasing and pre-phasing levels were evaluated using Illumina RTA software.
[0115] Example 2
[0116] Identification and screening of 9°N polymerase mutants for phasing / pre-phasing
[0117] Saturation mutagenesis screening was performed on residues in the 3' block pocket. Mutations to the two modified 9°N polymerase backbone sequences (SEQ ID NO: 29 and 31) were generated, cloned, expressed, and purified as generally described in Example 1.
[0118] Using the gel-based assay described above in Example 1, purified mutant polymerases were screened for burst kinetics and compared with control polymerases having the sequences listed in SEQ ID NO: 29 and 31. Among those mutants screened, a group of mutants including the following mutants were further screened for phasing / pre-phasing activity as described above in Example 1.
[0119] The screening results are summarized in the table below. As shown in the table, each of the above mutants showed unexpected and significant improvements in one or more aspects of phasing and pre-phasing compared to the control polymerases Pol957 or Pol955.
[0120] Mutation (name) SEQ ID NO: Is phasing reduced compared to the control? Comparison 29 - K477M 30 yes Comparison 31 - K477M 32 yes
[0121] Example 3
[0122] Screening for mutants of 9°N WT polymerase
[0123] As generally described in Example 1, the generation, cloning, expression, and purification of a mutation in the wild-type polymerase backbone sequence (SEQ ID NO:5) of the genus Thermococcus sp. 9°N-7 (9°N) produced polymerases having the amino acid sequences listed in SEQ ID NO:6-8.
[0124] The purified mutant polymerases were screened for burst kinetics using the gel-based assay described above in Example 1, and compared with a control polymerase having the sequence listed in SEQ ID NO:5. Of those mutants screened, a group was further screened for phasing / pre-phasing activity as generally described above in Example 1. Those polymerases having the following mutations showed improved phasing and / or pre-phasing activity compared to the control.
[0125]
[0126]
[0127] Example 4
[0128] Screening 9°N Exo - polymerase mutant
[0129] As generally described in Example 1, the generation, cloning, expression, and purification of 9°NExo - Mutations in the polymerase backbone sequence (SEQ ID NO:9) resulted in polymerases having amino acid sequences as listed in SEQ ID NO:10-12.
[0130] The purified mutant polymerases were screened for burst kinetics using the gel-based assays described above in Example 1, and compared with control polymerases having the sequence listed in SEQ ID NO:9. Of those mutants screened, a group was further screened for phasing / pre-phasing activity as generally described above in Example 1. Those polymerases having the following mutations showed improved phasing and / or pre-phasing activity compared to the control.
[0131]
[0132]
[0133] Example 5
[0134] Screening for mutants of the modified 9°N polymerase
[0135] Mutations in the modified 9°N polymerase backbone sequence (SEQ ID NO:13), which were generally described in Example 1, produced polymerases having the amino acid sequences listed in SEQ ID NO:14-16.
[0136] The purified mutant polymerases were screened for burst kinetics using the gel-based assay described above in Example 1, and compared with a control polymerase having the sequence listed in SEQ ID NO:13. Of those mutants screened, a group was further screened for phasing / pre-phasing activity as generally described above in Example 1. Those polymerases having the following mutations showed improved phasing and / or pre-phasing activity compared to the control.
[0137]
[0138] Example 6
[0139] Filter PfuExo - polymerase mutant
[0140] Sequence alignment analysis based on the 9°N polymerase backbone sequence (see...) Figure 1 As generally described in Example 1, the generation, cloning, expression, and purification of *P. fuga* Exo - A specific mutation in the polymerase backbone sequence (SEQ ID NO:17) produces a polymerase having the amino acid sequences listed in SEQ ID NO:18-20.
[0141] The purified mutant polymerases were screened for burst kinetics using the gel-based assay described above in Example 1, and compared with the control polymerase having the sequence listed in SEQ ID NO:17. Among those mutants screened, a group of mutants were further screened for phasing / pre-phasing activity as generally described above in Example 1. Those polymerases having the following mutations showed improved phasing and / or pre-phasing activity compared to the control.
[0142]
[0143] Example 7
[0144] Filtering KOD1 Exo - polymerase mutant
[0145] Sequence alignment analysis based on the 9°N polymerase backbone sequence (see...) Figure 1The generation, cloning, expression, and purification of Thermococcus kodakaraensis (KOD1) Exo, as generally described in Example 1, are also described. - A specific mutation in the polymerase backbone sequence (SEQ ID NO:21) produces a polymerase having the amino acid sequences listed in SEQ ID NO:22-24.
[0146] The purified mutant polymerases were screened for burst kinetics using the gel-based assay described above in Example 1, and compared with a control polymerase having the sequence listed in SEQ ID NO:21. Among those screened mutants, a group of mutants were further screened for phasing / pre-phasing activity as generally described above in Example 1. Those polymerases having the following mutations showed improved phasing and / or pre-phasing activity compared to the control.
[0147]
[0148] Example 8
[0149] Filter MMS2 Exo - polymerase mutant
[0150] Sequence alignment analysis based on the 9°N polymerase backbone sequence (see...) Figure 1 Based on homology with 9°N polymerase (see...) Figure 2 ), to identify M. maripaludis(MMS2)Exo - A specific mutation in the polymerase backbone sequence (SEQ ID NO:25). The mutant was generated, cloned, expressed, and purified as generally described in Example 1, resulting in a polymerase having the amino acid sequences listed in SEQ ID NO:26-28.
[0151] The purified mutant polymerases were screened for burst kinetics using the gel-based assay described above in Example 1, and compared with a control polymerase having the sequence listed in SEQ ID NO:25. Of those mutants screened, a group was further screened for phasing / pre-phasing activity as generally described above in Example 1. Those polymerases having the following mutations showed improved phasing and / or pre-phasing activity compared to the control.
[0152]
[0153]
[0154] Numerous publications, patents, and / or patent applications have been referenced throughout this application. The disclosures of these publications are incorporated herein by reference in their entirety.
[0155] The terminology used herein is open-ended and includes not only the cited element but also any other element.
[0156] Many embodiments have been described. However, it will be understood that various modifications can be made. Therefore, other embodiments are within the scope of the following claims.
Claims
1. A modified B-family archaea DNA polymerase, said modified B-family archaea DNA polymerase being selected from the amino acid sequences SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:29 or SEQ ID NO:31, differing in that a substitution mutation is present at Lys477 and one or more of Leu408, Tyr409 and Pro410 in the amino acid sequence SEQ ID NO:5 of the wild-type 9°N DNA polymerase, or a substitution mutation is present at one or more of Lys477 and one or more of Leu408, Tyr409 and Pro410 in SEQ ID NO:5, wherein said substitution mutation corresponds to Leu408Ala, Tyr409Ala, Pro410Ile and Lys477Met, and wherein, when used, said modified B-family archaea DNA polymerase exhibits polymerase activity.
2. A modified B-family archaeal DNA polymerase as described in claim 1, wherein the modified B-family archaeal DNA polymerase comprises an amino acid sequence of any one of the following: SEQ ID NO: 6-8, 10-12, 14-16, 18-20, 22-24, 26-28, 30 and 32.
3. A nucleic acid molecule encoding a modified B-family archaeal DNA polymerase as defined in any one of claims 1-2.
4. An expression vector comprising the nucleic acid molecule of claim 3.
5. A host cell comprising the expression vector of claim 4.
Citation Information
Patent Citations
Nucleic acid sequencing using microsphere arrays
US20050191698A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Polymerase enzymes and reagents for enhanced nucleic acid sequencing
US20080108082A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090026082A1