Engineered polymerases
Patent Information
- Application Number
- JP2023577945
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-03-25
- Filing Date
- 2022-06-17
- Publication Date
- 2025-06-23
AI Technical Summary
Current DNA polymerases struggle with incorporating nucleotide analogs containing chain-terminating moieties, such as 3'-O-azido groups, due to steric clashes and poor affinity, leading to inefficient nucleotide discrimination and incorporation, especially in sequencing technologies like SBS and SBB.
Engineered mutant polymerases from Candidatus Altiarchaeales with specific amino acid substitutions, such as D141A and E143A, exhibit improved thermostability, enhanced binding and incorporation of nucleotide analogs, and increased uracil resistance, forming stable ternary complexes with nucleotides, including those with 3'-O-azido groups.
The engineered polymerases enhance the efficiency and accuracy of nucleic acid sequencing by maintaining stable complexes and increasing the uptake rate of nucleotide analogs, improving signal-to-noise ratios in sequencing methods like SBS and SBB.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy, created on June 15, 2022, is named 52269WO_CRF_sequencelisting and is 2,585,266 bytes in size.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Patent Application No. 17 / 705,011, filed March 25, 2022, U.S. Patent Application No. 17 / 705,020, filed March 25, 2022, U.S. Patent Application No. 17 / 705,043, filed March 25, 2022, and U.S. Provisional Application No. 63 / 212,540, filed June 18, 2021, each of which is incorporated by reference herein in its entirety for all purposes.
[0003] Throughout this application, various publications, patents, and / or patent applications are referenced. The disclosures of the publications, patents, and / or patent applications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this disclosure pertains.
[0004] The present disclosure provides mutant polymerases engineered for improved thermostability and exhibiting improved binding and / or incorporation of nucleotide analogs, as well as improved uracil resistance. Exemplary nucleotide analogs include nucleotides containing 3' chain-terminating moieties. The mutant polymerases exhibit increased incorporation rates compared to wild-type polymerases. [Background technology]
[0005] Next-generation sequencing (NGS) technology has become a powerful tool for obtaining sequencing data used in molecular biology, taxonomy, agricultural science, medical diagnostics, and the development of new therapeutics. The present disclosure provides engineered polymerases that are useful for performing any nucleic acid sequencing method using labeled or unlabeled chain-terminating nucleotides, where the chain-terminating nucleotides contain a 3'-O-azido group (or a 3'-O-methyl azido group) or any other type of bulky blocking group at the 3' position of the sugar. For example, the engineered polymerases can be used to perform affinity sequencing (SBA) using labeled polyvalent molecules and unlabeled chain-terminating nucleotides. Furthermore, the engineered polymerases can be used to perform sequencing-by-synthesis (SBS) methods using labeled chain-terminating nucleotides and sequencing-by-binding (SBB) methods using unlabeled chain-terminating nucleotides.
[0006] The addition of a single nucleotide alone to a strand of DNA does not generate a signal sufficient for easy detection. Currently available SBS technologies overcome this problem by increasing the signal-to-noise ratio of nucleotide addition, coupled with a detection method sensitive enough to make accurate base calls. The most commercially successful platforms use monoclonal template DNA amplification in a spatially constrained matrix to generate distinct DNA islands containing multiple copies of the sequence to be interrogated. The result of this amplification is a "colony" of DNA copies, such that the addition of a single DNA base to all copies enriches the detection modality in a manner sufficient to overcome the signal-to-noise problem. Sequencing multiple, spatially constrained, identical copies of DNA further relies on a controlled step mechanism to ensure that one and only one nucleotide base can be added, ensuring that all copies within the DNA colony remain in the same position relative to each other (N, N+1, N+2, N+3, etc.).
[0007] The molecular engine required to carry out SBS is DNA polymerase. In vivo, this class of enzymes is involved in DNA replication and maintaining genome integrity. Under native conditions, DNA-dependent DNA polymerase (dDdP) catalyzes the addition of deoxynucleotide triphosphates (dNTPs) to DNA in a 5' to 3' direction, creating a phosphodiester bond between the 3' hydroxyl of the primer DNA terminus and the 5' α-phosphate of the input nucleotide. This chemical reaction occurs with high fidelity for correct Watson-Crick base pairing due to hydrogen bonding between the correct input dNTP and template base. This "correct" base pairing induces a conformational change in the enzyme that aligns the catalytic amino acid for efficient phosphodiester bond formation. The newly added dNTP also possesses a 3' OH that is used in the next round of catalysis to further extend the DNA strand.
[0008] To ensure that only a single dNTP is added to the growing strand of DNA per SBS cycle, reversibly terminated dNTPs are used. These bases contain a modification to the 3' hydroxyl of the dNTP, blocking subsequent rounds of incorporation. The most commercially successful reversible terminator is the 3' methyl azide, but other terminators, including 3' amino allyl and 3' oxyamine, have also been used. Each of these reversibly terminated dNTPs functions in the same way: once incorporated, the bulky 3' block inhibits the addition of the next nucleotide due to the absence of a 3' hydroxyl. Upon exposure to catalyst, the 3' block reacts to regenerate a 3' hydroxyl that can form a new phosphodiester bond during the next cycle. While effective, these bulky 3' modifications present a challenge for polymerases.
[0009] The evolutionary requirement for high-fidelity genome replication and stability is 4 ~10 7This has resulted in polymerases that incorporate only non-Watson-Crick base pairs per incorporation event. Polymerases also often need to discriminate between the vast excess of nucleotides in the cellular environment. Discrimination between nucleotides is typically achieved through a steric gate where the presence of the 2' hydroxyl sterically clashes with the amino acid side chain at the nucleotide binding site, selecting for nucleotide binding and catalysis. Furthermore, damage or modifications to the 3' hydroxyl of nucleotides are also sensed by the enzyme, as bases containing nonviable 3' hydroxyls can function as chain terminators that inhibit DNA synthesis. Discrimination of these undesirable bases results in inappropriate nucleotide substrates binding with weaker overall affinity, resulting in a slower rate of phosphodiester bond formation. 2 ~10 4 This occurs via a kinetic pathway that occurs orders of magnitude slower. This is due to a lack of induced fit to properly align the catalytic amino acids for bond formation. As a result, naturally evolved polymerases are poor at incorporating reversible chain terminator nucleotides. Summary of the Invention
[0010] The present disclosure provides mutant polymerases that have been engineered for improved thermostability, and that exhibit improved binding of nucleotide analogs and / or improved binding and incorporation of nucleotide analogs, as well as improved uracil tolerance. The engineered polymerases may be used in a variety of contexts and may have a variety of properties, as described in more detail below.
[0011] The present disclosure provides bound complexes (e.g., ternary complexes), each comprising a nucleotide. The present disclosure provides a plurality of ternary complexes, each comprising a nucleic acid duplex and a mutant or wild-type DNA polymerase bound to a nucleotide, the nucleic acid duplex comprising a nucleic acid template molecule hybridized to a nucleic acid primer, and in the ternary complex, the nucleotide is bound to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule. In some embodiments, in the ternary complex, the nucleotide is bound to the nucleic acid duplex and has not undergone polymerase-catalyzed incorporation, or the nucleotide is bound to the nucleic acid duplex and has undergone polymerase-catalyzed incorporation. In some embodiments, the wild-type DNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the mutant DNA polymerase comprises an amino acid sequence at least 85% identical to SEQ ID NO: 1.
[0012] In some embodiments, the mutant DNA polymerase is from the archaeon Candidatus Altiarchaeales and comprises the amino acid sequence of any one of SEQ ID NOs: 2-274, 288-375, or 385-397. In some embodiments, the mutant polymerase comprises the amino acid substitutions D141A and E143A (Asp141Ala and Glu143Ala). In some embodiments, the ternary complex remains stable without dissociating the mutant or wild-type polymerase from the nucleic acid duplex (or exhibits reduced dissociation), and the stable ternary complex exhibits a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. In some embodiments, the plurality of ternary complexes further comprises a plurality of non-catalytic divalent cations or a plurality of catalytic divalent cations. In some embodiments, the plurality of non-catalytic divalent cations comprises strontium, barium, and / or calcium. In some embodiments, the catalytic divalent cation comprises magnesium and / or manganese.
[0013] In some embodiments, the mutant polymerase from the archaeon Candidatus Altiarchaeales has at least 85% sequence identity, or at least 90% sequence identity, or at least 95% sequence identity, or at least 96% sequence identity, or at least 97% sequence identity, or at least 98% sequence identity, or at least 99% sequence identity, or at least 99.1% sequence identity, or at least 99.2% sequence identity, or at least 99.3 ... The mutant DNA polymerase comprises an amino acid sequence having 99.4% sequence identity, or at least 99.5% sequence identity, or at least 99.6% sequence identity, or at least 99.7% sequence identity, or at least 99.8% sequence identity, or a higher percent sequence identity, and the mutant DNA polymerase comprises an amino acid substitution at any one or any combination of two or more positions selected from the group consisting of Leu416, Tyr417, Pro418, Ala493, Arg515, Ile529, and Asn567. In some embodiments, the mutant polymerase comprises the amino acid substitutions D141A and E143A, which can confer exonuclease-minus activity. In some embodiments, the mutant polymerase exhibits desirable characteristics compared to a polymerase having a wild-type amino acid backbone sequence (e.g., SEQ ID NO: 1 or 391). For example, the mutant polymerase exhibits increased thermostability (Tm). In another example, the mutant polymerase exhibits an increased incorporation rate of nucleotide analogs containing chain-terminating moieties (e.g., blocking moieties) at the 2' and / or 3' sugar positions. In yet another example, the mutant polymerase exhibits increased uracil tolerance. One or more of the properties described in this paragraph may be present in any of the exemplary mutant polymerases in the various embodiments described herein. The properties described in this paragraph are referred to throughout this disclosure as "exemplary mutant polymerase properties."
[0014] In some embodiments, in a ternary complex comprising a nucleotide, the multiple immobilized polymerase complexes comprise nucleic acid template molecules having the same target sequence of interest or different target sequences of interest.
[0015] In some embodiments, in a ternary complex comprising a nucleotide, the nucleotide comprises an aromatic base, a five-carbon sugar, and one to ten phosphate groups, and the aromatic base of the nucleotide comprises adenine, guanine, cytosine, thymine, or uracil. The nucleotide comprises dATP, dGTP, dCTP, dTTP, or dUTP. The nucleotide can be labeled with a fluorophore. The nucleotide can lack a fluorophore.
[0016] In some embodiments, in a ternary complex comprising a nucleotide, the nucleotide comprises a chain-terminating moiety. In some embodiments, the chain-terminating moiety can be attached to the 3'-OH sugar position via a cleavable moiety. In some embodiments, the chain-terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide into a nascent chain during a primer extension reaction. In some embodiments, the chain-terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain-terminating moiety is removable / cleavable from the 3' sugar hydroxyl position to generate a nucleotide having a 3'-OH sugar group that can be extended with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group.In some embodiments, chain-terminating moieties alkyl, alkenyl, alkynyl, and allyl are cleavable / removable using tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or using 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ); chain-terminating moieties aryl and benzyl are cleavable / removable using HPd / C; chain-terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable / removable using thiol reagents including β-mercaptoethanol or dithiothritol (DTT); chain-terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide are cleavable / removable using phosphine reagents including tris(2-carboxyethyl)phosphine (TCEP), bis-sulfotriphenylphosphine (BS-TPP), or tri(hydroxypropyl)phosphine (THPP). The chain-terminating moieties amine, amide, keto, isocyanate, phosphate, thio, and disulfide can be cleaved / removed using 4-dimethylaminopyridine (4-DMAP); the chain-terminating moieties carbonate can be cleaved / removed using potassium carbonate (KCO) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH); and the chain-terminating moieties urea and silyl can be cleaved using tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride. In some embodiments, the chain-terminating moiety can be cleaved with nitrous acid. In some embodiments, the chain-terminating moiety can be cleaved using a solution containing nitrite, for example, a combination of nitrite and an acid, such as acetic acid, sulfuric acid, or nitric acid. In some further embodiments, the solution can include an organic acid. One or more of the properties described in this paragraph can be present in any exemplary chain-terminating moiety in various embodiments described herein. The features described in this paragraph are referred to throughout this disclosure as "chain-terminating moiety embodiments."
[0017] In some embodiments, in a ternary complex comprising a nucleotide, the nucleotide comprises a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, wherein the chain-terminating moiety comprises an azide, azido, or azidomethyl group. For example, in some embodiments, the chain-terminating moiety comprises a 3'-O-azido or 3'-O-azidomethyl group. In some embodiments, the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable using a phosphine compound, including a derivatized tri-alkylphosphine moiety, a derivatized tri-arylphosphine moiety, tris(2-carboxyethyl)phosphine (TCEP), bis-sulfotriphenylphosphine (BS-TPP), or tri(hydroxypropyl)phosphine (THPP), and the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable using 4-dimethylaminopyridine (4-DMAP). In some embodiments, in the system, the nucleotide analog comprises a chain-terminating moiety selected from the group consisting of 3'-deoxynucleotide, 2',3'-dideoxynucleotide, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino group, 3'-phosphorothioate, and 3-O-benzyl, or a derivative thereof. In some embodiments, chain-terminating moieties comprising one or more of a 3'-O-amino group, a 3'-O-aminomethyl group, a 3'-O-methylamino group, or derivatives thereof, can be cleaved with nitrous acid via a nitrous acid-based mechanism or using a solution containing nitrous acid.In some embodiments, chain-terminating moieties comprising one or more of a 3'-O-amino group, a 3'-O-aminomethyl group, a 3'-O-methylamino group, or derivatives thereof, can be cleaved using a solution containing nitrite. In some embodiments, for example, the nitrite can be combined with or contacted with an acid, such as acetic acid, sulfuric acid, or nitric acid. In some further embodiments, for example, the nitrite can be combined with or contacted with an organic acid, such as formic acid, acetic acid, propionic acid, butyric acid, isobutyric acid, etc. The phrase, when referring to subsets of groups, can also be stated as "azide-containing chain-terminating moieties," or "azido-containing chain-terminating moieties," or "azidomethyl-containing chain-terminating moieties," and still include the embodiments listed in this paragraph. The phrase "a chain-terminating moiety comprising an azide, azido, or azidomethyl group" is used throughout this disclosure to refer to any one or more of the chain-terminating moiety characteristics described in this paragraph.
[0018] In some embodiments, in a ternary complex containing a nucleotide, the wild-type or mutant DNA polymerase may or may not be fluorescently labeled. In some embodiments, the wild-type or mutant DNA polymerase comprises a fluorescently labeled DNA polymerase. In some embodiments, the wild-type or mutant DNA polymerase lacks a fluorophore. In some embodiments, the DNA polymerase comprises a fluorescently labeled DNA polymerase and the nucleotide lacks a fluorophore. In some embodiments, the DNA polymerase lacks a fluorophore and the nucleotide comprises a fluorescently labeled nucleotide. In some embodiments, the DNA polymerase comprises a fluorescently labeled DNA polymerase and the nucleotide comprises a fluorescently labeled nucleotide. One or more of the characteristics described in this paragraph may be present in any of the exemplary fluorophore molecules in the various embodiments described herein. The characteristics described in this paragraph are referred to throughout this disclosure as "fluorophore embodiments."
[0019] In some embodiments, in a ternary complex containing nucleotides, the nucleic acid template molecule can take various forms. For example, the nucleic acid template molecule comprises a linear nucleic acid molecule, a circular nucleic acid molecule, or a mixture of both linear and circular nucleic acid molecules. In some embodiments, the nucleic acid template molecule comprises a clonally amplified template molecule. In some embodiments, the nucleic acid template molecule comprises one copy of a target sequence of interest. In some embodiments, the nucleic acid template molecules in a plurality of nucleic acid template molecules comprise the same target sequence of interest or different target sequences of interest. In some embodiments, the nucleic acid template molecule comprises two or more tandem copies of a target sequence of interest (e.g., a concatemer). In some embodiments, the nucleic acid template molecule comprises at least one uridine nucleotide or lacks uridine nucleotides. One or more of the characteristics described in this paragraph can appear in any nucleic acid template in various embodiments described herein. The characteristics described in this paragraph are referred to throughout this disclosure as "nucleic acid template embodiments."
[0020] In some embodiments, in a ternary complex comprising a nucleotide, a plurality of ternary complexes are immobilized. For example, the ternary complex can be immobilized on a support or on a coating on a support. In some embodiments, the coating on the support comprises at least one hydrophilic polymer coating layer comprising a branched polyethylene glycol (PEG) having at least four branches, the coating having a water contact angle of 45 degrees or less. In some embodiments, the support comprises a functionalized polymer coating layer covalently bonded to at least a portion of the support via chemical groups on the support, an oligonucleotide primer grafted to the functionalized polymer coating, and a water-soluble protective coating on the primer and the functionalized polymer coating. In some embodiments, the density of the plurality of ternary complexes immobilized on the support is greater than 1 mm 2 10 per 2 ~10 6In some embodiments, the plurality of ternary complexes are immobilized at predetermined sites on a support. In some embodiments, the plurality of ternary complexes are immobilized at random sites on a support. In some embodiments, the plurality of immobilized ternary complexes are in fluid communication with each other, allowing a solution of reagents to be flowed over the support such that the plurality of immobilized ternary complexes on the support react in a massively parallel manner with a solution of reagents, the reagents comprising soluble primers, DNA polymerase, nucleotides, divalent cations, and / or buffer. One or more of the properties described in this paragraph may be present in any exemplary ternary complex in various embodiments described herein. The properties described in this paragraph are referred to throughout this disclosure as "immobilization embodiments."
[0021] The present disclosure provides a plurality of binding complexes (e.g., a plurality of ternary complexes), each comprising a multivalent molecule. A plurality of binding complexes or ternary complexes comprising multivalent molecules can comprise any of the same properties as those described above for binding complexes or ternary complexes. In some embodiments, a ternary complex comprising a multivalent molecule can take various forms. For example, in a ternary complex comprising multivalent molecules, each multivalent molecule in the plurality of multivalent molecules can comprise (a) a core and (b) a plurality of nucleotide arms comprising (i) a core-binding moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is linked to the plurality of nucleotide arms via their core-binding moieties, the spacer is linked to the linker, and the linker is linked to the nucleotide unit. In some embodiments, the core comprises a streptavidin-type or avidin-type moiety, and the core-binding moiety comprises biotin. In some embodiments, the linker comprises an aliphatic chain having 2 to 6 subunits or an oligoethylene glycol chain having 2 to 6 subunits. In some embodiments, the linker further comprises an aromatic moiety. An exemplary spacer is shown in Figure 16A (top), and an exemplary linker is shown in Figure 16A (bottom) and Figure 16B. An exemplary nucleotide arm is shown in Figure 15B. An exemplary multivalent molecule is shown in Figures 14A, 14B, and 15A. In some embodiments, the nucleotide unit comprises an aromatic base, a 5-carbon sugar, and 1 to 10 phosphate groups. In some embodiments, the linker is attached to the nucleotide unit via the base. In some embodiments, multiple nucleotide arms attached to the core have the same type of nucleotide unit, where the type of nucleotide unit is selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. In some embodiments, the multiple multivalent molecules comprise one type of multivalent molecule, where each multivalent molecule has the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.In some embodiments, the plurality of multivalent molecules comprises a mixture of any combination of two or more types of multivalent molecules, each type having a nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. One or more of the characteristics described in this paragraph may be present in any of the exemplary multivalent molecules in the various embodiments described in this disclosure. The characteristics described in this paragraph are referred to throughout this disclosure as "embodiments of multivalent molecules."
[0022] In some embodiments, in a ternary complex comprising multivalent molecules, the plurality of multivalent molecules comprise fluorescently labeled multivalent molecules. In some embodiments, the core of each fluorescently labeled multivalent molecule is bound to a fluorophore corresponding to the nucleotide unit bound to the nucleotide arm in a given multivalent molecule. In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or nucleotide base bound to a fluorophore, and the fluorophore bound to a given linker or nucleotide base corresponds to the nucleotide base of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, the multivalent molecule lacks a fluorophore.
[0023] In some embodiments, in a ternary complex comprising a multivalent molecule, at least one of the multivalent molecules in the plurality of multivalent molecules comprises a nucleotide unit having a chain-terminating moiety attached to the 3'OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, comprising any of the potential properties described above.
[0024] In some embodiments, in a ternary complex comprising a multivalent molecule, the wild-type or mutant DNA polymerase may or may not be fluorescently labeled, hi some embodiments, the wild-type or mutant DNA polymerase may comprise any of the above-described fluorophore embodiments.
[0025] In some embodiments, in a ternary complex comprising a multivalent molecule, the nucleic acid template molecule can include an embodiment of a nucleic acid template comprising any of the potential properties described above.
[0026] In some embodiments, ternary complexes containing multivalent molecules may be immobilized according to immobilization embodiments that include any of the potential properties described above.
[0027] In some embodiments, in a ternary complex comprising a multivalent molecule, the plurality of ternary complexes comprises at least a first and a second ternary complex comprising a single multivalent molecule that binds to the first and second ternary complexes to form a first affinity complex. In some embodiments, the first ternary complex comprises a first DNA polymerase (e.g., a first mutant or wild-type DNA polymerase) that binds to a first primer that hybridizes to a first portion of the concatemeric template molecule, and a first nucleotide unit of the single multivalent molecule binds to the first primer, thereby forming a first ternary complex. In some embodiments, the second ternary complex comprises a second DNA polymerase (e.g., a second mutant or wild-type DNA polymerase) that binds to a second primer that hybridizes to a second portion of the same concatemeric template molecule, and a second nucleotide unit of the single multivalent molecule binds to the second primer, thereby forming a second ternary complex. The first and second ternary complexes that bind to the multivalent molecule form a first affinity complex. In some embodiments, the first and / or second ternary complex remains stable without dissociation for a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0028] In some embodiments, in a ternary complex comprising a multivalent molecule, the first affinity complex further comprises at least a third and a fourth ternary complex in which a single multivalent molecule binds to the first, second, third, and fourth ternary complexes. In some embodiments, the third ternary complex comprises a third DNA polymerase (e.g., a third mutant or wild-type DNA polymerase) that binds to a third primer that hybridizes to a third portion of the concatemeric template molecule, and the third nucleotide unit of the single multivalent molecule binds to the third primer, thereby forming the third ternary complex. In some embodiments, the fourth ternary complex comprises a fourth DNA polymerase (e.g., a fourth mutant or wild-type DNA polymerase) that binds to a fourth primer that hybridizes to a fourth portion of the same concatemeric template molecule, and the fourth nucleotide unit of the single multivalent molecule binds to the fourth primer, thereby forming the fourth ternary complex. In some embodiments, the first, second, and third ternary complexes that bind to the multivalent molecule form the first affinity complex. In some embodiments, the first, second, third, and fourth ternary complexes that bind to the multivalent molecule form a first affinity complex, and in some embodiments, the third and / or fourth ternary complex remains stable without dissociation for a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0029] The present disclosure provides a nucleic acid sequencing method that uses a DNA polymerase from the archaeon Candidatus Altiarchaeales to form a bound complex (e.g., a ternary complex) containing nucleotides. The disclosure provides a method for nucleic acid sequencing, comprising: (a) contacting (i) a plurality of wild-type or mutant DNA polymerases with (ii) a plurality of nucleic acid duplexes, each comprising a nucleic acid template molecule hybridized to a nucleic acid primer, under conditions suitable to form a plurality of multiple polymerases, each comprising a wild-type or mutant DNA polymerase that binds to the nucleic acid duplex, wherein the plurality of wild-type DNA polymerases is 100% identical to SEQ ID NO: 1, or the plurality of mutant DNA polymerases comprises an amino acid sequence at least 85% identical to SEQ ID NO: 1 (e.g., at least 85% identical to any one of the amino acid sequences of SEQ ID NOs: 2-274, 288-375, or 385-397). (b) contacting a plurality of polymerase complexes with (iii) a plurality of nucleotides, and (iv) a plurality of catalytic divalent cations, wherein the contacting is performed under conditions suitable to form a plurality of ternary complexes, each comprising a wild-type or mutant DNA polymerase bound to a nucleic acid duplex and a nucleotide, wherein the nucleotide is bound to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule, and the conditions are suitable to promote polymerase-catalyzed incorporation of the nucleotide bound to the 3' end of the nucleic acid primer; (c) detecting the plurality of ternary complexes; and (d) identifying the plurality of incorporated nucleotides in the plurality of ternary complexes. In some embodiments, in each ternary complex of step (b), the nucleotide is bound to the nucleic acid duplex and has not undergone polymerase-catalyzed incorporation, or the nucleotide is bound to the nucleic acid duplex and has undergone polymerase-catalyzed incorporation.In some embodiments, the ternary complex remains stable (or exhibits reduced dissociation) without dissociating the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complex exhibits a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. In some embodiments, the catalytic divalent cation comprises magnesium and / or manganese. In some embodiments, the multiple polymerase complexes comprise nucleic acid template molecules having the same target sequence of interest or different target sequences of interest. In some embodiments, the mutant polymerase comprises the amino acid substitutions D141A and E143A.
[0030] In some embodiments, nucleic acid sequencing using a ternary complex comprising nucleotides includes contacting (a) (i) a plurality of wild-type or mutant DNA polymerases, and (ii) a plurality of nucleic acid duplexes, each comprising a nucleic acid template molecule hybridized to a nucleic acid primer, wherein the contacting is performed under conditions suitable to form a plurality of polymerase complexes, each comprising a wild-type or mutant DNA polymerase that binds to the nucleic acid duplex, wherein the plurality of wild-type DNA polymerases is 100% identical to SEQ ID NO: 1, or the plurality of mutant DNA polymerases comprises an amino acid sequence at least 85% identical to SEQ ID NOs: 2-274, 288-375, or 385-397. (b) contacting a plurality of composite polymerases with (iii) a plurality of nucleotides, and (iv) a plurality of non-catalytic divalent cations, the contacting being performed under conditions suitable to form a plurality of ternary complexes, each comprising a wild-type or mutant DNA polymerase bound to a nucleic acid duplex and a nucleotide, wherein the nucleotide is bound to the 3' end of a nucleic acid primer at a position opposite its complementary nucleotide in the nucleic acid template molecule, and the conditions are suitable to inhibit polymerase-catalyzed incorporation of the nucleotide bound to the 3' end of the nucleic acid primer; (c) detecting the plurality of ternary complexes; and (d) identifying the plurality of bound nucleotides in the plurality of ternary complexes. In some embodiments, the non-catalytic divalent cation comprises strontium, barium, and / or calcium. In some embodiments, the plurality of composite polymerases comprises nucleic acid template molecules having the same target sequence of interest or different target sequences of interest. In some embodiments, the mutant polymerase comprises the amino acid substitutions D141A and E143A.
[0031] In some embodiments, a nucleic acid sequencing method uses a ternary complex that can include a nucleotide unit. The nucleotide unit can include an aromatic base, a 5-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 to 10 phosphate groups), where the aromatic base of the nucleotide includes adenine, guanine, cytosine, thymine, or uracil. In some embodiments, the plurality of nucleotides includes one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. In some embodiments, the plurality of nucleotides includes a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is labeled with a fluorophore. In some embodiments, the plurality of nucleotides lacks a fluorophore label. One or more of the characteristics described in this paragraph can appear in any of the exemplary nucleotide units in the various embodiments described herein. The characteristics described in this paragraph are referred to throughout this disclosure as "exemplary nucleotide unit characteristics."
[0032] In some embodiments, in nucleic acid sequencing methods using ternary complexes containing nucleotides, the multiple DNA polymerases may or may not be fluorescently labeled and may include any of the above-described fluorophore embodiments. In some embodiments, the multiple DNA polymerases include multiple nucleotides.
[0033] In some embodiments, in a nucleic acid sequencing method using a ternary complex comprising nucleotides, at least one of the nucleotides in the plurality of nucleotides comprises a chain-terminating moiety and may comprise any of the above-described embodiments of a chain-terminating moiety. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the above-described potential characteristics.
[0034] In some embodiments, in nucleic acid sequencing methods using ternary complexes comprising nucleotides, the plurality of nucleic acid template molecules can include embodiments of nucleic acid templates comprising any of the above potential properties.
[0035] In some embodiments, in nucleic acid sequencing methods using a ternary complex comprising a nucleotide, this may be immobilized according to an embodiment of immobilization comprising any of the potential properties described above.
[0036] The present disclosure provides nucleic acid sequencing methods using binding complexes (eg, ternary complexes) that include multivalent molecules. The disclosure provides a method for nucleic acid sequencing, comprising: (a) contacting (i) a plurality of wild-type or mutant DNA polymerases and (ii) a plurality of nucleic acid duplexes, each comprising a nucleic acid template molecule hybridized to a nucleic acid primer, under conditions suitable to form a plurality of composite polymerases, each comprising a mutant DNA polymerase that binds to the nucleic acid duplex, wherein the plurality of wild-type DNA polymerases are 100% identical to SEQ ID NO:1, or the plurality of mutant DNA polymerases comprise an amino acid sequence that is at least 85% identical to SEQ ID NO:1 (e.g., at least 85% identical to any one of the amino acid sequences of SEQ ID NOs:2-274, 288-375, or 385-397); and (b) contacting the plurality of composite polymerases with (iii) a plurality of multivalent molecules, and (iv) a plurality of non-catalytic divalent cations, wherein the plurality of multivalent molecules each bind to a plurality of nucleotide arms. a nucleic acid duplex comprising a core and each nucleotide arm comprising a nucleotide unit, the contacting being performed under conditions suitable to form a plurality of ternary complexes, each comprising a mutant DNA polymerase that binds to the nucleic acid duplex and the multivalent molecule, wherein in the ternary complex, one nucleotide unit of the multivalent molecule binds to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule, the plurality of ternary complexes remaining stable without dissociation for a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second, the contacting being performed under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound nucleotide unit of the multivalent molecule; (c) detecting the plurality of ternary complexes; and (d) identifying the plurality of nucleotide units bound to the 3' end of the nucleic acid primer in the plurality of ternary complexes, thereby determining the sequence of the plurality of nucleic acid template molecules. In some embodiments, the mutant polymerase comprises amino acid substitutions D141A and E143A.
[0037] In yet another example, the mutant polymerase exhibits one or more of the properties of the exemplary mutant polymerases discussed above. Alternatively or additionally, the mutant polymerase exhibits increased uracil tolerance (e.g., any of SEQ ID NOs: 361, 362, 363, 364, 366, 367, 374, or 375, or any of SEQ ID NOs: 385-397). In some embodiments, the mutant polymerase comprises the amino acid substitutions D141A and E143A.
[0038] In some embodiments, in nucleic acid sequencing methods using a ternary complex comprising a multivalent molecule, the ternary complex remains stable (or exhibits reduced dissociation) without dissociating the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complex exhibits a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. In some embodiments, the plurality of non-catalytic divalent cations comprises strontium, barium, and / or calcium. In some embodiments, the plurality of polymerase complexes comprises nucleic acid template molecules having the same target sequence of interest or different target sequences of interest.
[0039] In some embodiments, in a nucleic acid sequencing method using a ternary complex comprising a multivalent molecule, in the ternary complex, the nucleotide units of the multivalent molecule are either bound to a nucleic acid duplex and have not undergone polymerase-catalyzed incorporation, or the nucleotide units are bound to a nucleic acid duplex and have undergone polymerase-catalyzed incorporation.
[0040] In some embodiments, in a method for nucleic acid sequencing using a ternary complex comprising multivalent molecules, the method further comprises forming an affinity complex, (1) contacting a plurality of wild-type or mutant DNA polymerases and a plurality of nucleic acid primers with different portions of a concatemeric nucleic acid template molecule to form at least first and second composite polymerases on the same concatemeric template molecule; and (2) contacting the plurality of multivalent molecules with at least first and second composite polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second composite polymerases, wherein at least a first nucleotide unit of the single multivalent molecule hybridizes to a first portion of the concatemeric template molecule, thereby binding to a first composite polymerase comprising a first primer, forming a first ternary complex, and a single (3) contacting, under conditions suitable for inhibiting polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second ternary complexes, wherein at least a second nucleotide unit of the multivalent molecule is bound to a second complex polymerase comprising a second primer that hybridizes to a second portion of the concatemeric template molecule, thereby forming a second ternary complex, such that the first and second ternary complexes bound to the same multivalent molecule form an affinity complex; (4) detecting the first and second ternary complexes on the same concatemeric template molecule; and (5) identifying the first nucleotide unit in the first ternary complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second ternary complex, thereby determining the sequence of the second portion of the concatemeric template molecule. In some embodiments, in the method for forming an affinity complex, the identifying in step (4) comprises identifying a first nucleotide unit attached to the 3' end of the first primer in the first ternary complex, thereby determining the sequence of a first portion of the concatemer template molecule, and identifying a second nucleotide unit attached to the 3' end of the second primer in the second ternary complex, thereby determining the sequence of a second portion of the concatemer template molecule.
[0041] In some embodiments, in nucleic acid sequencing methods using ternary complexes that include multivalent molecules, the multivalent molecules can include any of the embodiments of multivalent molecules that include any of the potential properties described above.
[0042] In some embodiments, in nucleic acid sequencing methods using ternary complexes comprising multivalent molecules, the multivalent molecules may or may not be fluorescently labeled and may comprise any of the fluorophore embodiments described above. In some embodiments, the core of an individual multivalent molecule is bound to a nucleotide unit that is bound to a nucleotide arm and may or may not be fluorescently labeled and may comprise any of the fluorophore embodiments described above. In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or nucleotide base that is bound to a fluorophore, and the fluorophore bound to a given linker or nucleotide base corresponds to the nucleotide base of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil).
[0043] In some embodiments, in a nucleic acid sequencing method using a ternary complex comprising multivalent molecules, at least one of the multivalent molecules in the plurality of multivalent molecules comprises a nucleotide unit having a chain-terminating moiety attached to the 3'OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, comprising any of the potential properties described above.
[0044] In some embodiments, in a nucleic acid sequencing method using a ternary complex comprising a multivalent molecule, the plurality of DNA polymerases comprises fluorescently labeled DNA polymerases. In some embodiments, the plurality of DNA polymerases lacks a fluorophore.
[0045] In some embodiments, in nucleic acid sequencing methods using ternary complexes comprising multivalent molecules, the plurality of nucleic acid template molecules can include embodiments of nucleic acid templates comprising any of the above potential properties.
[0046] In some embodiments, in nucleic acid sequencing methods using ternary complexes that include multivalent molecules, the multivalent molecules can be immobilized according to immobilization embodiments that include any of the potential properties described above.
[0047] The present disclosure provides a two-phase nucleic acid sequencing method using a polymerase, a polyvalent molecule, and nucleotides from the archaeon Candidatus Altiarchaeales. The disclosure provides a method for nucleic acid sequencing, comprising: (a) contacting (i) a plurality of first wild-type or mutant DNA polymerases with (ii) a plurality of nucleic acid duplexes, each comprising a nucleic acid template molecule hybridized to a nucleic acid primer, under conditions suitable to form a plurality of first composite polymerases, each comprising a first wild-type or mutant DNA polymerase that binds to the nucleic acid duplex, wherein the plurality of first wild-type DNA polymerases are selected from the group consisting of SEQ ID NO: or a plurality of first mutant DNA polymerases comprising an amino acid sequence that is 100% identical to SEQ ID NO: 1, or a plurality of first mutant DNA polymerases comprising an amino acid sequence that is at least 85% identical to SEQ ID NO: 1 (e.g., at least 85% identical to any one of the amino acid sequences of SEQ ID NOs: 2-274, 288-375, or 385-397); and (b) contacting the plurality of first composite polymerases with (iii) a plurality of multivalent molecules, and (iv) a plurality of non-catalytic divalent cations, wherein the plurality of multivalent molecules each comprise a plurality of nucleotides. a core attached to a nucleotide arm, each nucleotide arm comprising a nucleotide unit, the contacting being performed under conditions suitable to form a plurality of first ternary complexes (each comprising a first wild-type or mutant DNA polymerase bound to a nucleic acid duplex, and a multivalent molecule), wherein in the first ternary complex, a nucleotide unit of the multivalent molecule is attached to the 3' end of the nucleic acid primer at a position opposite a complementary nucleotide in the nucleic acid template molecule, the contacting being performed under conditions suitable to inhibit polymerase-catalyzed incorporation of the attached nucleotide unit of the multivalent molecule; (c) detecting the plurality of first ternary complexes and identifying the nucleotide unit attached to the 3' end of the nucleic acid primer, thereby determining the sequence of the nucleic acid template molecule; (d) dissociating the plurality of first wild-type or mutant polymerases and the plurality of first ternary complexes, thereby retaining the plurality of nucleic acid duplexes; and (e) contacting the retained nucleic acid duplexes of step (d) with: (i) a plurality of second wild-type or mutant DNA polymerases;(ii) a plurality of nucleotides; and (iii) a plurality of catalytic divalent cations, wherein the plurality of second wild-type DNA polymerases comprise amino acid sequences that are 100% identical to SEQ ID NO: 1, or the plurality of second mutant DNA polymerases comprise amino acid sequences that are at least 85% identical to SEQ ID NO: 1 (e.g., at least 85% identical to any one of the amino acid sequences of SEQ ID NOs: 2-274, 288-375, or 385-397), and the contacting in step (e) produces a plurality of second ternary complexes (each of which is a retained nucleotide of step (d)). (f) contacting a nucleic acid template molecule with a second wild-type or mutant DNA polymerase under conditions suitable to form a second ternary complex (comprising a nucleic acid duplex and a second wild-type or mutant DNA polymerase bound to a nucleotide), wherein in the second ternary complex, the nucleotide is bound to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule, and the conditions are suitable to promote polymerase-catalyzed incorporation of the nucleotide bound to the 3' end of the nucleic acid primer; and (f) detecting a plurality of second ternary complexes and identifying the incorporated nucleotide in the second ternary complex. In some embodiments, the detecting in step (f) is optional. In some embodiments, the identifying in step (f) is optional. In some embodiments, the mutant polymerase comprises amino acid substitutions D141A and E143A.
[0048] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, in a first ternary complex, a nucleotide unit of the multivalent molecule is bound to a nucleic acid duplex and is not undergoing polymerase-catalyzed incorporation, or a nucleotide unit is bound to a nucleic acid duplex and is undergoing polymerase-catalyzed incorporation. In some embodiments, in a second ternary complex, a nucleotide is bound to a nucleic acid duplex and is not undergoing polymerase-catalyzed incorporation, or a nucleotide is bound to a nucleic acid duplex and is undergoing polymerase-catalyzed incorporation.
[0049] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the first mutant DNA polymerase comprises one or more of the properties described above with respect to exemplary mutant polymerase properties. In yet another example, the first mutant polymerase exhibits increased uracil tolerance (e.g., any of SEQ ID NOs: 361, 362, 363, 364, 366, 367, 374, or 375, or any of SEQ ID NOs: 385-397).
[0050] In some embodiments, the first wild-type or mutant DNA polymerase comprises a fluorescently labeled DNA polymerase. In some embodiments, the first wild-type or mutant DNA polymerase lacks a fluorophore.
[0051] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the second mutant DNA polymerase comprises one or more of the properties described above with respect to exemplary mutant polymerase properties. In yet another example, the second mutant polymerase exhibits increased uracil tolerance (e.g., any of SEQ ID NOs: 361, 362, 363, 364, 366, 367, 374, or 375, or any of SEQ ID NOs: 385-397).
[0052] In some embodiments, the second wild-type or mutant DNA polymerase comprises a fluorescently labeled DNA polymerase. In some embodiments, the second mutant DNA polymerase lacks a fluorophore.
[0053] In some embodiments, in the two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the plurality of non-catalytic divalent cations in step (b) comprise strontium, barium, and / or calcium. In some embodiments, the plurality of catalytic divalent cations in step (e) comprise magnesium and / or manganese. In some embodiments, the plurality of first ternary complexes in step (b) remain stable without dissociating the mutant or wild-type polymerase from the nucleic acid duplex (or exhibit reduced dissociation), and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. In some embodiments, the plurality of second ternary complexes in step (e) remain stable (or exhibit reduced dissociation) without dissociating the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0054] In some embodiments, a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales includes: (g) removing the plurality of second wild-type or mutant DNA polymerases, retaining the plurality of nucleic acid duplexes of step (f); (h) contacting the retained nucleic acid duplexes of step (g) with a plurality of first wild-type or mutant DNA polymerases, a plurality of multivalent molecules, and a plurality of non-catalytic divalent cations, wherein the contacting is performed under conditions suitable for forming another plurality of first ternary complexes, and the contacting is performed under conditions suitable for inhibiting polymerase-catalyzed incorporation of bound nucleotide units of the multivalent molecules; (i) detecting the plurality of first ternary complexes formed in step (h) and identifying the nucleotide unit attached to the 3' end of the nucleic acid primer, thereby determining the sequence of the nucleic acid template molecule; and (j) detecting the plurality of first ternary complexes formed in step (h) and identifying the nucleotide unit attached to the 3' end of the nucleic acid primer, thereby determining the sequence of the nucleic acid template molecule. The method further includes (i) dissociating the plurality of first ternary complexes formed in step (h) by removing a plurality of the first wild-type or mutant polymerase and the plurality of multivalent molecules, thereby retaining a plurality of nucleic acid duplexes; (k) contacting the plurality of retained nucleic acid duplexes of step (j) with a plurality of second wild-type or mutant DNA polymerases, a plurality of nucleotides, and a plurality of catalytic divalent cations, the contacting being performed under conditions suitable for forming another plurality of second ternary complexes, the conditions being suitable for promoting polymerase-catalyzed incorporation of nucleotides attached to the 3' end of the nucleic acid primer; (l) detecting the plurality of second ternary complexes formed in step (k) and identifying the plurality of incorporated nucleotides in the second ternary complexes; and (m) repeating steps (g) through (l) at least once. In some embodiments, the detecting in step (l) is optional. In some embodiments, the identifying in step (l) is optional.In some embodiments, the plurality of first ternary complexes of step (h) remain stable without dissociating (or exhibit reduced dissociation) the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second. In some embodiments, the plurality of second ternary complexes of step (k) remain stable without dissociating (or exhibit reduced dissociation) the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0055] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the non-catalytic divalent cation comprises strontium, barium, and / or calcium, and the catalytic divalent cation comprises magnesium or manganese.
[0056] In some embodiments, a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales comprises the steps of: (1) contacting a plurality of wild-type or mutant DNA polymerases and a plurality of nucleic acid primers with different portions of a concatemeric nucleic acid template molecule to form at least first and second composite polymerases on the same concatemeric template molecule; and (2) contacting a plurality of multivalent molecules with at least first and second composite polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second composite polymerases, wherein at least a first nucleotide unit of the single multivalent molecule binds to the first composite polymerase comprising a first primer that hybridizes to a first portion of the concatemeric template molecule, thereby forming a first concatemeric ternary complex, and at least a second nucleotide unit of the single multivalent molecule binds to the second portion of the concatemeric template molecule. (3) contacting a second concatemeric polymerase with a second primer that hybridizes to a portion of the first and second concatemeric ternary complexes, thereby forming a second concatemeric ternary complex, under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second concatemeric ternary complexes, whereby first and second concatemeric ternary complexes that bind to the same multivalent molecule form an affinity complex; (4) detecting the first and second concatemeric ternary complexes on the same concatemeric template molecule; and (5) identifying the first nucleotide unit in the first concatemeric ternary complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second concatemeric ternary complex, thereby determining the sequence of the second portion of the concatemeric template molecule.In some embodiments, the identifying in step (4) comprises identifying a first nucleotide unit that binds to the 3' end of a first primer in the first concatemeric ternary complex, thereby determining the sequence of a first portion of the concatemeric template molecule, and identifying a second nucleotide unit that binds to the 3' end of a second primer in the second concatemeric ternary complex, thereby determining the sequence of a second portion of the concatemeric template molecule.
[0057] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, there can be a multivalent molecule that can include any of the embodiments of a multivalent molecule that includes any of the potential properties described above.
[0058] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the multivalent molecule lacks a fluorophore. In some embodiments, the multivalent molecule is labeled with a fluorophore. In some embodiments, the plurality of multivalent molecules in step (b) are fluorescently labeled multivalent molecules, and step (c) comprises detecting fluorescent signals from the plurality of first ternary complexes and identifying a nucleotide unit attached to the 3' end of the nucleic acid primer, thereby determining the sequence of the nucleic acid template molecule.
[0059] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from a Candidatus Altiarchaeales archaeon, at least one of the nucleotide units of the polyvalent molecule comprises a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, comprising any of the potential characteristics described above. In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from a Candidatus Altiarchaeales archaeon, the nucleotides of steps (e) and / or (k) comprise nucleotide units comprising one or more exemplary nucleotide unit characteristics discussed above. In some embodiments, the plurality of nucleotides of step (e) lacks a fluorophore. In some embodiments, at least one of the plurality of nucleotides of step (e) is labeled with a fluorophore.
[0060] In some embodiments, in the two-phase nucleic acid sequencing method using the polymerase from the archaeon Candidatus Altiarchaeales, the plurality of nucleotides in steps (e) and / or (k) comprise a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the potential properties described above.
[0061] In some embodiments, in the two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the plurality of nucleotides in step (e) comprise a chain-terminating moiety attached to the 3' OH sugar position via a cleavable moiety, and step (f) further comprises contacting the chain-terminating nucleotide incorporated into the nucleic acid primer with a cleavage agent to remove the chain-terminating moiety, thereby generating a plurality of nucleic acid primers having 3' extendable ends.
[0062] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the plurality of second wild-type or mutant DNA polymerases may or may not be fluorescently labeled and may comprise any of the above-described fluorophore embodiments. In some embodiments, the plurality of second wild-type or mutant DNA polymerases comprise a plurality of nucleotides, which may or may not be fluorescently labeled and may comprise any of the above-described fluorophore embodiments.
[0063] In some embodiments, a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales includes a nucleic acid template molecule, which may include an embodiment of a nucleic acid template comprising any of the potential properties described above.
[0064] In some embodiments, in a two-phase nucleic acid sequencing method using a polymerase from the archaeon Candidatus Altiarchaeales, the plurality of first composite polymerases are immobilized according to an immobilization embodiment comprising any of the potential characteristics described above.
[0065] The present disclosure provides recombinant mutant 9°N DNA polymerases. As used herein, "recombinant mutant 9°N DNA" can include any of the properties described in this and the following paragraphs. For example, recombinant mutant 9°N The DNA may comprise the backbone amino acid sequence of SEQ ID NO: 280 or 281 or 282 and may have at least one or any combination of two or more amino acid substitution mutations, including: (1) a leucine (L) at position 408 is substituted with serine (S), phenylalanine (F), tyrosine (Y), valine (V), glycine (G), threonine (T), alanine (A), isoleucine (I), phenylalanine (F), or methionine (M); (2) a tyrosine (Y) at position 409 is substituted with alanine (A), threonine (T), serine (S), glycine (G), valine (V), isoleucine (I), or tyrosine (Y); and (3) a proline (P) at position 410 is substituted with glycine (G), serine (S), valine (V), cis- (4) the alanine (A) at position 485 is substituted with serine (S) or valine (V); (5) the serine (S) at position 492 is substituted with glycine (G); (6) the lysine (K) at position 507 is substituted with leucine (L), treptophan (W), tyrosine (Y), proline (P), or phenylalanine (F); (7) the isoleucine at position 521 is substituted with histidine (H), threonine (T), valine (V), serine (S), glycine (G), alanine (A), leucine (L), or phenylalanine (F); and / or (8) the lysine (K) at position 559 is substituted with aspartic acid (D). In some embodiments, the mutant 9°N DNA polymerase further comprises the amino acid substitutions D141A and E143A.
[0066] In some embodiments, the recombinant mutant 9°N DNA polymerase comprises the backbone amino acid sequence of SEQ ID NO: 280, or 281, or 282 and has a combination of amino acid substitution mutations including: (i) a leucine (L) at position 408 is substituted with serine (S), phenylalanine (F), or tyrosine (Y); (ii) a tyrosine (Y) at position 409 is substituted with alanine (A), a proline (P) at position 410 is substituted with glycine (G), an alanine (A) at position 485 is substituted with serine (S), a lysine (K) at position 507 is substituted with leucine (L), an isoleucine at position 521 is substituted with histidine (H), and a lysine (K) at position 559 is substituted with aspartic acid (D). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 280 and having the substitution mutations L408S, Y409A, P418G, A485S, K507L, I521H, and K559D (SEQ ID NO: 376). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 281 and having the substitution mutations L408S, Y409A, P418G, A485S, K507L, I521H, and K559D (SEQ ID NO: 379). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 282 and having the substitution mutations L408S, Y409A, P418G, A485S, K507L, I521H, and K559D (SEQ ID NO: 382). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 280 and having the substitution mutations L408F, Y409A, P418G, A485S, K507L, I521H, and K559D (SEQ ID NO: 377). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 281 and having the substitution mutations L408F, Y409A, P418G, A485S, K507L, I521H, and K559D (SEQ ID NO: 380). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 282 and having the substitution mutations L408F, Y409A, P418G, A485S, K507L, I521H, and K559D (SEQ ID NO: 383). In some embodiments, the mutant 9°N DNA polymerase further comprises the amino acid substitutions D141A and E143A.
[0067] In some embodiments, the recombinant mutant 9°N DNA polymerase comprises the backbone amino acid sequence of SEQ ID NO: 280, or 281, or 282 and has a combination of amino acid substitution mutations including: (i) a leucine (L) at position 408 is substituted with serine (S), phenylalanine (F), or tyrosine (Y); (ii) a tyrosine (Y) at position 409 is substituted with alanine (A), a proline (P) at position 410 is substituted with glycine (G), an alanine (A) at position 485 is substituted with serine (S), a serine (S) at position 492 is substituted with glycine (G), a lysine (K) at position 507 is substituted with leucine (L), an isoleucine at position 521 is substituted with histidine (H), and a lysine (K) at position 559 is substituted with aspartic acid (D). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 280 and having the substitution mutations L408F, Y409A, P418G, A485S, S492G, K507L, I521H, and K559D (SEQ ID NO: 378). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 281 and having the substitution mutations L408F, Y409A, P418G, A485S, S492G, K507L, I521H, and K559D (SEQ ID NO: 381). A mutant 9°N DNA polymerase based on the backbone sequence of SEQ ID NO: 282 and having the substitution mutations L408F, Y409A, P418G, A485S, S492G, K507L, I521H, and K559D (SEQ ID NO: 384). In some embodiments, the mutant 9°N DNA polymerase further comprises the amino acid substitutions D141A and E143A.
[0068] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a nucleic acid template molecule and a nucleotide polymerization initiation site having a 3' extendable end. The nucleic acid template molecule can comprise any of the nucleic acid template embodiments, including any of the potential characteristics described above.
[0069] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a nucleic acid template molecule, a nucleotide polymerization initiation site having a 3′ extendable end, and at least one nucleotide, wherein the at least one nucleotide comprises a nucleotide unit comprising one or more exemplary nucleotide unit characteristics discussed above.
[0070] In some embodiments, the recombinant mutant 9°N DNA polymerase is part of a ternary complex that includes the recombinant mutant 9°N DNA polymerase bound to a nucleic acid template molecule that hybridizes to a nucleic acid primer, and at least one nucleotide bound to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule. In some embodiments, in the ternary complex, the nucleotide is bound to the nucleic acid duplex and has not undergone polymerase-catalyzed incorporation, or the nucleotide is bound to the nucleic acid duplex and has undergone polymerase-catalyzed incorporation.
[0071] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a nucleic acid template molecule, a nucleotide polymerization initiation site having a 3′ extendable end, and at least one nucleotide that is labeled with a fluorophore or at least one nucleotide that lacks a fluorophore label.
[0072] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a nucleic acid template molecule, a nucleotide polymerization initiation site having a 3' extendable end, and at least one nucleotide comprising a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the potential properties described above.
[0073] In some embodiments, the recombinant mutant 9°N DNA polymerase comprises a recombinant mutant 9°N DNA polymerase that may or may not be fluorescently labeled and may include any of the fluorophore embodiments described above. In some embodiments, the recombinant mutant 9°N DNA polymerase comprises at least one nucleotide that may or may not be fluorescently labeled.
[0074] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises at least one multivalent molecule, which may include any of the embodiments of multivalent molecules comprising any of the potential features described above.
[0075] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises at least one multivalent molecule, which may or may not be fluorescently labeled as described above, including at least one fluorescently labeled multivalent molecule. In some embodiments, at least one multivalent molecule comprises a core labeled with a fluorophore. In some embodiments, at least one multivalent molecule comprises one or more nucleotide arms having linkers and / or nucleotide units attached to the fluorophore. In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a multivalent molecule lacking a fluorophore.
[0076] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises at least one multivalent molecule comprising a nucleotide unit having a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the potential properties described above.
[0077] In some embodiments, the recombinant mutant 9°N DNA polymerase comprises a polymerase that may or may not be fluorescently labeled and may include any of the fluorophore embodiments described above. In some embodiments, the recombinant mutant 9°N DNA polymerase comprises a polymerase and further comprises at least one multivalent molecule that may or may not be fluorescently labeled and may include any of the fluorophore embodiments described above.
[0078] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a nucleic acid template molecule, a nucleotide polymerization initiation site having a 3′ extendable end, and a plurality of catalytic divalent cations that facilitate polymerase-catalyzed incorporation of nucleotides, wherein the catalytic divalent cations comprise magnesium and / or manganese.
[0079] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a nucleic acid template molecule, a nucleotide polymerization initiation site having a 3′ extendable end, and a plurality of non-catalytic divalent cations that inhibit polymerase-catalyzed nucleotide incorporation, wherein the non-catalytic divalent cations comprise strontium, barium, and / or calcium.
[0080] In some embodiments, the recombinant mutant 9°N DNA polymerase further comprises a plurality of mutant DNA polymerases that bind to a plurality of nucleic acid template molecules and a plurality of nucleotide polymerization initiation sites, which form a plurality of composite mutant DNA polymerases (each comprising a mutant DNA polymerase that binds to a nucleic acid duplex) when the nucleic acid duplex comprises nucleic acid template molecules hybridized to an oligonucleotide primer. In some embodiments, the plurality of composite mutant DNA polymerases are immobilized according to an immobilization embodiment comprising any of the potential characteristics described above.
[0081] The present disclosure provides a method for nucleic acid sequencing using a mutant 9°N DNA polymerase and nucleotides. The disclosure provides a method for nucleic acid sequencing, comprising: (a) contacting (i) a plurality of mutant 9°N DNA polymerases; and (ii) a plurality of nucleic acid duplexes, each comprising a nucleic acid template molecule hybridized to a nucleic acid primer, wherein the contacting comprises a plurality of multiplex polymerases, each comprising a mutant 9°N DNA polymerase that binds to the nucleic acid duplex. (b) contacting the plurality of polymerase complexes with (iii) a plurality of nucleotides, and (iv) a plurality of catalytic or non-catalytic divalent cations, wherein the contacting is performed under conditions suitable to form a plurality of ternary complexes, each comprising a mutant DNA polymerase bound to a nucleic acid duplex and a nucleotide, wherein the nucleotide is attached to the 3' end of a nucleic acid primer at a position opposite its complementary nucleotide in the nucleic acid template molecule, and wherein the conditions are suitable to promote polymerase-catalyzed incorporation of the nucleotide attached to the 3' end of the nucleic acid primer, or the conditions are suitable to inhibit polymerase-catalyzed incorporation of the nucleotide attached to the 3' end of the nucleic acid primer; (c) detecting the plurality of ternary complexes; and (d) identifying the plurality of incorporated nucleotides in the plurality of ternary complexes. The description of mutant polymerases described above (e.g., the paragraph above regarding mutant 9°N DNA polymerase) also applies herein.
[0082] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, a plurality of ternary complexes remain stable (or exhibit reduced dissociation) without dissociating the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0083] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, the nucleotides are bound to a nucleic acid duplex and are not undergoing polymerase-catalyzed incorporation, or the nucleotides are bound to a nucleic acid duplex and are undergoing polymerase-catalyzed incorporation.
[0084] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, the plurality of catalytic divalent cations comprises magnesium and / or manganese. In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, the plurality of non-catalytic divalent cations comprises strontium, barium, and / or calcium.
[0085] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, the plurality of mutant 9°N DNA polymerases comprises a plurality of fluorescently labeled polymerases, or the plurality of mutant 9°N DNA polymerases lack a fluorophore.
[0086] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, individual nucleotides in the plurality of nucleotides comprise nucleotide units that include one or more exemplary nucleotide unit characteristics discussed above.
[0087] In some embodiments, sequencing methods using mutant 9°N DNA polymerases and nucleotides, which may or may not be fluorescently labeled, may include any of the fluorophore embodiments described above.
[0088] In some embodiments, in a sequencing method using a mutant 9°N DNA polymerase and nucleotides, at least one of the nucleotides in the plurality of nucleotides comprises a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the potential properties described above.
[0089] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, the plurality of nucleic acid template molecules can include embodiments of nucleic acid templates that include any of the above potential properties.
[0090] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and nucleotides, the multiplexed polymerases are immobilized according to an immobilization embodiment that includes any of the potential characteristics described above.
[0091] The present disclosure provides sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, each of which has been previously described and those descriptions also apply herein.
[0092] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, multiple ternary complexes remain stable (or exhibit reduced dissociation) without dissociating the mutant or wild-type polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0093] In some embodiments, in sequencing methods using a mutant 9°N DNA polymerase and a multivalent molecule, in the ternary complex, the nucleotide units of the multivalent molecule are either bound to the nucleic acid duplex and have not undergone polymerase-catalyzed incorporation, or the nucleotide units are bound to the nucleic acid duplex and have undergone polymerase-catalyzed incorporation.
[0094] In some embodiments, in a sequencing method using a mutant 9°N DNA polymerase and a multivalent molecule, the method further comprises forming an affinity complex. The method for forming an affinity complex comprises steps (steps (1) to (4)) similar to those described above for the two-phase nucleic acid sequencing method using the polymerase from Candidatus Altiarchaeales, and further comprises forming an affinity complex, the steps comprising: (a) forming a plurality of mutant 9°N DNA polymerases; (b) contacting a DNA polymerase and a plurality of nucleic acid primers with different portions of a concatemeric nucleic acid template molecule to form at least first and second composite polymerases on the same concatemeric template molecule; and (b) contacting a plurality of multivalent molecules with at least the first and second composite polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second composite polymerases, wherein at least a first nucleotide unit of the single multivalent molecule hybridizes to a first portion of the concatemeric template molecule, thereby binding to a first composite polymerase comprising a first primer that hybridizes to the first portion of the concatemeric template molecule, thereby forming a first ternary complex, and at least a second nucleotide unit of the single multivalent molecule hybridizes to a second portion of the concatemeric template molecule. (c) contacting the first and second ternary complexes with a second complex containing a second primer that reacts with the first nucleotide unit and the second primer that reacts with the second nucleotide unit to form a second ternary complex under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second ternary complexes (respectively), such that first and second ternary complexes that bind to the same multivalent molecule form an affinity complex; (d) identifying the first nucleotide unit in the first ternary complex, thereby determining the sequence of a first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second ternary complex, thereby determining the sequence of a second portion of the concatemeric template molecule.In some embodiments, the identifying in step (d) comprises identifying a first nucleotide unit that binds to the 3' end of a first primer in the first ternary complex, thereby determining the sequence of a first portion of the concatemeric template molecule, and identifying a second nucleotide unit that binds to the 3' end of a second primer in the second ternary complex, thereby determining the sequence of a second portion of the concatemeric template molecule.
[0095] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, the non-catalytic divalent cations comprise strontium, barium, and / or calcium.
[0096] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, the individual multivalent molecules can include any of the embodiments of multivalent molecules that include any of the potential properties described above.
[0097] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, the plurality of multivalent molecules may or may not be fluorescently labeled and may comprise any of the fluorophore embodiments described above.
[0098] In some embodiments, in a sequencing method using a mutant 9°N DNA polymerase and a multivalent molecule, at least one of the multivalent molecules in the plurality of multivalent molecules comprises a nucleotide unit having a chain-terminating moiety attached to its 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, comprising any of the potential properties described above.
[0099] In some embodiments, sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, any of which may or may not be fluorescently labeled, may include any of the fluorophore embodiments described above.
[0100] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, the plurality of nucleic acid template molecules can include embodiments of nucleic acid templates that include any of the potential properties described above.
[0101] In some embodiments, in sequencing methods using mutant 9°N DNA polymerases and multivalent molecules, the multiplexed polymerases are immobilized according to an immobilization embodiment that includes any of the potential characteristics described above.
[0102] The present disclosure provides a two-phase nucleic acid sequencing method using a mutant 9°N DNA polymerase, a multivalent molecule, and nucleotides. The disclosure provides a nucleic acid sequencing method, comprising: (a) contacting (i) a plurality of first mutant 9°N DNA polymerases with (ii) a plurality of nucleic acid duplexes (each comprising a nucleic acid template molecule hybridized to a nucleic acid primer), the contacting being performed under conditions suitable to form a plurality of first composite polymerases (each comprising a first mutant 9°N DNA polymerase bound to the nucleic acid duplex); and (b) contacting the plurality of first composite polymerases with (iii) a plurality of multivalent molecules and (iv) a plurality of non-catalytic divalent cations, each of the plurality of multivalent molecules comprising a core bound to a plurality of nucleotide arms, each nucleotide arm comprising a nucleotide unit, the contacting being performed under conditions suitable to form a plurality of first ternary complexes, each comprising a first mutant 9°N DNA polymerase bound to the nucleic acid duplex. (c) detecting the plurality of first ternary complexes and identifying the nucleotide unit attached to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule, thereby determining the sequence of the nucleic acid template molecule; (d) dissociating the plurality of first ternary complexes by removing the plurality of first mutant 9°N polymerases and the plurality of multivalent molecules, thereby retaining the plurality of nucleic acid duplexes; and (e) contacting the retained nucleic acid duplexes of step (d) with (i) a plurality of second mutant 9°N polymerases. and (iii) a plurality of catalytic divalent cations, wherein the contacting in step (e) results in a plurality of second ternary complexes, each of which contains a second mutant 9°N bound to the retained nucleic acid duplex of step (d) and a nucleotide.(f) detecting the plurality of second ternary complexes and identifying the incorporated nucleotides in the second ternary complexes. In some embodiments, the detecting step (f) is optional. In some embodiments, the identifying step (f) is optional. In some embodiments, the non-catalytic divalent cation comprises strontium, barium, and / or calcium, and the catalytic divalent cation comprises magnesium or manganese.
[0103] In some embodiments, the two-phase nucleic acid sequencing method using a mutant 9°N polymerase includes: (g) removing the plurality of second mutant 9°N DNA polymerases and retaining the plurality of nucleic acid duplexes of step (f); and (h) reacting the retained nucleic acid duplexes of step (g) with the plurality of first mutant 9°N DNA polymerases. contacting a plurality of nucleic acid template molecules with a DNA polymerase, a plurality of multivalent molecules, and a plurality of non-catalytic divalent cations, wherein the contacting is performed under conditions suitable to form another plurality of first ternary complexes, and wherein the contacting is performed under conditions suitable to inhibit polymerase-catalyzed incorporation of bound nucleotide units of the multivalent molecules; (i) detecting the plurality of first ternary complexes formed in step (h) and identifying a nucleotide unit attached to the 3' end of the nucleic acid primer, thereby determining the sequence of the nucleic acid template molecule; (j) dissociating the plurality of first ternary complexes formed in step (h) by removing the plurality of first mutant 9°N polymerases and the plurality of multivalent molecules, thereby retaining the plurality of nucleic acid duplexes; and (k) separating the plurality of retained nucleic acid duplexes of step (j) from a plurality of second mutant 9°N polymerases. The method further includes contacting a DNA polymerase, a plurality of nucleotides, and a plurality of catalytic divalent cations, wherein the contacting is performed under conditions suitable for forming another plurality of second ternary complexes, the conditions being suitable for promoting polymerase-catalyzed incorporation of a nucleotide attached to the 3' end of the nucleic acid primer; (l) detecting the plurality of second ternary complexes formed in step (k) and identifying the plurality of incorporated nucleotides in the second ternary complexes; and (m) repeating steps (g) through (l) at least once. In some embodiments, the detecting in step (l) is optional. In some embodiments, the identifying in step (l) is optional. In some embodiments, the non-catalytic divalent cation comprises strontium, barium, and / or calcium, and the catalytic divalent cation comprises magnesium or manganese.
[0104] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, either of them can be mutated as described above.
[0105] In some embodiments, in the two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of first mutant 9°N DNA polymerases form a plurality of first ternary complexes in step (b) that remain stable (or exhibit reduced dissociation) without dissociating the first mutant 9°N DNA polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0106] In some embodiments, in the two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of second mutant 9°N DNA polymerases form a plurality of first ternary complexes in step (e) that remain stable (or exhibit reduced dissociation) without dissociating the first mutant 9°N DNA polymerase from the nucleic acid duplex, and the stable ternary complexes exhibit a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0107] In some embodiments, a two-phase nucleic acid sequencing method using a mutant 9°N polymerase comprises, in a first ternary complex, a nucleotide unit of a multivalent molecule that is bound to a nucleic acid duplex and not undergoing polymerase-catalyzed incorporation, or a nucleotide unit that is bound to a nucleic acid duplex and undergoing polymerase-catalyzed incorporation. In some embodiments, in a second ternary complex, a nucleotide unit that is bound to a nucleic acid duplex and not undergoing polymerase-catalyzed incorporation, or a nucleotide unit that is bound to a nucleic acid duplex and undergoing polymerase-catalyzed incorporation.
[0108] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the method further comprises forming an affinity complex as described above for a polymerase from Candidatus Altiarchaeales, and comprises the same steps (steps (1)-(4)), including: (a) contacting a plurality of wild-type or mutant DNA polymerases and a plurality of nucleic acid primers with different portions of a concatemeric nucleic acid template molecule to form at least first and second composite polymerases on the same concatemeric template molecule; and (b) contacting a plurality of multivalent molecules with at least first and second composite polymerases on the same concatemeric template molecule under conditions suitable for binding of a single multivalent molecule from the plurality of multivalent molecules to the first and second composite polymerases, wherein at least a first nucleotide unit of the single multivalent molecule binds to a first composite polymerase comprising a first primer that hybridizes to a first portion of the concatemeric template molecule, thereby forming a first concatemeric ternary complex, and at least a second nucleotide unit of the single multivalent molecule binds to a first composite polymerase comprising a first primer that hybridizes to a first portion of the concatemeric template molecule, thereby forming a first concatemeric ternary complex. The method includes contacting the nucleotide units with a second composite polymerase comprising a second primer that hybridizes to a second portion of the concatemeric template molecule, thereby forming a second concatemeric ternary complex, under conditions suitable to inhibit polymerase-catalyzed incorporation of the bound first and second nucleotide units in the first and second concatemeric ternary complexes, such that first and second concatemeric ternary complexes bound to the same multivalent molecule form an affinity complex; (c) detecting the first and second concatemeric ternary complexes on the same concatemeric template molecule; and (d) identifying the first nucleotide unit in the first concatemeric ternary complex, thereby determining the sequence of the first portion of the concatemeric template molecule, and identifying the second nucleotide unit in the second concatemeric ternary complex, thereby determining the sequence of the second portion of the concatemeric template molecule.In some embodiments, the identifying in step (d) comprises identifying a first nucleotide unit that binds to the 3' end of a first primer in the first concatemeric ternary complex, thereby determining the sequence of a first portion of the concatemeric template molecule, and identifying a second nucleotide unit that binds to the 3' end of a second primer in the second concatemeric ternary complex, thereby determining the sequence of a second portion of the concatemeric template molecule.
[0109] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of first mutant 9°N DNA polymerases comprises a plurality of fluorescently labeled first mutant 9°N DNA polymerases. In some embodiments, the plurality of first mutant 9°N DNA polymerases lacks a fluorophore. In some embodiments, the plurality of second mutant 9°N DNA polymerases comprises a plurality of fluorescently labeled first mutant 9°N DNA polymerases. In some embodiments, the plurality of second mutant 9°N DNA polymerases lacks a fluorophore.
[0110] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the multivalent molecule may include any of the embodiments of a multivalent molecule that includes any of the potential properties described above.
[0111] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of multivalent molecules comprises a plurality of fluorescently labeled multivalent molecules. In some embodiments, the core of each multivalent molecule in the plurality is bound to a fluorophore corresponding to the nucleotide unit bound to the nucleotide arm. In some embodiments, at least one of the nucleotide arms of the multivalent molecule comprises a linker and / or a nucleotide base bound to a fluorophore, and the fluorophore bound to a given linker or nucleotide base corresponds to the nucleotide base of the nucleotide arm (e.g., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, the plurality of multivalent molecules in step (b) are fluorescently labeled multivalent molecules, and step (c) comprises detecting fluorescent signals from the plurality of first ternary complexes and identifying the nucleotide unit bound to the 3' end of the nucleic acid primer, thereby determining the sequence of the nucleic acid template molecule. In some embodiments, the plurality of multivalent molecules lacks a fluorophore.
[0112] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, at least one of the nucleotide units of a plurality of one or more multivalent molecules comprises a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the potential properties described above.
[0113] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, each nucleotide in the plurality of nucleotides comprises a nucleotide unit comprising one or more exemplary nucleotide unit characteristics discussed above.
[0114] In some embodiments, in the two-phase nucleic acid sequencing method using the mutant 9°N polymerase, at least one of the nucleotides in the plurality of nucleotides comprises a fluorescently labeled nucleotide. In some embodiments, the plurality of nucleotides lacks a fluorophore label.
[0115] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of second mutant 9°N DNA polymerases may or may not be fluorescently labeled and may comprise any of the fluorophore embodiments described above. In some embodiments, the plurality of polymerases comprises one or more nucleotides, any of which may or may not be fluorescently labeled and may comprise any of the fluorophore embodiments described above.
[0116] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, at least one of the nucleotides in the plurality of nucleotides comprises a chain-terminating moiety attached to the 3'-OH sugar position via a cleavable moiety, which may comprise any of the chain-terminating moiety embodiments described above. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group, including any of the potential properties described above.
[0117] In some embodiments, in the two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of nucleotides in step (e) comprises a chain-terminating moiety attached to the 3'OH sugar position via a cleavable moiety, and step (f) further comprises contacting the chain-terminating nucleotide incorporated into the nucleic acid primer with a cleavage agent to remove the chain-terminating moiety, thereby generating a plurality of nucleic acid primers having 3' extendable ends.
[0118] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of nucleic acid template molecules can include embodiments of nucleic acid templates that include any of the above potential properties.
[0119] In some embodiments, in a two-phase nucleic acid sequencing method using a mutant 9°N polymerase, the plurality of first composite polymerases are immobilized according to an immobilization embodiment comprising any of the potential characteristics described above.
[0120] This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication containing color figure(s) will be provided by the U.S. Patent and Trademark Office upon request and payment of the necessary fee.
[0121] The novel advantages and features of the compositions and methods disclosed herein are set forth with particularity in the appended claims. A better understanding of the features and advantages of the compositions and methods of the present disclosure can be obtained by reference to the following detailed description that sets forth illustrative embodiments and the accompanying drawings. [Brief explanation of the drawings]
[0122] [Figure 1] 1 is a graph comparing the rates of product formation for purified wild-type and mutant DNA polymerases from the archaeon Candidatus Altiarchaeales in the presence of various nucleotide concentrations of 3′-methylazido-dCTP. The graph shows data for mutant polymerases having the amino acid sequences of SEQ ID NOs: 105, 129, 130, 134, 141, 149, 151, and 155.
[0123] [Figure 2] 1 is a graph showing the relative percentage incorporation of 3'-methylazido nucleotides by mutants of Bst polymerase (eg, DNA polymerase I from Geobacillus stearothermophilus).
[0124] [Figure 3A]Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3B] Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3C] Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3D]Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3E] Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3F] Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3G]Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1. [Figure 3H] Table 1 (eight sheets presented as Figures 3A-3H) lists the relative incorporation activity of wild-type (SEQ ID NO: 1) and mutant variants (SEQ ID NOs: 2-157) of DNA polymerase from the archaeon Candidatus Altiarchaeales in incorporating 3' methyl azidonucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases (SEQ ID NOs: 2-157) listed in Table 1 contain the substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 1.
[0125] [Figure 4A] Table 2 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 158-255) for incorporation of 3' methyl azido nucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases listed in Table 2 (SEQ ID NOS: 158-255) contain substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 2. [Figure 4B]Table 2 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 158-255) for incorporation of 3' methyl azido nucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases listed in Table 2 (SEQ ID NOS: 158-255) contain substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 2. [Figure 4C] Table 2 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 158-255) for incorporation of 3' methyl azido nucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases listed in Table 2 (SEQ ID NOS: 158-255) contain substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 2. [Figure 4D] Table 2 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 158-255) for incorporation of 3' methyl azido nucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases listed in Table 2 (SEQ ID NOS: 158-255) contain substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 2. [Figure 4E]Table 2 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 158-255) for incorporation of 3' methyl azido nucleotides at the N+1 position of elongating polynucleotide chains at 42°C. The variants are present in lysates removed from expression strains. All of the mutant polymerases listed in Table 2 (SEQ ID NOS: 158-255) contain substitution mutations D141A and E143A, even if the genotype does not have these substitution mutations listed in Table 2.
[0126] [Figure 5A] Table 3 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 288-375, 385-390, and 394-397) for incorporation of a 3'-methyl azidonucleotide at the N+1 position of a growing polynucleotide chain at 42°C. The variants are present in lysates removed from expression strains. Mutant polymerases having amino acid sequences SEQ ID NOS: 353, 354, and 386 are truncation mutants, and the cleavage site is indicated with an asterisk (*). All of the mutant polymerases listed in Table 3 (SEQ ID NOS: 288-375 and 385-390) contain substitution mutations D141A and E143A, even if the genotype does not contain these substitution mutations as listed in Table 3. [Figure 5B]Table 3 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 288-375, 385-390, and 394-397) for incorporation of a 3'-methyl azidonucleotide at the N+1 position of a growing polynucleotide chain at 42°C. The variants are present in lysates removed from expression strains. Mutant polymerases having amino acid sequences SEQ ID NOS: 353, 354, and 386 are truncation mutants, and the cleavage site is indicated with an asterisk (*). All of the mutant polymerases listed in Table 3 (SEQ ID NOS: 288-375 and 385-390) contain substitution mutations D141A and E143A, even if the genotype does not contain these substitution mutations as listed in Table 3. [Figure 5C] Table 3 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 288-375, 385-390, and 394-397) for incorporation of a 3'-methyl azidonucleotide at the N+1 position of a growing polynucleotide chain at 42°C. The variants are present in lysates removed from expression strains. Mutant polymerases having amino acid sequences SEQ ID NOS: 353, 354, and 386 are truncation mutants, and the cleavage site is indicated with an asterisk (*). All of the mutant polymerases listed in Table 3 (SEQ ID NOS: 288-375 and 385-390) contain substitution mutations D141A and E143A, even if the genotype does not contain these substitution mutations as listed in Table 3. [Figure 5D]Table 3 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 288-375, 385-390, and 394-397) for incorporation of a 3'-methyl azidonucleotide at the N+1 position of a growing polynucleotide chain at 42°C. The variants are present in lysates removed from expression strains. Mutant polymerases having amino acid sequences SEQ ID NOS: 353, 354, and 386 are truncation mutants, and the cleavage site is indicated with an asterisk (*). All of the mutant polymerases listed in Table 3 (SEQ ID NOS: 288-375 and 385-390) contain substitution mutations D141A and E143A, even if the genotype does not contain these substitution mutations as listed in Table 3. [Figure 5E] Table 3 lists the relative incorporation activity of mutant variants of DNA polymerase from the archaeon Candidatus Altiarchaeales (SEQ ID NOS: 288-375, 385-390, and 394-397) for incorporation of a 3'-methyl azidonucleotide at the N+1 position of a growing polynucleotide chain at 42°C. The variants are present in lysates removed from expression strains. Mutant polymerases having amino acid sequences SEQ ID NOS: 353, 354, and 386 are truncation mutants, and the cleavage site is indicated with an asterisk (*). All of the mutant polymerases listed in Table 3 (SEQ ID NOS: 288-375 and 385-390) contain substitution mutations D141A and E143A, even if the genotype does not contain these substitution mutations as listed in Table 3.
[0127] [Figure 6A]Table 4 lists various amino acid substitution mutations in the DNA polymerase from the archaeon Candidatus altiarchaeales (relative to SEQ ID NO: 1), as well as the equivalent amino acid substitution mutations in 9°N DNA polymerase (relative to SEQ ID NO: 280), VENT DNA polymerase (relative to SEQ ID NO: 283), DEEP VENT DNA polymerase (relative to SEQ ID NO: 284), Geobacillus stearothermophilus DNA polymerase (relative to SEQ ID NO: 275), Pfu DNA polymerase (relative to SEQ ID NO: 285), and Pyrococcus abyssi DNA polymerase (relative to SEQ ID NO: 286). [Figure 6B] Table 4 lists various amino acid substitution mutations in the DNA polymerase from the archaeon Candidatus altiarchaeales (relative to SEQ ID NO: 1), as well as the equivalent amino acid substitution mutations in 9°N DNA polymerase (relative to SEQ ID NO: 280), VENT DNA polymerase (relative to SEQ ID NO: 283), DEEP VENT DNA polymerase (relative to SEQ ID NO: 284), Geobacillus stearothermophilus DNA polymerase (relative to SEQ ID NO: 275), Pfu DNA polymerase (relative to SEQ ID NO: 285), and Pyrococcus abyssi DNA polymerase (relative to SEQ ID NO: 286). [Figure 6C] Table 4 lists various amino acid substitution mutations in the DNA polymerase from the archaeon Candidatus altiarchaeales (relative to SEQ ID NO: 1), as well as the equivalent amino acid substitution mutations in 9°N DNA polymerase (relative to SEQ ID NO: 280), VENT DNA polymerase (relative to SEQ ID NO: 283), DEEP VENT DNA polymerase (relative to SEQ ID NO: 284), Geobacillus stearothermophilus DNA polymerase (relative to SEQ ID NO: 275), Pfu DNA polymerase (relative to SEQ ID NO: 285), and Pyrococcus abyssi DNA polymerase (relative to SEQ ID NO: 286).
[0128] [Figure 7A] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and 9°N DNA polymerase (to SEQ ID NO: 280) from the archaeon Candidatus altiarchaeales. [Figure 7B] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and 9°N DNA polymerase (to SEQ ID NO: 280) from the archaeon Candidatus altiarchaeales.
[0129] [Figure 8A] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and VENT DNA polymerase (to SEQ ID NO: 283) from the archaeon Candidatus altiarchaeales. [Figure 8B] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and VENT DNA polymerase (to SEQ ID NO: 283) from the archaeon Candidatus altiarchaeales. [Figure 8C] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and VENT DNA polymerase (to SEQ ID NO: 283) from the archaeon Candidatus altiarchaeales.
[0130] [Figure 9A] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and DEEP VENT DNA polymerase (to SEQ ID NO: 284) from the archaeon Candidatus altiarchaeales. [Figure 9B]1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and DEEP VENT DNA polymerase (to SEQ ID NO: 284) from the archaeon Candidatus altiarchaeales.
[0131] [Figure 10A] 1 is an amino acid sequence alignment between the wild-type DNA polymerase from the archaeon Candidatus altiarchaeales (to SEQ ID NO: 1) and Geobacillus stearothermophilus DNA polymerase (to SEQ ID NO: 275). [Figure 10B] 1 is an amino acid sequence alignment between the wild-type DNA polymerase from the archaeon Candidatus altiarchaeales (to SEQ ID NO: 1) and Geobacillus stearothermophilus DNA polymerase (to SEQ ID NO: 275).
[0132] [Figure 11A] 1 is an amino acid sequence alignment between wild-type DNA polymerase from the archaeon Candidatus altiarchaeales (to SEQ ID NO: 1) and Pfu DNA polymerase (to SEQ ID NO: 285). [Figure 11B] 1 is an amino acid sequence alignment between wild-type DNA polymerase from the archaeon Candidatus altiarchaeales (to SEQ ID NO: 1) and Pfu DNA polymerase (to SEQ ID NO: 285).
[0133] [Figure 12A] 1 is an amino acid sequence alignment between the wild-type DNA polymerase from the archaeon Candidatus altiarchaeales (to SEQ ID NO: 1) and the Pyrococcus abyssi polymerase (to SEQ ID NO: 286). [Figure 12B]1 is an amino acid sequence alignment between the wild-type DNA polymerase from the archaeon Candidatus altiarchaeales (to SEQ ID NO: 1) and the Pyrococcus abyssi polymerase (to SEQ ID NO: 286).
[0134] [Figure 13A] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and RB69 polymerase (to SEQ ID NO: 287) from the archaeon Candidatus altiarchaeales. [Figure 13B] 1 is an amino acid sequence alignment between wild-type DNA polymerase (to SEQ ID NO: 1) and RB69 polymerase (to SEQ ID NO: 287) from the archaeon Candidatus altiarchaeales.
[0135] [Figure 14A] FIG. 1 is a schematic diagram of a multivalent molecule comprising a general core attached to multiple nucleotide arms.
[0136] [Figure 14B] FIG. 1 is a schematic diagram of a multivalent molecule comprising a dendrimer core attached to multiple nucleotide arms.
[0137] [Figure 15A] 1 is a schematic diagram of a multivalent molecule comprising a core bound to multiple nucleotide arms, the nucleotide arms comprising biotin, a spacer, a linker, and a nucleotide unit.
[0138] [Figure 15B] FIG. 1 is a schematic diagram of a nucleotide arm comprising a core-binding moiety, a spacer, a linker, and a nucleotide unit.
[0139] [Figure 16A]The chemical structures of exemplary spacers and various exemplary linkers are shown, including an 11-atom linker, a 16-atom linker, a 23-atom linker, and an N3 linker.
[0140] [Figure 16B] 1 shows the chemical structures of various exemplary linkers, including linkers 1-9.
[0141] [Figure 17A] 1 shows the chemical structures of various exemplary linkers that link / bond nucleotide units.
[0142] [Figure 17B] 1 shows the chemical structures of various exemplary linkers that link / bond nucleotide units.
[0143] [Figure 17C] 1 shows the chemical structures of various exemplary linkers that link / bond nucleotide units.
[0144] [Figure 18] The chemical structure of an exemplary nucleotide arm is shown. In this example, the nucleotide unit is connected to the linker via a propargylamine bond at the 5-position of the pyrimidine base or the 7-position of the purine base. This nucleotide arm represents an exemplary biotinylated nucleotide arm.
[0145] [Figure 19] 1 is a bar graph comparing the rates of product formation for purified wild-type and mutant DNA polymerases from the archaeon Candidatus Altiarchaeales in the presence of 3' methylazido-dCTP. The graph shows data for mutant polymerases having the amino acid sequences of SEQ ID NOs: 39, 297, 27, 164, or 225.
[0146] [Figure 20]1 is a series of graphs showing the results of primer extension reactions on polony immobilized on a flow cell using an engineered polymerase (e.g., SEQ ID NO: 27), and the length of the extension product was monitored by capillary electrophoresis. DETAILED DESCRIPTION OF THE INVENTION
[0147] Definition: The headings provided herein are not limitations on the various aspects of the disclosure, which aspects can be understood by reference to the specification as a whole.
[0148] Unless otherwise defined, technical and scientific terms used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. In general, terms related to molecular biology, nucleic acid chemistry, protein chemistry, genetics, microbiology, transgenic cell production, and hybridization techniques described herein are well known and commonly used in the art. The techniques and procedures described herein are generally performed according to conventional methods well known in the art and as described in various general and more specific references cited and discussed throughout the specification. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). See also Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992). The nomenclature associated with the specification and the experimental procedures and techniques described herein are those well known and commonly used in the art.
[0149] Unless otherwise required by context herein, singular terms shall include plurals and plural terms shall include the singular. The singular forms "a," "an," and "the," as well as the singular use of any word, include plural referents unless expressly and unambiguously limited to one referent.
[0150] The use of alternative language (eg, "or") is understood to mean either one or both of the alternatives, or any combination thereof.
[0151] The term "and / or" as used herein is intended to imply specific disclosure of each of the specified features or components, regardless of the presence or absence of others. For example, the term "and / or" used in phrases such as "A and / or B" herein is intended to include "A and B," "A or B," "A" (A only), and "B" (B only). Similarly, the term "and / or" used in phrases such as "A, B, and / or C" is intended to encompass each of the following embodiments: "A, B, and C," "A, B, or C," "A or C," "A or B," "B or C," "A and B," "B and C," "A and C," "A" (A only), "B" (B only), and "C" (C only).
[0152] As used in this specification and the appended claims, the terms "comprising," "including," "having," and "containing," as well as grammatical variants thereof as used herein, are intended to be open-ended, such that one or more items in a list do not exclude other items that may be substituted for or added to the listed items. When embodiments are described herein using "comprising" language, it is understood that similar embodiments otherwise described in the sense of "consisting of" and / or "consisting essentially of" are also provided.
[0153] As used herein, the terms "about" and "approximately" refer to a value or composition that falls within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which depends in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "approximately" can mean within one or more standard deviations, according to practice in the art. Alternatively, "about" or "approximately" can mean a range of up to 10% (i.e., ±10%) or more, depending on the limitations of the measurement system. For example, about 5 mg can include any number between 4.5 mg and 5.5 mg. Furthermore, particularly with respect to biological systems or processes, the term can mean up to an order of magnitude, or up to five times the value. When a particular value or composition is provided in this disclosure, unless otherwise specified, the meaning of "about" or "approximately" should be assumed to be within an acceptable error range for that particular value or composition. Also, where ranges and / or sub-ranges of values are provided, the ranges and / or sub-ranges may include the endpoints of the ranges and / or sub-ranges.
[0154] The terms "peptide," "polypeptide," and "protein," as well as other related terms used herein, are used interchangeably and refer to a polymer of amino acids and are not limited to any particular length. Polypeptides can contain natural and unnatural amino acids. Polypeptides include recombinant or chemically synthesized forms. Polypeptides also include precursor molecules that have not yet undergone post-translational modifications, such as proteolytic cleavage, ribosomal skipping cleavage, hydroxylation, methylation, lipidation, acetylation, sumoylation, ubiquitination, glycosylation, phosphorylation, and / or disulfide bond formation. These terms encompass natural and artificial proteins, protein fragments, and polypeptide analogs of protein sequences (such as muteins, variants, chimeric proteins, and fusion proteins), as well as proteins that are post-translationally or otherwise covalently or non-covalently modified.
[0155] The terms "polymerase" and variants thereof, as used herein, include any enzyme capable of catalyzing the polymerization of nucleotides (including analogs thereof) into nucleic acid strands. Typically, although not necessarily, such nucleotide polymerization can occur in a template-dependent manner. Typically, a polymerase contains one or more active sites at which catalysis of nucleotide binding and / or nucleotide polymerization can occur. In some embodiments, a polymerase contains other enzymatic activities, such as, for example, 3' to 5' exonuclease activity or 5' to 3' exonuclease activity. In some embodiments, a polymerase has strand displacement activity. Polymerases include, but are not limited to, naturally occurring polymerases and any subunits thereof and truncations, mutant polymerases, variant polymerases, recombinant, fused, or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, as well as any analogs, derivatives, or fragments thereof (e.g., catalytically active fragments) that retain the ability to catalyze nucleotide polymerization. In some embodiments, a polymerase can be isolated from a cell or produced using recombinant DNA technology or chemical synthesis methods. In some embodiments, the polymerase can be expressed in a prokaryotic, eukaryotic, viral, or phage organism. In some embodiments, the polymerase can be a post-translationally modified protein or fragment thereof. The polymerase can be derived from a prokaryotic, eukaryotic, viral, or phage organism. Polymerases include DNA-guided DNA polymerases and RNA-guided DNA polymerases.
[0156] As used herein, the term "fidelity" refers to the accuracy of DNA polymerization by a template-dependent DNA polymerase. The fidelity of a DNA polymerase is typically measured by the error rate (the frequency of incorporating an incorrect nucleotide, i.e., a nucleotide that is not complementary to the template nucleotide). The accuracy or fidelity of DNA polymerization is maintained by both the polymerase activity and the 3'-5' exonuclease activity of the DNA polymerase.
[0157] As used herein, the term "bound complex" refers to a complex formed by binding together a nucleic acid duplex, a polymerase, and a free nucleotide or nucleotide unit of a multivalent molecule, where the nucleic acid duplex includes a nucleic acid template molecule hybridized to a nucleic acid primer. In the bound complex, the free nucleotide or nucleotide unit may or may not be bound to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide in the nucleic acid template molecule. A "ternary complex" is an example of a bound complex formed by binding together a nucleic acid duplex, a polymerase, and a free nucleotide or nucleotide unit of a multivalent molecule, where the free nucleotide or nucleotide unit is bound to the 3' end of the nucleic acid primer (as part of the nucleic acid duplex) at a position opposite the complementary nucleotide in the nucleic acid template molecule.
[0158] The term "duration" and related terms refer to the length of time a binding complex remains stable without dissociating any of its components, including the nucleic acid template and nucleic acid primer, polymerase, and nucleotide units or free (e.g., unbound) nucleotides of a multivalent molecule. The nucleotide units or free nucleotides can be complementary or non-complementary to nucleotide residues in the template molecule. The nucleotide units or free nucleotides can be bound to the 3' end of the nucleic acid primer at a position opposite the complementary nucleotide residue in the nucleic acid template molecule. Duration indicates the stability of the binding complex and the strength of the binding interaction. Duration can be measured by observing the onset and / or duration of the binding complex, for example, by observing a signal from a labeled component of the binding complex. For example, labeled nucleotides or labeled reagents comprising one or more nucleotides can be present in the binding complex, thus allowing a signal from the label to be detected during the duration of the binding complex. One exemplary label is a fluorescent label. The bound complex (e.g., ternary complex) remains stable until exposed to conditions that cause dissociation of interactions between the polymerase, template molecule, primer, and / or any of the nucleotide units or nucleotides. For example, dissociation conditions include contacting the bound complex with any one or any combination of detergent, EDTA, and / or water. In some embodiments, the bound complex remains stable without dissociation for a duration of greater than 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 second.
[0159] The terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms used herein, are used interchangeably and refer to a polymer of nucleotides and are not limited to any particular length. Nucleic acids include recombinant and chemically synthesized forms. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA produced using nucleotide analogs (e.g., peptide nucleic acids and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, where the nucleotides comprise natural or non-natural bases and / or sugars. Nucleic acids comprise naturally occurring internucleoside linkages, such as phosphodiester linkages. Nucleic acids comprise non-natural internucleoside linkages, including phosphorothioate, phosphorothiolate, or peptide nucleic acid (PNA) linkages. In some embodiments, nucleic acids comprise one type of polynucleotide or a mixture of two or more different types of polynucleotides.
[0160] As used herein, the term "primer" and related terms refer to an oligonucleotide, either natural or synthetic, that can hybridize to a DNA and / or RNA polynucleotide template and form a double-stranded molecule. Primers can be of any length but typically range from 4 to 50 nucleotides. A typical primer includes a 5' end and a 3' end. The 3' end of a primer can contain a 3' OH moiety that functions as a nucleotide polymerization initiation site in a polymerase-mediated primer extension reaction. Alternatively, the 3' end of a primer can lack a 3' OH moiety or can contain a terminal 3' blocking group that inhibits nucleotide polymerization in a polymerase-mediated reaction. Any one or more nucleotides along the length of the primer can be labeled with a detectable reporter moiety. Primers can be in solution (e.g., soluble primers) or immobilized on a support (e.g., capture primers).
[0161] The terms "template nucleic acid," "template polynucleotide," "target nucleic acid," "target polynucleotide," "template strand," and other variations refer to a nucleic acid strand that serves as a base nucleic acid molecule for generating a complementary nucleic acid strand. The sequence of the template nucleic acid can be partially or completely complementary to the sequence of the complementary strand. Template nucleic acids can be obtained from naturally occurring sources, recombinant, or chemically synthesized to contain any type of nucleic acid analog. Template nucleic acids can be linear, circular, or in other forms. Template nucleic acids can be isolated in any form, including chromosomes, genomes, organelles (e.g., mitochondria, chloroplasts, or ribosomes), recombinant molecules, cloned, amplified, RNA such as cDNA, precursor mRNA or mRNA, oligonucleotides, total genomic DNA obtained from fresh, frozen, paraffin-embedded tissue, needle biopsies, cell-free circulating DNA, or any type of nucleic acid library. Template nucleic acid molecules can be isolated from any source, including prokaryotes, eukaryotes (e.g., humans, plants, and animals), fungi, and viruses; cells; tissues; bodily fluids, including normal or diseased cells or tissues, blood, urine, serum, lymph, tumors, saliva, anal and vaginal secretions, amniotic fluid samples, sweat, and semen; environmental samples; culture samples; or synthetic nucleic acid molecules prepared using recombinant molecular biology or chemical synthesis methods. Template nucleic acids can be subjected to nucleic acid analysis, including sequencing and compositional analysis.
[0162] When used in reference to nucleic acid molecules, the terms "hybridize" or "hybridizing" or "hybridization" or other related terms refer to hydrogen bonding between two different nucleic acids to form a double-stranded nucleic acid. Hybridization also includes hydrogen bonding between two different regions of a single nucleic acid molecule to form a self-hybridizing molecule having a double-stranded region. Hybridization can involve Watson-Crick or Hoogstein binding to form a duplex double-stranded nucleic acid or a double-stranded region within a nucleic acid molecule. A double-stranded nucleic acid, or two different regions of a single nucleic acid, can be fully complementary or partially complementary. Complementary nucleic acid strands need not hybridize to each other over their entire length. Complementary base pairing can be standard AT or CG base pairing or can be other forms of base pairing interactions. A double-stranded nucleic acid can contain mismatched base-paired nucleotides.
[0163] The term "nucleotide" and related terms refer to a molecule comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose), and at least one phosphate group. Canonical or non-canonical nucleotides are consistent with the use of the terms. In some embodiments, the phosphate comprises a monophosphate, diphosphate, or triphosphate, or the corresponding phosphate analogs. In some embodiments, a nucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 phosphate groups. The term "nucleoside" refers to a molecule comprising an aromatic base and a sugar.
[0164] Nucleotides (and nucleosides) typically contain a heterocyclic base containing a substituted or unsubstituted nitrogen-containing parent heteroaromatic ring commonly found in nucleic acids, including naturally occurring, substituted, modified, or engineered variants, or analogs thereof. The base of a nucleotide (or nucleoside) can form Watson-Crick and / or Hoogstein hydrogen bonds with an appropriate complementary base. Exemplary bases include, for example, 2-aminopurine, 2,6-diaminopurine, adenine (A), ethenoadenine, N, N-aminopurine ... 6 -Δ 2 -Isopentenyladenine (6iA), N 6 -Δ 2 -Isopentenyl-2-methylthioadenine (2ms6iA), N 6 -Methyladenine, guanine (G), isoguanine, N 2 -dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine and O 6 -methylguanine; 7-deaza-purines such as 7-deazadenine (7-deaza-A) and 7-deazaguanine (7-deaza-G); cytosine (C), 5-propynylcytosine, isocytosine, thymosin (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O 4 Examples of bases include, but are not limited to, purines and pyrimidines such as methylthymine, uracil (U), 4-thiouracil (4sU), and 5,6-dihydrouracil (dihydracyl; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; nebularine; inosine; hydroxymethylcytosine; 5-methylcytosine (methycytosine); base (Y); and methylated, glycated, and acylated base moieties. Additional exemplary bases can be found in Fasman, 1989, in "Practical Handbook of Biochemistry and Molecular Biology," pp. 385-394, CRC Press, Boca Raton, Fla.
[0165] Nucleotides (and nucleosides) typically include a sugar moiety, such as a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez, et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez, et al., 1997 Bioorganic & Medicinal Chemistry Letters vol. 7:3013-3016), and another sugar moiety (Joeng, et al., 1993 J. Med. Chem. 36:2627-2638; Kim, et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124 and U.S. Pat. No. 5,558,991). Sugar moieties include ribosyl, 2'-deoxyribosyl, 3'-deoxyribosyl, 2',3'-dideoxyribosyl, 2',3'-didehydrodideoxyribosyl, 2'-alkoxyribosyl, 2'-azidoribosyl, 2'-aminoribosyl, 2'-fluoribosyl, 2'-mercaptoriboxyl, 2'-alkylribosyl, 3'-alkoxyribosyl, 3'-azidoribosyl, 3'-aminoribosyl, 3'-fluoribosyl, 3'-mercaptoriboxyl, 3'-alkylribosylcarbosyl, acyclic, or other modified sugars.
[0166] In some embodiments, the nucleotide comprises a chain of one, two, or three phosphorus atoms, typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramido linkage. In some embodiments, the nucleotide is an analog having a phosphorus chain linked together with an intervening O, S, NH, methylene, or ethylene atom. In some embodiments, the phosphorus atom in the chain comprises a substituted side group comprising O, S, or BH3. In some embodiments, the chain comprises a phosphate group substituted with an analog comprising a phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite group.
[0167] When used in reference to nucleic acids, the terms "extend," "extending," "extension," and other variants refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation involves the polymerization of one or more nucleotides to the terminal 3'OH terminus of a nucleic acid chain, resulting in the elongation of the nucleic acid chain. Nucleotide incorporation can be performed using natural nucleotides and / or nucleotide analogs. Typically, although not necessarily, nucleotide incorporation occurs in a template-dependent manner. Any suitable method for extending a nucleic acid molecule can be used, including primer extension catalyzed by DNA polymerase or RNA polymerase.
[0168] The terms "reporter moiety," "reporter moieties," or related terms refer to a compound that generates or causes to be generated a detectable signal. A reporter moiety is sometimes referred to as a "label." Any suitable reporter moiety can be used, including luminescent, photoluminescent, electroluminescent, bioluminescent, chemiluminescent, fluorescent, phosphorescent, chromophore, radioisotope, electrochemical, mass spectrometry, Raman, hapten, affinity tag, atom, or enzyme. The reporter moiety generates a detectable signal due to a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzymatic activity, or a proximity event). A proximity event includes the proximity, association, or binding of two reporter moieties to each other. It is well known to those skilled in the art to select reporter moieties so that each absorbs excitation radiation and / or fluoresces at wavelengths that are distinguishable from other reporter moieties, allowing the presence of different reporter moieties to be monitored in the same or different reactions. Two or more different reporter moieties can be selected that have spectrally distinct emission profiles or that have minimally overlapping spectral emission profiles. The reporter moiety can be linked (e.g., operably linked) to a nucleotide, a nucleoside, a nucleic acid, an enzyme (e.g., a polymerase or a reverse transcriptase), or a support (e.g., a surface).
[0169] Reporter moieties (or labels) include fluorescent labels or fluorophores. Exemplary fluorescent moieties that can function as fluorescent labels or fluorophores include fluorescein and fluorescein derivatives, such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynapthofluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamidofluorescein, fluorescein maleimide, SAMSA-fluorescein, fluorescein thiosemicarbazide, carbohydrazinomethylthioacetyl-aminofluorescein, rhodamine and rhodamine derivatives. Conductors such as TRITC, TMR, Lissamine rhodamine, Texas Red, rhodamine B, rhodamine 6G, rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, Lissamine rhodamine B sulfonyl chloride, Lissamine rhodamine B sulfonyl hydrazine, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY 530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503, Cascade Blue and derivatives such as Cascade Blue acetyl azide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanines and derivatives such as indolium-based cyanine dyes, benzo-indolium-based cyanine dyes, pyridium-based cyanine dyes, thiozolium-based cyanine dyes, quinolinium-based cyanine dyes, imidazolium-based cyanine dyes, Cy3, Cy5, lanthanide chelates and derivatives such as BCPDA, TBP, TMT,Examples of suitable dyes include, but are not limited to, BHHCT, BCOT, europium chelates, terbium chelates, Alexa Fluor dyes, DyLight dyes, Atto dyes, LightCycler Red dyes, CAL Flour dyes, JOE and their derivatives, Oregon Green dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, Malachite green, stilbenes, DEG dyes, NR dyes, near-infrared dyes, and others known in the art, such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition, Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999), or Hermanson, Bioconjugate Techniques, 2nd Edition, or any derivatives thereof, or any combination thereof. Cyanine dyes can exist in either sulfonated or non-sulfonated forms and consist of two indolenine, benzo-indolium, pyridinium, thiozolium, and / or quinolinium groups separated by a polymethine bridge between the two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3 (1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indol-2-ylidene). indolium or 1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-2-(3-{1-[6-(2,5-dioxopyrrolidin-1-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene}prop-1-en-1-yl)-3,3-dimethyl-3H-indolium-5-sulfonate),Cy5(1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indolin-2-ylidene)penta-1,3-dien-1-yl)-3,3-dimethyl- 3H-Indol-1-ium or 1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-2-((1E,3E)-5-((E)-1-(6-((2,5-dioxopyrrolidin-1-yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-sulfoindolin-2-ylidene)penta-1,3-diene Cy7 (1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium or 1-(5-carboxypentyl)-2-[(1E,3E,5E,7Z)-7-(1-ethyl-5-sulfo-1,3-dihydro-2H-indol-2-ylidene)hepta-1,3,5-trien-1-yl]-3H-indolium-5-sulfonate), where "Cy" stands for "cyanine" and the first number identifies the number of carbon atoms between the two indolenine groups. Cy2, which is an oxazole derivative rather than an indolenine, and benzo-derivatized Cy3.5, Cy5.5, and Cy7.5 are exceptions to this rule.
[0170] In some embodiments, the reporter moieties can be FRET pairs, allowing for multiple classifications under a single excitation and imaging step. As used herein, FRET can include excitation exchange (Förster) transfer or electron exchange (Dexter) transfer.
[0171] The terms "linked," "joined," "attached," and variants thereof include any type of fusion, bond, adhesion, or association between any combination of compounds or molecules that is stable enough to withstand use in a particular procedure. Procedures can include, but are not limited to, nucleotide transient binding, nucleotide incorporation, deblocking, washing, removal, flow, detection, imaging, and / or identification. Such linkages can include, for example, covalent, ionic, hydrogen, dipole-dipole, hydrophilic, hydrophobic, or affinity binding, bonds or associations involving van der Waals forces, mechanical bonds, etc. In some embodiments, such linkages occur intramolecularly, such as, for example, linking the ends of a single- or double-stranded linear nucleic acid molecule together to form a circular molecule. In some embodiments, such linkages can occur between different molecular combinations or between molecules and non-molecules, including, but not limited to, linkages between nucleic acid molecules and solid surfaces, linkages between proteins and detectable reporter moieties, linkages between nucleotides and detectable reporter moieties, etc. Some examples of linkages can be found, for example, in Hermanson, G., "Bioconjugate Techniques", Second Edition (2008), Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998), Aslam, M., Dent, A., "Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences", London: Macmillan (1998).
[0172] As used herein, the terms "operably linked" and "operably associated" or related terms refer to the juxtaposition of components. Juxtaposed components can be covalently linked. For example, two nucleic acid components can be enzymatically ligated, where the linkage joining the two components consists of a phosphodiester linkage. A first nucleic acid component and a second nucleic acid component can be linked together, where the first nucleic acid component can confer a function to the second nucleic acid component. For example, a linkage between a primer binding sequence and a sequence of interest forms a nucleic acid library molecule having a portion capable of binding to a primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or nucleic acid sequence of interest) can be ligated into a vector, where the linkage allows for expression or function of the transgene sequence included in the vector. In some embodiments, the transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects expression of the transgene. In some embodiments, the vector includes at least one host cell regulatory sequence, including a promoter sequence, an enhancer, a transcription and / or translation initiation sequence, a transcription and / or translation termination sequence, a polypeptide secretion signal sequence, etc. In some embodiments, host cell regulatory sequences control the level, timing, and / or location of expression of the transgene.
[0173] In some embodiments, the support is solid, semi-solid, or a combination of both. In some embodiments, the support is porous, semi-porous, non-porous, or any combination of porous. In some embodiments, the support can be substantially planar, concave, convex, or any combination thereof. In some embodiments, the support can be cylindrical, for example, comprising a capillary or the inner surface of a capillary.
[0174] In some embodiments, the surface of the support can be substantially smooth, hi some embodiments, the support can have a regular or irregular texture, including ridges, etchings, pores, a three-dimensional scaffold, or any combination thereof.
[0175] In some embodiments, the support comprises beads having any shape, including spherical, hemispherical, cylindrical, barrel, toroidal, disc-shaped, rod-shaped, conical, triangular, cubic, polygonal, tubular, or wire-shaped.
[0176] The support can be made of any material, including, but not limited to, glass, fused silica, silicon, polymer (e.g., polystyrene (PS), macroporous polystyrene (MPPS), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET)), or any combination thereof. Various compositions of both glass and plastic substrates are contemplated.
[0177] In some embodiments, the surface of the support is coated with one or more compounds to create a passivation layer on the support. In some embodiments, the support comprises a low non-specific binding surface, which allows for improved nucleic acid hybridization and amplification performance on the support. Generally, the support may comprise a covalently or non-covalently bonded, low-chemically modified layer, such as a silane layer, a polymer film, and one or more layers of one or more covalently or non-covalently bonded oligonucleotides that can be used to immobilize multiple nucleic acid template molecules on the support.
[0178] In some embodiments, the degree of hydrophilicity (or "wettability" with aqueous solutions) of a surface coating can be assessed by measuring water contact angles, for example, where a small drop of water is placed on the surface and the contact angle with the surface is measured, for example, using an optical tensiometer. In some embodiments, a static contact angle can be determined. In some embodiments, an advancing or receding contact angle can be determined. In some embodiments, the water contact angle of a hydrophilic, low-binding support disclosed herein can range from about 0 degrees to about 30 degrees. In some embodiments, the water contact angle of a hydrophilic, low-binding support disclosed herein can be 50 degrees, 40 degrees, 30 degrees, 25 degrees, 20 degrees, 18 degrees, 16 degrees, 14 degrees, 12 degrees, 10 degrees, 8 degrees, 6 degrees, 4 degrees, 2 degrees, or 1 degree or less. In many cases, the contact angle is 40 degrees or less. One of ordinary skill in the art will recognize that the surface of a given hydrophilic, low-binding support of the present disclosure can exhibit a water contact angle having a value anywhere within this range.
[0179] The present disclosure provides a plurality (e.g., two or more) of nucleic acid templates immobilized on a support. In some embodiments, the immobilized plurality of nucleic acid templates have the same sequence or different sequences. In some embodiments, individual nucleic acid template molecules within the plurality of nucleic acid templates are immobilized at different sites on the support. In some embodiments, two or more individual nucleic acid template molecules within the plurality of nucleic acid templates are immobilized at sites on the support. In some embodiments, the support comprises a plurality of sites arranged in an array. The term "array" refers to a support comprising a plurality of sites in a predetermined arrangement on the support to form a series of sites. The sites may be discrete and separated by interstitial regions. In some embodiments, the predetermined sites on the support can be arranged in one dimension, in rows or columns, or in two dimensions, in rows and columns. In some embodiments, the plurality of predetermined sites are arranged on the support in an organized manner. In some embodiments, the plurality of predetermined sites are arranged in any organized pattern, including linear, hexagonal, grid, patterns with reflection symmetry, patterns with rotational symmetry, etc. The pitch between different pairs of sites can be the same or different. In some embodiments, the support is 1 mm thick to form a nucleic acid template array. 2 Approximately 10 per 2 ~10 15 In some embodiments, the support may have nucleic acid template molecules immobilized at a plurality of sites with a surface density of at least 10 sites. 2 at least 10 sites 3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites 9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13at least 10 sites 14 at least 10 sites 15 In some embodiments, the substrate comprises a plurality of predetermined sites (e.g., 10 or more sites) on the substrate, the sites being in a predetermined arrangement on the substrate. 2 ~10 15 In some embodiments, the nucleic acid templates are immobilized at a plurality of predetermined sites (e.g., 10 or more sites) to form a nucleic acid template array. In some embodiments, the nucleic acid templates are immobilized at a plurality of predetermined sites by hybridization to the immobilized surface capture primers, or the nucleic acid templates are covalently attached to the surface capture primers. In some embodiments, the nucleic acid templates are immobilized at a plurality of predetermined sites, e.g., 10 2 ~10 15 Nucleic acid templates immobilized at individual or more sites. In some embodiments, the nucleic acid templates immobilized at multiple sites on a support comprise linear or circular nucleic acid template molecules, or a mixture of both linear and circular molecules. In some embodiments, the immobilized nucleic acid templates are clonally amplified to generate nucleic acid polonies immobilized at multiple predetermined sites. In some embodiments, each immobilized nucleic acid template molecule comprises one copy of the target sequence of interest or comprises a concatemer having two or more tandem copies of the target sequence of interest.
[0180] In some embodiments, a support comprising a plurality of sites in a random arrangement on the support is referred to herein as a support having randomly arranged sites thereon. The arrangement of the randomly arranged sites on the support is not predetermined. The plurality of randomly arranged sites are arranged on the support in a disordered and / or unpredictable manner. In some embodiments, the support comprises at least 10 2 at least 10 sites 3 at least 10 sites 4 at least 10 sites 5 at least 10 sites 6 at least 10 sites 7 at least 10 sites 8 at least 10 sites9 at least 10 sites 10 at least 10 sites 11 at least 10 sites 12 at least 10 sites 13 at least 10 sites 14 at least 10 sites 15 In some embodiments, the substrate comprises a plurality of randomly arranged sites (e.g., 10 sites, 100 sites, or more), which are randomly arranged on the substrate. 2 ~10 15 In some embodiments, the nucleic acid template is immobilized at a plurality of randomly spaced sites (e.g., 10 or more sites) to form a nucleic acid-templated support. In some embodiments, the nucleic acid template is immobilized at a plurality of randomly spaced sites by hybridization to an immobilized surface capture primer, or the nucleic acid template is covalently attached to the surface capture primer. In some embodiments, the nucleic acid template is immobilized at a plurality of randomly spaced sites, e.g., 10 or more sites. 2 ~10 15 Nucleic acid templates immobilized at multiple sites or more. In some embodiments, nucleic acid templates immobilized at multiple sites on a support may include embodiments of nucleic acid templates comprising any of the potential properties listed above, and in some embodiments, the immobilized nucleic acid templates are clonally amplified to generate nucleic acid polonies immobilized at multiple randomly positioned sites.
[0181] In some embodiments, with respect to nucleic acid template molecules immobilized at predetermined or random sites on a support, the multiple immobilized nucleic acid template molecules on the support are in fluid communication with each other, allowing solutions of reagents (e.g., enzymes, including polymerases, polyvalent molecules, nucleotides, divalent cations, and / or buffers, etc.) to be flowed over the support, thereby allowing the multiple immobilized nucleic acid template molecules on the support to react with the reagents in a massively parallel manner. In some embodiments, the fluid communication of the multiple immobilized nucleic acid template molecules can be used to perform nucleotide binding assays and / or nucleotide polymerization reactions (e.g., primer extension or sequencing) on the multiple immobilized nucleic acid template molecules, as well as detection and imaging for massively parallel sequencing. In some embodiments, the term "immobilized" and related terms refer to nucleic acid molecules or enzymes (e.g., polymerases) that are bound to the support in predetermined or random locations, whether the nucleic acid molecules or enzymes are directly bound to the support via covalent or non-covalent interactions, or the nucleic acid molecules or enzymes are bound to a coating on the support.
[0182] As used herein, the term "sequencing" and its variants typically involve obtaining sequence information from a nucleic acid strand by determining the identity of at least some nucleotides (including their nucleobase components) within a nucleic acid template molecule. In some embodiments, "sequencing" a given region of a nucleic acid molecule involves identifying each and every nucleotide within the sequenced region, although in some embodiments, "sequencing" also includes methods in which the identities of only some nucleotides within a region are determined, while the identities of some nucleotides remain unclear or are incorrectly determined. Any suitable sequencing method can be used. In exemplary embodiments, sequencing can include label-free or ion-based sequencing methods. In some embodiments, sequencing can include labeled or dye-containing nucleotide or fluorescence-based nucleotide sequencing methods. In some embodiments, sequencing can include polony-based sequencing or bridge-sequencing methods. In some embodiments, sequencing can include massively parallel sequencing platforms using sequence-by-synthesis, sequence-by-hybridization, or sequence-by-ligation procedures. Examples of massively parallel sequence-by-synthesis procedures include polony sequencing, pyrosequencing (e.g., U.S. Pat. Nos. 7,211,390, 7,244,559, and 7,264,929 from 454 Life Sciences), chain terminator sequencing (e.g., U.S. Pat. No. 7,566,537 from Illumina, Bentley 2006 Current Opinion Genetics and Development 16:545-552, and Bentley, et al., 2008 Nature 456:53-59), ion-sensitive sequencing (e.g., from Ion Torrent), probe-anchor ligation sequencing (e.g., Complete Genomics), DNA nanoball sequencing, and nanopore DNA sequencing. Examples of single molecule sequencing include Heliscope single molecule sequencing and single molecule real-time (SMRT) sequencing.Examples of sequencing by hybridization include SOLiD sequencing (e.g., from Life Technologies, WO2006 / 084132). Examples of sequencing by binding include Omniome sequencing (e.g., U.S. Pat. No. 10,246,744).
[0183] Engineered polymerases The present disclosure provides compositions, including mutant polymerases with amino acid substitutions and / or truncated amino acid sequences, nucleic acids encoding the mutant polymerases, and systems and kits comprising the mutant polymerases. Further provided herein are methods using the mutant polymerases, including methods for binding nucleic acid duplexes, methods for binding complementary nucleotides or methods for binding multivalent molecules with complementary nucleotide units, methods for incorporating complementary nucleotides or methods for incorporating complementary nucleotide units, methods for extending primers, and methods for nucleic acid sequencing, which employ any of the mutant polymerases described herein. The mutant polymerases are engineered to exhibit desired characteristics, including increased incorporation of nucleotide analogs, compared to wild-type polymerases. In one embodiment, the mutant polymerases comprise polypeptides or fragments thereof derived from directed evolution of recently identified novel B-family and A-family polymerases, and the mutant polymerases exhibit improved specificity while maintaining high discrimination for correct Watson-Crick base pairing. One exemplary polymerase enzyme is derived from the archaeal species Candidatus altiarchaeales (e.g., SEQ ID NO: 1 or 391), which was first identified in 2018. This enzyme contains less than 42% sequence identity with 9°N, Pfu, VENT, DEEP VENT, and Pyrococcus abyssi, demonstrating extreme evolutionary divergence, which may be further evidenced by the fact that the 9°N, Pfu, VENT, DEEP VENT, and Pyrococcus abyssi polymerases are thermostable at high temperatures (e.g., above 95°C), while the Candidatus altiarchaeales archaeal polymerase is thermostable at low temperatures (e.g., below 75°C), with an optimal catalytic temperature of approximately 68°C. Another exemplary polymerase enzyme exhibiting activity at temperatures below 75°C is derived from Geobacillus stearothermophilus (e.g., SEQ ID NO: 275).Candidatus altiarchaeales archaeal polymerase exhibits nucleotide binding and incorporation activity in a temperature range of about 25-50°C, or about 45-75°C, or about 65-75°C. Candidatus altiarchaeales archaeal polymerase exhibits optimal nucleotide binding and incorporation activity in a temperature range of about 65-75°C. Candidatus altiarchaeales archaeal polymerase is a moderately thermostable polymerase (e.g., a mesothermal polymerase). Engineered polymerases having a Candidatus altiarchaeales archaeal sequence backbone with one or more mutations can be used to perform nucleotide binding, nucleotide unit binding, nucleotide incorporation, nucleotide unit incorporation, and / or nucleic acid sequencing reactions in a temperature range of about 25-50°C, or about 45-75°C, or about 50-65°C, or about 50-60°C. For example, thermostable polymerases such as 9°N, Pfu, VENT, DEEP VENT, and Pyrococcus abyssi polymerases are suitable for use in PCR reactions where the typical cycling step is performed at temperatures above or equal to 90-95° C. One of skill in the art will understand that the thermostable polymerases described herein may not be suitable for use in nucleotide binding, nucleotide incorporation, and / or nucleic acid sequencing reactions performed at lower temperature ranges, such as, for example, about 25-50° C. or about 45-75° C.
[0184] Polymerases include DNA polymerases, RNA polymerases, template-independent polymerases, reverse transcriptases, or other enzymes capable of catalyzing the incorporation of nucleotides. Archaeal polymerases are often derived from thermophilic organisms and thus may represent a class of thermostable or thermotolerant enzymes. Thus, polypeptide backbones derived from archaeal polymerases provide desirable protein engineering targets to further enhance the incorporation of reversible terminator nucleotides (removable chemical groups that prevent nucleic acid elongation), for applications that can be improved by the application of enzymes with increased thermostability or other resistance to degradation due to repeated exposure to high temperatures, changes in buffer conditions, etc.
[0185] We have made the surprising discovery that engineered polymerases having a Candidatus altiarchaeales archaeal sequence backbone and containing one or more mutations exhibit increased incorporation rates of nucleotide analogs compared to wild-type polymerases. Compared to wild-type Candidatus altiarchaeales archaeal polymerases, some of the engineered polymerases exhibit one or more desirable characteristics, including increased binding affinity for nucleotide analogs with a 3' chain terminator, improved ability to incorporate dATP nucleotides opposite uracil-containing template molecules (e.g., uracil-resistant mutant polymerases), improved ability to bind to complementary nucleotide units of multivalent molecules, and increased thermal stability up to about 75°C. We also show that engineered polymerases based on a Geobacillus stearothermophilus backbone (e.g., Bst polymerase) and containing mutant sequences exhibit improved incorporation of nucleotide analogs.
[0186] The present disclosure provides engineered polymerases that are useful for performing any nucleic acid sequencing method using labeled or unlabeled chain-terminating nucleotides, where the chain-terminating nucleotides contain a 3'-O-azido group (or a 3'-O-methyl azido group) or any other type of bulky blocking group at the 3' position of the sugar. For example, the engineered polymerases can be used to perform affinity sequencing (SBA) using labeled polyvalent molecules and unlabeled chain-terminating nucleotides. Furthermore, the engineered polymerases can be used to perform sequencing-by-synthesis (SBS) methods using labeled chain-terminating nucleotides and sequencing-by-binding (SBB) methods using unlabeled chain-terminating nucleotides.
[0187] DNA affinity sequencing (SBA) ideally requires (a) the detection of n+1 bases, two or more copies of a target nucleic acid sequence, two or more primer nucleic acid molecules complementary to one or more regions of the target nucleic acid sequence, and two or more polymerases, where the polymeric nucleotide conjugates contain two or more nucleotide moieties, and the two or more polymerases are contacted with the composition under conditions sufficient to allow the formation of a multivalent binding complex between the polymeric nucleotide conjugates and the two or more copies of the target nucleic acid sequence in the composition, and the detection substrate is then washed away; and (b) structural modifications ("blocking groups") of unlabeled nucleotides are required to ensure the incorporation of a single nucleotide, but prevent the incorporation of any additional nucleotides into the polynucleotide chain, so that only a single incorporation occurs. The blocking groups must then be removable under reaction conditions that do not interfere with the integrity of the DNA being sequenced. The sequencing cycle can then continue with the N+1 detection of the next multivalent polymerase-conjugated DNA complex, and so on. To be of practical application, the affinity step requires both (a) a stable substrate that persists long enough to allow imaging for more than 30 seconds, and (b) a stepping step in which the entire process should consist of high-yielding, highly specific chemical and enzymatic steps to facilitate multiple sequencing cycles.
[0188] DNA sequencing by synthesis (SBS) ideally requires controlled (i.e., one at a time) incorporation of the correct complementary nucleotide opposite the oligonucleotide to be sequenced. This allows for accurate sequencing by adding nucleotides in multiple cycles, such that each nucleotide residue is sequenced one at a time, thus preventing uncontrolled sequential incorporation. The incorporated nucleotide is read using the appropriate label attached to it before removal of the label moiety and the next sequencing round. To ensure that only a single incorporation occurs, a structural modification ("blocking group") of the sequencing nucleotide is required to ensure incorporation of a single nucleotide but then prevent the incorporation of any additional nucleotides into the polynucleotide chain. The blocking group must then be removable under reaction conditions that do not interfere with the integrity of the DNA to be sequenced. The sequencing cycle can then proceed with the incorporation of the next blocked, labeled nucleotide. For practical applications, the entire process should consist of high-yield, highly specific chemical and enzymatic steps to facilitate multiple sequencing cycles.
[0189] Sequencing binding (SBB) includes the steps of: (a) sequentially contacting a primed template nucleic acid with at least two separate mixtures under ternary complex-stabilizing conditions, each of the at least two separate mixtures comprising a polymerase and nucleotides, whereby the sequential contacting results in a primed template nucleic acid contacted with nucleotide analogs of first, second, and third base types in the template under ternary complex-stabilizing conditions; (b) examining the at least two separate mixtures to determine whether a ternary complex has formed; and (c) examining the at least two separate mixtures to determine whether a ternary complex has formed. The method involves identifying a next correct nucleotide, where the next correct nucleotide is identified as a cognate of the first, second, or third base type if a ternary complex is detected in step (b), and the next correct nucleotide is presumed to be a nucleotide cognate of the fourth base type based on the absence of a ternary complex in step (b); (d) adding the next correct nucleotide to the primed template nucleic acid after step (b), thereby producing an extension primer; and (e) repeating steps (a) through (d) for the primed template nucleic acid containing the extension primer.
[0190] The polypeptides described herein include, but are not limited to, polypeptides having enzymatic activity, such as polymerase activity, and are often described as families. Often, the polymerase is a DNA polymerase, RNA polymerase, template-independent polymerase, reverse transcriptase, or other enzyme capable of nucleotide binding and nucleotide incorporation (e.g., primer extension). Many DNA polymerases are known in the art, and such enzymes are optionally mutated to generate the compositions described herein. Members of a DNA polymerase family are often defined in terms of polymerase activity, active site structure, domain homology / function, or sequence homology to other known DNA polymerase family members. For example, DNA polymerases include, but are not limited to, E. coli DNA polymerase I, E. coli DNA polymerase II, or other members of the DNA polymerase family. Known thermostable DNA polymerases include Taq polymerase, Pfu polymerase, and 9°N polymerase, or other members of the DNA polymerase family. Wild-type DNA polymerases are or can be obtained from any number of origins, such as eukaryotic, prokaryotic, or viral origins, and in some embodiments, for purposes of this disclosure, are or can be obtained from archaeal origins. In some embodiments, a polymerase comprising the amino acid sequence of any of SEQ ID NOs: 1-274, 288-375, and 385-397 is a member of the DNA polymerase family.
[0191] Further provided herein is a polypeptide comprising a sequence having at least 85% identity to SEQ ID NO: 1 or 391 and at least one mutation at positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567 of the polypeptide sequence numbered according to the residues of SEQ ID NO: 1 or 391. In some cases, the polypeptides described herein comprise a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to SEQ ID NO: 1 or 391. In some cases, the polypeptides described herein include sequences having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to SEQ ID NO: 1 or 391, as well as sequences having at least one mutation at positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567 according to the numbering of SEQ ID NO: 1. In some embodiments, the present invention relates to at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or a mixture of any of SEQ ID NOs: 1 or 2-268 or 288-375 or 385-397. Disclosed herein are polypeptides having 99.8% or greater identity to the polypeptides of SEQ ID NOs: 1, 393, or 391 and having at least one mutation at a position similar to one or more of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567 of SEQ ID NOs: 1, 393, or 391.
[0192] Further provided herein is a polypeptide comprising a sequence having at least 85% identity to SEQ ID NO: 1 or 391 and at least one mutation at positions G403, H405, D406, R414, S415, L416, Y417, P418, R468, A493, K495, N499, M501, Y502, F507, R515, I529, and / or N567 of the polypeptide sequence numbered according to the residues of SEQ ID NO: 1 or 391. In some cases, the polypeptides described herein comprise a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to SEQ ID NO: 1 or 391. In some cases, the polypeptide also has at least one mutation at positions G403, H405, D406, R414, S415, L416, Y417, P418, R468, A493, K495, N499, M501, Y502, F507, R515, I529, and / or N567 according to the numbering of SEQ ID NO: 1, 393, or 391. In some embodiments, the nucleic acid sequence is at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identical to SEQ ID NO: 1 or 2-268 or 288-375 or 385-397. and at least one mutation at a position analogous to one or more of positions G403, H405, D406, R414, S415, L416, Y417, P418, R468, A493, K495, N499, M501, Y502, F507, R515, I529, and / or N567 of SEQ ID NO: 1 or 391.In some embodiments, the variant polypeptides comprise a variant polypeptide at position(s) D9, Y10, I11, E14, E27, F37, M41, H48, P45, L50, K51, Q54, K58, K61, I63, I68, E73, D75, E77, M84, Q87, V91, G96, E102, K105, V107, A115, E116, L124, P126, N132, M142, R170, E179, D191, E205, K225, V239, R256, I272, E291, D305, E308, E310, E328, I343, T344, S372, E434, Y448, V465, R467, G474, N480, R483, D488, A498, S500, M501 , Y502, R509, E516, S520, K538, F539, D560, V564, M565, A568, E569, D573, K574, S57 7, further comprising at least one mutation at E578, E581, M583, K610, T618, D636, N657, T675, K682, V689, E700, N705, S717, E730, S746, E758, K762, G763, L764, G765, K766, Q767, and / or F773.
[0193] The present disclosure provides compositions and methods comprising mutant polypeptides related to polymerase enzymes that exhibit increased ability to bind and discriminate between nucleotide analogs and improved incorporation of nucleotide analogs compared to wild-type polymerases. Nucleotide analogs include, for example, nucleotides containing a chain-terminating group attached to the 2' or 3' sugar position. The chain-terminating group may include an azide, azido, or azidomethyl group, or another type of chain-terminating group. The engineered DNA polymerases exhibit increased incorporation rates of nucleotide analogs compared to wild-type polymerases having the amino acid sequence of SEQ ID NO: 1 or 391. The data presented in Table 1 (Figures 3A-3H) provide a number of exemplary mutant polymerases that exhibit increased incorporation rates of nucleotide analogs.
[0194] The present disclosure provides compositions and methods comprising mutant polymerase enzymes that have increased thermostability compared to the wild-type polymerase having the amino acid sequence of SEQ ID NO: 1 or 391.
[0195] The present disclosure provides compositions and methods comprising mutant polymerase enzymes that can be used for sequencing uracil-containing nucleic acid template molecules. The mutant polymerases can exhibit uracil resistance, having an increased ability to incorporate dATP into the 3' end of a nucleic acid primer at a position opposite the uracil base in the nucleic acid template molecule. The mutant polymerases can also bind to adenine-bearing nucleotide units of multivalent molecules at a position opposite the uracil base in the nucleic acid template molecule. Exemplary mutant polymerases exhibiting uracil resistance include the amino acid sequence of any one of SEQ ID NOs: 361, 362, 363, 364, 366, 367, 374, or 375, or any of SEQ ID NOs: 385-397.
[0196] Mutations in the polymerases described herein variously involve one or more changes to amino acid residues present in the polypeptide. Additions, substitutions, or deletions are all examples of mutations used to generate variant polypeptides. In some embodiments, a substitution involves replacing one amino acid with an alternative amino acid, which differs from the original amino acid in terms of size, shape, conformation, or chemical structure. In some embodiments, the mutation is conservative or non-conservative. Conservative mutations involve the substitution of an amino acid with an amino acid having similar chemical properties. Additions often involve the insertion of one or more amino acids at the N-terminus, C-terminus, or internal position of the polypeptide. In some cases, additions comprise fusion polypeptides, in which one or more additional polypeptides are attached to the polypeptide. In some embodiments, such additional polypeptides comprise domains with additional activities or sequences with additional functions (e.g., improving expression, aiding in purification, improving solubility, binding to solid supports, or other functions). Often, the polypeptides described herein comprise one or more non-amino acid groups. Fusion polypeptides optionally include amino acids or other chemical linkers connecting one or more proteins. Any number of mutations can be introduced into a polypeptide or portion of a polypeptide described herein, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more than 50 mutations.
[0197] In some embodiments, entire domains (portions of a polypeptide with a defined function) are added, deleted, or substituted for domains from other polypeptides. Exemplary domains include DNA / RNA binding domains, nucleotide binding domains, nuclease domains, subcellular localization domains such as nuclear localization domains, or other domains. In some embodiments, the disclosed methods and compositions include attachment of domains that function as spacers or labels and / or provide for attachment of linkers such as SNAP tags, avidin moieties, streptavidin moieties, epitope tags, fluorescent proteins, affinity tags, metal binding (i.e., His6 (SEQ ID NO: 398) or polyhistidine tags). In some embodiments, one or more mutations are present in the catalytic site or binding domain. For example, the polypeptide comprises the nucleic acid binding domain disclosed in SEQ ID NO: 1. The domain optionally contains a DNA or RNA binding site, for example, comprising residues at positions 354-773 of SEQ ID NO: 1, or alternatively, residues 249-253, 392-396, 598-601, 613-615, 673-676, and / or 681-683. Such sites can be found in similar positions after alignment of other sequences to SEQ ID NO: 1 (see, e.g., Figures 7A-7B, 8A-8C, 9A-9B, 10A-10B, 11A-11B, 12A-12B, and 13A-13B). In some embodiments, the domain comprises a polymerase domain comprising residues at positions 1-135, 136-347, 348-454, 455-504, 505-624, or 625-773 of SEQ ID NO: 1, or a positional equivalent thereof (see, e.g., Table 4 in Figures 6A-6C, and Figures 7A-7B, 8A-8C, 9A-9B, 10A-10B, 11A-11B, 12A-12B, and 13A-13B). In some cases, the polypeptide comprises an active site.The active site of the polypeptide may comprise residues D412, D547, and / or D549 of SEQ ID NO: 1, or positional equivalents thereof (see, e.g., Table 4 in Figures 6A-6C, and Figures 7A-7B, 8A-8C, 9A-9B, 10A-10B, 11A-11B, 12A-12B, and 13A-13B). Such sites are often found at analogous positions in other domains (e.g., identified by aligning two or more sequences for comparison), and polypeptides containing such domains are consistent with the methods and compositions described herein.
[0198] As used herein, the term "surrounding" an amino acid residue or sequence position has its normal meaning in the art to include and incorporate modifications such as substitutions, deletions, insertions, or post-translational modifications at residues from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more residues away from the named residue, i.e., residues N- or C-terminal to the named residue. In some contexts, as will be understood by those of skill in the art, more than 12 residues or sequence positions N- or C-terminal from the named residue can be considered to "surround" the named residue based on sequence or structural (i.e., three-dimensional) context.
[0199] It is understood that the residue substitutions or modifications described herein may also incorporate or include non-standard amino acids known in the art, including, but not limited to, hydroxyproline, N-formylmethionine, selenomethionine, selenocysteine, phosphotyrosine, phosphohistidine, and the like. The mutations, modifications, truncations, substitutions, and the like described herein may be made by any method known in the art, particularly in the art of molecular biology and / or protein engineering. Such methods may include site-directed mutagenesis using mutagenic and / or partially degenerate primers, in vitro gene assembly, gene editing (e.g., by CRISPR or related methods), and the like. The mutant or engineered proteins described herein may be further expressed, isolated, and / or purified by any such means known in the art. Relevant methods are described in Green, M. and Sambrook, J., Molecular Cloning: A Laboratory Manual (4th Edition), which is incorporated herein by reference in its entirety, particularly with respect to its disclosure of methods for modifying, introducing, and expressing recombinant, modified, altered, and engineered gene sequences, and methods for extracting, isolating, and / or purifying engineered proteins.
[0200] The polypeptides disclosed herein have been shown to function as nucleotide polymerases that exhibit greater thermostability, higher incorporation rates of 3'-O-azidomethyl-derivatized nucleosides, and / or increased uracil tolerance compared to wild-type enzymes. The polypeptides disclosed herein can be used to extend nucleic acids during replication or synthesis, or can be trapped at nucleotide addition sites, for example, by the use of non-incorporable or blocked nucleotides, or can be used under conditions in the absence of required salts or cofactors. The polypeptides disclosed herein can be utilized, for example, in polynucleotide sequencing applications, such as sequencing by synthesis and sequencing by binding applications.
[0201] The present disclosure provides engineered DNA polymerases that include an amino acid sequence backbone of a Family B polymerase, typically including replicative polymerases that exhibit improved incorporation of nucleotide analogs. Examples of Family B polymerases include Family B archaeal DNA polymerases and Phi29 polymerase. In some embodiments, the engineered DNA polymerase includes a Family B archaeal DNA polymerase, which may be selected from Thermococcus, Pyrococcus, Methanococcus, or Candidatus. In some embodiments, the engineered DNA polymerase includes an amino acid sequence backbone from Candidatus Altiarchaeales archaeal DNA polymerase, 9°N polymerase (including THERMINATOR polymerase), VENT polymerase, DEEP VENT polymerase, Pfu polymerase, Pyrococcus abyssi polymerase, or RB69 polymerase. In some embodiments, the engineered DNA polymerase can be based on the amino acid sequence backbone of a family A type polymerase, including Geobacillus (eg, Geobacillus stearothermophilus).
[0202] Engineered DNA polymerases can be designed and prepared by introducing one or more mutations into the amino acid sequence of a DNA polymerase of interest (e.g., a wild-type or mutant polymerase), and the phenotype of the resulting engineered polymerase can be determined. Any one or any combination of two or more mutation sites can be transferred from one type of polymerase to a positionally equivalent site in a second type of polymerase. For example, any one or any combination of two or more mutation sites from the Candidatus Altiarchaeales archaeal DNA polymerase can be introduced into the positionally equivalent site of 9°N polymerase (including THERMINATOR polymerase), VENT polymerase, DEEP VENT polymerase, Pfu polymerase, Pyrococcus abyssi polymerase, and / or RB69 polymerase (see, e.g., Table 4 in Figures 6A-6C). Mutations include any one or any combination of two or more amino acid substitutions, insertions, deletions, and / or truncations.
[0203] Functional equivalents of residues include one or more amino acid residues that occupy a similar position in a sequence (e.g., sequence alignment) and / or three-dimensional structure of an enzyme (e.g., DNA polymerase) and perform substantially the same function as a known amino acid residue in a known enzyme. Functionally equivalent amino acid substitutions include one or more amino acid residues at a particular position in a basic polypeptide that have the same functional role in another polypeptide. Functionally equivalent amino acid substitutions include any one or any combination of conservative and / or non-conservative amino acid substitutions. Table 4 in Figures 6A-6C lists examples of amino acid residues at sites in Candidatus Altiarchaeales archaeal DNA polymerases and functionally equivalent amino acid positions, such as 9°N DNA polymerase (for SEQ ID NO: 280 or 281), THERMINATOR (for SEQ ID NO: 282), VENT DNA polymerase (for SEQ ID NO: 283), DEEP VENT DNA polymerase (for SEQ ID NO: 284), Pfu DNA polymerase (for SEQ ID NO: 285), Pyrococcus abyssi DNA polymerase (for SEQ ID NO: 286), and Geobacillus stearothermophilus DNA polymerase (for SEQ ID NO: 275).
[0204] A wild-type polypeptide sequence is often the starting point for protein or enzyme engineering to generate mutant polypeptides. In some embodiments, a mutant polypeptide differs from a wild-type polypeptide by at least one amino acid residue. Often, a mutant polypeptide differs from the closest wild-type polypeptide by at least one amino acid residue. In some embodiments, a mutant polypeptide differs from a wild-type polypeptide by at least two amino acid residues. In some embodiments, a mutant polypeptide differs from a wild-type polypeptide by at least three, four, five, or at least six amino acid residues. Often, the wild-type sequence is the closest wild-type sequence identified by aligning polypeptides containing at least one mutation in the wild-type sequence. In some embodiments, the wild-type polypeptide sequence comprises the sequence of a naturally occurring polypeptide.
[0205] Amino acid substitution refers to replacing an amino acid residue at a selected position in a polypeptide with a different amino acid that has similar or different biochemical properties, such as similar size, shape, conformation, chemical structure, charge, and / or hydrophobicity. Amino acid substitutions can be conservative or non-conservative. In some embodiments, the amino acid residue at a selected position in a polypeptide can be replaced with an amino acid having a polar side chain. Examples of amino acids having polar side chains include arginine, asparagine, aspartic acid, glutamine, glutamic acid, histidine, lysine, serine, and threonine. In some embodiments, the amino acid residue at a selected position in a polypeptide can be replaced with an amino acid having a non-polar side chain. Examples of amino acids having non-polar side chains include alanine, cysteine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, tyrosine, and valine. In some embodiments, the amino acid residue at a selected position in a polypeptide can be replaced with an amino acid having a hydrophobic side chain. Examples of amino acids having hydrophobic side chains include glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tyrosine, and tryptophan. In some embodiments, amino acid residues at selected positions within a polypeptide can be replaced with amino acids having uncharged side chains. Examples of amino acids having uncharged side chains include glycine, serine, cysteine, asparagine, glutamine, tyrosine, and threonine. In some embodiments, amino acid residues at selected positions within a polypeptide can be replaced with amino acids having positively charged side chains. Examples of amino acids having positively charged side chains include arginine, histidine, and lysine. In some embodiments, amino acid residues at selected positions within a polypeptide can be replaced with amino acids having negatively charged side chains. Examples of amino acids having negatively charged side chains include aspartic acid and glutamic acid.
[0206] The present disclosure provides a method for the preparation of a ribonucleotide-binding fragment comprising the backbone amino acid sequence of SEQ ID NO: 1 or 391, and including the steps of F22, E26, V34, A35, K52, K58, I72, E79, M84, C104, I112, C130, E150, S162, C269, I272, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q491, Q52, Q53, Q54, Q55, Q56, Q57, Q58, Q59, Q60, Q61, Q62, Q63, Q64, Q65, Q66, Q67, Q68, Q69, Q70, Q71, Q72, Q73, Q74, Q75, Q76, Q77, Q78, Q79, Q80, Q81, Q82, Q83, Q84, Q85, Q86, Q87, Q88, Q89, Q90, Q91, Q92, Q93, Q94, Q95, Q96, Q97, Q98, Q99, Q104, Q112, Q130, Q150, Q162, Q170, Q181, Q190, Q191, Q201, Q212, Q221, Q222, Q231, Q241, Q242, Q253, Q261, Q26 Mutant polymerases are provided from the archaeon Candidatus Altiarchaeales having amino acid substitution mutations at one or more positions, including 492, A493, L494, K495, L496, N499, M501, Y502, F507, C514, R515, C517, T522, I529, N567, E569, S577, R608, K610, L611, D622, K633, V651, D653, T660, A669, Q673, T683, R697, S717, R723, I750, L751, and / or E760. In some embodiments, the amino acid substitution mutations can also include positions D141 and E143.In some embodiments, the mutant polymerase comprises a mutant polymerase comprising a polypeptide sequence numbered according to the residues of SEQ ID NO: 1, 393, or 391 at positions D9, Y10, I11, E14, E27, F37, M41, P45, H48, L50, K51, Q54, K61, I63, I68, E73, D75, E77, M84, Q87, V91, G96, E102, K105, V107, A115, E116, L124, P126, N132, M142, R170, E179, D191, E205, K225, V239, R256, I272, E291, D305, E308, E310, E328, I330, I340, I350, I360, I372, I380, I391, I405, I410, I425, I430, I440, I450, I460, I472, I480, I491, I505, I510, I520, I530, I540, I550, I560, I570, I580, I591, I610, I620, I630, I640, I650, I660, I671, I680, I730, I750, I760, I771, I780, I791, I820, I830, I840, I850, I860, I871, I880, I891, I900, I910, I920, I930, I940 43, T344, S372, E434, Y448, V465, R467, G474, N480, R483, D488, A498, S500, Y50 2, R509, E516, S520, K538, F539, D560, V564, M565, A568, D573, K574, E578, E581 , M583, D680, T618, D636, N657, T675, E680, K682, V689, E700, N705, S717, E730, S746, E758, K762, G763, L764, G765, K766, Q767, and / or F773. In some embodiments, the mutant polymerase further comprises additional amino acids at the C-terminus comprising the sequence QGSYT (single letter code) set forth in SEQ ID NO: 391.
[0207] In some embodiments, the mutant polymerase from the archaeon Candidatus Altiarchaeales is selected from the group consisting of F22L, E26G, V34I, A35E, A35T, K58M, I72F, E79K, E79V, M84T, C104S, I112F, C130S, C130R, E150K, S162S, C269S, C269V, I272N, P335Q, G355S, E370D, D381Y, G403A, H405P, H405S, D406H, R414S, S415A, S415G, S415V, L416V, L416G, L416T, L416A, L416S , L416I, L416F, L416Y, L416M, Y417T, Y417S, Y417G, Y417A, Y417V, Y417 I, P418S, P418G, P418V, P418C, P418K, P418I, P418T, P418A, D439S, S44 0N, S443N, C450S, R468A, R468V, R468S, R468K, R468H, R468G, K473A, V4 89I, Q492R, 492C, Q492F, Q492A, Q492G, A493V, A493S, L494V, K495G, K4 95A, K495Q, K495S, K495V, L496A, L496G, L496S, L496R, L496H, L496N, L 496I, L496M, L496C, L496Y, N499G, N499A, N499S, N499V, M501I, Y502T, Y502V, Y502S, Y502R, Y502G, Y502N, Y502A, Y502Q, Y502P, Y502H, Y502F , F507S, C514S, R515L, R515W, R515Y, R515P, R515F, C517S, T522S, T522 and / or E760G.In some embodiments, the amino acid substitution mutations can also include D141A and E143A. In some embodiments, the mutant polymerase from the archaeon Candidatus altiarchaeales is D9N, Y10F, I11F, E14K, E14G, E27R, F37S, M41L, P45S, H48R, L50P, K51R, Q54L, K58R, K61E, I63T, I63V, I68V, E73K, D75N, E77G, M84L, Q87H, V91Q, G96S, E102R, K105R, V107I , A115V, E116G, L124Q, P126S, N132S, M142L, R170H, E179R, D191G, E205K, K225E, V239I, R256H, R256K, I 272V, E291R, D305N, E308R, E310R, E328Q, I343V, T344I, S372N, E434R, Y448H, V465M, R467C, G474S, G474 D, N480I, R483H, D488N, A498G, S500G, M501V, Y502F, R509H, E516G, S520N, S520G, K538R, F539Y, D560G, D560E, V564I, M565V, A568V, D573N, K574R, E578N, E581G, M583K, T618A, D636G, N657D, T675A, E680D, K68 In some embodiments, the mutant polymerase further comprises any one or any combination of two or more amino acid substitutions (according to the numbering of SEQ ID NO: 1, 393, or 391) including 2I, V689A, E700K, N705D, N705S, S717N, E730R, S746C, E758R, K762R, G763V, G763S, L764W, G765A, K766S, Q767K, and / or F773S. In some embodiments, the mutant polymerase further comprises additional amino acids at the C-terminus comprising the sequence QGSYT (single letter code) set forth in SEQ ID NO: 391.
[0208] Segments or portions of larger polypeptides are further described herein. Optionally, the segment has catalytic activity, such as nucleotide incorporation and nucleic acid extension activity, particularly in the context of a reverse transcriptase domain or polymerase domain described herein. Described herein are polypeptides comprising any one of the segments from any subset of SEQ ID NOs: 1-391 and at least one additional residue (e.g., a +1 residue) at the N- or C-terminus.
[0209] In some embodiments, both the N-terminus and C-terminus have at least one additional residue, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, or more than 100 additional residues. Described herein are polypeptides that include additional residues such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NO: 1, such as an adjacent N-terminal aspartic acid, an adjacent C-terminal arginine, or a combination thereof, which account for single mutated residues or other residues that contribute to a polypeptide comprising SEQ ID NO: 1. Described herein are polypeptides that include an additional residue, such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NOS: 1, such as an adjacent N-terminal glutamine, an adjacent C-terminal histidine, or a combination thereof, which accounts for a single mutated or other residue that contributes to a polypeptide comprising SEQ ID NOS: 1. Described herein are polypeptides that include an additional residue, such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NOS: 1, such as an adjacent N-terminal valine, an adjacent C-terminal cysteine, or a combination thereof, which accounts for a single mutated or other residue that contributes to a polypeptide comprising SEQ ID NOS: 1. Described herein are polypeptides that contain additional residues, such as any of SEQ ID NOs: 1 and 2-391 (+1 residue), or residues identified by alignment of any of SEQ ID NOs: 1 and 2-391 to SEQ ID NO: 1, such as an adjacent N-terminal threonine, an adjacent C-terminal cysteine, or a combination thereof, which account for single mutated residues or other residues that contribute to a polypeptide comprising SEQ ID NO: 1.Described herein are polypeptides that include additional residues, such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NOS: 1, such as an adjacent N-terminal aspartic acid, an adjacent C-terminal leucine, or a combination thereof, which account for a single mutated or other residue that contributes to a polypeptide comprising SEQ ID NOS: 1. Described herein are polypeptides that include additional residues, such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NOS: 1, such as an adjacent N-terminal aspartic acid, an adjacent C-terminal arginine, or a combination thereof, which account for a single mutated or other residue that contributes to a polypeptide comprising SEQ ID NOS: 1. Described herein are polypeptides that include additional residues, such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NOS: 1, such as an adjacent N-terminal threonine, an adjacent C-terminal threonine, or a combination thereof, which account for a single mutated or other residue that contributes to a polypeptide comprising SEQ ID NOS: 1. Described herein are polypeptides that include additional residues, such as any of SEQ ID NOS: 1 and 2-391 (+1 residue), or a residue identified by alignment of any of SEQ ID NOS: 1 and 2-391 to SEQ ID NOS: 1, such as an adjacent N-terminal threonine, an adjacent C-terminal asparagine, or a combination thereof, which account for a single mutated or other residue that contributes to a polypeptide comprising SEQ ID NOS: 1. Described herein are polypeptides that contain additional residues, such as any of SEQ ID NOs: 1 and 2-391 (+1 residue), or residues identified by alignment of any of SEQ ID NOs: 1 and 2-391 to SEQ ID NO: 1, such as an adjacent N-terminal threonine, an adjacent C-terminal serine, or a combination thereof, which account for single mutated residues or other residues that contribute to a polypeptide comprising SEQ ID NO: 1.
[0210] The disclosure provides polymerases from Geobacillus stearothermophilus (e.g., SEQ ID NOS:275-279). In some embodiments, the disclosure provides one or more polypeptides having one or more mutations, e.g., substitutions, deletions, or insertions, at or around positions 314, 332, 334, 368, 381, 385, 417, 434, 454, 471, 528, 601, 615, 635, 649, 654, 655, 656, 657, 658, 659, 665, 680, 682, 702, 706, 707, 710, 714, 758, 760, and / or 829 of SEQ ID NO:275. Further provided are polypeptides having at least 85% identity to SEQ ID NO:275 and comprising a sequence having at least one mutation at positions 314, 332, 334, 368, 381, 385, 417, 434, 454, 471, 528, 601, 615, 635, 649, 654, 655, 656, 657, 658, 659, 665, 680, 682, 702, 706, 707, 710, 714, 758, 760, and / or 829 of the polypeptide sequence numbered according to the residues of SEQ ID NO:275. In some cases, the polypeptides described herein comprise a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to SEQ ID NO:275. In some cases, the polypeptides described herein may comprise a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to SEQ ID NO: 275, as well as .... 275 numbering and at least one mutation at positions 314, 332, 334, 368, 381, 385, 417, 434, 454, 471, 528, 601, 615, 635, 649, 654, 655, 656, 657, 658, 659, 665, 680, 682, 702, 706, 707, 710, 714, 758, 760, and / or 829.
[0211] In some embodiments, disclosed herein are polypeptides having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to any of SEQ ID NOs: 1-257, 288-375, and 385-397, and having at least one mutation at a position similar to one or more of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, and / or 529 of SEQ ID NO: 1.
[0212] Further provided herein is a polypeptide having at least 85% identity to SEQ ID NO:275, comprising a sequence having at least one or more mutations at positions R615, Y654, S655, Q656, 1657, E658, L659, D680, H682, R702, K706, A707, F710, Y714, H829, D314, 1332, 1334, K368, K381, 1385, K417, K434, 1454, D471, 1528, K601, K635, 1649, 1665, K758, and / or K760 of the polypeptide sequence numbered according to the residues of SEQ ID NO:275. In some cases, the polypeptides described herein comprise a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, or 99.8% identity to SEQ ID NO:275.
[0213] Exemplary polypeptide variants described herein are listed in Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), and Table 3 (FIGS. 5A-5E). In some embodiments, a polypeptide described herein has a sequence that has at least 85% identity to SEQ ID NO: 1 or 391 and further exhibits at least one of the mutations shown in Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E). In some embodiments, the polypeptides described herein have a sequence that has at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8% or more sequence identity or substantial identity to any of SEQ ID NOs: 1 or 2-274 or 288-375 or 385-397, and further exhibits at least one of the mutations set forth in Table 1 (Figures 3A-3H), Table 2 (Figures 4A-4E), or Table 3 (Figures 5A-5E). Additional polypeptides contemplated and disclosed herein include a DNA polymerase domain having at least one mutation at a position similar to at least one of the positions in Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E), up to and including all of the positions shown in Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E), optionally resulting in a polypeptide having one or more of the mutations shown in Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E) at the homologous position.
[0214] Tables 1, 2, and 3 show the relative incorporation activity of various mutant variants relative to the wild-type (SEQ ID NO: 1) DNA polymerase from the archaeon Candidatus Altiarchaeales in different symbols. The symbol "0" indicates that, based on experimental data, the mutant variant has insignificant incorporation activity or no significant enhancement in incorporation activity compared to the wild-type. The symbol "+" indicates that, based on experimental data, the mutant variant has some enhancement in incorporation activity compared to the wild-type. The symbol "++" indicates that, based on experimental data, the mutant variant has a high enhancement in incorporation activity compared to the wild-type.
[0215] In some embodiments, one or more mutant variants exhibit an increased average incorporation rate that is at least 5-fold greater than the average incorporation rate of a wild-type polymerase having the amino acid sequence of SEQ ID NO: 1. In some embodiments, one or more mutant variants exhibit an increased average incorporation rate that is at least 10-fold greater than the average incorporation rate of a wild-type polymerase having the amino acid sequence of SEQ ID NO: 1. In some embodiments, one or more mutant variants exhibit an increased average incorporation rate that is at least 20-fold greater than the average incorporation rate of a wild-type polymerase having the amino acid sequence of SEQ ID NO: 1. In some embodiments, one or more mutant variants exhibit an increased average incorporation rate that is at least 50-fold greater than the average incorporation rate of a wild-type polymerase having the amino acid sequence of SEQ ID NO: 1.
[0216] The present disclosure provides polymerases from the archaeon Candidatus Altiarchaeales that are mutated in certain domains to improve binding and / or nucleotide analog incorporation. See, e.g., SEQ ID NOs:269-274.
[0217] For example, an N-terminal domain comprising the amino acid sequence of SEQ ID NO: 269 can be mutated at one or more positions Y10, K58, V91, C104, and / or C130. In some embodiments, the domain comprising the amino acid sequence of SEQ ID NO: 269 comprises any one or any combination of two or more amino acid substitutions Y10F, Y10A, Y10V, Y10I, Y10L, Y10M, Y10W, K58M, V91Q, V91A, V91I, V91L, V91M, V91F, V91Y, V91W, V91S, V91T, V91N, C104S, C130S, and / or C130R.
[0218] An exonuclease domain comprising the amino acid sequence of SEQ ID NO: 270 can be mutated at one or more positions D141, E143, C269, and / or P335Q. In some embodiments, the domain comprising the amino acid sequence of SEQ ID NO: 270 comprises any one or any combination of two or more amino acid substitutions D141A, E143A, C269S, and / or P335. In some embodiments, the amino acid substitution mutations can include D141A and E143A to knock out 3' to 5' exonuclease activity (e.g., proofreading activity).
[0219] In some embodiments, a palm (1) domain (e.g., the first palm domain) comprising the amino acid sequence of SEQ ID NO: 271 can be mutated at one or more positions G355, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, and / or C450. In some embodiments, a domain (e.g., the first palm domain) comprising the amino acid sequence of SEQ ID NO: 271 can be mutated at two or more amino acid substitutions G355S, E370, D381Y, G403A, H405P, H405S, D406H, R414S, S415A, L416V, L416G, L416T, L416A, L416S, L416I, L416J, L416K ... 16F, L416Y, L416M, Y417T, Y417S, Y417G, Y417A, Y417V, Y417I, P418S, P418G, P418V, P418C, P418K, P418I, P418T, P418A, D439S, S440N, S443N, and / or C450S.
[0220] In some embodiments, a finger domain (e.g., a finger domain) comprising the amino acid sequence of SEQ ID NO: 272 can be mutated at one or more positions R468, K473, V489, Q492, A493, L494, K495, N499, M501, and / or Y502. In some embodiments, a domain (e.g., a finger domain) comprising the amino acid sequence of SEQ ID NO: 272 can be mutated at two or more amino acid substitutions R468A, R468V, R468S, R468K, R468H, R468G, K473A, V489I, Q492R, Q492C, Q492F, Q492A, Q492G, A493V, A493S, L494V , K495G, K495A, K495Q, K495S, K495V, N499G, N499A, N499S, N499V, M501I, Y502T, Y502V, Y502S, Y502R, Y502G, Y502N, Y502A, Y502Q, Y502P, Y502H, and / or Y502F. In some embodiments, a domain (e.g., a finger domain) comprising the amino acid sequence of SEQ ID NO: 272 further comprises any one or any combination of two or more amino acid substitutions G474S, G474D, N480I, R483H, D488N, A498G, S500G, M501V, and / or Y502F.
[0221] In some embodiments, the palm(2) domain (e.g., the second palm domain) comprising the amino acid sequence of SEQ ID NO: 273 can be mutated at one or more positions and / or F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, and / or L611. In some embodiments, a domain (e.g., the second palm domain) comprising the amino acid sequence of SEQ ID NO: 273 comprises any one or any combination of two or more amino acid substitutions F507S, C514S, R515L, R515W, R515Y, R515P, R515F, C517S, I529H, I529T, I529V, I529S, I529G, I529A, I529L, I529F, N567D, E569G, S577I, R608K, K610E, L611S and / or D622. In some embodiments, the domain (e.g., the second palm domain) comprising the amino acid sequence of SEQ ID NO: 273 further comprises any one or any combination of two or more amino acid substitutions E516G, S520N, S520G, K538R, F539Y, D560G, D560E, V564I, M565V, A568V, D573N, K574R, E578N, E581G, M583K, and / or D622T.
[0222] In some embodiments, a thumb domain (e.g., thumb domain) comprising the amino acid sequence of SEQ ID NO: 274 can be mutated at one or more positions V651, D653, A669, Q673, E680, R697, S717, R723, I750, and / or E760. In some embodiments, a domain (e.g., thumb domain) comprising the amino acid sequence of SEQ ID NO: 274 comprises any one or any combination of two or more amino acid substitutions V651M, D653G, E680D, A669D, Q673I, R697G, S717G, R723H, I750V, and / or E760G. In some embodiments, a domain (e.g., a thumb domain) comprising the amino acid sequence of SEQ ID NO: 274 further comprises any one or any combination of two or more amino acid substitutions: D636G, T675A, K682I, V689A, N705D, S717N, E730R, S746C, and / or E758R.
[0223] The present disclosure provides polymerases from the archaeon Candidatus Altiarchaeales that are mutated at two or more positions to increase the rate of incorporation of nucleotide analogs. In some embodiments, the mutant polymerase from the archaeon Candidatus Altiarchaeales comprises the amino acid sequence of SEQ ID NO: 1 or 391 with one or more amino acid substitution mutations selected from the group consisting of L416, Y417, P418, A493, and / or I529 (see, e.g., Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E)). In some embodiments, the amino acid substitution mutation at position L416 comprises a nonpolar amino acid or a polar, uncharged amino acid. In some embodiments, the amino acid substitution mutation at position L416 comprises valine, glycine, threonine, alanine, serine, isoleucine, leucine, phenylalanine, tyrosine, or methionine. In some embodiments, the amino acid substitution mutation at position Y417 comprises a nonpolar amino acid or a polar, uncharged amino acid. In some embodiments, the amino acid substitution mutation at position Y417 comprises threonine, serine, glycine, alanine, valine, isoleucine, or tyrosine. In some embodiments, the amino acid substitution mutation at position P418 comprises a polar, uncharged amino acid, a nonpolar amino acid, or a positively charged amino acid. In some embodiments, the amino acid substitution mutation at position P418 comprises serine, glycine, valine, cysteine, lysine, isoleucine, threonine, or proline. In some embodiments, the amino acid substitution mutation at position A493 comprises a nonpolar amino acid or a polar, uncharged amino acid. In some embodiments, the amino acid substitution mutation at position A493 comprises valine or serine. In some embodiments, the amino acid substitution mutation at position I529 comprises a positively charged amino acid, a polar, uncharged amino acid, or a nonpolar amino acid. In some embodiments, the amino acid substitution mutation at position 1529 comprises histidine, threonine, valine, serine, glycine, alanine, leucine, phenylalanine.In some embodiments, the mutant polymerase from a Candidatus Altiarchaeales archaeon comprises the amino acid sequence of SEQ ID NO: 1 or 391 with amino acid substitution mutations at positions L416S, Y417A, P418G, A493S, and I529H. In some embodiments, the mutant polymerase from a Candidatus Altiarchaeales archaeon comprises the amino acid sequence of SEQ ID NO: 1 with amino acid substitution mutations at positions L416F, Y417A, P418G, A493S, and I529H. In some embodiments, the amino acid substitution mutations can also include D141A and E143A.
[0224] In some embodiments, a polypeptide according to the present disclosure may include a single point mutation (see, e.g., Table 1 (Figures 3A-3H), Table 2 (Figures 4A-4E), or Table 3 (Figures 5A-5E)). In some embodiments, the single point mutation in SEQ ID NO: 1 or 391 is G403A, H405P, H405S, D406H, R414S, S415A, S415G, S415V, L416V, L416G, L416T, L416A, L416S, L416I, L416L, Y417T, Y417S, Y417G, Y417A, Y417V, Y417I, P418S, P418G, P418V, P418C, P418K, P418I, P418T, R468A, R468V, R468S, R468K, R468H , R468G, A493V, K495G, K495A, K495Q, K495S, K495V, N499G, N499A, N499S, N499V, S500G, M501I, Y502T, Y502V, Y502S, Y502R, Y502G, Y502N, Y502A, Y502Q, Y502P, Y502H, Y502F, F507S, I529H, I529T, I529V, I529S, I529G, I529A, I529L, and / or I529F, or any combination thereof. In some embodiments, the single point mutation in SEQ ID NO: 1 or 391 may include one or more of Y502V, Y502S, Y502Q, Y502P, Y502H, Y502G, Y502, Y502A, S415V, S415G, S415A, R468V, R468S, R468K, R468H, R468G, R468A, R414S, L416V, L416S, L416A, I529V, I529S, I529H, I529G, I529A, I529H, or any combination thereof.
[0225] In some embodiments, a polypeptide according to the present disclosure comprises the amino acid sequence of SEQ ID NO: 1 or 391 and includes multiple mutations, e.g., see Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E), such as G403A, H405P, H405S, D406H, R414S, S415A, S415G, S415V, L416V, L416G, L416T, L416A, L416S, L416I, L416L, Y417T, Y417S, Y417G, Y417A, Y417V, Y417I, P418S, P418G, P418V, P418C, P418K, P418L ... 8I, P418T, R468A, R468V, R468S, R468K, R468H, R468G, A493V, K495G, K495A, K495 Q, K495S, K495V, N499G, N499A, N499S, N499V, M501I, Y502T, Y502V, Y502S, Y502R , Y502G, Y502N, Y502A, Y502Q, Y502P, Y502H, Y502F, F507S, I529H, I529T, I529V, I529S, I529G, I529A, I529L, and / or I529F, or any combination thereof.
[0226] In some embodiments, a polypeptide according to the present disclosure comprises the amino acid sequence of SEQ ID NO: 1 or 391 and has a double mutation (see, e.g., Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E)), such as Y417V_P418V, Y417V_P418S, Y417V_P418A, Y417T_P418K, Y417S_P418S, Y417S_P418G, Y417S _P418A, Y417G_P418V, Y417G_P418S, Y417G_P418G, Y417G_P418C, Y417G_P418A, Y417A_P418V, Y417A_P418S, Y417 A_P418G, Y417A_P418A, S415V_Y417S, S415V_Y417A, S415V_P418V, S415V_P418S, S415V_P418G, S415V_L416V, S41 5V_L416S, S415G_Y417V, S415G_Y417S, S415G_P418V, S415G_P418S, S415G_P418G, S415G_P418A, S415G_L416V, S4 15G_L416S, S415G_L416G, S415G_L416A, S415A_Y417G, S415A_Y417A, S415A_P418G, S415A_L416S, L416V_P418S, L 416V_P418G, L416S_Y417G, L416S_Y417A, L416S_P418V, L416S_P418G, L416G_Y417G, L416G_Y417A, L416G_P418G, L416G_P418A, L416A_Y417S, L416A_P418S, L416A_P418G, and / or L416A_P418A, or any combination thereof.
[0227] In some embodiments, a polypeptide according to the present disclosure comprises the amino acid sequence of SEQ ID NO: 1 or 391 and may have a triple mutation. In some exemplary embodiments, the triple mutation (see, e.g., Table 1 (FIGS. 3A-3H), Table 2 (FIGS. 4A-4E), or Table 3 (FIGS. 5A-5E)) is Y417V_P418V_Y502S, Y417V_P418G_Y502G, Y417V_P418A_Y502R, Y417S_P418V_Y502R, Y417S_P418G_Y502S, Y417G_P418G_Y502V, Y417G_P418A_Y502N, Y417G_P418A_Y502G, Y417A_P418S _Y502R, G403A_H405S_D406H, A493V_K495V_N499S, A493V_K495V_N499G, A493V_K495V_N499A, A493V_K495S_N499V, A493V_K495S_N499S, A493V_K495S_N499G, A493V_K495Q_N499V, A493V_K495Q_N499G, A493V_K495Q_N499A, A493V_K495G_N499V, A493V_K495G_N499S, A493V_K 495G_N499G, A493V_K495G_N499A, A493V_K495A_N499G, A493V_K495A_N499A, A493V_K495S_N499G, L416A_Y417A_P418A, L416A_Y417A_P4 18G, L416A_Y417A_P418I, L416A_Y417S_P418A, L416A_Y417S_P418G, L416A_Y417S_P418S, L416G_Y417G_P418G, L416I_Y417A_P418G, L41 6I_Y417A_P418S, L416I_Y417G_P418A, L416I_Y417I_P418V, L416S_Y417A_P418G, L416S_Y417G_P418A, L416T_Y417A_P418A, L416T_Y417G_P418A, L416V_Y417A_P418A, L416V_Y417G_P418G, L416I_Y417S_P418S, and / or L416V_Y417V_P418G, or any combination thereof.
[0228] In some embodiments, a polypeptide according to the present disclosure comprises the amino acid sequence of SEQ ID NO: 1 or 391 and may have a quadruple mutation (see, e.g., (Figures 3A-3H), Table 2 (Figures 4A-4E), or Table 3 (Figures 5A-5E)). In some exemplary embodiments, the quadruple mutation is H405P_A493V_K495G_N499A, A493V_K495S_N499A_Y502T, A493V_K495Q_N499G_F507S, A493V_K495G_N499G_M501I, L416A_Y417A_P418A_I529L, L416A_Y417A_P418A_I529H, L416A_Y417A_P418S_I529H, L416A_Y417G_P418A_I529H, L4 16A_Y417G_P418G_I529H, L416A_Y417G_P418S_I529H, L416G_Y417T_P418S_I529H, L416I_Y417A_P418A_I529L, L416I_Y417A_P41 8G_I529F, L416I_Y417A_P418S_I529H, L416I_Y417G_P418A_I529L, L416I_Y417S_P418G_I529H, L416L_Y417Y_P418P_I529H, L416S _Y417A_P418G_I529H, L416S_Y417A_P418G_I529F, L416S_Y417A_P418T_I529H, L416S_Y417G_P418G_I529H, L416S_Y417G_P418V_ I529H, L416T_Y417A_P418A_I529H, L416T_Y417A_P418G_I529H, L416T_Y417G_P418A_I529H, L416V_Y417A_P418A_I529H, L416V_Y 417A_P418G_I529S, L416V_Y417A_P418G_I529T, L416V_Y417A_P418G_I529H, L416V_Y417A_P418S_I529, L416V_Y417G_P418A_I529H, L416V_Y417G_P418G_I529H, L416V_Y417G_P418S_I529H, and / or L416V_Y417T_P418S_I529H, or any combination thereof.
[0229] In some embodiments, the disclosed compositions and methods include one or more mutations disclosed herein that can affect the thermostability of the enzyme and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at or surrounding one or more of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof. In some embodiments, the mutation(s) may include substitution of the residue with one of the 19 other naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to one of skill in the art.
[0230] 3' or 5' modified substrates disclosed herein. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 403 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) may include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) may include the substitution G403A. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 405 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) may include the substitution H405P or H405S. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 406 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) may include the substitution D406H. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 414 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) may include the substitution R414S. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 415 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at a position or location surrounding thereof. In some embodiments, the mutation(s) may include the substitution S415A, S415G, or S415V.In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 416 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitutions L416V, L416G, L416T, L416A, or L416S. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 417 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitutions Y417T, Y417S, Y417G, Y417A, Y417V, or Y417I. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 418 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitutions P418S, P418G, P418V, P418C, P418K, P418I, or P418T. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 468 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitutions R468A, R468V, R468S, R468K, R468H, or R468G. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 493 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at a position or location surrounding thereof. In some embodiments, the mutation(s) may include the substitution A493V or A493S.In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 495 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitution K495G, K495A, K495Q, K495S, or K495V. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 499 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitution N499G, N499A, N499S, or N499V. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 501 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitution M501I. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 502 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitution Y502T, Y502V, Y502S, Y502R, Y502G, Y502N, Y502A, Y502Q, Y502P, Y502H, or Y502F. In some embodiments, the mutation(s) may include one or more substitutions, deletions, or insertions at position 507 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at a position or location surrounding thereof. In some embodiments, the mutation(s) may include the substitution F507S.In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 515 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitutions R515L, R515W, R515Y, R515P, or R515F. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 529 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding thereof. In some embodiments, the mutation(s) may comprise the substitutions I529H, I529T, I529V, I529S, I529G, I529A, I529L, or I529F. In some embodiments, the mutation(s) may comprise one or more substitutions, deletions, or insertions at position 567 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) may comprise the substitution N567D. In some embodiments, the mutation(s) may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them. In some embodiments, the disclosed compositions and methods include one or more mutations disclosed herein that can affect the thermostability of the enzyme and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at or surrounding positions or locations of SEQ ID NO: 1, 393, or 391, or any combination thereof, or homologs or orthologs thereof.In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to one of skill in the art. In some embodiments, the mutation(s) can include the substitution H405P or H405S. In some embodiments, the mutation can be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0231] In some embodiments, the disclosed compositions and methods include one or more mutations disclosed herein that can affect the thermostability of an enzyme and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 406 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) can include the substitution D406H. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0232] In some embodiments, the disclosed compositions and methods include one or more mutations disclosed herein that can affect the thermostability of an enzyme and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 414 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) can include the substitution R414S. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0233] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 415 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or unnatural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions S415A, S415G, or S415V. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0234] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 416 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or non-natural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions L416V, L416G, L416T, L416A, L416S, L416I, or L416L. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0235] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 417 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or unnatural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions Y417T, Y417S, Y417G, Y417A, Y417V, Y417I, or Y417Y. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 418, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0236] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 418 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include substitutions of the residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-naturally occurring amino acid known to one of skill in the art. In some embodiments, the mutation(s) may include substitutions P418S, P418G, P418V, P418C, P418K, P418I, P418T, or P418P. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 468, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0237] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 468 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or non-natural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions R468A, R468V, R468S, R468K, R468H, or R468G. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 493, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0238] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 493 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or unnatural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions A493V or A493S. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 495, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0239] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 495 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or unnatural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions K495G, K495A, K495Q, K495S, or K495V. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 499, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0240] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 499 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or non-natural amino acids known to those of skill in the art. In some embodiments, the mutation(s) can include substitutions N499G, N499A, N499S, or N499V. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 501, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0241] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 501 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) can include the substitution M501I. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 502, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0242] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 502 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at or surrounding positions or locations. In some embodiments, the mutation(s) can include substitutions of the residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to one of skill in the art. In some embodiments, the mutation(s) may include substitutions Y502T, Y502V, Y502S, Y502R, Y502G, Y502N, Y502A, Y502Q, Y502P, Y502H, or Y502F. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 507, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0243] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 507 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) can include the substitution F507S. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 515, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0244] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 515 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) can include the substitution R515L, R515W, R515Y, R515P, or R515F. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 529, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0245] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 529 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to one of skill in the art. In some embodiments, the mutation(s) may include the substitutions I529H, I529T, I529V, I529S, I529G, I529A, I529L, or I529F. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, and / or 567, or any combination thereof, or at positions or locations surrounding them.
[0246] In some embodiments, the disclosed compositions and methods include one or more mutations that can affect the thermostability of an enzyme as disclosed herein and / or its ability to accept modified substrates, e.g., 3'- or 5'-modified substrates. In some embodiments, the mutation(s) can include one or more substitutions, deletions, or insertions at position 567 of SEQ ID NO: 1, 393, or 391, or any combination thereof, or a homolog or ortholog thereof, or at positions or locations surrounding them. In some embodiments, the mutation(s) can include a substitution of that residue with one of the 20 naturally occurring amino acids (i.e., W, I, M, P, F, G, A, V, L, H, E, R, K, D, N, Y, C, S, T, or Q) or a non-natural amino acid known to those of skill in the art. In some embodiments, the mutation(s) can include the substitution N567D. In some embodiments, the mutations may be combined with one or more mutations at other positions, such as one or more substitutions, deletions, or insertions at any of positions 403, 405, 406, 414, 415, 416, 417, 418, 468, 493, 495, 499, 501, 502, 507, 515, and / or 529, or any combination thereof, or at positions or locations surrounding them.
[0247] In some embodiments, the methods and compositions provide polymerase variants with increased thermostability and increased resistance to the incorporation of non-canonical nucleotides, particularly 3'-blocked nucleotides. In some embodiments, the disclosed methods and compositions provide one or more polypeptides having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to SEQ ID NOs: 2-274, or 288-375, or 385-397. The disclosed methods and compositions include one or more polypeptides having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to any of SEQ ID NOs: 2-274, or 288-375, or 385-397.In some embodiments, the methods and compositions of the disclosure provide a method for detecting 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 99.8%, at least 99.9 ... and R615K, Y654A, Y654D, Y654E, Y654F, Y654G, S655A, S655G, S655V, Q656A, Q656G, Q656N, Q656S, Q656V, I657A, I657B, I657C, I657D, I657E, I657F, I657G, I657H ... 57G, I657S, I657V, E658A, E658D, E658G, E658S, E658V, L659A, L659G, L659P, L659S, L659V, D680A, D680G, D680I, D 680L, D680N, D680S, D680V, H682A, H682G, H682N, H682Q, H682S, H682V, R702A, R702G, R702H, R702K, R702S, R702V, K and / or H829G, or any combination thereof.
[0248] In some embodiments, the disclosed methods and compositions provide polymerase variants with increased resistance to the incorporation of non-standard nucleotides, such as 3'-blocked nucleotides, and particularly enhanced thermostability. In some embodiments, the disclosed methods and compositions provide polymerase variants that have one or more polypeptides having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to SEQ ID NOs: 2-274, 288-375, or 385-397. The disclosed methods and compositions include one or more polypeptides having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to any of SEQ ID NOs: 2-274, or 288-375, or 385-397.In some embodiments, the methods and compositions of the disclosure provide a method for detecting a nucleotide sequence encoding a nucleotide sequence of any of SEQ ID NOs: 2-274, 288-375, or 385-397 that is 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80% or more of a nucleotide sequence encoding a nucleotide sequence of any of SEQ ID NOs: 2-274, 288-375, or 385-397. , at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity and having one or more mutations selected from D314E, I332L, I334L, K368R, K381R, I385L, K417R, K434R, I454L, D471E, I528L, K601R, K635R, I649L, I665L, K758R, and / or K760R, or any combination thereof.
[0249] In some embodiments, the methods and compositions provide polymerase variants with increased thermostability and increased resistance to the incorporation of non-canonical nucleotides, particularly 3'-blocked nucleotides. In some embodiments, the disclosed methods and compositions provide one or more polypeptides having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to SEQ ID NOs: 2-274, or 288-375, or 385-397. The disclosed methods and compositions include one or more polypeptides having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to any of SEQ ID NOs: 2-274, or 288-375, or 385-397.In some embodiments, the disclosed methods and compositions comprise a nucleic acid sequence having 100%, at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99%, at least 98%, at least 97%, at least 95%, at least 90%, at least 85%, at least 80%, at least 75%, at least 70%, at least 65%, at least 60%, at least 55%, or at least 50% sequence identity to any of SEQ ID NOs: 2-274, or 288-375, or 385-397, and R615K, Y654A, Y654D, Y654E, Y654F, Y654G, S655A, S655G, S655V, Q656A, Q656B, Q656C, Q656D, Q656E, Q656F, Q656G, Q656H ... G, Q656N, Q656S, Q656V, I657A, I657G, I657S, I657V, E658A, E658D, E658G, E658S, E658V, L659A, L659G, L6 59P, L659S, L659V, D680A, D680G, D680I, D680L, D680N, D680S, D680V, H682A, H682G, H682N, H682Q, H682S, H 682V, R702A, R702G, R702H, R702K, R702S, R702V, K706H, K706K, K706R, A707G, A707S, A707T, F710A, F710D, F710E, F710G, F710Q, F710S, F710T, F710V, Y714A, Y714D, Y714E, Y714F, Y714G, Y714S, Y714W, H829A, H829G The present invention also includes one or more polypeptides having one or more mutations selected from D314E, I332L, I334L, K368R, K381R, I385L, K417R, K434R, I454L, D471E, I528L, K601R, K635R, I649L, I665L, K758R, and / or K760R, or any combination thereof.
[0250] The present disclosure provides a polymerase from the archaeon Candidatus Altiarchaeales, a truncated polypeptide that exhibits increased thermostability, eg, a polymerase having the amino acid sequence of SEQ ID NO: 353 or 354.
[0251] The present disclosure provides polymerases from the archaeon Candidatus Altiarchaeales that are mutated at two or more positions to increase the rate of incorporation of nucleotide analogs. In some embodiments, the mutant polymerase from the archaeon Candidatus Altiarchaeales comprises the amino acid sequence of SEQ ID NO: 1 or 391 with one or more amino acid substitution mutations selected from the group consisting of L416, Y417, P418, A493, and / or I529, and also comprises one or more amino acid substitution mutations selected from the group consisting of K58, R515, N567, E569, S577, K610, and / or S717. In some embodiments, the amino acid substitution mutation at position K58 comprises a polar, uncharged amino acid. In some embodiments, the amino acid substitution mutation at position K58 comprises a methionine. In some embodiments, the amino acid substitution mutation at position R515 comprises a nonpolar amino acid or a polar, uncharged amino acid. In some embodiments, the amino acid substitution mutation at position R515 comprises leucine, tryptophan, tyrosine, proline, or phenylalanine. In some embodiments, the amino acid substitution mutation at position N567 comprises a negatively charged amino acid. In some embodiments, the amino acid substitution mutation at position N567 comprises aspartic acid. In some embodiments, the amino acid substitution mutation at position E569 comprises a nonpolar amino acid. In some embodiments, the amino acid substitution mutation at position E569 comprises glycine. In some embodiments, the amino acid substitution mutation at position S577 comprises a nonpolar amino acid. In some embodiments, the amino acid substitution mutation at position S577 comprises isoleucine. In some embodiments, the amino acid substitution mutation at position K610 comprises a negatively charged amino acid. In some embodiments, the amino acid substitution mutation at position K610 comprises glutamic acid. In some embodiments, the amino acid substitution mutation at position S717 comprises a nonpolar amino acid. In some embodiments, the amino acid substitution mutation at position S717 comprises glycine. In some embodiments, the amino acid substitution mutations can also include D141A and E143A.
[0252] The present disclosure provides polymerases from the archaeon Candidatus Altiarchaeales that are mutated at two or more positions to increase the rate of nucleotide analog incorporation compared to a wild-type polymerase comprising SEQ ID NO: 1 or 391. In some embodiments, the mutant polymerase exhibits increased thermostability compared to a wild-type polymerase having the amino acid sequence of SEQ ID NO: 1 or 391. For example, the mutant polymerase exhibits increased thermostability in a temperature range of about 25-50°C or about 45-75°C. In some embodiments, the mutant polymerase comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8% identical, or has a higher level of sequence identity, to any of SEQ ID NOs: 1 or 2-274, 288-375, or 385-397. The mutant polymerase can include any of the characteristics described in this paragraph. For example, in some embodiments, a mutant polymerase from the archaeon Candidatus Altiarchaeales comprises an amino acid sequence having at least 85% sequence identity, at least 90% sequence identity, or at least 95% sequence identity, or at least 96% sequence identity, or at least 97% sequence identity, or at least 98% sequence identity, or at least 99.8%, at least 99.7%, at least 99.6%, at least 99.5%, at least 99.4%, at least 99.3%, at least 99.2%, at least 99.1%, at least 99% sequence identity, or higher percent sequence identity to SEQ ID NO: 1, 393, or 391, and the mutant DNA polymerase comprises an amino acid substitution at any one or any combination of two or more positions selected from the group consisting of Leu416, Tyr417, Pro418, Ala493, Arg515, Ile529, and Asn567. In some embodiments, the mutant polymerase comprises the amino acid substitutions D141A and E143A, which can confer exonuclease-minus activity.In some embodiments, the mutant polymerase exhibits desirable characteristics compared to a polymerase having a wild-type amino acid backbone sequence (e.g., SEQ ID NO: 1 or 391). For example, the mutant polymerase exhibits increased thermostability (Tm). In another example, the mutant polymerase exhibits an increased incorporation rate of nucleotide analogs containing chain-terminating moieties (e.g., blocking moieties) at the 2' and / or 3' sugar positions. In yet another example, the mutant polymerase exhibits increased uracil tolerance. One or more of the properties described in this paragraph may be present in any of the exemplary mutant polymerases in the various embodiments described herein. The properties described in this paragraph are referred to throughout this disclosure as "exemplary mutant polymerase properties."
[0253] The present disclosure provides polymerases from the archaeon Candidatus Altiarchaeales that are mutated at one or more positions to increase the rate of incorporation of nucleotide analogs. In some embodiments, the mutant polymerase from the archaeon Candidatus Altiarchaeales comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or more sequence identity to any one of the amino acid sequences set forth in SEQ ID NOs: 2-274, 288-375, or 385-397.
[0254] The present disclosure provides archaeal family B DNA polymerases, including 9°N DNA polymerases and THERMINATOR polymerases, mutated at one or more positions. In some embodiments, the mutant 9°N and THERMINATOR polymerases mutate at one or more positions Y7, T55, V106, D132, I264, Y291, P328, S348, L352, K363, E374, G395, W397, D398, R406, S407, L408, Y409, P410, Y431, D432, P435, C442, R460, R465, Y481, R484, A485, I486, K487, I488, N491, In some embodiments, the one or more amino acid substitutions of SEQ ID NO: 280 comprise at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO: 280, 281, or 282, and the one or more amino acid substitution mutations at F493, Y494, Y499, C506, K507, C509, I521, K559, K561, P569, E600, K602, I603, D614, V643, E645, V661, Q665, R689, R709, I715, I744, and / or D754. In some embodiments, the one or more amino acid substitutions of SEQ ID NO: 280 comprise the amino acid sequence of any one of SEQ ID NOs: 1-274, 288-375, or 385-397, respectively. Positions Y7, K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A49 in Altiarchaeal polymerases 3, L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760 are positionally equivalent to the amino acid substitutions.In some embodiments, the positionally equivalent amino acid positions of the Candidatus Altiarchaeales archaeal polymerase and 9°N polymerase, VENT polymerase, DEEP VENT polymerase, Geobacillus stearothermophilus polymerase, Pfu polymerase, and Pyrococcus abyssi polymerase are listed in Table 4 shown in Figures 6A-6C. See also the sequence alignments in Figures 7A-7B, 8A-8C, 9A-9B, 10A-10B, 11A-11B, 12A-12B, and 13A-13B.
[0255] The present disclosure provides archaeal family B DNA polymerases, including 9°N DNA polymerases mutated at one or more positions. In some embodiments, the mutant 9°N DNA polymerases mutate at one or more positions Y7, T55, V106, D132, I264, Y291, P328, S348, L352, K363, E374, G395, W397, D398, R406, S407, L408, Y409, P410, Y431, D432, P435, C442, R460, R465, Y481, R484, A485, I486, K487, I488, N491, F493, Y494, Y4 In some embodiments, the one or more amino acid substitution mutations in SEQ ID NO:280 include at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:280, including one or more amino acid substitution mutations at positions 99, C506, K507, C509, I521, K559, K561, P569, E600, K602, I603, D614, V643, E645, V661, Q665, D672, R689, R709, I715, I744, and / or D754. In some embodiments, the amino acid at position 129 can be methionine or alanine. In some embodiments, the one or more amino acid substitutions in SEQ ID NO:280 include at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:280, including one or more amino acid substitution mutations at positions 99, C506, K507, C509, I521, K559, K561, P569, E600, K602, I603, D614, V643, E645, V661, Q665, D672, R689, R709, I715, I744, and / or D754. In some embodiments, the one or more amino acid substitutions in SEQ ID NO:280 include at least one Candidatus strain comprising the amino acid sequence of any one of SEQ ID NOs:1-274, 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760 are positionally equivalent to the amino acid substitutions.In some embodiments, the positionally equivalent amino acid positions of the Candidatus Altiarchaeales archaeal polymerase and the 9°N polymerase are listed in Table 4, shown in Figures 6A-6C. See also the sequence alignments in Figures 7A-7C.
[0256] The present disclosure provides archaeal family B DNA polymerases, including 9°N DNA polymerases mutated at one or more positions. In some embodiments, the mutant 9°N DNA polymerases mutate at one or more positions Y7, T55, V106, D132, I264, Y291, P328, S348, L352, K363, E374, G395, W397, D398, R406, S407, L408, Y409, P410, Y431, D432, P435, C442, R460, R465, Y481, R484, A485, I486, K487, I488, N491, F493, Y494, Y4 In some embodiments, the one or more amino acid substitution mutations in SEQ ID NO:281 include at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:281, including one or more amino acid substitution mutations at positions 99, C506, K507, C509, I521, K559, K561, P569, E600, K602, I603, D614, V643, E645, V661, Q665, D672, R689, R709, I715, I744, and / or D754. In some embodiments, the amino acid at position 129 can be methionine or alanine. In some embodiments, the one or more amino acid substitutions in SEQ ID NO:281 include at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:281, including one or more amino acid substitution mutations at positions 99, C506, K507, C509, I521, K559, K561, P569, E600, K602, I603, D614, V643, E645, V661, Q665, D672, R689, R709, I715, I744, and / or D754. In some embodiments, the one or more amino acid substitutions in SEQ ID NO:281 include at least one Candidatus strain comprising the amino acid sequence of any one of SEQ ID NOs:1-274, 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760 are positionally equivalent to the amino acid substitutions.In some embodiments, the positionally equivalent amino acid positions of the Candidatus Altiarchaeales archaeal polymerase and the 9°N polymerase are listed in Table 4, shown in Figures 6A-6C. See also the sequence alignment in Figures 7A-7B.
[0257] The present disclosure provides archaeal family B DNA polymerases, including a THERMINATOR DNA polymerase mutated at one or more positions. In some embodiments, the mutant THERMINATOR polymerase mutates at one or more positions Y7, T55, V106, D132, I264, Y291, P328, S348, L352, K363, E374, G395, W397, D398, R406, S407, L408, Y409, P410, Y431, D432, P435, C442, R460, R465, Y481, R484, A485, I486, K487, I488, N491 , F493, Y494, Y499, C506, K507, C509, I521, K559, K561, P569, E600, K602, I603, D614, V643, E645, V661, Q665, D672, R689, R709, I715, I744, and / or D754. In some embodiments, the amino acid at position 129 can be methionine or alanine. In some embodiments, the one or more amino acid substitutions in SEQ ID NO:282 are selected from the group consisting of Candidatus spp. comprising the amino acid sequence of any one of SEQ ID NOs: 1-274, 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760 are positionally equivalent to the amino acid substitutions.In some embodiments, the positionally equivalent amino acid positions of the Candidatus Altiarchaeales archaeal polymerase and the THERMINATOR polymerase correspond to the positions of the 9°N polymerase listed in Table 4 shown in Figures 6A-6C. See also the sequence alignment in Figures 7A-7B.
[0258] The present disclosure provides archaeal family B DNA polymerases, including VENT DNA polymerases mutated at one or more positions. In some embodiments, the mutant VENT DNA polymerases mutate at positions Y7, K61, V106, D132, V266, G293, P330, S350, L354, A365, E376, G398, W400, E401, R409, S410, L411, Y412, P413, Y434, D435, P438, C445, R463, K468, Y484, R487, A488, I489, K490, L491, N494, I496, Y1035, Y1040, S1047, In some embodiments, the one or more amino acid substitutions of SEQ ID NO:283 comprise at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:283, and the one or more amino acid substitution mutations at K1048, C1050, I1062, K1490, K1492, S1500, E1531, R1533, I1534, D1545, V1574, D1576, V1592, Q1596, D1604, R1620, K1640, I1646, I1675, and / or D1685. In some embodiments, the one or more amino acid substitutions of SEQ ID NO:283 comprise the amino acid sequence of any one of SEQ ID NOs:1-274, or 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760. In some embodiments, the positionally equivalent amino acid positions in the Candidatus Altiarchaeales archaeal polymerase and the VENT polymerase are listed in Table 4, shown in Figures 6A-6C.See also the sequence alignments in Figures 8A-8C.
[0259] The present disclosure provides archaeal family B DNA polymerases, including DEEP VENT DNA polymerases mutated at one or more positions. In some embodiments, the mutant DEEP VENT DNA polymerases mutate at one or more positions Y7, K61, V106, D132, I264, Y291, P328, S348, L352, E363, E374, G396, W398, E399, R407, S408, L409, Y410, P411, Y432, D433, P436, C443, R461, R466, Y482, R485, A486, I487, K488, I489, N492, I494, Y1032, Y1037, C1044, In some embodiments, the one or more amino acid substitutions of SEQ ID NO:284 comprise at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:284, and the one or more amino acid substitution mutations are at K1045, C1047, I1059, K1097, L1099, A1107, E1138, K1140, I1141, D1152, V1181, E1183, V1199, Q1203, E1210, R1227, P1247, I1253, I1282, and / or D1292. In some embodiments, the one or more amino acid substitutions of SEQ ID NO:284 comprise the amino acid sequence of any one of SEQ ID NOs:1-274, or 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760. In some embodiments, the positionally equivalent amino acid positions in the Candidatus Altiarchaeales archaeal polymerase and the DEEP VENT polymerase are listed in Table 4, shown in Figures 6A-6C.See also the sequence alignments in Figures 9A-9C.
[0260] The present disclosure provides archaeal family B DNA polymerases, including Pfu DNA polymerases mutated at one or more positions. In some embodiments, the mutant Pfu DNA polymerase is mutated at one or more positions Y7, K61, V106, E150, I264, Y291, P328, S348, L352, E363, E374, G396, W398, E399, R407, S408, L409, Y410, P411, Y432, D433, P436, C443, R461, T466, Y482, K485, A486, I487, K488, L489, N492, F494, Y495, Y5 In some embodiments, the one or more amino acid substitutions of SEQ ID NO:285 comprise at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:285, and the one or more amino acid substitution mutations at positions 00, C507, K508, C510, I522, K560, L562, S570, E601, K603, V604, D615, V644, E646, A662, Q666, E673, K690, P710, I716, I745, and / or D755. In some embodiments, the one or more amino acid substitutions of SEQ ID NO:285 comprise the amino acid sequence of any one of SEQ ID NOs:1-274, or 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760. In some embodiments, the positionally equivalent amino acid positions in the Candidatus Altiarchaeales archaeal polymerase and the Pfu polymerase are listed in Table 4, shown in Figures 6A-6C. See also the sequence alignments in Figures 11A-11B.
[0261] The present disclosure provides archaeal family B DNA polymerases, including Pyrococcus abyssi DNA polymerases mutated at one or more positions. In some embodiments, the mutant Pyrococcus abyssi DNA polymerase is mutated at one or more positions Y7, K61, V106, N132, I264, Y291, P328, S348, L352, E363, E374, G396, W398, E399, R407, S408, L409, Y410, P411, Y432, D433, P436, C443, R461, K466, Y482, R485, A486, I487, K488, I489, N492, Y494, Y495, Y5 In some embodiments, the one or more amino acid substitutions of SEQ ID NO:286 comprise at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:286, and the one or more amino acid substitution mutations at positions 00, C507, K508, C510, I522, K559, L561, S569, E600, K602, I603, D614, V643, E645, V661, Q665, E672, K689, P709, I715, I744, and / or D754. In some embodiments, the one or more amino acid substitutions of SEQ ID NO:286 comprise the amino acid sequence of any one of SEQ ID NOs:2-274, or 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760. In some embodiments, the positionally equivalent amino acid positions in the Candidatus Altiarchaeales archaeal polymerase and the Pyrococcus abyssi polymerase are listed in Table 4, shown in Figures 6A-6C.See also the sequence alignment in Figures 12A-12B.
[0262] The present disclosure provides archaeal family A DNA polymerases, including Geobacillus stearothermophilus DNA polymerases mutated at one or more positions. In some embodiments, the mutant Geobacillus stearothermophilus DNA polymerase is mutated at one or more positions G10, D59, A97, Q123, S240, G277, E325, G342, R343, E349, A359, G384, E386, L387, L394, L395, L396, A397, A398, E420, A421, S424, K431, K450, W455, K476, Q479, P480, L481, A482, A483, A486, M488, E489, V4 In some embodiments, the one or more amino acid substitutions of SEQ ID NO:275 comprise at least 80%, 85%, 90%, 95%, or higher percent sequence identity to SEQ ID NO:275, and the one or more amino acid substitution mutations at positions 93, G504, S505, L507, N527, N573, L575, S585, E632, R634, K635, D647, I689, H691, A707, G711, D718, P744, Y772, S799, V828, and / or E840. In some embodiments, the one or more amino acid substitutions of SEQ ID NO:275 comprise the amino acid sequence of any one of SEQ ID NOs:1-274, or 288-375, or 385-397, respectively. Positions of Altiarchaeal polymerase: K58, C104, C130, C269, R296, P335, G355, R359, E370, D381, G403, H405, D406, R414, S415, L416, Y417, P418, D439, S440, S443, C450, R468, K473, V489, Q492, A493 , L494, K495, L496, N499, S500, M501, Y502, F507, C514, R515, C517, I529, N567, E569, S577, R608, K610, L611, D622, V651, D653, A669, Q673, R697, S717, R723, I750, and E760 are positionally equivalent to the amino acid substitutions.In some embodiments, the positionally equivalent amino acid positions of the Candidatus Altiarchaeales archaeal polymerase and the Geobacillus stearothermophilus polymerase are listed in Table 4, shown in Figures 6A-6C. See also the sequence alignment in Figures 10A-10B.
[0263] The present disclosure provides a polymerase operably linked to a detectable reporter moiety. Any of the polymerases described herein can be labeled with a detectable reporter moiety, including polymerases having a wild-type or mutant amino acid sequence backbone of any of the polymerases described herein, including Candidatus Altiarchaeales archaeal DNA polymerase (e.g., any of SEQ ID NOS: 1-274, or 288-375, or 385-397), Geobacillus DNA polymerase (e.g., SEQ ID NOS: 275-279), 9°N DNA polymerase (e.g., SEQ ID NOS: 280 and 281), THERMINATOR DNA polymerase (e.g., SEQ ID NO: 282), VENT DNA polymerase (e.g., SEQ ID NO: 283), DEEP VENT DNA polymerase (e.g., SEQ ID NO: 284), Pfu DNA polymerase (e.g., SEQ ID NO: 285), Pyrococcus abyssi DNA polymerase (e.g., SEQ ID NO: 286), and RB69 DNA polymerase (e.g., SEQ ID NO: 287). In some embodiments, the detectable reporter moiety generates a detectable signal due to a chemical or physical change (e.g., heat, light, electricity, pH, salt concentration, enzymatic activity, or a proximity event such as FRET). In some embodiments, the detectable reporter moiety comprises a luminescent moiety, a fluorescent moiety, or a quencher. In some embodiments, the detectable moiety comprises a fluorescent moiety that acts as a FRET donor or acceptor. The detectable reporter moiety can be attached to the polymerase at the N-terminus, C-terminus, or any internal position. The detectable reporter moiety is attached to the polymerase in a manner that does not interfere with the ability of the polymerase to bind to a nucleic acid template molecule, a nucleic acid primer, or a nucleotide. The detectable reporter moiety is attached to the polymerase in a manner that does not interfere with the catalytic activity of the polymerase, including the incorporation of nucleotides.
[0264] The present disclosure provides recombinant fusion polypeptides comprising any of the DNA polymerases described herein operably linked to any one or any combination of two or more exogenous amino acid sequences for affinity purification, cleavage, or solubilization. In some embodiments, the recombinant fusion polypeptides include any of the wild-type and mutant polymerases described herein, as well as polymerases having a substitution mutation at a site that is a positionally equivalent mutation site to the site shown in Table 4 (FIGS. 6A-6C) described herein, including Candidatus Altiarchaeales archaeal DNA polymerase (e.g., any of SEQ ID NOS: 1-274, 288-375, or 385-397), Geobacillus DNA polymerase (e.g., SEQ ID NOS: 275-279), 9°N DNA polymerase (e.g., SEQ ID NOS: 280 and 281), THERMINATOR DNA polymerase (e.g., SEQ ID NO: 282), VENT DNA polymerase (e.g., SEQ ID NO: 283), DEEP VENT DNA polymerase (e.g., SEQ ID NO: 284), Pfu DNA polymerase (e.g., SEQ ID NO: 285), Pyrococcus abyssi DNA polymerase (e.g., SEQ ID NO: 286), and RB69. Included are polymerases having the amino acid sequence backbone sequence of a DNA polymerase (e.g., SEQ ID NO: 287).
[0265] In some embodiments, the recombinant fusion polypeptide comprises any of the wild-type and mutant polymerases described herein operably linked at their N-terminus and / or C-terminus(s) to at least one affinity purification tag sequence(s), wherein the affinity purification tag sequence(s) comprise a histidine tag (e.g., a hexa-histidine tag (SEQ ID NO: 398)), a FLAG tag, a T7 tag, a Strep II tag, an S tag (e.g., derived from pancreatic ribonuclease A), an HA tag (e.g., derived from human influenza hemagglutinin protein), and / or a c-Myc tag.
[0266] In some embodiments, the recombinant fusion polypeptide comprises any of the wild-type and mutant polymerases described herein operably linked at their N-terminus and / or C-terminus(s) to at least one polypeptide cleavage sequence, or the polypeptide cleavage sequence can be positioned between the affinity tag sequence and the N-terminus or C-terminus of the polymerase sequence. In some embodiments, the polypeptide cleavage sequence can be recognized and cleaved by a protease or reducing conditions. In some embodiments, the polypeptide cleavage sequence comprises a thrombin cleavage sequence, a TEV cleavage sequence (e.g., from tobacco etch virus, including AcTEV and ProTEV), a factor Xa cleavage sequence, an enterokinase cleavage sequence, and a SUMO cleavage sequence (e.g., small ubiquitin-like modification, including Ulp1, Senp2, and SUMOstar).
[0267] In some embodiments, the recombinant fusion polypeptide comprises any of the wild-type and mutant polymerases described herein operably linked at the N-terminus and / or C-terminus(s) to at least one exogenous amino acid sequence for improving solubilization, including maltose-binding protein (MBP), small ubiquitin-like modifier (SUMO), and glutathione S-transferase (GST).
[0268] Polymerase-containing systems The present disclosure provides systems comprising one or more mutant polymerases and at least one nucleic acid template molecule having a self-priming 3' end. In some embodiments, the one or more mutant polymerases may or may not bind to the at least one nucleic acid template molecule having a self-priming 3' end. In some embodiments, the self-priming 3' end of the template molecule provides an initiation site for nucleotide polymerization. In some embodiments, the mutant polymerase comprises the properties of one or more of the exemplary mutant polymerases discussed above.
[0269] The present disclosure provides a system comprising one or more mutant polymerases, at least one nucleic acid template molecule, and at least one nucleic acid primer. In some embodiments, the one or more mutant polymerases may or may not bind to the at least one nucleic acid template molecule and the at least one nucleic acid primer. In some embodiments, the primer provides an initiation site for nucleotide polymerization. In some embodiments, the primer comprises a 3' extendable end for a polymerase-catalyzed nucleotide incorporation reaction, or the primer comprises a 3' non-extendable end. In some embodiments, the nucleic acid template molecule comprises at least one uridine nucleotide or lacks uridine nucleotides. In some embodiments, the mutant polymerase comprises the properties of one or more of the exemplary mutant polymerases discussed above.
[0270] In some embodiments, the system comprises one or more mutant polymerases bound to a nucleic acid duplex, each comprising a nucleic acid template hybridized to a nucleic acid primer, thereby forming a multiplex polymerase. In some embodiments, the primer provides an initiation site for nucleotide polymerization. In some embodiments, the mutant polymerase binds to a nucleic acid template molecule having a self-priming 3' end to form a multiplex polymerase lacking a separate primer molecule. In some embodiments, the nucleic acid template molecule contains at least one uridine nucleotide or lacks uridine nucleotides. In some embodiments, the mutant polymerase comprises the properties of one or more of the exemplary mutant polymerases discussed above.
[0271] In some embodiments, a system comprises one or more mutant polymerases, at least one nucleic acid template molecule, and an initiation site for nucleotide polymerization, wherein the mutant polymerase is in solution, the nucleic acid template molecule is in solution, and the initiation site (e.g., a primer) is in solution. In some embodiments, a system comprises one or more mutant polymerases, at least one nucleic acid template molecule, and an initiation site for nucleotide polymerization, wherein the system comprises any combination of mutant polymerases in solution, nucleic acid template molecules in solution or immobilized on a support, and initiation sites (e.g., primers) in solution or immobilized on a support. In some embodiments, a system comprises one or more mutant polymerases, at least one nucleic acid template molecule, and an initiation site for nucleotide polymerization, wherein the system comprises any combination of mutant polymerases in solution, nucleic acid template molecules in solution or immobilized on a support, and initiation sites (e.g., primers) in solution or immobilized on a support.
[0272] In some embodiments of the system, the mutant polymerase exhibits increased thermostability compared to a wild-type polymerase having the amino acid sequence of SEQ ID NO: 1 or 391. For example, the mutant polymerase exhibits increased thermostability in a temperature range of about 25-50°C or about 45-75°C.
[0273] In some embodiments of the system, the mutant polymerase exhibits an increased rate of incorporation of a nucleotide analog compared to a wild-type polymerase comprising SEQ ID NO: 1 or 391, wherein the nucleotide analog comprises a chain-terminating moiety (e.g., a blocking moiety) at the 2' and / or 3' sugar position.
[0274] In some embodiments, the system includes one or more mutant polymerases and a plurality of nucleic acid duplexes, each including a nucleic acid template hybridized to a nucleic acid primer. In some embodiments, the one or more polymerases and nucleic acid duplexes further include a plurality of nucleotides. The one or more mutant polymerases may or may not bind to the nucleic acid duplex. The one or more mutant polymerases may or may not bind to one of the nucleotides. In some embodiments, one or more mutant polymerases bind to a nucleic acid duplex including a nucleic acid template hybridized to a nucleic acid primer, thereby forming a multiplex polymerase, and the system further includes a plurality of nucleotides. In some embodiments, the mutant polymerase includes characteristics of one or more exemplary mutant polymerases discussed above.
[0275] In some embodiments of the system, a nucleotide can bind to the composite polymerase without being incorporated, hi some embodiments, a complementary nucleotide can bind to the composite polymerase without undergoing polymerase-catalyzed incorporation, forming a ternary complex in which the complementary nucleotide is attached to the 3' end of the primer at a position opposite the complementary nucleotide in the template strand.
[0276] In some embodiments of the system, at least one nucleotide of the plurality of nucleotides comprises a base, a sugar, and at least one phosphate group. For example, the nucleotide unit may comprise an aromatic base, a 5-carbon sugar (e.g., ribose or deoxyribose), and one or more phosphate groups (e.g., 1 to 10 phosphate groups), wherein the aromatic base of the nucleotide comprises adenine, guanine, cytosine, thymine, or uracil. In some embodiments, the plurality of nucleotides comprises one type of nucleotide selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. In some embodiments, the plurality of nucleotides comprises a mixture of any combination of two or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. In some embodiments, at least one of the nucleotides in the plurality of nucleotides is labeled with a fluorophore. In some embodiments, the plurality of nucleotides lacks a fluorophore label. One or more of the characteristics described in this paragraph may be present in any exemplary nucleotide unit in various embodiments described herein. The properties described in this paragraph are referred to throughout this disclosure as "exemplary nucleotide unit properties."
[0277] In some embodiments, the system comprises at least one nucleotide of the plurality of nucleotides comprising a chain of one, two, or three phosphorus atoms, typically linked to the 5' carbon of the sugar moiety via an ester or phosphoramide linkage. In some embodiments, at least one nucleotide of the plurality is an analog having a phosphorus chain linked together with an O, S, NH, methylene, or ethylene atom interposed therebetween. In some embodiments, the phosphorus atom in the chain comprises a substituted side group comprising O, S, or BH3. In some embodiments, the chain comprises a phosphate group substituted with an analog comprising a phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoramidite group.
[0278] In some embodiments, in the system, at least one nucleotide of the plurality of nucleotides comprises a nucleotide analog having a chain-terminating moiety (e.g., a blocking moiety) at the sugar 2' position, the sugar 3' position, or the sugar 2' and 3' positions. In some embodiments, the chain-terminating moiety can inhibit polymerase-catalyzed incorporation of a subsequent nucleotide unit or free nucleotide in a nascent strand during a primer extension reaction. In some embodiments, the chain-terminating moiety is attached to the 3' sugar hydroxyl position, where the sugar comprises a ribose or deoxyribose sugar moiety. In some embodiments, the chain-terminating moiety is removable / cleavable from the 3' sugar hydroxyl position to generate a nucleotide having a 3' OH sugar group that can be extended with a subsequent nucleotide in a polymerase-catalyzed nucleotide incorporation reaction. In some embodiments, the chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the chain-terminating moiety is cleavable / removable from the nucleotide, for example, by reacting the chain-terminating moiety with a chemical, a pH change, light, or heat. In some embodiments, the chain-terminating moieties alkyl, alkenyl, alkynyl, and aryl are cleavable using tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzo-quinone (DDQ). In some embodiments, the chain-terminating moieties aryl and benzyl are cleavable using HPd / C. In some embodiments, chain-terminating moieties amine, amide, keto, isocyanate, phosphate, thio, disulfide can be cleaved using phosphines or thiol groups, including β-mercaptoethanol or dithiothritol (DTT). In some embodiments, chain-terminating moieties carbonate can be cleaved using potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH).In some embodiments, the chain-terminating moieties urea and silyl are cleavable using tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride. One or more properties of nucleotide analogs containing chain-terminating moieties described in this paragraph may appear in any exemplary nucleotide analog in various embodiments described herein. The properties of nucleotide analogs described in this paragraph are referred to throughout this disclosure as "exemplary nucleotide analog properties." Similarly, one or more chain-terminating moiety properties described in this paragraph may appear in any exemplary chain-terminating moiety in various embodiments described herein. The phrase "chain-terminating moiety embodiments" is used throughout this disclosure to refer to any of the one or more chain-terminating moiety properties described in this paragraph.
[0279] In some embodiments, in the system, at least one nucleotide of the plurality of nucleotides comprises a terminator nucleotide analog having a chain-terminating moiety (e.g., a blocking moiety) at the 2' sugar position, the 3' sugar position, or the 2' and 3' sugar positions. In some embodiments, the chain-terminating moiety comprises an azide, azido, or azidomethyl group. In some embodiments, the chain-terminating moiety comprises a 3'-O-azido or 3'-O-azidomethyl group. In some embodiments, the chain-terminating azide, azido, and azidomethyl groups are cleavable / removable using a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bissulfotriphenylphosphine (BS-TPP), or tri(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP). In some embodiments, in the system, the nucleotide analog comprises a chain-terminating moiety selected from the group consisting of 3'-deoxynucleotide, 2',3'-dideoxynucleotide, 3'-methyl, 3'-azido, 3'-azidomethyl, 3'-O-azidoalkyl, 3'-O-ethynyl, 3'-O-aminoalkyl, 3'-O-fluoroalkyl, 3'-fluoromethyl, 3'-difluoromethyl, 3'-trifluoromethyl, 3'-sulfonyl, 3'-malonyl, 3'-amino, 3'-O-amino, 3'-sulfhydral, 3'-aminomethyl, 3'-ethyl, 3'butyl, 3'-tertbutyl, 3'-fluorenylmethyloxycarbonyl, 3'tert-butyloxycarbonyl, 3'-O-alkylhydroxylamino group, 3'-phosphorothioate, and 3-O-benzyl, or a derivative thereof. One or more of the properties described in this paragraph may appear in any chain-terminating moiety that includes an azide, azido, or azidomethyl group in various embodiments described herein.The phrase "a chain-terminating moiety comprising an azide, azido, or azidomethyl group" is used throughout this disclosure to refer to any one or more of the chain-terminating moiety characteristics described in this paragraph.
[0280] In some embodiments, the system includes a plurality of nucleotides lacking a detectable reporter moiety, e.g., a fluorophore. In some embodiments, the system includes a plurality of nucleotides labeled with a detectable reporter moiety. The detectable reporter moiety comprises a fluorophore. In some embodiments, the fluorophore is attached to the nucleotide base. In some embodiments, the fluorophore is attached to the nucleotide base by a linker that is cleavable / removable from the base.
[0281] In some embodiments, in the system, the cleavable linker on the base comprises a cleavable moiety including an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group. In some embodiments, the cleavable linker on the base is cleavable / removable from the base by reacting the cleavable moiety with a chemical, a pH change, light, or heat. In some embodiments, the cleavable moieties alkyl, alkenyl, alkynyl, and aryl are cleavable using tetrakis(triphenylphosphine)palladium(0) (Pd(PPh3)4) with piperidine or 2,3-dichloro-5,6-dicyano-1,4-benzoquinone (DDQ). In some embodiments, the cleavable moieties aryl and benzyl are cleavable using HPd / C. In some embodiments, amine, amide, keto, isocyanate, phosphate, thio, and disulfide cleavable moieties can be cleaved using phosphines or thiol groups, including β-mercaptoethanol or dithiothritol (DTT). In some embodiments, carbonate cleavable moieties can be cleaved using potassium carbonate (KCO) in MeOH, triethylamine in pyridine, or Zn in acetic acid (AcOH). In some embodiments, urea and silyl cleavable moieties can be cleaved using tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0282] In some embodiments, in the system, the cleavable linker on the base comprises a cleavable moiety comprising an azide, azido, or azidomethyl group. In some embodiments, the cleavable moieties azide, azido, and azidomethyl groups are cleavable / removable using a phosphine compound. In some embodiments, the phosphine compound comprises a derivatized trialkylphosphine moiety or a derivatized triarylphosphine moiety. In some embodiments, the phosphine compound comprises tris(2-carboxyethyl)phosphine (TCEP), bissulfotriphenylphosphine (BS-TPP), or tri(hydroxypropyl)phosphine (THPP). In some embodiments, the cleaving agent comprises 4-dimethylaminopyridine (4-DMAP).
[0283] In some embodiments, in the system, the cleavable linker on the chain-terminating moiety (e.g., sugar 2' and / or sugar 3' position) and the base have the same or different cleavable moieties. In some embodiments, the detectable reporter moiety linked to the chain-terminating moiety (e.g., sugar 2' and / or sugar 3' position) and the base are chemically cleavable / removable using the same chemical agent. In some embodiments, the detectable reporter moiety linked to the chain-terminating moiety (e.g., sugar 2' and / or sugar 3' position) and the base are chemically cleavable / removable using different chemical agents.
[0284] In some embodiments, the system comprises one or more mutant polymerases and a nucleic acid duplex, each comprising a nucleic acid template hybridized to a nucleic acid primer. In some embodiments, the one or more polymerases and nucleic acid duplex further comprise a plurality of multivalent molecules. The one or more mutant polymerases may or may not be bound to the nucleic acid duplex. The one or more mutant polymerases may or may not be bound to one or more of the multivalent molecules. In some embodiments, one or more mutant polymerases bind to a nucleic acid duplex comprising a nucleic acid template hybridized to a nucleic acid primer, thereby forming a multiplex polymerase, and the system further comprises a plurality of multivalent molecules. In some embodiments, the mutant polymerase comprises the properties of one or more of the exemplary mutant polymerases discussed above.
[0285] In some embodiments of the system, at least one multivalent molecule of the plurality of multivalent molecules comprises (a) a core and (b) a plurality of nucleotide arms, the plurality of nucleotide arms comprising (i) a core-binding moiety, (ii) a spacer comprising a PEG moiety, (iii) a linker, and (iv) a nucleotide unit, wherein the core is bound to the plurality of nucleotide arms, the spacer is bound to the linker, and the linker is bound to the nucleotide unit. In some embodiments, the nucleotide unit comprises a base, a sugar, and at least one phosphate group, and the linker is bound to the nucleotide unit via the base. An exemplary spacer is shown in Figure 16A (top). Various exemplary linkers are shown in Figures 16A (bottom) and 16B. Examples of various linkers linking / binding to nucleotide units are shown in Figures 17A-17C, where the 5-position of a pyrimidine base or the 7-position of a purine base is bound to the linker via a propargylamine bond (see also Figure 18). In some embodiments, the core comprises a streptavidin- or avidin-type moiety, and the core-binding moiety comprises biotin. In some embodiments, the linker comprises an aliphatic chain having 2 to 6 subunits or an oligoethylene glycol chain having 2 to 6 subunits. In some embodiments, the linker further comprises an aromatic moiety. In some embodiments, the linker comprises an aliphatic chain or an oligoethylene glycol chain, and both linker chains have 2 to 6 subunits. In some embodiments, the linker also comprises an aromatic moiety. An exemplary spacer is shown in Figure 16A (top), and exemplary linkers are shown in Figures 16A (bottom) and 16B. An exemplary nucleotide arm is shown in Figure 15B. An exemplary multivalent molecule is shown in Figures 14A, 14B, and 15A. In some embodiments, the nucleotide unit comprises an aromatic base, a 5-carbon sugar, and 1 to 10 phosphate groups. In some embodiments, the linker is attached to the nucleotide unit via the base. In some embodiments, multiple nucleotide arms attached to the core have the same type of nucleotide unit, and the type of nucleotide unit is selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP.In some embodiments, the plurality of multivalent molecules comprises one type of multivalent molecule, and each multivalent molecule of the plurality has the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP. In some embodiments, the plurality of multivalent molecules comprises a mixture of any combination of two or more types of multivalent molecules having nucleotide units selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP. One or more of the characteristics described in this paragraph may be present in any exemplary multivalent molecule in the various embodiments described herein. The characteristics described in this paragraph are referred to throughout this disclosure as "embodiments of multivalent molecules."
[0286] In some embodiments of the system, the nucleotide arms are designed so that the nucleotide units of the nucleotide arms can interact with a polymerase enzyme in a manner similar to that of free nucleotides. The nucleotide units of the nucleotide arms can bind to a polymerase complexed with a nucleic acid template and a nucleic acid primer (e.g., nucleotide conjugate). The nucleotide units can also dissociate from the polymerase complex and rebind to the same polymerase complex or bind to a different polymerase complex in close proximity to the multivalent molecule. Because the multivalent molecule contains multiple nucleotide arms, the nucleotide units of a single multivalent molecule can simultaneously bind to multiple polymerase complexes. The multivalent molecule effectively increases the local concentration of nucleotides, which can enhance the signal in the nucleotide conjugation reaction.
[0287] In some embodiments of the system, a nucleotide unit of the multivalent molecule can bind to the multiple polymerase without being incorporated, hi some embodiments, a complementary nucleotide unit of the multivalent molecule can bind to the multiple polymerase without undergoing polymerase-catalyzed incorporation, with the complementary nucleotide attached to the 3' end of the primer at a position opposite the complementary nucleotide in the template strand.
[0288] In some embodiments of the system, a nucleotide unit of a multivalent molecule can bind to a composite polymerase and undergo primer extension by incorporation into the 3' end of an extendable primer (e.g., complexed with a polymerase), resulting in primer extension. If the nucleotide unit contains a sugar 3' OH, it can incorporate a subsequent nucleotide into the nascent extended primer. If the nucleotide unit contains a sugar 3' OH substituted with a blocking group, it blocks the incorporation of a subsequent nucleotide into the nascent extended primer strand. The nucleotide unit (of the multivalent molecule) can be attached to the 3' end of the primer at a position opposite the complementary nucleotide in the template strand. The nucleotide unit can undergo nucleotide incorporation in a polymerase-catalyzed reaction, thereby extending the primer by one nucleotide.
[0289] In some embodiments of the system, the core units of the multivalent molecules can be labeled with a detectable reporter moiety (e.g., a fluorophore) in a manner that allows for differentiation between different multivalent molecules carrying different types of nucleotide units. For example, the core unit of a first multivalent molecule is labeled with a first fluorophore, and the first multivalent molecule comprises multiple nucleotide arms with dGTP nucleotide units. The core unit of a second multivalent molecule is labeled with a second fluorophore (different from the first fluorophore), and the second multivalent molecule comprises multi...
Claims
1. An engineered polymerase, having an amino acid sequence that is at least 85% identical to the amino acid sequence of SEQ ID NO: 1 and having amino acid substitutions D141A and E143A, wherein the engineered polymerase has an increased ability to incorporate chain-terminating nucleotide analogs as compared to the wild-type polymerase having the amino acid sequence of SEQ ID NO: 1, and said engineered polymerase comprising said amino acid sequence.
2. The engineered polymerase according to claim 1, wherein said amino acid substitution further comprises an additional substitution at one or more amino acid positions selected from L416, Y417, P418, and I529.
3. Said additional substitution is L416S, L416I, L416A, L416V, or L416G; Y417A, Y417T, Y417G, or Y417S; P418A, P418G, or P418S; and / or I529H, I529T, or I529L The engineered polymerase according to claim 2.
4. The engineered polymerase according to claim 1, wherein said amino acid substitution further comprises a first additional substitution at a first amino acid position L416 and a second additional substitution at a second amino acid position Y417.
5. The engineered polymerase according to claim 4, wherein said amino acid substitution further comprises a third additional substitution at a third amino acid position P418 and a fourth additional substitution at a fourth amino acid position I529.
6. The engineered polymerase according to claim 1, wherein the amino acid substitution further comprises one or more additional substitutions at one or more amino acid positions selected from amino acid positions Y10, E14, H48, D75, C104, A115, C130, C269, D305, E328, L416, Y417, P418, A493, S500, M501, R515, I529, N567, A568, E569, S577, K610, N657, E700, N705, S717, K762, G763, L764, G765, K766, Q767, and M768.
7. The engineered polymerase according to claim 1, wherein the amino acid sequence is at least 90% identical to one of the amino acid sequences selected from SEQ ID NOs: 12, 14, 18, 20, 27, 39, 66, 76, 85, 86, 114, 117, 118, 124, 128, 129, 130, 138, 163, 164, 189, 194, 207, 225, 233, 235, 289, 291, 310, 319, 323, 333, 353, 346, 361, 364, 362, 366, 392, and 393.
8. The engineered polymerase according to claim 1, wherein the amino acid substitution further comprises additional substitutions within the palm domain specified by the amino acid sequence of SEQ ID NO: 271 or SEQ ID NO:
273.
9. The engineered polymerase according to claim 1, wherein the amino acid substitution removes the 3' - 5' exonuclease activity of the engineered polymerase.
10. The engineered polymerase shows increased thermal stability as compared to the wild - type polymerase having the amino acid sequence of SEQ ID NO: 1, shows increased incorporation rate as compared to the wild - type polymerase having the amino acid sequence of SEQ ID NO: 1, shows increased incorporation rate of nucleotide analogs containing a chain - terminating moiety at the 2' or 3' position of the sugar moiety as compared to the wild - type polymerase having the amino acid sequence of SEQ ID NO: 1, and / or showing increased uracil resistance as compared to said wild-type polymerase having the amino acid sequence of SEQ ID NO: 1 The engineered polymerase according to claim 1.
11. The engineered polymerase according to claim 1, wherein said increased average incorporation rate is at least 5-fold the average incorporation rate of said wild-type polymerase having the amino acid sequence of SEQ ID NO:
1.
12. A composition comprising: one or more of the engineered polymerases according to claim 1; one or more nucleic acid template molecules; and one or more molecules comprising a nucleotide polymerization initiation site having a 3'-extendable end.
13. The composition according to claim 12, wherein said one or more nucleic acid template molecules are linear nucleic acid molecules or circular nucleic acid molecules.
14. Said one or more nucleic acid template molecules are clonally amplified template molecules; and / or a concatemer having one copy of a target sequence of interest or two or more tandem copies of a target sequence of interest The composition according to claim 12.
15. The composition according to claim 12, wherein said one or more molecules comprising a nucleotide polymerization initiation site comprise nucleic acid primers that hybridize to a portion of said nucleic acid template molecule, or wherein said one or more molecules comprising a nucleotide polymerization initiation site comprise self-priming end portions of said nucleic acid template molecule.
16. Said one or more engineered polymerases, said one or more nucleic acid template molecules, and said one or more nucleotide polymerization initiation sites form one or more complex polymerases, each complex polymerase The composition according to claim 12, comprising the engineered polymerase bound to a nucleic acid double strand, wherein the double strand comprises one of the nucleic acid template molecules that hybridizes to a nucleic acid primer.
17. The one or more complex polymerases further comprise a multivalent molecule, and the multivalent molecule (a) a core, and (b) a plurality of nucleotide arms, each of which (i) a core-binding moiety, (ii) a spacer, (iii) a linker, and (iv) a nucleotide unit, and the plurality of nucleotide arms, the composition according to claim 12.
18. The core binds to each of the nucleotide arms via the core-binding moiety, the core-binding moiety binds to the spacer, the spacer binds to the linker, and the linker binds to the nucleotide unit, the composition according to claim 17.
19. The plurality of nucleotide arms bind to the core via the core-binding moiety, each of the nucleotide arms bound to the core has the same type of nucleotide unit, and the nucleotide unit comprises dATP, dGTP, dCTP, dTTP, or dUTP, the composition according to claim 17.
20. The one or more multivalent molecules comprise one type of multivalent molecule, and each type of the multivalent molecule has the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP, or comprise two or more types of multivalent molecules, and each type of the multivalent molecule has the same type of nucleotide unit selected from the group consisting of dATP, dGTP, dCTP, dTTP, and dUTP, the composition according to claim 17.
21. The composition according to claim 16, wherein the one or more composite polymerases further comprise a plurality of nucleotides, and the nucleotides among the plurality of nucleotides comprise aromatic bases, pentose sugars, and 1 to 10 phosphate groups.
22. The composition according to claim 21, wherein the plurality of nucleotides comprise one or more types of nucleotides selected from the group consisting of dATP, dGTP, dCTP, dTTP, and / or dUTP.
23. The composition according to claim 21, wherein at least one of the nucleotides among the plurality of nucleotides comprises a removable chain-terminating moiety attached to the 3'-carbon position of the sugar group, and the removable chain-terminating moiety comprises an alkyl group, an alkenyl group, an alkynyl group, an allyl group, an aryl group, a benzyl group, an azide group, an azido group, an O-azidomethyl group, an amine group, an amide group, a keto group, an isocyanate group, a phosphate group, a thio group, a disulfide group, a carbonate group, a urea group, or a silyl group, and the removable chain-terminating moiety is cleavable using a chemical compound to generate an extendable 3'-OH moiety on the sugar group.
24. The one or more composite polymerases further comprise a plurality of non-catalytic divalent cations that inhibit nucleotide incorporation by polymerase catalysis, and the non-catalytic divalent cations comprise strontium or barium, or further comprise a plurality of catalytic divalent cations that promote nucleotide incorporation by polymerase catalysis, and the catalytic divalent cations comprise magnesium or manganese, The composition according to claim 16.
25. The one or more composite polymerases further comprise first and second binding complexes, (1) The first binding complex binds to a first nucleic acid primer, a first engineered polymerase, and a first portion of a concatemer template molecule, thereby forming the first binding complex and including a first multivalent molecule, wherein a first nucleotide unit of the multivalent molecule binds to the first engineered polymerase, (2) The second binding complex binds to a second nucleic acid primer, a second engineered polymerase, and a second portion of the same concatemer template molecule, thereby forming the second binding complex and including the first multivalent molecule, wherein a second nucleotide unit of the first multivalent molecule binds to the second engineered polymerase, The composition according to claim 16, wherein the first and second binding complexes include the same first multivalent molecule and form an affinity complex.
26. A method for forming one or more composite polymerases, comprising: contacting one or more engineered polymerases with (i) one or more nucleic acid template molecules and (ii) one or more nucleic acid primers to form the one or more composite polymerases, wherein at least one of the composite polymerases includes an engineered polymerase bound to a nucleic acid double strand, and the nucleic acid double strand includes a nucleic acid template molecule that hybridizes to a nucleic acid primer, The method, wherein the one or more engineered polymerases include an amino acid sequence that is at least 85% identical to the amino acid sequence of SEQ ID NO: 1 and has substitutions Asp141Ala and Glu143Ala.
27. A method for performing nucleic acid sequencing, comprising: (a) contacting a first set of engineered polymerases with (i) a plurality of nucleic acid template molecules and (ii) a plurality of nucleic acid primers, wherein the contacting is performed under conditions suitable for the engineered polymerase to bind to the nucleic acid template molecule and the nucleic acid primer, thereby forming a first set of composite polymerases, and each of the first set of composite polymerases includes the engineered polymerase bound to a nucleic acid double strand, the nucleic acid double strand includes one of the nucleic acid template molecules that hybridizes to one of the nucleic acid primers, the contacting, wherein the engineered polymerase comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of SEQ ID NO: 1 and has amino acid substitutions Asp141Ala and Glu143Ala; (b) contacting the first set of composite polymerases with a plurality of multivalent molecules to form a first set of multivalent binding complexes, each multivalent molecule of the plurality of multivalent molecules includes a core that binds to a plurality of nucleotide arms, each nucleotide arm binds to a nucleotide unit, the contacting is binding complementary nucleotide units of the multivalent molecule to at least two of the first set of composite polymerases, thereby forming a first set of multivalent binding complexes; the contacting, which is performed under suitable conditions to inhibit the incorporation of complementary nucleotides of each multivalent molecule into the nucleic acid primers of the first set of multivalent binding complexes; (c) detecting the first set of multivalent binding complexes; (d) identifying the nucleotide bases of the complementary nucleotides within the first set of multivalent binding complexes, thereby determining the sequence of the nucleic acid template molecule.