Engineered DNA polymerases that reduce artifact formation
By fusing thioredoxin binding domain (TBD) and thioredoxin (TRX) to DNA polymerases, the engineered enzymes effectively reduce stutter artifacts in PCR, enhancing the accuracy of genetic analysis.
Patent Information
- Application Number
- JP2025551150
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-03
- Filing Date
- 2024-03-04
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional PCR methods result in stutter artifacts due to strand slippage, complicating the analysis of short tandem repeat (STR) profiles and microsatellite instability, which affects the accuracy and reliability of genetic analysis.
Engineering DNA polymerases by fusing or conjugating a thioredoxin binding domain (TBD) and thioredoxin (TRX) to the DNA polymerase domain to reduce stutter artifacts.
The engineered DNA polymerases exhibit a significant reduction in stutter tendency, improving the accuracy and reliability of genetic analysis by minimizing stutter artifacts.
Smart Images

Figure 2026507239000061 
Figure 2026507239000062 
Figure 2026507239000063
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 488,035, filed March 2, 2023, and U.S. Provisional Patent Application No. 63 / 488,416, filed March 3, 2023, both of which are incorporated herein by reference.
[0002] Sequence Listing The text of the computer-readable sequence listing submitted herewith (entitled "PRMG_41353_601_SequenceListing.xml", created March 4, 2024, file size 353,296 bytes) is incorporated herein by reference in its entirety.
[0003] Provided herein are compositions and systems comprising a DNA polymerase domain, a thioredoxin binding domain (TBD), and thioredoxin (TRX), wherein one or both of the TRX and TBD are fused or otherwise conjugated to the DNA polymerase domain. The TRX or TBD may also be provided as separate entities (e.g., binary systems) in the systems herein. The DNA polymerase / TBD / TRX compositions and systems herein are engineered to reduce stutter and / or produce fewer stutter artifacts. Kits comprising the DNA polymerase / TBD / TRX compositions and systems herein, and methods for using the same, are also within the scope of the present specification. [Background technology]
[0004] Microsatellites, or short tandem repeats (STRs), consist of tandemly repeated DNA sequence motifs 1–8 nucleotides in length. They are widely distributed and abundant in eukaryotic genomes and are often highly polymorphic due to variation in the number of repeat units.
[0005] Forensic short tandem repeat (STR) profiling relies on accurately determining the number of repetitive DNA sequences at a given genomic locus, with each repeat unit typically consisting of 3–6 base pairs. Traditional polymerase chain reaction (PCR) methods result in a collection of amplicons, including products with erroneous insertions or deletions of repetitive sequences in a phenomenon known as strand slippage or "stutter." These stutter products can complicate the analysis of STR profiles and potentially mask the contribution of trace DNA in STR profiles derived from multiple individuals.
[0006] Microsatellite instability (MSI) is a well-established biomarker that often indicates susceptibility to cancer development and can be found in a wide range of solid tumors. MSI provides genetic evidence of impaired DNA mismatch repair, known to be one of the most frequently mutated gene sets in cancer. MSI may also predict Lynch syndrome. MSI results in the addition or deletion of nucleotides during DNA replication, which are then inherited by daughter cells. Mononucleotide repeats are particularly susceptible to these types of MSI-induced errors. While these aberrant insertions or deletions can be detected by PCR-based assays, stutter artifacts (particularly problematic when amplifying mononucleotide repeat sequences) significantly compromise the sensitivity of such tests.
[0007] Stutter signals differ from PCR products representing genomic alleles by a multiple of the size of the repeat unit. For dinucleotide repeat loci, the dominant stutter signal is generally two bases shorter than the genomic allele signal, accompanied by additional side products that are four and six bases shorter. The multiple signal patterns observed for each allele complicate interpretation, particularly when the two alleles from an individual are close in size (e.g., in medical and genetic mapping applications) or when the DNA sample contains a mixture from two or more individuals (e.g., in forensic applications). Such confusion is greatest in mononucleotide microsatellite genotyping, where both the genomic and stutter fragments exhibit single-nucleotide spacing.
[0008] There is a need in the art to develop PCR reaction conditions that minimize or eliminate stutter so that genetic analysis can be more accurate and reliable. Summary of the Invention
[0009] Provided herein are compositions and systems comprising a DNA polymerase domain, a thioredoxin binding domain (TBD), and thioredoxin (TRX), wherein one or both of the TRX and TBD are fused or otherwise conjugated to the DNA polymerase domain. The TRX or TBD may also be provided as separate entities (e.g., binary systems) in the systems herein. The DNA polymerase / TBD / TRX compositions and systems herein are engineered to reduce stutter and produce fewer stutter artifacts. Kits comprising the DNA polymerase / TBD / TRX compositions and systems herein, and methods for using the same, are also within the scope of the present specification. In certain embodiments, the present disclosure provides (1) a chimera of a DNA polymerase, a thioredoxin binding domain, and thioredoxin; (2) a chimera of 0.1 to 2000 (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 60 Provided are (1) a chimera of a DNA polymerase and a thioredoxin binding domain in the presence of TRX at a TRX:TBD ratio of 0, 700, 800, 900, 1000, 1500, 2000, or a range therebetween (e.g., 0.1 to 800, 0.6 to 600, etc.), and / or (2) a chimera of a thermostable DNA polymerase and a thioredoxin binding domain in the presence of a thioredoxin binding domain.
[0010] In some embodiments, provided herein are DNA polymerase systems comprising (a) a DNA polymerase domain, (b) a thioredoxin binding domain (TBD), and (c) a thioredoxin (TRX) domain. In some embodiments, the DNA polymerase system is capable of synthesizing a DNA product from deoxynucleotide triphosphates in the presence of template DNA under appropriate reaction conditions. In some embodiments, the DNA polymerase system exhibits a reduced tendency to stutter compared to a DNA polymerase comprising the DNA polymerase domain in the absence of the TBD and / or TRX. In some embodiments, the DNA polymerase system exhibits at least a 10% (e.g., 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, 99%) reduction in stutter tendency compared to a DNA polymerase comprising the DNA polymerase domain in the absence of the TBD and / or TRX. In some embodiments, the DNA polymerase system exhibits at least 10% (e.g., 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, 99%) fewer stutter artifacts compared to a DNA polymerase comprising a DNA polymerase domain in the absence of a TBD and / or a TRX. In some embodiments, the DNA polymerase system comprises a DNA polymerase domain conjugated to a TBD and / or a TRX. In some embodiments, the DNA polymerase system comprises a DNA polymerase domain genetically fused to one or both of a TBD and / or a TRX. In some embodiments, the DNA polymerase system comprises a genetic fusion of a DNA polymerase domain, a TBD, and a TRX. In some embodiments, one of the TBD and the TRX is not conjugated to other components of the system. In some embodiments, the system comprises a free TBD and a DNA polymerase domain conjugated or genetically fused to a TBD. In some embodiments, the system comprises a free TBD and a DNA polymerase domain conjugated or genetically fused to a TRX.In some embodiments, the system comprises a DNA polymerase domain, a TRX, and a TBD conjugated or genetically fused to each other.
[0011] In some embodiments, provided herein are chimeric DNA polymerases with reduced stutter propensity (e.g., by at least 10% (e.g., 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, 99%)), the chimeric DNA polymerase comprising a genetic fusion of (a) a DNA polymerase domain, (b) a thioredoxin binding domain (TBD), and (c) a thioredoxin (TRX) domain. In some embodiments, the DNA polymerase domain is thermophilic. In some embodiments, the DNA polymerase domain is derived from a naturally occurring thermophilic DNA polymerase. In some embodiments, the naturally occurring thermophilic DNA polymerase is selected from the group consisting of Thermus aquaticus DNA polymerase, Thermus thermophilus DNA polymerase, Thermus flavus DNA polymerase, Thermotoga neapolitana polymerase, and Geobacillus stearothermophilus DNA polymerase. In some embodiments, the DNA polymerase domain is derived from a Family A DNA polymerase. In some embodiments, the DNA polymerase domain comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 1. In some embodiments, the DNA polymerase domain further comprises an internal amino acid sequence insertion.In some embodiments, the DNA polymerase domain comprises an N-terminal portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 14 and a C-terminal portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 12, wherein the N-terminal portion and the C-terminal portion are separated by an internal amino acid sequence insert. In some embodiments, the internal amino acid sequence insert comprises a TBD. In some embodiments, the TBD is derived from the thioredoxin binding domain of T3 or T7 bacteriophage DNA polymerase. In some embodiments, the TBD comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 15. In some embodiments, the TBD is derived from the thioredoxin binding domain of a Klebsiella pneumoniae, Salmonella enterica, or Aeromonas hydrophila phage DNA polymerase. In some embodiments, the TBD comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 101-103. In some embodiments, the TBD sequence is internal to the DNA polymerase domain sequence. In some embodiments, the TRX domain is derived from Escherichia coli thioredoxin. In some embodiments, the TRX domain constitutes at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to SEQ ID NO: 16, 17, or 107.In some embodiments, the TRX domain is derived from Alishewanella jeotgali or Thiococcus pfennigii thioredoxin. In some embodiments, the TRX domain comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 93 or 94. In some embodiments, the TRX domain comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 51-53. In some embodiments, the TRX sequence is fused to the N-terminus or C-terminus of the DNA polymerase domain. In some embodiments, the TRX sequence is fused to the DNA polymerase domain by a linker of 1-300 amino acids (e.g., 1, 2, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, or more, or any range or length therebetween). In some embodiments, the linker is a flexible linker. In some embodiments, 50-100% (e.g., 50%, 60%, 70%, 80%, 90%, 100%, or any range therebetween) of the linker are glycine and serine residues. For example, the linker can include one or more repeating GS units, one or more repeating GSAT units, etc. In some embodiments, the linker includes a rigid linker segment. In some embodiments, the rigid segment includes one or more EAAAK peptide segments. In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 22-49.In some embodiments, the linker comprises the sequence GS(24)-CASSIDYKRISRMPSKIMDAVIDTLNICKLANCE-GS(24), GS(24)-CASSIDYKRISRMPAVLADAVIDTLNICKLANCE-GS(24), etc., from Table 13.
[0012] In some embodiments, provided herein are compositions comprising a DNA polymerase domain conjugated to a thioredoxin binding domain (TBD). In some embodiments, the DNA polymerase domain is genetically fused to the thioredoxin binding domain (TBD). In some embodiments, the DNA polymerase domain is derived from a Family A DNA polymerase (e.g., Taq polymerase, Tne polymerase, etc.). In some embodiments, the DNA polymerase domain is thermophilic. In some embodiments, the DNA polymerase domain is derived from a naturally occurring thermophilic DNA polymerase. In some embodiments, the naturally occurring thermophilic DNA polymerase is selected from the group consisting of Thermus aquaticus DNA polymerase, Thermus thermophilus DNA polymerase, Thermus flavus DNA polymerase, Thermotoga neapolitana polymerase, and Geobacillus stearothermophilus DNA polymerase. In some embodiments, the DNA polymerase domain comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 1. In some embodiments, the DNA polymerase domain comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to one or more of SEQ ID NOs: 2-12 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 in any combination in any order). In some embodiments, the DNA polymerase domain comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity with SEQ ID NOs: 14 and 12.In some embodiments, the DNA polymerase domain includes a portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 14, excluding SEQ ID NO: 13. In some embodiments, the TBD sequence is internal to the DNA polymerase domain sequence. In some embodiments, the DNA polymerase domain comprises an N-terminal portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 14 and a C-terminal portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 12, wherein the N-terminal portion and the C-terminal portion are separated by a TBD. In some embodiments, the TBD is derived from the thioredoxin binding domain of T3 or T7 bacteriophage DNA polymerase. In some embodiments, the TBD constitutes at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 15. In some embodiments, the TBD is derived from the thioredoxin binding domain of Klebsiella pneumoniae, Salmonella enterica, or Aeromonas hydrophila phage DNA polymerase. In some embodiments, the TBD constitutes at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 101-103.In some embodiments, the composition further comprises thioredoxin (TRX), wherein the thioredoxin is present in the composition at an 800 molar excess or less relative to TBD (e.g., 5x, 10x, 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, 100x, 150x, 200x, 250x, 300x, 400x, 500x, 600x, 700x, 800x, or ranges therebetween). In some embodiments, TRX is derived from Escherichia coli thioredoxin. In some embodiments, TRX shares at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 16, 17, or 107. In some embodiments, TRX is derived from Thiococcus pfennigii thioredoxin. In some embodiments, TRX shares at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity with SEQ ID NO:93. In some embodiments, TRX is derived from Alishewanella jeotgali thioredoxin. In some embodiments, TRX shares at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity with SEQ ID NO:94. In some embodiments, TRX comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity with SEQ ID NOs: 51-53.
[0013] In some embodiments, TRX is a fusion with an additional polypeptide sequence. In some embodiments, the additional polypeptide sequence is a DNA-binding protein, an amino acid sequence capable of binding to DNA, a protein associated with a DNA replication site, a TBD, and / or a DNA polymerase. In some embodiments, the additional polypeptide sequence is fused to TRX by a linker peptide or polypeptide. In some embodiments, the linker peptide or polypeptide is 1-300 amino acids in length (e.g., 1, 2, 5, 10, 20, 50, 100, 150, 200, 250, 300, or a range or value therebetween). In some embodiments, any linker described herein may be found to be used in such embodiments.
[0014] In some embodiments, thioredoxin is present in the composition at a TRX:TBD ratio of 0.1 to 2000 (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, or ranges therebetween (e.g., 0.1-800, 0.6-600)). In some embodiments, fusion proteins are provided that have at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to one of SEQ ID NOs: 28-34.
[0015] In some embodiments, as used herein, a segment comprises a DNA polymerase domain corresponding to SEQ ID NO: 1 and has (a) at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to SEQ ID NOs: 2, 4, 6, 8, 10, and 12; (b) (i) at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to SEQ ID NOs: 3, 5, 7, 9, and 11; , 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween), or (ii) a segment in which all or a portion of the sequence in SEQ ID NO: 1 corresponding to one or more of SEQ ID NOs: 3, 5, 7, 9, and 11 is replaced with a heterologous sequence selected from a TBD, a TRX, and a TIS (TBD / TRX interacting sequence). In some embodiments, the DNA polymerases herein comprise a TBD with at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 15. In some embodiments, the TBD is located at the C-terminus, the N-terminus, inserted within one of SEQ ID NOs: 3, 5, 7, 9, and 11, and / or replaces all or a portion of one of SEQ ID NOs: 3, 5, 7, 9, and 11. In some embodiments, the DNA polymerase herein comprises a TRX having at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 16, 17, or 107. In some embodiments, the TRX is located at the C-terminus, the N-terminus, inserted within one of SEQ ID NOs: 3, 5, 7, 9, and 11, and / or replaces all or a portion of one of SEQ ID NOs: 3, 5, 7, 9, and 11.In some embodiments, the DNA polymerases herein comprise a TIS having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 18-21. In some embodiments, the TIS is located at the C-terminus, the N-terminus, inserted within one of SEQ ID NOs: 3, 5, 7, 9, and 11, and / or replaces all or a portion of one of SEQ ID NOs: 3, 5, 7, 9, and 11. In some embodiments, the exonuclease domain of SEQ ID NO: 13 is removed from the sequence corresponding to SEQ ID NO: 1.
[0016] In some embodiments, provided herein are DNA polymerases comprising a sequence having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of the following: (a) (SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12), (b) (SEQ ID NO: 2)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (c) (SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (d) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(one of SEQ ID NOs:18-21)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12); (e) (SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (f) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(SEQ ID NO:7)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107), (g) (SEQ ID NO:2)-(one of SEQ ID NOs:18-21)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(SEQ ID NO:7)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107); (h) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(one of SEQ ID NOs:18-21)-(SEQ ID NO:6)-(SEQ ID NO:7)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107); (i) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(one of SEQ ID NOs:18-21)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107); (j) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(SEQ ID NO:7)-(SEQ ID NO:8)-(one of SEQ ID NOs:18-21)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107); (k) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (l) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (m) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (n) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12), and (o) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12).
[0017] In some embodiments, provided herein are DNA polymerases comprising a sequence having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to one of SEQ ID NOs: 22-49.
[0018] In some embodiments, provided herein is a DNA polymerase comprising: (a) one DNA polymerase domain, one TBD, and one TRX; (b) one DNA polymerase domain, one TBD, and two or more TRXs; (c) one DNA polymerase domain, two or more TBDs, and one TRX; (d) one exonuclease-deficient DNA polymerase domain, one TBD, and one TRX; (e) one exonuclease-deficient DNA polymerase domain, one TBD, and two or more TRXs; (f) one exonuclease-deficient DNA polymerase domain, two or more TBDs, and one TRX; (g) one DNA polymerase domain, one TBD, one TRX, and one TIS; (h) one DNA polymerase domain, one TBD, two or more TRXs, and one TIS; (i) one DNA polymerase domain, two or more TBDs, one TRX, and one TIS; (j) one exonuclease-deficient DNA polymerase domain, one TBD, one TRX, and one TIS; (k) one exonuclease-deficient DNA polymerase domain, one TBD, two or more TRXs, and one TIS; or (l) an exonuclease-deficient DNA polymerase domain, two or more TBDs, a TRX, and a TIS. In some embodiments, the DNA polymerase domain has at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 1, or an ordered combination of eight or more (e.g., 8, 9, 10, 11) of SEQ ID NOs: 2-12. In some embodiments, the TBD has at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 15. In some embodiments, the TRX has at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 16, 17, or 107. In some embodiments, the TIS has at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 18-21. In some embodiments, the exonuclease-deficient DNA polymerase domain lacks all or a portion of SEQ ID NO: 13.
[0019] In some embodiments, provided herein are reaction mixtures comprising a composition, a DNA polymerase, or a DNA polymerase system, and sufficient amplification reagents to amplify a DNA target sequence. In some embodiments, the amplification reagents include one or more of oligonucleotide primers, deoxynucleotide triphosphates, magnesium chloride, a buffer, water, and a template DNA comprising the DNA target sequence. In some embodiments, the DNA target sequence comprises one or more short tandem repeats (STRs). In some embodiments, the STRs comprise repeating units of 1-50 nucleotides ranging in length from 10-500 nucleotides. In some embodiments, the reaction mixture further comprises a reducing agent. In some embodiments, the reducing agent is a thiol reducing agent or a non-thiol reducing agent. In some embodiments, the reducing agent is dithiothreitol (DTT) or tris(2-carboxyethyl)phosphine (TCEP).
[0020] In some embodiments, provided herein are methods for amplifying a DNA target sequence, comprising exposing a reaction mixture to polymerase chain reaction temperature cycling conditions described herein.
[0021] In some embodiments, provided herein are thioredoxin (TRX) polypeptides that share at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to positions 29-37, 60-77, and 89-98 of SEQ ID NO: 16, and the TRX polypeptides are capable of binding to a TRX binding domain (TBD) having the amino acid sequence of SEQ ID NO: 15. In some embodiments, the TRX polypeptides share at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to positions 29-37, 60-77, and 89-98 of SEQ ID NO: 16. In some embodiments, the TRX polypeptides share 100% sequence identity to positions 29-37, 60-77, and 89-98 of SEQ ID NO: 16. In some embodiments, the TRX polypeptide comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 16. In some embodiments, the TRX polypeptide comprises 50-60% sequence identity to SEQ ID NO: 16. In some embodiments, the TRX polypeptide is 100-120 amino acids in length. In some embodiments, the TRX polypeptide has a 3D fold threshold of 0.8 or greater (e.g., 0.8, 0.85, 0.90, 0.95, 1.0, or ranges therebetween) relative to TRX in protein database model 6N7W. In some embodiments, the TRX polypeptide has an instability score of less than 40 (e.g., <35, <30, <25, <20, etc.).
[0022] In some embodiments, provided herein are thioredoxin (TRX) polypeptides capable of binding to a TRX binding domain (TBD) and having (i) a 3D fold threshold of 0.8 or greater (e.g., 0.8, 0.85, 0.90, 0.95, 1.0, or a range therebetween) relative to TRX in protein database model 6N7W, and / or (ii) an instability score of less than 40 (e.g., <35, <30, <25, <20, etc.). In some embodiments, the TRX polypeptide comprises 100% sequence similarity to positions 29-37, 60-77, and 89-98 of SEQ ID NO: 16. In some embodiments, the TRX polypeptide is 100-120 amino acids in length. In some embodiments, the TRX polypeptide is longer than 120 amino acids in length (e.g., 125, 130, 140, 150, 175, 200, 250, 300, 400, 500, or more).
[0023] In some embodiments, provided herein is a thioredoxin (TRX) polypeptide capable of binding to a TRX binding domain (TBD), wherein the root mean square deviation (RMSD) calculated for the alpha carbons of at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) of the amino acid residues corresponding to amino acids 29-37, 60-77, and 89-98 of SEQ ID NO: 16, for a 3D molecular structure (e.g., calculated by ESMFold), relative to TRX in protein database model 6N7W is 3.0 Å or less (e.g., 3.0 Å, 2.8 Å, 2.6 Å, 2.4 Å, 2.2 Å, 2.0 Å, 1.8 Å, 1.6 Å, 1.4 Å, 1.2 Å, 1.0 Å, 0.8 Å, 0.6 Å, 0.4 Å, 0.2 Å, or less, or ranges or values therebetween). In some embodiments, the alpha carbon RMSD of the TBD interacting residues of TRX is 3.0 Å or less to protein database model 6N7W. In some embodiments, the TBD interacting residues of TRX have at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence similarity to SEQ ID NO: 16. In some embodiments, the TBD interacting residues of TRX have at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 16.
[0024] In some embodiments, provided herein are DNA polymerase systems that include: (a) a DNA polymerase domain that constitutes at least 40% sequence identity to a Family A DNA polymerase; (b) a thioredoxin binding domain (TBD) that has at least 50% sequence identity to a naturally occurring phage-derived TBD; and (c) a thioredoxin (TRX) domain (e.g., a TRX described herein) that is capable of binding to the TBD.
[0025] In some embodiments, provided herein are DNA polymerase systems comprising: (a) a first polypeptide comprising (i) a DNA polymerase domain, (ii) a thioredoxin binding domain (TBD), and (iii) a thioredoxin (TRX) domain; and (b) a second polypeptide comprising (i) a DNA polymerase domain, and (ii) a TBD. [Brief explanation of the drawings]
[0026] [Figure 1A] Example electropherogram of a PowerPlex® Fusion STR multiplex amplified with standard Taq. [Figure 1B] Example electropherogram of a PowerPlex® Fusion STR multiplex amplified with Taq-TBD. [Figure 1C] Example electropherogram of a PowerPlex® Fusion STR multiplex amplified with Taq-TBD with 80x molar excess of TRX. [Figure 1D] Representative loci illustrating stutter reduction between standard Taq (top) and Taq-TBD with TRX (bottom). Stutter peaks are indicated by orange arrows. [Figure 1E] Comparison of mean stutter frequencies at the indicated loci between Taq and Taq-TBD in the presence of 80x molar excess of TRX. [Figure 1F] Quantification of average peak heights using Taq-TBD in combination with the indicated amounts of TRX. [Figure 1G] Example electropherograms illustrating the D2S1338 locus amplified with standard Taq or with Taq-TBD in combination with the indicated molar excess of TRX. The blue arrows point to the obvious stutter peaks. Note that the peak height decreases as the molar ratio of TRX decreases, as indicated by the y-axis scale. Reducing TRX reduces signal amplitude, so we focused on well-separated single loci. [Figure 1H] Quantification of stutter frequency for the D2S1338 locus for standard Taq or for Taq-TBD used in combination with the indicated molar excess of TRX. Reducing TRX reduces signal amplitude, so we focused on well-separated single loci. [Figure 2] (A) Example electropherogram demonstrating the effect of DTT concentration ([DTT]) on a PowerPlex® Fusion STR multiplex amplified with the Taq-TBD plus TRX system. As [DTT] was lowered, amplification efficiency decreased, particularly for larger loci. In addition, stuttering increased as [DTT] was lowered; this effect was seen for the smallest locus in the indicated dye channel (D8S1179, blue arrow). (B) Quantification of stuttering observed at various concentrations of DTT. [Figure 3] Average peak amplitude (A) and stutter percentage (B) of PowerPlex® Fusion STR multiplexes amplified with the indicated polymerases. Taq-TBD reactions were amplified in the presence of modified thioredoxin (C36S) with or without a reducing agent (DTT). Significant reduction in stutter with the use of chimeric polymerases and modified thioredoxin was only observed with modified TRX in the presence of a reducing agent. [Figure 4]Average amplicon peak heights (A) and stutter percentages (B) for constructs with thioredoxin terminally linked. Thioredoxin (TRX) was linked via gene fusion to the C-terminus (Taq-TBD-TRX), N-terminus (TRX-Taq-TBD), or both termini (TRX-Taq-TBD-TRX) of the Taq-TBD construct. These Taq-TBD / TRX chimeras were used to amplify PowerPlex® Fusion STR multiplexes, and the resulting amplification products were analyzed by capillary electrophoresis. Compared to Taq as a control, the chimeric constructs were able to amplify STR multiplexes without the need for exogenous TRX, and importantly, all constructs still demonstrated significantly reduced stutter. (C) Example electropherograms of PowerPlex® Fusion STR multiplexes amplified with the TRX-Taq-TBD enzyme. [Figure 5] Example electropherograms of the MONO-27 locus amplified using control Taq (top) or N-terminally linked TRX-Taq-TBD constructs (bottom) from DNA obtained from matched normal (left) and tumor (center) tissue samples. In the TRX-Taq-TBD amplified tumor sample, a novel tumor allele diagnostic of frequent MSI is clearly visible as a new local maximum (red arrow). [Figure 6A]A variation of the Taq-TBD construct was generated. This is an exonuclease domain deletion variant. The stuttering properties of this construct were investigated by amplifying a Promega PowerPlex® Fusion multiplex and analyzing it by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentage was determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the example shown, significant stutter reduction was demonstrated compared to a matched control condition (i.e., a multiplex amplified with standard Taq). [Figure 6B] A variation of the Taq-TBD construct was generated. This is a cysteine-mutated variant. The stuttering properties of this construct were investigated by amplifying a Promega PowerPlex® Fusion multiplex and analyzing it by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentage was determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the example shown, significant stutter reduction was demonstrated compared to a matched control condition (i.e., a multiplex amplified with standard Taq). [Figure 6C]A variation of the TRX-Taq-TBD construct was generated. This is a point mutation (I677T) variant that converts the T3 TBD sequence to a T7 sequence. The stuttering properties of this construct were investigated by amplifying a Promega PowerPlex® Fusion multiplex and analyzing it by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the example shown, significant stutter reduction was demonstrated compared to a matched control condition (i.e., a multiplex amplified with standard Taq). [Figure 6D] Variations of the TRX-Taq-TBD construct were generated, featuring linker variants of varying lengths used to connect TRX to the Taq-TBD construct at the N-terminus. The stutter properties of this construct were investigated by amplifying a Promega PowerPlex® Fusion multiplex and analyzing it by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the examples shown, significant stutter reduction was demonstrated compared to matched control conditions (i.e., multiplexes amplified with standard Taq). [Figure 6E]Variations of the TRX-Taq-TBD construct were generated, featuring linker variants of varying lengths used to connect TRX to the Taq-TBD construct at the C-terminus. The stuttering properties of this construct were investigated by amplifying a Promega PowerPlex® Fusion multiplex and analyzing it by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the examples shown, significant stutter reduction was demonstrated compared to matched control conditions (i.e., multiplexes amplified with standard Taq). # indicates that dropout (i.e., no allele peak) was observed at the indicated locus for that condition. [Figure 6F] Variations of the TRX-Taq-TBD construct were generated, consisting of one to three tandemly repeated TRX motifs linked to the N-terminus of Taq-TBD. The stuttering properties of these constructs were investigated by amplifying Promega PowerPlex® Fusion multiplexes and analyzing them by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the examples shown, significant stutter reduction was demonstrated compared to matched control conditions (i.e., multiplexes amplified with standard Taq). [Figure 6G]Variations of the TRX-Taq-TBD construct were generated. These were variants of the tandem TBD motif inserted into the TRX-Taq-TBD construct with and without exogenous TRX supplementation (approximately 160x molar ratio). The stuttering properties of this construct were investigated by amplifying a Promega PowerPlex® Fusion multiplex and analyzing by capillary electrophoresis. Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only for loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). In the examples shown, significant stutter reduction was demonstrated compared to matched control conditions (i.e., multiplexes amplified with standard Taq). [Figure 7] Example of STR data analyzed by next-generation sequencing (NGS). (A) Example histogram showing STR allele calls at loci D18S51, D21S11, D22S1045, and DYS481 for a sample amplified with standard Taq enzyme. Gray bars represent unknown alleles (Unk) or sequences below the filter setting (minimum 10 reads or 1.5% of the total reads at the locus). Blue bars indicate allele calls. Orange arrows indicate stutter alleles. (B) Example histogram showing STR allele calls at loci D18S51, D21S11, D22S1045, and DYS481 for a sample amplified with Taq-TBD plus TRX. Gray bars represent unknown alleles (Unk) or sequences below the filter setting (minimum 10 reads or 1.5% of the total reads at the locus). Blue bars indicate allele calls. Orange arrows point to stutter alleles. (C) Bar graphs representing stutter at each locus as a percentage of the associated main peak. [Figure 8]The TBD-interacting sequence (TIS, see SEQ ID NO: 20) was inserted into the Trx-Taq-TBD construct (TRX-Taq-TIS-TBD, see SEQ ID NO: 48) and examined by amplifying control DNA on a Promega PowerPlex® Fusion multiplex and analyzing by capillary electrophoresis. Amplification with Taq was also performed in parallel. The stutter artifact percentage was quantified by dividing the amplitude of the stutter allele peak by the amplitude of the corresponding allele peak. Stutter artifacts were determined only at loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap). (A) Overall polymerase efficiency, as measured by the average peak height across all loci, was similar for TRX-Taq-TBD and TRX-Taq-TIS-TBD. (B) However, select loci that were poorly amplified with TRX-Taq-TBD appeared to be better amplified with the TRX-Taq-TIS-TBD construct, as measured by locus mean peak height. (C) Importantly, the formation of stutter artifacts was similarly reduced with either TRX-Taq-TBD or TRX-Taq-TIS-TBD compared to control Taq amplification. [Figure 9A] Development of a medium-throughput lysate screen. Clarified lysate containing TRX-Taq-TBD (SEQ ID NO: 35) was used to amplify primers targeting the D22S1045 STR locus (top), the DYS481 STR locus (middle), or both loci (bottom). [Figure 9B] Development of a medium-throughput lysate screen. The degree of backstutter (i.e., minus one repeat unit) at each locus was determined as a percentage of the major allele height for one variant containing TRX-Taq-TBD and the H914F substitution from duplex PCR amplification targeting both DYS481 and D22S1045. [Figure 9C]Development of a medium-throughput lysate screen. Example electropherogram obtained when this medium-throughput lysate screen was used to assay lysate containing Taq-TBD (SEQ ID NO: 50) together with TRX (SEQ ID NO: 16), supplied separately as a purified protein. Clarified lysate containing Taq (SEQ ID NO: 1) was included as a control condition. [Figure 9D] Development of a medium-throughput lysate screen. Using this medium-throughput lysate screen, stutter quantification results were obtained when lysate containing Taq-TBD (SEQ ID NO: 50) was assayed together with TRX (SEQ ID NO: 16), supplied separately as a purified protein. Clarified lysate containing Taq (SEQ ID NO: 1) was included as a control condition. [Figure 9E] Development of a medium-throughput lysate screen. Example electropherogram obtained when this medium-throughput lysate screen was used to assay a lysate containing TRX (SEQ ID NO: 16) together with Taq-TBD (SEQ ID NO: 50), supplied separately as a purified protein. Purified Taq was included as a control condition. [Figure 9F] Development of a medium-throughput lysate screen. Stutter quantification results obtained when this medium-throughput lysate screen was used to assay lysates containing TRX (SEQ ID NO: 16) together with Taq-TBD (SEQ ID NO: 50) supplied separately as a purified protein. Purified Taq was included as a control condition. [Figure 10A] Experimentally determined 3D structural model (PDB model 6N7W) showing the TBD (turquoise) and TRX (gray / purple). Residues in TRX, highlighted in purple, were identified as potential interaction motifs and fixed for the generated AI model. [Figure 10B]PDB model 6N7W (showing the TBD (turquoise) and TRX (gray / purple) with the highlighted interaction motif (purple)) overlaid with the predicted structures (yellow, blue, and green) of the three AI-generated sequences tested in the lysate duplex screen. [Figure 10C] Example electropherograms of DYS481 and D22S1045 alleles and stutter peaks obtained from amplification using the indicated cell lysates in a duplex screen. Taq is ATG6964 (SEQ ID NO: 1), TRX-Taq-TBD is ATG7346 (SEQ ID NO: 35), AI sequence 1 is ATG8280 (SEQ ID NO: 51), AI sequence 2 is ATG8279 (SEQ ID NO: 52), and AI sequence 3 is ATG8281 (SEQ ID NO: 53). [Figure 10D] Quantification of backstutter (minus one repeat unit). N=3 for all conditions. Bars represent mean ± standard deviation. X indicates failure to amplify the indicated allele peak. [Figure 11] Reduced stutter is observed with mixed assemblies of TRX-Taq-TBD and Taq-TBD monomers. (A) Example electropherograms obtained when the indicated loci were amplified using purified TRX-Taq-TBD (SEQ ID NO: 35), Taq-TBD (SEQ ID NO: 50), or the indicated mixture of the two (expressed as the ratio of TRX-Taq-TBD:Taq-TBD) in the duplex PCR amplification system of Example 7 (Figure 9). (B) Quantification of backstutter (i.e., minus one repeat). N=3 for all conditions except TRX-Taq-TBD, where n=6. Bars represent the mean ± standard deviation. DETAILED DESCRIPTION OF THE INVENTION
[0027] definition Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the embodiments described herein, certain preferred methods, compositions, devices, and materials are described herein. However, before the present materials and methods are described, it should be understood that this invention is not limited to the specific molecules, compositions, methodologies, or protocols described herein, as these may vary through routine experimentation and optimization. It should also be understood that the terminology used herein is for the purpose of describing particular versions or embodiments only, and is not intended to limit the scope of the embodiments described herein.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. However, in case of conflict, the present specification, including definitions, will prevail. Therefore, with respect to the embodiments described herein, the following definitions will apply:
[0029] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly requires otherwise. Thus, for example, reference to "a domain" is a reference to one or more domains and equivalents thereof known to those skilled in the art, and so forth.
[0030] As used herein, the term "and / or" includes any and all combinations of the listed items, including any and all other combinations of the listed items. For example, "A, B, and / or C" includes A, B, C, AB, AC, BC, and ABC, each of which should be considered individually set forth by the statement "A, B, and / or C."
[0031] As used herein, the term "comprising" and its linguistic variants mean that the recited feature(s), element(s), method step(s), etc. are present, and that the presence of additional feature(s), element(s), method step(s), etc. is not excluded. Conversely, the term "consisting of" and its linguistic variants mean that the recited feature(s), element(s), method step(s), etc. are present, and that any unrecited feature(s), element(s), method step(s), etc., except for impurities ordinarily associated therewith, are excluded. The phrase "consisting essentially of" means the recited feature(s), element(s), method step(s), etc., and any additional feature(s), element(s), method step(s), etc. that do not materially affect the basic nature of the composition, system, or method. Many embodiments herein are described using the open-ended term "comprising." Such embodiments include multiple restrictive "consisting of" and / or "consisting essentially of" embodiments, which may alternatively be claimed or described using language such as "consisting of" and / or "consisting essentially of."
[0032] As used herein, the term "system" refers to a collection of compositions that are brought together in any suitable manner (e.g., physically associated, in the same fluid (e.g., reaction mixture, cell lysate, etc.), in the same body (e.g., cell), packaged together (e.g., in a kit), etc.) for a particular purpose.
[0033] As used herein, the term "sample" is used in its broadest sense. In one sense, a sample is intended to include specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from animals (including humans) and can include fluids, solids, tissues, and gases. Biological samples include blood products such as plasma and serum. A sample can also refer to a cell lysate or purified forms of enzymes, peptides, and / or polypeptides described herein. A cell lysate can include cells lysed with a lysing agent or a lysate such as a rabbit reticulocyte or wheat germ lysate. A sample can also include a cell-free expression system. Environmental samples include environmental materials such as surface material, soil, water, crystals, and industrial samples. However, these examples should not be construed as limiting the types of samples applicable to the present invention.
[0034] As used herein, the term "DNA polymerase" refers to an enzyme that can catalyze the synthesis of DNA molecules from nucleoside triphosphate building blocks using a template DNA molecule to guide the sequence of nucleotide species to be added. Natural DNA polymerases have a highly conserved structure among polymerases within the same class, and the "DNA polymerase domain" or "catalytic domain" varies little between species. The DNA polymerase domain resembles a right hand and contains "thumb," "fingers," and "palm" subdomains. DNA polymerases may also contain additional domains that confer various functionalities (e.g., exonuclease domain(s), thioredoxin-binding domain, TIS, etc.). DNA polymerases are divided into seven families based on their sequence homology and tertiary structure. These include families A, B, C, D, X, Y, and RT. Polymerase family A includes Pol I (encoded by the polA gene), the most abundant and ubiquitous DNA polymerase in prokaryotes, as well as various thermostable DNA polymerases (such as Thermus aquaticus DNA polymerase, Thermus thermophilus DNA polymerase, Thermus flavus DNA polymerase, Thermotoga neapolitana DNA polymerase, and Geobacillus stearothermophilus DNA polymerase) and certain bacteriophage DNA polymerases (such as T7 bacteriophage DNA polymerase and T3 bacteriophage DNA polymerase). In addition to the catalytic domain, family A polymerases contain a 3' to 5' exonuclease domain.
[0035] The terms "DNA polymerase activity," "synthetic activity," and "polymerase activity" are used interchangeably and refer to the ability of a DNA polymerase to synthesize new DNA strands by incorporation of deoxynucleotide triphosphates.
[0036] As used herein, the terms "Taq DNA polymerase" or "Taq" refer to the DNA polymerase of SEQ ID NO: 1, unless otherwise indicated.
[0037] The term "genomic DNA," as used herein, refers to any DNA ultimately derived from the DNA of a genome. The term includes, for example, DNA cloned in a heterologous organism, total genomic DNA, and partial genomic DNA (e.g., isolated DNA from a single chromosome). DNA detected, analyzed, isolated, etc., according to embodiments herein can be single-stranded or double-stranded. For example, single-stranded DNA can be obtained from bacteriophage, bacteria, or genomic DNA fragments. Double-stranded DNA can be obtained from any one of several different sources, such as DNA containing tandem repeat sequences, including phage libraries, cosmid libraries, and bacterial genomes or plasmid DNA, and DNA isolated from the genomic DNA of any eukaryotic organism, including humans. In some embodiments, the DNA is obtained from human genomic DNA. Any one of several different sources of human genomic DNA can be used, including medical or forensic samples such as blood, semen, vaginal swabs, tissue, hair, saliva, urine, and bodily fluid mixtures. Such samples may be fresh, stale, dried, and / or partially decomposed. Samples may be taken from evidence at a crime scene.
[0038] As used herein, the terms "slippage," "slippage," and "stutter" refer to the skipping or rereading of a few nucleotides (e.g., 1-8 nucleotides) in a template DNA strand by a DNA polymerase, resulting in a deletion or duplication of nucleotides in the resulting complementary product strand. In forward stutter, a few nucleotides (e.g., 1-8 nucleotides) in the template strand are read twice by the polymerase, resulting in a product strand containing a duplication of the sequence complementary to the reread nucleotides. In backward stutter, a few nucleotides (e.g., 1-8 nucleotides) in the template strand are skipped, resulting in a product strand containing a deletion of the sequence complementary to the skipped nucleotides. Stutter typically occurs at a very low rate with most template sequences, but is more common when the template strand contains a repetitive sequence of 1-8 nucleotides (e.g., tandem repeats).
[0039] As used herein, the term "tandem repeat" ("simple tandem repeat") refers to a DNA sequence pattern in which a sequence of one or more nucleotides is repeated, with the repeats immediately adjacent to one another. Typically, short repeat sequences (e.g., 1-8 nucleotides) span DNA segments of 10-500 (e.g., 10, 20, 50, 100, 200, 300, 400, 500, or any range therebetween) nucleotides in length, while tandem repeats are longer (e.g., 9-50 nucleotides), spanning DNA segments of 500, 750, 1000 nucleotides, or longer. Repeats of short sequences (e.g., 1-8 nucleotides) are sometimes referred to herein as "short tandem repeats" ("STRs") or "microsatellites." A repeat of a single nucleotide is referred to as a "mononucleotide repeat" (e.g., "AAAAA"), a repeat of two nucleotides is referred to as a "dinucleotide repeat" (e.g., "ACACACAC"), and a repeat of three nucleotides is referred to as a "trinucleotide repeat" (e.g., "AGCAGCAGCAGC"), and so forth.
[0040] As used herein, the term "compound repeat" refers to two or more adjacent simple repeats (i.e., simple tandem repeats with different sequences).
[0041] As used herein, the term "complex repeat" refers to several repeat blocks of variable unit length, plus variable intervening sequences.
[0042] As used herein, the term "complex hypervariable repeat" includes a large number of non-consensus alleles that can vary in both size and sequence (e.g., SE33).
[0043] STR types (e.g., simple, compound, complex, complex hypervariable, etc.) are described, for example, in Chapter 5 (p100) of "Advanced Topics in Forensic DNA Typing: Methodology" by John M. Butler (2012), Academic Press (incorporated by reference in its entirety).
[0044] Tandem repeats used in forensic analysis ("forensic STRs") can be simple or complex repeats. In some embodiments, the type of STR (simple or complex) is not distinguished during forensic analysis.
[0045] As used herein, the term "stutter artifact" refers to a DNA product that has an insertion or deletion of a nucleotide or series of nucleotides as a result of stuttering. In analysis of a DNA product, a stutter artifact will typically appear as a minor signal (e.g., an insertion or deletion) paired with a major signal (e.g., occurring without stuttering). Stutter artifacts result from strand mismatching due to slippage during DNA replication both in vivo and in vitro (see, e.g., Levinson and Gutman (1987), Mol. Biol. Evol., 4(3):203-221, and Schlotterer and Tautz (1992), Nucleic Acids Research 20(2):211-215, which are incorporated by reference in their entireties). Such artifacts become particularly apparent when DNA containing any such repetitive sequences is amplified in vitro using amplification methods such as the polymerase chain reaction (PCR), as any minor fragments present in the sample or arising during polymerization will be amplified along with the major fragment.
[0046] As used herein, the term "backstutter" refers to a stutter artifact caused by the absence of exactly one repeat unit.
[0047] As used herein, the term "stutter propensity" refers to the likelihood that a given set of reaction conditions will cause stutter and / or stutter artifacts. For example, if a particular DNA polymerase produces fewer stutter artifacts than a control, the stutter propensity for that DNA polymerase is low. If a particular template sequence (e.g., tandem repeats) increases the incidence of stutter, the stutter propensity for that template is increased.
[0048] As used herein, the term "template strand" or "template DNA" refers to the DNA sequence that is read by a DNA polymerase during DNA replication or synthesis. The term "product strand" or "product DNA" refers to the DNA sequence that is synthesized during DNA replication. If stuttering occurs when replicating the template strand, the resulting stutter artifact will be present in the product strand.
[0049] As used herein, the term "primer" refers to an oligonucleotide that can hybridize to a template DNA and serve as a starting point for DNA synthesis by a DNA polymerase. A primer can be single-stranded or double-stranded. A primer can be perfectly complementary to a sequence in the template DNA, or it can have one or more mismatches or non-Watson-Crick pairings, provided that the primer is capable of hybridizing to the template under amplification conditions. A primer is said to be "capable of hybridizing to a DNA molecule" if it can anneal to the DNA molecule. That is, the primer shares some degree of complementarity with the DNA molecule. The degree of complementarity can be, but need not be, perfect (i.e., the primer need not be 100% complementary to the DNA molecule). Any primer that can anneal to a template DNA molecule and support primer extension along the template DNA molecule under the reaction conditions employed can hybridize to the DNA molecule.
[0050] As used herein, the terms "complementary" or "complementarity" are used in reference to nucleotide sequences related by the base-pairing rules. For example, the sequence 5'"AGT"3' is complementary to the sequence 3'"TCA"5'. Complementarity can be "partial," in which only a portion of the nucleic acid bases match according to the base-pairing rules. Alternatively, there can be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions, as well as in detection methods that rely on nucleic acid hybridization.
[0051] As used herein, the term "polymerase chain reaction" ("PCR") refers to the method described, for example, in U.S. Pat. Nos. 4,683,195, 4,889,818, and 4,683,202 (all of which are incorporated herein by reference). These patents describe a method for increasing the concentration of a target sequence segment in a genomic DNA mixture without cloning or purification. This process for amplifying a target sequence involves introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target sequence, followed by precise sequential temperature cycling in the presence of a DNA polymerase (e.g., Taq polymerase). The two primers are complementary to their corresponding strands of the double-stranded target sequence. To allow amplification, the mixture is denatured, and the primers are then annealed to their complementary sequences within the target molecule. After annealing, the primers are extended by a polymerase, resulting in the formation of a new pair of complementary strands. The denaturation, primer annealing, and polymerase extension steps can be repeated many times (i.e., denaturation, annealing, and extension constitute one "cycle," and there can be many "cycles") to enrich the amplified segment of the desired target sequence. The length of the amplified segment of the desired target sequence is determined by the relative positions of the primers, and therefore, this length is a controllable parameter. Because of the repetitive nature of the process, the method is referred to as a "polymerase chain reaction" (hereinafter "PCR"). Because the desired amplified segments of the target sequence become the predominant sequences in the mixture (in terms of concentration), they are said to be "PCR amplified."
[0052] Using PCR, it is possible to amplify single copies of specific target sequences in genomic DNA to detectable levels by several different methodologies (i.e., hybridization with a labeled probe, incorporation of a biotinylated primer followed by detection with an avidin-enzyme conjugate, incorporation of labeled deoxynucleotide triphosphates, etc.). In addition to genomic DNA, any oligonucleotide sequence can be amplified with an appropriate set of primer molecules. In particular, the amplified segments created by the PCR process itself are themselves effective templates for subsequent PCR amplification.
[0053] As used herein, the term "fusion protein" refers to a chimeric protein that includes two or more peptide / polypeptide moieties that originate or are derived from different sources.
[0054] As used herein, the term "modifier" refers to any peptide or polypeptide sequence that is fused to a peptide, polypeptide, or protein of interest to confer functionality. Non-limiting examples of modifiers include His tags, HaloTags, streptavidin, antibodies, epitopes, FLAG tags, etc.
[0055] As used herein, the terms "conjugated," "linked," or linguistic variations thereof refer to the connection of two moieties via a covalent or non-covalent bond. Conjugation or linking may require a direct covalent bond, or any suitable linking agent may be employed, such as, for example, a peptide linker, a non-peptide linker, a chemical cross-linker, etc.
[0056] As used herein, the term "peptide" refers to a short polymer of amino acids linked together by peptide bonds. In contrast to other amino acid polymers (e.g., proteins, polypeptides, etc.), peptides are about 50 amino acids or less in length. Peptides can contain natural amino acids, unnatural amino acids, amino acid analogs, and / or modified amino acids. Peptides can be subsequences of naturally occurring proteins or unnatural (artificial) sequences.
[0057] As used herein, a "conservative" amino acid substitution refers to the substitution of an amino acid in a peptide or polypeptide with another amino acid having similar chemical properties (such as size or charge). For purposes of this disclosure, each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) alanine (A) and glycine (G), 2) aspartic acid (D) and glutamic acid (E), 3) asparagine (N) and glutamine (Q), 4) arginine (R) and lysine (K), 5) isoleucine (I), leucine (L), methionine (M), and valine (V), 6) phenylalanine (F), tyrosine (Y), and tryptophan (W), 7) serine (S) and threonine (T), and 8) cysteine (C) and methionine (M).
[0058] Naturally occurring residues can be divided into classes based on the properties of their common side chains, for example, polar positively charged (histidine (H), lysine (K), and arginine (R)), polar negatively charged (aspartic acid (D), glutamic acid (E)), polar neutral (serine (S), threonine (T), asparagine (N), glutamine (Q)), nonpolar aliphatic (alanine (A), valine (V), leucine (L), isoleucine (I), methionine (M)), nonpolar aromatic (phenylalanine (F), tyrosine (Y), tryptophan (W)), proline and glycine, and cysteine. As used herein, a "semi-conservative" amino acid substitution refers to the replacement of an amino acid in a peptide or polypeptide with another amino acid within the same class.
[0059] In some embodiments, unless otherwise specified, conservative or semi-conservative amino acid substitutions may also include non-naturally occurring amino acid residues that have similar chemical properties to the natural residues. These non-natural residues are typically incorporated by chemical peptide synthesis rather than synthesis in biological systems. These include, but are not limited to, peptidomimetics and other inverted or reversed amino acid moieties. Embodiments herein may, in some embodiments, be limited to natural amino acids, non-natural amino acids, and / or amino acid analogs.
[0060] Non-conservative substitutions may require that a member of one class be exchanged for a member of another class.
[0061] As used herein, the term "sequence identity" refers to the extent to which two polymer sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have the same sequence composition of monomer subunits. The term "sequence similarity" refers to the extent to which two polymer sequences (e.g., peptides, polypeptides, nucleic acids, etc.) differ only by conservative and / or semi-conservative amino acid substitutions. "Percent sequence identity" (or "percent sequence similarity") is calculated by: (1) comparing two optimally aligned sequences over a comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window, etc.); (2) determining the number of positions containing identical (or similar) monomers (e.g., identical amino acids found in both sequences, similar amino acids found in both sequences) to obtain the number of matched positions; (3) dividing the number of matched positions by the total number of positions within the comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); and (4) multiplying the result by 100 to obtain the percent sequence identity or percent sequence similarity. For example, if peptides A and B are both 20 amino acids long and all amino acids except one position are identical, then peptide A and peptide B have 95% sequence identity. If the amino acids at non-identical positions share the same biophysical properties (e.g., both are acidic), then peptide A and peptide B will have 100% sequence similarity. As another example, if peptide C is 20 amino acids long and peptide D is 15 amino acids long, and 14 of the 15 amino acids in peptide D are identical to some amino acids in peptide C, then peptides C and D will have 70% sequence identity, but peptide D will have 93.3% sequence identity over the optimal comparison window of peptide C. For purposes of calculating "percent sequence identity" (or "percent sequence similarity") herein, any gap in the aligned sequences is treated as a mismatch at that position.
[0062] Any peptide described herein as having a particular percent sequence identity or sequence similarity (e.g., at least 70%) with a reference sequence can also be expressed as having the maximum number of substitutions (or terminal deletions) relative to that reference sequence. For example, a sequence "having at least 70% sequence identity with SEQ ID NO:X" can have up to three substitutions relative to SEQ ID NO:X (if SEQ ID NO:X is 10 amino acids in length), and therefore can also be expressed as "having no more than three substitutions relative to SEQ ID NO:X." Furthermore, a sequence "having at least 80% sequence similarity with SEQ ID NO:X" can have zero, one, or two non-conservative substitutions relative to SEQ ID NO:X, and therefore can also be expressed as "having no more than two non-conservative substitutions relative to SEQ ID NO:X."
[0063] As used herein, the term "root mean square deviation (RMSD)" refers to a commonly used quantitative measure of similarity between pairs of superimposed atomic coordinates. RMSD values are expressed in angstroms (Å) and are calculated by:
number
[0064] As used herein, the term "closely homologous 3D structures" refers to a pair of polypeptides or domains or subdomains thereof where the alpha carbon RMSD between the two is less than 3 Å.
[0065] As used herein, the term "3D fold threshold" refers to the TM score calculated using TMAlign v 20170708 (https: / / bioweb.pasteur.fr / packages / pack@TM-align@20170708, Y. Zhang, J. Skolnick, TM-align - A protein structure alignment algorithm based on TM-score, Nucleic Acids Research, 33 2302-2309 (2005) (incorporated by reference in its entirety)). In some embodiments, a 3D fold threshold greater than 0.8 (e.g., 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, or higher) indicates a high degree of 3D structural identity. For calculating the 3D fold threshold for polypeptides herein, the 3D molecular structure can be calculated using ESMFold (Zeming Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023) (incorporated by reference in its entirety)).
[0066] As used herein, the term "instability score" refers to a quantitative prediction of a protein's in vivo stability based on its primary sequence (Guruprasad et al. Protein Engineering, Design and Selection, Volume 4, Issue 2, December 1990, Pages 155-161, which is incorporated by reference in its entirety).
[0067] Provided herein are compositions and systems comprising a DNA polymerase domain, a thioredoxin binding domain (TBD), and thioredoxin (TRX), wherein one or both of the TRX and TBD are fused or otherwise conjugated to the DNA polymerase domain. The TRX or TBD may also be provided as separate entities (e.g., binary systems) in the systems herein. The DNA polymerase / TBD / TRX compositions and systems herein are engineered to reduce stutter and / or produce fewer stutter artifacts. Kits comprising the DNA polymerase / TBD / TRX compositions and systems herein, and methods for using the same, are also within the scope of the present specification.
[0068] Although T3 and T7 bacteriophage DNA polymerases are structurally similar to Taq DNA polymerase, these phage polymerases contain an additional domain called the "thioredoxin binding domain" (TBD). Phage propagation requires binding of host thioredoxin (TRX) to the TBD, which greatly enhances the processivity of these phage polymerases. Previous publications have described grafting the T3 TBD onto thermostable Taq polymerase (Davidson et al., 2003, incorporated by reference in its entirety) to form a functional chimeric polymerase. This chimera demonstrated increased processivity in the presence of very high concentrations of TRX, similar to studies performed with T7 DNA polymerase. Typically, increasing processivity does not reduce stutter formation (Verheij, S., Harteveld, J., and Sijen, T. (2012) Forensic Science International: genetics, Vol. 6, pp. 167-175, incorporated by reference in its entirety). The need for a large molar excess of TRX relative to chimeric Taq-TBD (Davidson et al., incorporated by reference in its entirety) significantly limits the usefulness of this approach when applied to conventional PCR techniques. Specifically, the solubility and volume limitations to achieve this molar ratio while maintaining stability throughout temperature cycling limit the commercial viability and practical application of this approach. To overcome this limitation, experiments were conducted during development of the embodiments herein through multiple approaches, such as by genetically fusing one or more thioredoxins to TBD-modified Taq polymerases (e.g., internally or at the N- or C-terminus). These modifications ameliorate the need for an excess of exogenous thioredoxin for reduced stutter and robust polymerase activity.Experiments described herein demonstrate that STR multiplexes amplified with these gene fusion constructs exhibit approximately 10-50% fewer stutter artifacts than multiplexes amplified with unmodified Taq, demonstrating that the apparent stutter formation depends on the specific locus being amplified. Experiments conducted during the development of the embodiments herein also demonstrate that a 0.6-160-fold molar excess of free thioredoxin in the presence of Taq-TBD (without genetically fused thioredoxin) is sufficient to amplify STR multiplexes with stutter reduction comparable to that of gene fusions. This amount of thioredoxin can be delivered in a sufficiently concentrated stock solution compatible with conventional PCR approaches. Additional experiments conducted during the development of the embodiments herein demonstrate the utility of other Taq, TBD, and / or TRX constructs, for example, that may be found to be useful for stutter reduction.
[0069] In addition to the TBD, T3 and T7 DNA polymerases contain a "TBD / TRX-interacting sequence" (TIS), which is believed to interact with one or more of the TBD, TRX, catalytic domain, and / or template DNA to enhance aspects of DNA synthesis. In some embodiments, all (e.g., SEQ ID NO: 18) or a portion (e.g., one of SEQ ID NOs: 19-21 or a portion of SEQ ID NO: 18) of the TIS is fused to or inserted within a polymerase described herein to enhance one or more aspects of DNA synthesis.
[0070] In some embodiments, the polymerases (or polymerase-containing systems) herein reduce the formation of stutter products when amplifying highly repetitive sequences (e.g., STR multiplex). In some embodiments, the polymerases (or polymerase-containing systems) herein reduce stutter artifacts (e.g., because the polymerases or systems herein have a reduced tendency to stutter relative to Taq or other polymerases). For example, in some embodiments, the polymerases (or polymerase-containing systems) herein reduce stutter artifacts (e.g., have a reduced tendency to stutter) compared to polymerases containing only a DNA polymerase domain (e.g., Taq polymerase of SEQ ID NO: 1). In some embodiments, the polymerases (or polymerase-containing systems) herein exhibit fewer stutter artifacts (e.g., 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% fewer, or ranges therebetween) compared to a polymerase that includes only a DNA polymerase domain (e.g., the Taq polymerase of SEQ ID NO: 1). In some embodiments, the polymerases (or polymerase-containing systems) herein exhibit reduced stutter propensity (e.g., 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or ranges therebetween) compared to polymerases comprising only a DNA polymerase domain (e.g., Taq polymerase of SEQ ID NO: 1). In some embodiments, the polymerases (or polymerase-containing systems) herein exhibit fewer stutter artifacts, for example, when amplifying highly repetitive sequences (e.g., STR multiplexes, mononucleotide repeats, etc.).
[0071] In some embodiments, provided herein are systems and compositions comprising a DNA polymerase domain, thioredoxin (TRX), and a thioredoxin binding domain (TBD). In some embodiments, the systems and compositions herein further comprise one or more additional components, such as a linker to all or a portion of the heterologous polymerase domain (e.g., capable of interacting with the TBD and / or TRX). In some embodiments, the systems and compositions herein further comprise a portion of a heterologous polymerase (e.g., T3, T7, etc.), such as a portion of the T3 or T7 exonuclease domain (e.g., all or a portion of the TIS (e.g., SEQ ID NOs: 18-21)).
[0072] In some embodiments, two or more of the various components of the compositions and systems herein are provided as a fusion (e.g., a single polypeptide). In some embodiments, all of the components of the compositions herein (e.g., polymerase domain, TRX, TBD, etc.) are provided as a single fusion polypeptide. In some embodiments, one or more of the various components of the compositions and systems herein are provided as separate polypeptides (e.g., not fused to one or more of the other components). In some embodiments, either the TBD or TRX (or both) are fused or otherwise conjugated to the DNA polymerase domain. In some embodiments, the components of the systems herein may be provided as two, three, or more different polypeptides. In some embodiments, the components of the compositions herein may be provided as a single polypeptide.
[0073] DNA polymerase domain In some embodiments, the polymerase herein comprises a DNA polymerase domain. As defined herein, a DNA polymerase domain is a polypeptide capable of catalyzing DNA synthesis under appropriate conditions. In some embodiments, the DNA polymerase domain of the compositions or systems herein comprises sequence homology (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%) with all or a portion of a DNA polymerase enzyme (e.g., a Family A DNA polymerase (e.g., Taq polymerase, Tne polymerase, etc.)). In some embodiments, the compositions (or system components) herein comprise a DNA polymerase domain having sequence homology to all or a portion of a DNA polymerase enzyme, and various other components (e.g., TRX, TBD, TIS or portions thereof, linker) inserted within the sequence of the DNA polymerase enzyme, replacing a portion of the sequence of the DNA polymerase enzyme (e.g., 1-50 amino acids (e.g., 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any range therebetween), or fused (directly or via one or more linkers) to the N- or C-terminus of the DNA polymerase enzyme. In certain embodiments, for example, in exonuclease-deficient polymerases, regions of the polymerase domain between 100-300 amino acids (e.g., 235 amino acids) in size may be deleted or replaced with alternative domains or components. In some embodiments, homology modeling and tertiary structure analysis are utilized to identify regions of the DNA polymerase enzyme sequence that are suitable sites for insertion or substitution with components of the compositions herein (e.g., TRX, TBD, TIS, linker, etc.).
[0074] In some embodiments, jFATCAT pairwise structural alignments between pdb files 1TAQ and 1T7P were used to determine suitable sites for insertion of or substitution with components of the compositions herein (e.g., TRX, TBD, TIS, linkers, etc.). In some embodiments, primary sequence homology within the alpha helical or flexible domains of Taq DNA polymerase and T7 DNA polymerase was used to determine suitable sites for insertion of or substitution with components of the compositions herein (e.g., TRX, TBD, TIS, linkers, etc.).
[0075] In some embodiments, the polymerases herein comprise a DNA polymerase domain derived from a naturally occurring or known DNA polymerase. In some embodiments, the DNA polymerase domain is derived from a Family A DNA polymerase, such as Thermus aquaticus DNA polymerase (SEQ ID NO: 1), T7 DNA polymerase, DNA polymerase I, DNA polymerase γ, Tne polymerase, and DNA polymerase θ. In some embodiments, the DNA polymerase domain is a chimera of two or more different Family A DNA polymerases. In some embodiments, the DNA polymerase domain is derived from a naturally occurring thermophilic DNA polymerase. In some embodiments, the naturally occurring thermophilic DNA polymerase is selected from the group consisting of Thermus aquaticus DNA polymerase, Thermus thermophilus DNA polymerase, Thermus flavus DNA polymerase, Thermotoga neapolitana polymerase, and Geobacillus stearothermophilus DNA polymerase. In some embodiments, the DNA polymerase domain comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or ranges therebetween) sequence identity to SEQ ID NO: 1, while in other embodiments, a functional DNA polymerase domain may comprise less than 70% (e.g., <60%, <50%, <40%, or less) sequence identity to SEQ ID NO: 1. In some embodiments, the DNA polymerase domain maintains DNA synthesis function as well as one or more additional properties derived from DNA polymerases (e.g., thermostability). In some embodiments, polymerases with proofreading activity, polymerases with no (or negligible) proofreading activity, polymerases with exonuclease activity (e.g., 3'→5', 5'→3', etc.), polymerases with no exonuclease activity, hot start polymerases, non-hot start polymerases, etc. are used as the basis for a DNA polymerase domain.Examples of DNA polymerases from which the DNA polymerase domain is derived include HotStarTaq DNA polymerase (QIAGEN catalog number 203203), AmpliTaq Gold® DNA Polymerase (Applied Biosystems catalog number N8080241), KAPA Taq DNA Polymerase, KAPA Taq HotStart DNA Polymerase (KAPA BIOSYSTEMS catalog number BK1000), Pfu DNA polymerase (Thermo Scientific catalog number EP0501), Klentaq1 (DNA POLYMERASE TECHNOLOGY, Inc., St. Louis, Mo., catalog number 100), PHUSION DNA polymerase (PHUSION High Fidelity DNA polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA polymerase (M0535S, New England BioLabs, Inc.)), Q5® DNA polymerase, etc. Polymerase (such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.)), T4 DNA polymerase (M0203S, New England BioLabs, Inc.), and Sequenase Version 2.0 DNA polymerase (ThermoFisher Scientific catalog number 70775Y200UN).
[0076] In some embodiments, a DNA polymerase domain herein is defined with reference to the Taq DNA polymerase sequence of SEQ ID NO: 1. In certain embodiments, the DNA polymerase domain of a DNA polymerase herein has at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity to SEQ ID NO: 1. In some embodiments, the DNA polymerase domain constitutes a truncated form of 1-50 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, or any range therebetween) at the C-terminus and / or N-terminus compared to SEQ ID NO: 1. In some embodiments, the DNA polymerase domain comprises conservative or non-conservative substitutions to SEQ ID NO: 1. In some embodiments, the DNA polymerase domain of a DNA polymerase herein has at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity to a portion of SEQ ID NO: 1.
[0077] In some embodiments, the DNA polymerase may share at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more, or any range therebetween) sequence identity with one or more of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12. In some embodiments, the DNA polymerase shares at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more, or any range therebetween) sequence identity with each of SEQ ID NOs: 2, 4, 6, 8, 10, and 12 (in that order, optionally interspersed with one or more of SEQ ID NOs: 3, 5, 7, 9, 11 and / or other sequences). In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or ranges therebetween) sequence identity with each of one or more of SEQ ID NOs: 2, 4, 6, 8, 10, and 12 (in that order) and SEQ ID NOs: 3, 5, 7, 9, and 11 (arranged in that order).
[0078] In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity with each of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, and 12 (in that order). In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted between SEQ ID NOs: 10 and 12.
[0079] In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity with each of SEQ ID NOs: 2, 4, 5, 6, 7, 8, 9, 10, and 12 (in that order). In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted between SEQ ID NOs: 2 and 4 and / or between SEQ ID NOs: 10 and 12.
[0080] In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity with each of SEQ ID NOs: 2, 3, 4, 6, 7, 8, 9, 10, and 12 (in that order). In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted between SEQ ID NOs: 4 and 6 and / or between SEQ ID NOs: 10 and 12.
[0081] In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity with each of SEQ ID NOs: 2, 3, 4, 5, 6, 8, 9, 10, and 12 (in that order). In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted between SEQ ID NOs: 6 and 8 and / or between SEQ ID NOs: 10 and 12.
[0082] In some embodiments, the DNA polymerase comprises at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or any range therebetween) sequence identity with each of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 10, and 12 (in that order). In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted between SEQ ID NOs: 8 and 10 and / or between SEQ ID NOs: 10 and 12.
[0083] In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted into SEQ ID NO: 3. In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted into SEQ ID NO: 5. In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted into SEQ ID NO: 7. In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted into SEQ ID NO: 9. In some embodiments, a heterologous amino acid sequence (e.g., TIS, TRX, TBD, etc.) is inserted into SEQ ID NO: 11.
[0084] In some embodiments, the DNA polymerase domain of a DNA polymerase herein has overall homology to SEQ ID NO: 1 (as described in the preceding paragraph), but all or a portion of one or more of SEQ ID NOs: 3, 5, 7, 9, and 11 is replaced by a heterologous insert sequence. For example, a position corresponding to SEQ ID NOs: 3, 5, 7, and / or 9 (or a portion thereof) may be replaced by all or a portion of a TIS of a DNA polymerase (e.g., a T7 polymerase TIS (e.g., SEQ ID NOs: 18-21 or a portion or variant thereof), a T3 polymerase TIS, etc.). In some embodiments, a position corresponding to SEQ ID NO: 11 may be replaced by all or a portion of a TBD (e.g., SEQ ID NO: 15 or a portion or variant thereof) or a TRX (e.g., SEQ ID NO: 16, 17, or 107 or a portion or variant thereof). In some embodiments, other sequences within the DNA binding domain may be deleted or replaced with heterologous sequences (e.g., TBD, TRX, TIS, other portions of other polymerases, etc.), provided that the DNA polymerase domain maintains DNA synthesis catalytic activity. In some embodiments, only a portion of one or more of SEQ ID NOs: 3, 5, 7, 9, and 11 is replaced with a heterologous insert sequence. In such embodiments, a portion(s) of one or more of SEQ ID NOs: 3, 5, 7, 9, and 11 remains in place within the DNA polymerase domain.
[0085] In some embodiments, the DNA polymerase domain comprises an internal amino acid sequence insertion. In some embodiments, the DNA polymerase domain comprises an N-terminal portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 14 and a C-terminal portion having at least 40% (e.g., 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 12, wherein the N-terminal portion and the C-terminal portion are separated by the internal amino acid sequence insertion. In some embodiments, the internal amino acid sequence insertion comprises a TBD.
[0086] In some embodiments, the DNA polymerase domain (e.g., SEQ ID NO: 1) of a DNA polymerase (or polymerase-containing system) herein comprises an exonuclease domain (SEQ ID NO: 13). In some embodiments, the DNA polymerase domain is truncated by deletion of the exonuclease domain (SEQ ID NO: 13).
[0087] In some embodiments, the DNA polymerase domain herein comprises one or more substitutions with respect to the reference DNA polymerase sequence.For example, when using Taq DNA polymerase as the base sequence of the DNA polymerase domain, the DNA polymerase domain can comprise one or more substitutions with respect to SEQ ID NO: 1 (for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, or ranges or values therebetween). Exemplary substitutions include the H914 substitution in Table 1 (position 914 relative to SEQ ID NO:35), the A913 substitution in Table 2 (position 913 relative to SEQ ID NO:35), the R915 substitution in Table 3 (position 915 relative to SEQ ID NO:35), and various mutations of Taq DNA polymerase in Table 4, but substitutions relative to SEQ ID NO:1 or another base DNA polymerase sequence in the DNA polymerase domain are not limited to these positions or substitutions.
[0088] In some embodiments, the DNA polymerase domain is based on Tne DNA polymerase, Tfl DNA polymerase, Taq DNA polymerase, or chimeras thereof (see, eg, Tables 20 and 21).
[0089] Experiments conducted during the development of the embodiments herein demonstrate that Family A DNA polymerases with a wide variety of sequences and substitutions at a wide variety of positions throughout the sequence find use as the DNA polymerase domain of the embodiments herein. The DNA polymerase domain of the constructs herein is not limited to the sequence of a particular DNA polymerase.
[0090] Thioredoxin binding domain (TBD) In some embodiments, a DNA polymerase (or polymerase-containing system) herein comprises a thioredoxin-binding domain. In some embodiments, the DNA polymerase domain of a DNA polymerase herein comprises a TBD fused (e.g., directly or via one or more linkers) to the N-terminus or C-terminus of the DNA polymerase domain. In some embodiments, the DNA polymerase domain comprises a TBD inserted within the DNA polymerase domain. In some embodiments, the TBD is inserted at a position corresponding to or adjacent to an amino acid position within a sequence provided herein (e.g., SEQ ID NO: 1 or a sequence having at least 50% sequence identity thereto). In other embodiments, the TBD is inserted within or replaces all or a portion of an amino acid sequence provided herein (e.g., SEQ ID NO: 3, 5, 7, 9, 11, or any suitable region of SEQ ID NO: 1, or a sequence having at least 50% sequence identity thereto) that is replaced by the TBD. In some embodiments, the TBD is fused or inserted into a DNA polymerase domain position that maintains all or part of the catalytic function or other functional properties of the DNA polymerase domain. In some embodiments, the TBD is inserted into the thumb domain (e.g., SEQ ID NO: 11) of the DNA polymerase domain (e.g., SEQ ID NO: 1). In some embodiments, all or part of the thumb domain (e.g., SEQ ID NO: 11) of the DNA polymerase domain (e.g., SEQ ID NO: 1) herein is replaced by the TBD. In some embodiments, the system includes a TBD that is not fused or otherwise conjugated to the DNA polymerase domain (e.g., in a binary system in which a TRX is fused / conjugated to a DNA polymerase domain).
[0091] In embodiments where the DNA polymerase domain corresponds to a Family A DNA polymerase without a high degree of sequence identity to SEQ ID NO: 1, the TBD is inserted at a position that retains all or part of the catalytic activity of the DNA polymerase.
[0092] In some embodiments, the systems herein include a TBD that is not fused to the DNA polymerase domain. In such embodiments, the DNA polymerase domain is fused or otherwise conjugated to at least one TRX. In some embodiments, the presence of a TBD in the same system as a DNA polymerase domain fused / conjugated to a TRX reduces the stuttering tendency of the DNA polymerase domain compared to a system lacking the TBD and / or TRX. In some embodiments, a TRX is fused or otherwise conjugated to the DNA polymerase domain. In some embodiments, a TRX is fused or otherwise conjugated to a TBD. In some embodiments, a TBD is conjugated (e.g., covalently or non-covalently) to the DNA polymerase domain, but is not fused.
[0093] In some embodiments, the TBD of a DNA polymerase (or system comprising a DNA polymerase) herein is derived from the thioredoxin binding domain of T3 or T7 bacteriophage DNA polymerase. In some embodiments, the TBD comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 15. In some embodiments, the TBD comprises a truncated form of SEQ ID NO: 15 by 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus. In some embodiments, the TBD includes up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions (e.g., conservative or non-conservative) relative to SEQ ID NO: 15.
[0094] In some embodiments, the TBD (e.g., a sequence derived from a T3 or T7 TBD) of a DNA polymerase (or system comprising a DNA polymerase) herein can include substitutions (such as the exemplary substitutions in Tables 6 and 7) relative to a reference sequence (e.g., SEQ ID NO: 15). In some embodiments, the TBD includes substitutions at one or more of T489, R506, T535, E537, E548, and S555 (relative to SEQ ID NO: 35), such as those listed in Table 8. Other substitutions relative to the reference TBD (e.g., SEQ ID NO: 15) are within the scope of this specification.
[0095] In some embodiments, the TBD of a DNA polymerase (or system comprising a DNA polymerase) herein is derived from the thioredoxin binding domain of Salmonella enterica phage DNA polymerase. In some embodiments, the TBD comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO:101. In some embodiments, the TBD comprises a truncated form of SEQ ID NO:101 by 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus. In some embodiments, the TBD includes up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions (e.g., conservative or non-conservative) relative to SEQ ID NO: 101.
[0096] In some embodiments, the TBD of a DNA polymerase (or system comprising a DNA polymerase) herein is derived from the thioredoxin binding domain of Aeromonas hydrophila phage DNA polymerase. In some embodiments, the TBD comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 102. In some embodiments, the TBD comprises a truncated form of SEQ ID NO: 102 by 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus. In some embodiments, the TBD includes up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions relative to SEQ ID NO: 102.
[0097] In some embodiments, the TBD of a DNA polymerase (or system comprising a DNA polymerase) herein is derived from the thioredoxin binding domain of Klebsiella pneumoniae phage DNA polymerase. In some embodiments, the TBD comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 103. In some embodiments, the TBD comprises a truncated form of SEQ ID NO: 103 by 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus. In some embodiments, the TBD includes up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions relative to SEQ ID NO: 103.
[0098] In some embodiments, a DNA polymerase (or a system comprising a DNA polymerase) herein can comprise two or more (e.g., 2, 3, 4, 5, or more) TBDs. In some embodiments, the TBDs are fused or conjugated to different positions in the DNA polymerase domain. In some embodiments, two or more TBD sequences (e.g., identical TBD sequences (e.g., having at least 50% sequence identity to SEQ ID NO: 15), different TBD sequences) are fused or conjugated to the DNA polymerase domain in tandem (e.g., one after the other). In some embodiments, two or more TBDs are comprised in a monomeric DNA polymerase polypeptide. In other embodiments, two or more TBDs are comprised in separate polypeptides in a DNA polymerase binary system (e.g., TBD / Pol-TRX, TBD-Pol-TRX / TBD, TBD-Pol-TRX / TBD-Pol, etc.).
[0099] In some embodiments, the TBD is fused or conjugated to the TRX. In some embodiments, the system comprises a TBD and a TRX, and the TBD and TRX are fused or conjugated (e.g., directly, via one or more linkers, through an interacting partner, etc.) in a manner that facilitates binding of the TBD to the TRX (and subsequently reduces the stutter tendency of an associated DNA polymerase domain (e.g., bound to one or both of the TBD or TRX in the same system). In embodiments in which the TRX and TBD are fused or otherwise conjugated (e.g., directly or via a linker), one or both of the TBD and / or TRX are fused or otherwise conjugated (e.g., directly or via a linker) to the DNA polymerase domain.
[0100] In some embodiments, a free TBD is provided (e.g., in a binary system comprising a DNA polymerase domain fused / conjugated to TRX). In some embodiments, the free TBD is not fused or conjugated to the DNA polymerase domain or TRX. In some embodiments, addition of free TBD to a system comprising a suitable DNA polymerase domain fused or otherwise linked to TRX reduces stutter compared to the DNA polymerase domain in the absence of TRX and / or free TBD. In some embodiments, the binary system comprises a first polypeptide comprising a TBD (e.g., TBD, TBD-Pol, TBD-Pol-TRX, etc.) and a second polypeptide comprising a TRX (e.g., TRX, TRX-Pol, TBD-Pol-TRX, etc.).
[0101] Thioredoxin (TRX) In some embodiments, a DNA polymerase (or polymerase-containing system) herein comprises thioredoxin. In some embodiments, a DNA polymerase domain of a polymerase herein comprises thioredoxin (TRX) fused to the N-terminus or C-terminus of the DNA polymerase domain or inserted within the DNA polymerase domain. In some embodiments, TRX is fused or inserted at a position in the DNA polymerase domain that can maintain all or a portion of the catalytic activity or other functional properties of the DNA polymerase or TRX. In some embodiments, a DNA polymerase domain comprises TRX inserted within the DNA polymerase domain. In some embodiments, TRX is inserted at a position corresponding to or adjacent to an amino acid position in a sequence provided herein (e.g., SEQ ID NO: 1 or a sequence having at least 50% sequence identity thereto). In other embodiments, TRX is inserted into or replaces all or part of an amino acid sequence corresponding to all or part of a sequence provided herein (e.g., SEQ ID NO: 3, 5, 7, 9, 11, or any suitable region of SEQ ID NO: 1, or a sequence having at least 40% sequence identity thereto).
[0102] In some embodiments, the TRX of the DNA polymerase (or system comprising a DNA polymerase) herein is derived from E. coli thioredoxin. In some embodiments, the TRX domain comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 16, 17, or 107. In some embodiments, the TRX comprises a truncated form of 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus relative to SEQ ID NO: 16, 17, or 107. In some embodiments, the TRX comprises up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions relative to SEQ ID NO: 16, 17, or 107. In some embodiments, the TRX (e.g., a sequence derived from E. coli TRX) of a DNA polymerase (or a system comprising a DNA polymerase) herein can comprise substitutions (such as the exemplary substitutions in Tables 9 and 10) relative to a reference sequence (e.g., SEQ ID NO: 16, 17, or 107). In some embodiments, the TRX comprises a substitution at E31 (relative to SEQ ID NO: 35), such as those listed in Table 11. Other substitutions to the reference TRX (eg, SEQ ID NOs: 16, 17, 51-53, 93, 94, 107, etc.) are within the scope herein.
[0103] In some embodiments, the TRX of a DNA polymerase (or a system comprising a DNA polymerase) herein is derived from Alishewanella jeotgali thioredoxin. In some embodiments, the TRX domain comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 94. In some embodiments, the TRX comprises a truncated form of SEQ ID NO: 94 by 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus. In some embodiments, the TRX comprises up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions relative to SEQ ID NO:94.
[0104] In some embodiments, the TRX of a DNA polymerase (or a system comprising a DNA polymerase) herein is derived from Thiococcus pfennigii thioredoxin. In some embodiments, the TRX domain comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 93. In some embodiments, the TRX comprises a truncated form of SEQ ID NO: 93 by 1-20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any range therebetween) at the C-terminus and / or N-terminus. In some embodiments, the TRX comprises up to 30 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range therebetween) substitutions relative to SEQ ID NO:93.
[0105] In some embodiments, the TRX of the polymerase and / or polymerase system herein includes an engineered TRX that is functionally and / or structurally based on a known TRX polypeptide(s), but has divergent sequences with low sequence identity, such as the exemplary engineered TRX sequences in Table 25. In some embodiments, the TRX is engineered via conventional methods such as random mutagenesis, directed mutagenesis, and other techniques that alter amino acid sequence in a directed (e.g., rational) or undirected (e.g., random) manner. In other embodiments, the engineered TRX is generated by maintaining the 3D structure of all or a portion of a reference TRX (e.g., SEQ ID NO: 16). For example, the 3D structure is the portion of TRX that contacts the TBD (e.g., in PDB 6N7W).
[0106] Experiments conducted during the development of embodiments herein demonstrate that TRXs from E. coli, T. pfennigii, and A. jeotgali can all function to reduce stutter in the DNA polymerase systems described herein. These TRXs exhibit 69-76% overall sequence identity between each other, but much higher sequence identity between their TBD-interacting subdomains, ranging from 77.8% to 100%: TBD-interacting subdomain 1 of SEQ ID NO: 16 has 100% identity to E. coli TRX, 100% identity to T. pfennigii TRX, and 88.9% identity to A. jeotgali TRX; TBD-interacting subdomain 2 of SEQ ID NO: 16 has 100% identity to E. coli TRX, 77.8% identity to T. pfennigii TRX, and 83.3% identity to A. jeotgali TRX; and · TBD-interacting subdomain 3 has 100% identity to E. coli TRX, 80% identity to T. pfennigii TRX, and 90% identity to A. jeotgali TRX.
[0107] As described in Example 21, TRX polypeptides were engineered using AI-assisted protein sequence design, protein structure prediction, and protein structure alignment software and methods. The identities and 3D fold structures of selected residues within the putative TBD-TRX binding interface (residues 29-37 (TBD-interacting subdomain 1), 60-77 (TBD-interacting subdomain 2), and 89-98 (TBD-interacting subdomain 3)) were fixed, and candidate sequences predicted to fold such that the TBD-interacting residues were presented in the same 3D configuration were generated. For the 1,000 candidate TRX molecules generated, the alpha-carbon RMSD for the TBD-interacting residues between the 3D models of those sequences and 6N7W was 0.01. The average stutter-reducing interactions between TRX and TBD were 0.86 Å to 2.82 Å, with averages of 1.15 Å and 1.10 Å. Three TRXs engineered by this process were tested for their ability to function to reduce stutter. Despite having less than 55% sequence identity to E. coli, these three TRX sequences (SEQ ID NOs: 51-53) were able to reduce stutter in our test system. These experiments indicate that structurally similar representation of TBD-interacting residues is sufficient to provide a stutter-reducing interaction between TRX and TBD.
[0108] The 3D molecular structures of SEQ ID NOS: 51-53 were calculated using ESMFold (Zeming Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023) (incorporated by reference in its entirety)). The RMSDs of the "TBD-interacting residues" (residues 29-37, 60-77, and 89-98 relative to SEQ ID NOS: 16) in each of the 3D molecular structures calculated using ESMFold for SEQ ID NOS: 51-53 were calculated with respect to the molecular structure of PDB entry 6N7W (Gao et al. (2019) Science 363 (6429) (incorporated by reference in its entirety)). The resulting RMSDs for the TBD-interacting residues of the three engineered TRXs were 1.0 Å to 1.1 Å. RMSD was calculated using the "superimpose Proteins" plugin tool (docs.nanome.ai / plugins / superimpose.html#instructions, incorporated by reference in its entirety) in Nanome Version 1.24 (Bennie S, Maritan M, Gast J, Loschen M, Gruffat D, Bartolotta R, Hessenauer S, Leija E, McCloskey SA Virtual and Mixed Reality Platform for Molecular Design & Drug Discovery - Nanome Version 1.24.5th Workshop on Molecular Graphics and Visual Analysis of Molecular Data, 2023; 2023 / 06 / 12, The Eurographics Association, incorporated by reference in its entirety).
[0109] In some embodiments, provided herein are TRX polypeptides having a predicted 3D molecular structure (e.g., predicted using ESMFold (Zeming Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023) (incorporated by reference in its entirety)) in which the alpha-carbon RMSD of the TBD-interacting residues (residues 29-37, 60-77, and 89-98 relative to SEQ ID NO: 16) relative to PDB 6N7W is 3 Å or less (e.g., 3 Å, 2.8 Å, 2.6 Å, 2.4 Å, 2.2 Å, 2.0 Å, 1.8 Å, 1.6 Å, 1.4 Å, 1.2 Å, 1.0 Å, 0.8 Å, 0.6 Å, 0.4 Å, 0.2 Å, or less, or any value or range therebetween). In some embodiments, the alpha carbon RMSD relative to PDB 6N7W for at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) of the TBD interacting residues (residues 29-37, 60-77, and 89-98 relative to SEQ ID NO: 16) is 3 Å or less (e.g., 3 Å, 2.8 Å, 2.6 Å, 2.4 Å, 2.2 Å, 2.0 Å, 1.8 Å, 1.6 Å, 1.4 Å, 1.2 Å, 1.0 Å, 0.8 Å, 0.6 Å, 0.4 Å, 0.2 Å, or less, or any value or range therebetween). In some embodiments, the alpha carbon RMSD of TBD interacting subdomain 1 (residues 29-37 relative to SEQ ID NO: 16) relative to PDB 6N7W is 3 Å or less (e.g., 3 Å, 2.8 Å, 2.6 Å, 2.4 Å, 2.2 Å, 2.0 Å, 1.8 Å, 1.6 Å, 1.4 Å, 1.2 Å, 1.0 Å, 0.8 Å, 0.6 Å, 0.4 Å, 0.2 Å, or less, or any value or range therebetween).In some embodiments, the alpha carbon RMSD of TBD interacting subdomain 2 (residues 60-77 relative to SEQ ID NO: 16) relative to PDB 6N7W is 3 Å or less (e.g., 3 Å, 2.8 Å, 2.6 Å, 2.4 Å, 2.2 Å, 2.0 Å, 1.8 Å, 1.6 Å, 1.4 Å, 1.2 Å, 1.0 Å, 0.8 Å, 0.6 Å, 0.4 Å, 0.2 Å, or less, or any value or range therebetween). In some embodiments, the alpha carbon RMSD of TBD interacting subdomain 3 (residues 89-98 relative to SEQ ID NO: 16) relative to PDB 6N7W is 3 Å or less (e.g., 3 Å, 2.8 Å, 2.6 Å, 2.4 Å, 2.2 Å, 2.0 Å, 1.8 Å, 1.6 Å, 1.4 Å, 1.2 Å, 1.0 Å, 0.8 Å, 0.6 Å, 0.4 Å, 0.2 Å, or less, or any value or range therebetween).
[0110] In some embodiments, the TBD-interacting residues of TRX (residues 29-37, 60-77, and 89-98 relative to SEQ ID NO: 16) have at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16. In some embodiments, the TBD-interacting residues of TRX (residues 29-37, 60-77, and 89-98 relative to SEQ ID NO: 16) have at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16.
[0111] In some embodiments, TBD-interacting subdomain 1 of TRX (residues 29-37 relative to SEQ ID NO: 16) has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16. In some embodiments, TBD-interacting subdomain 1 of TRX (residues 29-37 relative to SEQ ID NO: 16) has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16.
[0112] In some embodiments, TBD-interacting subdomain 2 of TRX (residues 60-77 relative to SEQ ID NO: 16) has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16. In some embodiments, TBD-interacting subdomain 2 of TRX (residues 60-77 relative to SEQ ID NO: 16) has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16.
[0113] In some embodiments, TBD-interacting subdomain 3 of TRX (residues 89-98 relative to SEQ ID NO: 16) has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16. In some embodiments, TBD-interacting subdomain 3 of TRX (residues 89-98 relative to SEQ ID NO: 16) has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to the TBD-interacting residues of SEQ ID NO: 16.
[0114] In some embodiments, the engineered TRX may be shorter or longer than the TRX of SEQ ID NO: 16, provided that the TBD-interacting residues (e.g., residues having structural and / or sequence identity or similarity to the TRX of SEQ ID NO: 16) are in a closely homologous 3D structure. The length of the TRX can be about 75 to 500 residues or more (e.g., 75, 100, 125, 150, 175, 200, 250, 300, 400, 500, or more).
[0115] In some embodiments, the engineered TRX constitutes a 3D fold threshold of greater than 0.8 (e.g., 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, or higher) relative to PDB 6N7W, indicating a high degree of 3D structural identity.
[0116] In some embodiments, the TRX of a polymerase or system herein comprises at least 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to one of SEQ ID NOs: 51, 52, or 53. In some embodiments, the TRX comprises structural elements of the TRX and / or the ability to reduce the tendency for stuttering in the polymerase system.
[0117] In some embodiments, the TRX sequence is fused to the N-terminus or C-terminus of the DNA polymerase domain. In some embodiments, the TRX sequence is fused to the DNA polymerase domain by a linker of 1-300 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, or a range therebetween (e.g., 30-70 amino acids in length)). The linker can be of any suitable peptide / polypeptide sequence, including, but not limited to, those in Table 13. In some embodiments, the linker is a flexible linker. For example, in some embodiments, 50-100% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) of the linker are glycine and serine residues, although the linker can be composed of any suitable amino acids. In some embodiments, the linker comprises a sequence having at least 40% sequence identity to an exemplary linker in Table 13 or 14. In some embodiments, the linker is a rigid linker and / or comprises a rigid segment. For example, the linker can comprise one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, or more) EAAK peptide segments or other peptides that can introduce rigidity into the linker. Certain embodiments herein are not limited by the identity of the linker.
[0118] In some embodiments, the TRX sequence is not fused to a DNA polymerase and / or a TBD. In some embodiments, free TRX can be fused to one or more peptide or polypeptide modifiers of 1-100 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or any range therebetween). In some embodiments, free TRX comprises a peptide or polypeptide modifier fused to the C-terminus or N-terminus of the TRX sequence. Examples of modifiers include, but are not limited to, His tags, HaloTag, streptavidin, antibodies, epitopes, FLAG tags, etc. In some embodiments, free TRX is conjugated (e.g., non-genetically linked) to a peptide, polypeptide, or non-peptide (e.g., small molecule, solid surface, etc.) by any suitable conjugation method, such as, for example, click chemistry, thiol-maleimide linkage, cysteine-maleimide-cysteine conjugation, etc.
[0119] Conjugation and Linkers Provided herein are systems comprising various components (e.g., DNA polymerase domain(s), TBD(s), TRX(s), TIS, etc.). In some embodiments, two or more components (one of which is a DNA polymerase domain) are conjugated, fused, or otherwise physically connected to one another. For example, in certain embodiments herein, a DNA polymerase domain is genetically fused to a TBD and / or a TRX to form a chimeric DNA polymerase. In some embodiments, any of the components described herein can be fused in a manner consistent with the present disclosure to result in a DNA polymerase and / or polymerase system within the scope of the present disclosure. However, the present disclosure is not limited to genetic fusion of components (e.g., including a DNA polymerase domain) into a single polypeptide. In some embodiments, components can be covalently or non-covalently conjugated or linked (e.g., directly or via one or more linkers) via any suitable conjugation system.
[0120] In the case of fusion of two or more peptide or polypeptide components, the components (e.g., DNA polymerase domain(s), TBD(s), TRX(s), TIS(s), etc.) can be directly fused (e.g., one component is inserted into the other, the C-terminus of one component is fused to the N-terminus of the second component, one component replaces a portion of the other, etc.) or can be indirectly fused (e.g., via a linker segment). For example, in some embodiments, the DNA polymerase domain and the TBD are connected by a linker. In some embodiments, the DNA polymerase domain and the TRX are connected by a linker. In some embodiments, the TBD and the TRX are connected by a linker. In some embodiments, the components herein (e.g., DNA polymerase domain(s), TBD(s), TRX(s), TIS(s), etc.) are connected to an additional component (e.g., an antibody, an affinity molecule, a DNA-binding protein, etc.) by a linker. In some embodiments involving genetic fusion of two or more components, the linker is a peptide or polypeptide linker. In some embodiments, the linker is of a length suitable to allow the components to properly interact with each other, increasing the local concentration of one component relative to another, and / or allowing the components to retain their activity or function (e.g., allowing the TBD to function in the chimeric polymerase in a manner similar to the TBD of T3 or T7 DNA polymerase). In some embodiments, the TBD sequence is fused to the DNA polymerase domain by a linker of, for example, 1-300 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, or any value or range therebetween (e.g., 4-10 amino acids, 30-70 amino acids in length, etc.)).In some embodiments, the TRX sequence is fused to the DNA polymerase domain by a linker of, for example, 1-300 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, or any value or range therebetween (e.g., 4-10 amino acids, 30-70 amino acids in length, etc.)). In some embodiments, the TBD sequence is fused to the TRX by a linker of, for example, 1-300 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, or any value or range therebetween (e.g., 4-10 amino acids, 30-70 amino acids in length, etc.)). In some embodiments, two tandem elements (e.g., two TRXs, two TBDs, etc.) are fused by a linker of, for example, 1-300 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, or any value or range therebetween (e.g., 4-10 amino acids, 30-70 amino acids in length, etc.)). In some embodiments, the components herein (e.g., DNA polymerase domain(s), TBD(s), TRX(s), TIS(s), etc.) and additional elements (e.g., antibodies, affinity molecules, DNA binding proteins, etc.) are fused by a linker of, for example, 1-300 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, or any value or range therebetween (e.g., 4-10 amino acids, 30-70 amino acids in length, etc.)).In some embodiments, two or more linker segments are provided within the polypeptides herein and / or link two components.
[0121] Exemplary linkers for connecting any suitable elements described herein are shown in Tables 12-14. In some embodiments, a linker having at least 60% (e.g., 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%) identity to an exemplary linker in Tables 12-14 is provided, which connects two components herein (e.g., DNA polymerase domain(s), TBD(s), TRX(s), TIS(s), etc.).
[0122] In some embodiments, two components of the DNA polymerase and / or DNA polymerase system herein are conjugated by disulfide bond formation between the components (e.g., TRX and TBD), chemical linkage (e.g., via click chemistry), the use of protein and / or chemical tags, etc.
[0123] The two components (e.g., DNA polymerase domain, TBD(s), TRX(s), TIS, etc.) can be linked by any means known in the art, including covalent and non-covalent interactions. In some embodiments, the first component can be linked to the second component enzymatically or chemically. In some embodiments, the first component can be linked to the second component via ligation. In other embodiments, the first component can be linked to the second component via an affinity binding pair (e.g., biotin and streptavidin). In some cases, the first component can be linked to the second component via a non-natural amino acid (e.g., via a covalent interaction with a non-natural amino acid).
[0124] In some embodiments, a first component can be linked to a second component via a SpyCatcher-SpyTag interaction. The SpyTag peptide forms an irreversible covalent bond to the SpyCatcher protein via a spontaneous isopeptide linkage, thereby providing a genetically encoded strategy for creating peptide interactions that withstand force and harsh conditions (Zakeri et al., 2012, Proc. Natl. Acad. Sci. 109:E690-697; Li et al., 2014, J. Mol. Biol. 426:309-317). The binding factor can be expressed as a fusion protein containing the SpyCatcher protein. In some embodiments, the SpyCatcher protein is added to the N- or C-terminus of a component of the DNA polymerase system herein. The SpyTag peptide can be linked to a second component using standard conjugation chemistry (Hermanson, Bioconjugate Techniques, (2013) Academic Press).
[0125] In some embodiments, an enzyme-based strategy is used to link the first component to the second component. For example, the first component can be linked to the second component using formylglycine (FGly)-generating enzyme (FGE). In one example, a protein (e.g., SpyLigase) is used to link the first component to the second component (Fierer et al., Proc Natl Acad Sci USA. 2014; 111(13):E1176-E1181).
[0126] In other embodiments, the first component can be linked to the second component via a SnoopTag-SnoopCatcher peptide-protein interaction. The SnoopTag peptide forms an isopeptide bond with the SnoopCatcher protein (Veggiani et al., Proc. Natl. Acad. Sci. USA, 2016, 113:1202-1207). The first component can be expressed as a fusion protein containing the SnoopCatcher protein. In some embodiments, the SnoopCatcher protein is added to the N- or C-terminus of the component. The SnoopTag peptide can be linked to the second component using standard conjugation chemistry.
[0127] In yet another embodiment, the first component can be linked to the second component via a HaloTag® protein fusion tag and its chemical ligand. HaloTag is a modified haloalkane dehalogenase designed to covalently bind to a synthetic ligand (the HaloTag® ligand) (Los et al., 2008, ACS Chem. Biol. 3:373-382). The synthetic ligand contains a chloroalkane linker attached to a variety of molecules. The covalent bond between the HaloTag and the chloroalkane linker is highly specific, occurs rapidly under physiological conditions, and is essentially irreversible.
[0128] In some cases, the first component can be linked to the second component by enzymatic attachment (conjugation), such as sortase-mediated labeling (see, e.g., Antos et al., Curr Protoc Protein Sci. (2009) CHAPTER 15:Unit-15.3, International Patent Publication No. WO2013003555). Sortase enzymes catalyze transpeptidation reactions (see, e.g., Falck et al., Antibodies (2018) 7(4):1-19). In some embodiments, the first component is modified with or attached to one or more N- or C-terminal glycine residues.
[0129] In some embodiments, the first component can be linked to the second component using cysteine bioconjugation. In some embodiments, the first component can be linked to the second component using π-TIS-mediated cysteine bioconjugation (see, for example, Zhang et al., Nat Chem. (2016) 8(2):120-128). In some cases, the first component can be linked to the second component using 3-arylpropiolonitrile (APN)-mediated tagging (see, for example, Koniev et al., Bioconjug Chem. 2014;25(2):202-206).
[0130] Other mechanisms (e.g., click chemistry, antibody conjugation, etc.) for linking components of the systems herein (e.g., DNA polymerase domain, TBD(s), TRX(s), TIS, etc.) are within the scope of this disclosure.
[0131] Chimeras, systems, and methods In some embodiments, provided herein are chimeric DNA polymerases comprising a first DNA polymerase domain fused to a second heterologous (e.g., non-native) sequence relative to the DNA polymerase domain. In some embodiments, the DNA polymerase domain may be fused (or otherwise conjugated) to two or more heterologous sequences. In some embodiments, one or more heterologous sequences may be inserted within the DNA polymerase domain or may replace an amino acid segment of the underlying sequence of the DNA polymerase domain (e.g., SEQ ID NO: 1).
[0132] In some embodiments, provided herein are compositions comprising a chimeric DNA polymerase with reduced stuttering propensity, the chimeric DNA polymerase comprising (a) a DNA polymerase domain, (b) a thioredoxin binding domain (TBD), and (c) a thioredoxin (TRX) domain. In some embodiments, the chimeric DNA polymerase with reduced stuttering propensity comprises a sequence having at least 60% (e.g., 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 22-27 and 35-49.
[0133] Some embodiments herein involve chimeric DNA polymerases that include a DNA polymerase domain (e.g., based on the DNA polymerase domain of SEQ ID NO: 1) with one or more insertions, substitutions, N-terminal or C-terminal additions, or deletions. For reference, the DNA polymerase domain sequence of SEQ ID NO: 1 can be divided into 11 segments: N-terminal segment (SEQ ID NO: 2), insertion site A (SEQ ID NO: 3), internal segment 1 (SEQ ID NO: 4), insertion site B (SEQ ID NO: 5), internal segment 2 (SEQ ID NO: 6), insertion site C (SEQ ID NO: 7), internal segment 3 (SEQ ID NO: 8), insertion site D (SEQ ID NO: 9), internal segment 4 (SEQ ID NO: 10), thumb insertion site (SEQ ID NO: 11), and C-terminal segment (SEQ ID NO: 12). Each of the insertion sites (AD and thumb) represents a portion of the DNA polymerase domain that, in certain embodiments, is replaced with or is an insertion site for a heterologous sequence (e.g., TIS, TBD, TRX). All or part of the insertion site can be replaced with a heterologous sequence. Alternatively, the entire insertion site can be retained, with the heterologous sequence inserted between two amino acids of the insertion site. Each of the internal segments (C-terminal, 1-4, and N-terminal) in certain embodiments represents a portion of the remaining DNA polymerase domain, and no heterologous segment is inserted or substituted therein. In some embodiments, the internal segments can be locations for various substitutions, deletions, additions, etc., for the purpose of enhancing the properties of the polymerase. Any of the above sequences or combinations thereof can include various substitutions to enhance one or more properties of the systems herein.
[0134] In certain embodiments, the insertion site AD is a position for insertion of or substitution with a TIS described herein. In some embodiments, the DNA polymerase comprises one or more of the insertion sites AD containing an insertion or substitution (all or part of the insertion site) with a TIS (e.g., a sequence having greater than 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 18-21). In some embodiments, the insertion site AD is a position for insertion of or substitution with a TBD or TRX described herein. In some embodiments, the DNA polymerase comprises one or more of insertion sites AD that contain an insertion or substitution (all or part of the insertion site) in a TBD or TRX (e.g., a sequence having greater than 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 15-17).
[0135] In some embodiments, the thumb insertion site is a location for insertion of or substitution with a TIS described herein. In some embodiments, the thumb insertion site is a location for insertion of or substitution with a TBD or TRX described herein. In some embodiments, the DNA polymerase comprises a thumb insertion site that contains an insertion or substitution (all or part of the insertion site) with a TBD or TRX (e.g., a sequence having greater than 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 15-16). In some embodiments, the DNA polymerase comprises a thumb insertion site that contains an insertion or substitution (all or part of the insertion site) in a TIS (e.g., a sequence having greater than 50% (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 18-21).
[0136] In some embodiments, provided herein are DNA polymerase systems that include a first DNA polymerase domain and a second heterologous (e.g., non-native to the DNA polymerase domain) sequence, wherein the DNA polymerase domain and the heterologous sequence are not fused as a single polypeptide.
[0137] In some embodiments, the TRX sequence is fused to a DNA polymerase domain or TBD via a linker that allows both intramolecular interactions between the TRX and TBD on the same protein monomer and intermolecular interactions between the TRX and TBD on different protein monomers. In other embodiments, the TRX sequence is fused to a DNA polymerase domain or TBD via a linker that allows only intramolecular interactions between the TRX and TBD on the same protein monomer. In some embodiments, the TRX sequence is fused to a DNA polymerase domain or TBD via a linker that allows only intermolecular interactions between the TRX and TBD on different protein monomers.
[0138] In some embodiments, the TBD sequence is fused to a DNA polymerase domain or TRX via a linker that allows both intramolecular interactions between the TBD and TRX on the same protein monomer and intermolecular interactions between the TBD and TRX on different protein monomers. In other embodiments, the TBD sequence is fused to a DNA polymerase domain or TRX via a linker that allows only intramolecular interactions between the TBD and TRX on the same protein monomer. In some embodiments, the TBD sequence is fused to a DNA polymerase domain or TRX via a linker that allows only intermolecular interactions between the TBD and TRX on different protein monomers.
[0139] In some embodiments, the TRX sequence is not fused to a DNA polymerase domain or TBD. In some embodiments, the TRX sequence is fused to another protein or peptide that interacts with DNA, the Pol-TBD, and / or the TRX-Pol-TBD. In some embodiments, the TRX sequence is fused to another protein or peptide. In some embodiments, free TRX can be fused to one or more peptide or polypeptide modifiers that are 1-200 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 160, 165, 170, 180, 185, 190, 195, 200, or a range therebetween). In some embodiments, free TRX comprises a peptide or polypeptide modifier fused to the C-terminus or N-terminus of the TRX sequence. Examples of modifiers include, but are not limited to, His tags, HaloTag, streptavidin, antibodies, epitopes, FLAG tags, etc. In some embodiments, free TRX is conjugated (e.g., non-genetically linked) to a peptide, polypeptide, or non-peptide (e.g., small molecule, solid surface, etc.) by any suitable conjugation method, such as, for example, click chemistry, thiol-maleimide linkage, cysteine maleimide-cysteine conjugation, etc. In some embodiments, a chemical reaction is utilized that increases the local concentration of TRX relative to the TBD, which is higher than would otherwise be achieved in a simple binary system.
[0140] In some embodiments, the TBD sequence is not fused to a DNA polymerase domain or TRX. In some embodiments, the TBD sequence is fused to another protein or peptide that interacts with DNA, Pol-TRX, and / or TRX-Pol-TBD. In some embodiments, the TBD sequence is fused to another protein or peptide. In some embodiments, the free TBD may be fused to one or more peptide or polypeptide modifiers that are 1-200 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 160, 165, 170, 180, 185, 190, 195, 200 or a range therebetween). In some embodiments, the free TBD comprises a peptide or polypeptide modifier fused to the C-terminus or N-terminus of the TBD sequence. Examples of modifiers include, but are not limited to, His tags, HaloTag, streptavidin, antibodies, epitopes, FLAG tags, etc. In some embodiments, free TRX is conjugated (e.g., non-genetically linked) to a peptide, polypeptide, or non-peptide (e.g., small molecule, solid surface, etc.) by any suitable conjugation method, such as, for example, click chemistry, thiol-maleimide linkage, cysteine maleimide-cysteine conjugation, etc. In some embodiments, a chemical reaction is utilized that increases the local concentration of TBD relative to TRX, which is higher than would otherwise be achieved in a simple binary system.
[0141] In some embodiments, provided herein are compositions comprising: (a) a fusion protein comprising (i) a DNA polymerase domain and (ii) a thioredoxin binding domain (TBD), and (b) free thioredoxin (TRX). In some embodiments, the free thioredoxin is present in the composition at a TRX:TBD ratio of 0.1 to 2000 (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, or a range therebetween (e.g., 0.1-800, 0.6-600, etc.)). In some embodiments, the fusion protein comprises a sequence having at least 60% (e.g., 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, 100%, or a range therebetween) sequence identity to SEQ ID NOs: 28-34.
[0142] In some embodiments, provided herein are compositions comprising: (a) a fusion protein comprising (i) a DNA polymerase domain, (ii) a thioredoxin binding domain (TBD), and (iii) thioredoxin (TRX); and (b) a fusion protein comprising (i) free thioredoxin or (ii) TRX and a DNA polymerase. In some embodiments, (a) and (b) are present in the composition at a ratio of 1:100 to 100:1 (e.g., 1:100, 1:80, 1:60, 1:40, 1:20, 1:10, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 10:1, 20:1, 40:1, 60:1, 80:1, 100:1).
[0143] In some embodiments, provided herein are compositions comprising: (a) a fusion protein comprising (i) a DNA polymerase domain, (ii) a thioredoxin binding domain (TBD), and (iii) thioredoxin (TRX); and (b) (i) free TBD or (ii) a TBD and DNA polymerase fusion. In some embodiments, (a) and (b) are present in the composition at a ratio of 1:100 to 100:1 (e.g., 1:100, 1:80, 1:60, 1:40, 1:20, 1:10, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 10:1, 20:1, 40:1, 60:1, 80:1, 100:1).
[0144] In some embodiments, a DNA polymerase domain herein is a fusion of portions of two or more DNA polymerases (e.g., naturally occurring sequences (e.g., portions of Taq and Tfl DNA polymerases), engineered sequences, etc.). In some embodiments, the DNA polymerase domain is a chimeric DNA polymerase.
[0145] In some embodiments, the polymerases (or polymerase-containing systems) herein find use in any system (e.g., amplification reactions) in which a DNA polymerase (e.g., a thermostable DNA polymerase, such as Taq polymerase, etc.) would otherwise find use. In some embodiments, the polymerases (or polymerase-containing systems) herein find use in PCR reactions, multiplex amplification, STR amplification, sequencing applications (e.g., Sanger, NGS), MSI-related technologies, etc.
[0146] In some embodiments, any PCR conditions disclosed herein or any standard PCR conditions can be used with the polymerases and subsystems described herein. Any PCR conditions can be used in any of the methods herein to amplify a target nucleic acid. In some embodiments, provided herein are kits or reaction mixtures comprising a chimeric DNA polymerase or fusion protein described herein and amplification reagents sufficient to amplify a DNA target sequence. In some embodiments, the amplification reagents include one or more of oligonucleotide primers, deoxynucleotide triphosphates, magnesium, ethylenediaminetetraacetic acid (EDTA), a buffer, water, and template DNA comprising a DNA target sequence.
[0147] In some embodiments, the kit or reaction mixture further comprises a reducing agent. In some embodiments, the reducing agent is a thiol reducing agent or a non-thiol reducing agent. In some embodiments, the reducing agent is dithiothreitol (DTT) or tris(2-carboxyethyl)phosphine (TCEP).
[0148] In some embodiments, the DNA target sequence comprises one or more short tandem repeats (STRs). In some embodiments, the STRs comprise repeating units of 1-8 nucleotides, ranging in length from 10-500 nucleotides (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, or any range therebetween). In some embodiments, the tandem repeats comprise repeating units of 1-50 nucleotides (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any range therebetween) ranging in length up to 1000 nucleotides.
[0149] In some embodiments, the reaction volume includes ethylenediaminetetraacetic acid (EDTA), magnesium, tetramethylammonium chloride (TMAC), or any combination thereof. In some embodiments, the concentration of TMAC is 20-80 mM (e.g., 25-70 mM, 30-60 mM, 30-40 mM, 40-50 mM, 50-60 mM, or 60-70 mM), inclusive. In some embodiments, the concentration of magnesium (e.g., magnesium from magnesium chloride) is 1-10 mM (e.g., 1-8 mM, 1-5 mM, 1-3 mM, 3-5 mM, 3-6 mM, or 5-8 mM), inclusive. In some embodiments, the concentration of available magnesium (i.e., magnesium not bound by phosphate groups on dNTPs, primers, or template nucleic acids, or by carboxylic acid groups on magnetic or other beads, if present) (the concentration of magnesium available for binding to the polymerase and not expected to bind to molecules other than the polymerase) is 0.5-10 mM (e.g., 1-8 mM, 1-5 mM, 1-3 mM, 3-5 mM, 3-6 mM, 4-6 mM, or 5-8 mM), inclusive. In some embodiments, EDTA is used to reduce the amount of magnesium available as a polymerase cofactor, since high magnesium concentrations can lead to PCR errors, such as amplification of non-target nucleic acids. In some embodiments, EDTA reduces the concentration of available magnesium to 1-5 mM (e.g., 3-5 mM).
[0150] In some embodiments, the pH is 6.0 to 9.0 (e.g., 6.0 to 6.8, 6.8 to 7.5, 7.5 to 8.8, 8 to 8.3, or 8.3 to 8.5), inclusive. In some embodiments, Tris is used at a concentration of, for example, 10 to 100 mM (e.g., 10 to 25 mM, 25 to 50 mM, 50 to 75 mM, or 25 to 75 mM), inclusive. In some embodiments, Tris at any of these concentrations is used at a pH of 7.5 to 8.5.
[0151] In some embodiments, a combination of KCl and (NH4)2SO4 is used, such as 50-150 mM KCl and 10-90 mM (NH4)2SO4, inclusive. In some embodiments, the concentration of KCl is 0-30 mM, 50-100 mM, or 100-150 mM, inclusive. In some embodiments, the concentration of (NH4)2SO4 is 10-50 mM, 50-90 mM, 10-20 mM, 20-40 mM, 40-60 mM, or 60-80 mM (NH4)2SO4, inclusive. In some embodiments, ammonium [NH4. + ] concentration is 0-160 mM (0-50, 50-100, or 100-160 mM, etc.) (endpoints included).
[0152] In some embodiments, a crowding agent such as polyethylene glycol (PEG, such as PEG 8,000) or glycerol is used. In some embodiments, the amount of PEG (e.g., PEG 8,000) is 0.1 to 20% (e.g., 0.5 to 15%, 1 to 10%, 2 to 8%, or 4 to 8%), inclusive. In some embodiments, the amount of glycerol is 0.1 to 20% (e.g., 0.5 to 15%, 1 to 10%, 2 to 8%, or 4 to 8%), inclusive. In some embodiments, the crowding agent allows for either a lower polymerase concentration to be used and / or a shorter annealing time to be used. In some embodiments, the crowding agent improves the uniformity of direct oxide reduction (DOR) and / or reduces dropouts (undetected alleles).
[0153] In some embodiments, 5 to 2000 units / mL (units per mL of reaction volume) (e.g., 5 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900, 900 to 1000, 1000 to 1500, or 1500 to 2000 units / mL, inclusive) of polymerase are used. One unit is defined as the amount of enzyme required to catalyze the incorporation of 10 nanomoles of dNTP into acid-insoluble material at 74°C in 30 minutes.
[0154] In some embodiments, hot-start PCR is used to reduce or prevent polymerization prior to PCR thermal cycling. Exemplary hot-start PCR methods include initial DNA polymerase inhibition or physical separation of reaction components until the reaction mixture reaches a higher temperature. In some embodiments, the enzyme is spatially separated from the reaction mixture by wax, which melts as the reaction reaches a higher temperature. In some embodiments, sustained release of magnesium is used. Because DNA polymerase requires magnesium ions for activity, magnesium is chemically separated from the reaction by binding to a chemical compound and released into solution only at higher temperatures. In some embodiments, non-covalent binding of an inhibitor is used. In this method, a peptide, antibody, or aptamer non-covalently binds to the enzyme at low temperatures and inhibits its activity. After incubation at elevated temperatures, the inhibitor is released and the reaction begins. In some embodiments, a cold-sensitive Taq polymerase is used, such as a modified DNA polymerase that has little activity at low temperatures. In some embodiments, chemical modifications are used. In this method, a molecule is covalently attached to the side chain of an amino acid in the active site of the DNA polymerase. The molecule is released from the enzyme by incubating the reaction mixture at elevated temperature. Release of the molecule activates the enzyme. In some embodiments, the amount of template nucleic acid (e.g., RNA or DNA sample) is 20-5,000 ng (e.g., 20-200, 200-400, 400-600, 600-1,000, 1,000-1,500, or 2,000-3,000 ng), inclusive. In some embodiments, the reaction contains 0.2 ng / mL to 2 μg / mL of template DNA (e.g., 0.2 ng / mL, 0.5 ng / mL, 1 ng / mL, 2 ng / mL, 5 ng / mL, 10 ng / mL, 20 ng / mL, 50 ng / mL, 100 ng / mL, 500 ng / mL, 1 μg / mL, 2 μg / mL, or ranges therebetween).
[0155] In some embodiments, provided herein are methods for amplifying a DNA target sequence, comprising exposing a reaction mixture comprising a chimeric DNA polymerase or fusion protein herein and amplification reagents to PCR temperature cycling conditions.
[0156] In some embodiments, exemplary PCR thermocycling conditions include 95° C. for 10 minutes (hot start), 20 cycles of 96° C. for 30 seconds, 65° C. for 15 seconds, and 72° C. for 30 seconds, followed by 72° C. for 2 minutes (final extension), followed by a hold at 4° C. In some embodiments, PCR thermocycling conditions include 95° C. for 10 minutes (hot start), 25 cycles of 96° C. for 30 seconds, 65° C. for 20 seconds, and 72° C. for 30 seconds, followed by 72° C. for 2 minutes (final extension), followed by a hold at 4° C. In some embodiments, an exemplary set of PCR thermocycling conditions includes 95° C. for 10 minutes, 15 cycles (95° C. for 30 seconds, 65° C. for 1 minute, 60° C. for 5 minutes, 65° C. for 5 minutes, and 72° C. for 30 seconds), followed by 72° C. for 2 minutes. In some embodiments, an exemplary set of PCR thermocycling conditions includes 96°C for 1 minute, 30 cycles (94°C for 10 seconds, 59°C for 30 seconds, 72°C for 1 minute), a final 60°C for 10 minutes, and a 4°C hold. In other embodiments, an exemplary set of PCR thermocycling conditions includes 96°C for 1 minute, 30 cycles (94°C for 10 seconds, 59°C for 30 seconds), a final 60°C for 10 minutes, and a 4°C hold. In some embodiments, PCR thermocycling is used with the following exemplary reaction conditions: 100 mM KCl, 50 mM (NH4)2SO4, 3 mM MgCl2, 7.5 nM of each primer in the library, 50 mM TMAC, and 7 ul of template DNA in a final volume of 20 ul (pH 8.1). In some embodiments, other reaction conditions understood in the art are utilized. [Example]
[0157] Example 1 Taq-TBD + free thioredoxin system to reduce stutter TBDs derived from T3 or T7 bacteriophage DNA polymerase were selected for insertion into Taq DNA polymerase. The TBD insertion site is within or near the thumb domain, with or without adjacent linker sequences. Once cloned, the Taq-TBD construct and TRX are expressed and purified. Expression and purification of Taq-TBD and TRX can be performed independently from the same expression vector or from separate expression vectors in the same or mixed culture. A reduction in amplicon stutter can be observed when PCR is performed in the presence of cell lysate. Cell lysate enriched for Taq-TBD and / or thioredoxin can be used as a means to reduce stutter. Certain components introduced or already present in the cell lysate during purification may mask activity and / or amplicon detection.
[0158] Experiments conducted during the development of embodiments herein demonstrate that hot start enhances Taq-TBD activity and the formation of specific amplicons with reduced stutter. In this process, polymerase activity is inhibited to prevent the formation of nonspecific and undesired amplicons before PCR begins. Hot start can be achieved using conjugation of chemicals, proteins, or antibodies to the polymerase enzyme or to other elements that interact with the enzyme. Alternatively, hot start can be achieved thermally, temporally, and / or spatially by adding critical PCR components that are separated by space and / or time and interact only at the start of PCR. Once the reagents are prepared, they can then be subjected to conditions that alleviate the inhibitory effects of the hot start elements, allowing PCR to proceed.
[0159] The presence and amount of thioredoxin are important for polymerase activity and reduced stuttering. In one experiment, a titration was performed in which the concentration of thioredoxin was increased stepwise while maintaining the concentration of Taq-TBD. The signal increased with increasing concentrations of TRX relative to Taq-TBD (up to approximately 160x) (Figure 1F) and decreased with decreasing concentrations of TRX relative to Taq-TBD (Figure 1G). Compared to amplification using Taq or Taq-TBD in the absence of TRX, stuttering was reduced at all TRX concentrations relative to Taq-TBD tested (Figure 1H), but was no longer significantly reduced above a certain threshold of approximately 10-20 molar excess of TRX relative to Taq-TBD.
[0160] In the experiments performed here, stuttering was reduced even at a molar ratio of thioredoxin of only 0.6 (Fig. 1A-H).
[0161] The redox state of the reaction mixture is important for polymerase activity and stutter reduction. Certain reducing agents can be added to the PCR suspension to ensure that thioredoxin remains in a reduced state. In one example, a thiol reducing agent (DTT) was added at various concentrations to a multiplex PCR suspension (Figure 2). The resulting amplicons varied in quantity and total number of stuttered products, with increased stutter and decreased yield associated with lower reducing agent concentrations. Increasing the reducing agent concentration increased yield and reduced stutter, but the response diminished above a threshold concentration. A similar effect was observed when a non-thiol reducing agent was titrated into the reaction suspension.
[0162] The active site of all thioredoxins contains two cysteine residues that are involved in maintaining the redox balance inside the cell. The two native cysteine residues are repeatedly oxidized and reduced due to their ability to form transient disulfide bonds. Mutation of one or both cysteine residues impairs this ability. However, the effects of thiol-deficient thioredoxin variants on stutter and amplicon formation in the presence of Taq-TBD polymerase are relatively unchanged. Thus, efficient PCR amplification and reduced stuttering of repetitive sequences do not require the presence of redox-active thioredoxin (Figure 3). However, despite the elimination of redox-active thioredoxin, the reduction of stuttering still requires the presence of a reducing agent.
[0163] The requirement for a large molar excess of thioredoxin relative to Taq-TBD to reduce stutter is unclear but can be explained by several possibilities, including: (1) the thermal cycling process weakens the interaction between bound thioredoxin and TBD, causing thioredoxin to dissociate from the polymerase complex at higher temperatures; (2) thioredoxin unfolds / denatures at the selected cycling parameters; and / or (3) thioredoxin is required to dissociate in order to bind and amplify another template DNA strand.
[0164] Example 2 Chimera of Taq-TBD and thioredoxin to reduce stutter Covalently linking thioredoxin to Taq-TBD polymerase increases the local concentration of thioredoxin near the TBD binding site without adding exogenous thioredoxin to the reaction. This can be achieved by various strategies, such as disulfide bond formation and / or chemical cross-linking. A chimera of Taq-TBD and thioredoxin connected by a linker is a single polypeptide. The linker can be rigid, flexible, or intermediate and can be composed of the same amino acid residue, a repeating sequence of residues, or a random sequence of residues. The linker can flank an internal insertion sequence or be attached to either the amino or carboxy terminus. There can be more than one covalently linked thioredoxin moiety per Taq-TBD protein. The length of the linker depends on the region within the protein into which it is inserted. Several covalently linked thioredoxin and Taq-TBD chimeras were investigated for their performance in multiplex PCR on DNA repetitive sequences. All were able to amplify the target sequence and exhibited significantly reduced stutter compared to Taq, however, the activity and amount of stutter artifacts produced varied across the chimeras evaluated (Figure 4).
[0165] Example 3 Chimera of Taq-TBD and thioredoxin to reduce stutter in the detection of microsatellite instability PCR-based microsatellite instability (MSI) detection involves examining repetitive DNA sequences, typically mononucleotide repeats, in specific genomic regions. Detection of novel alleles resulting from length variations in repeat sequences in tumor tissue (not present in corresponding normal tissue) can indicate a diagnosis of high-frequency MSI. Stutter artifacts introduced during PCR can significantly complicate this type of analysis by masking the appearance of these novel tumor alleles. While novel tumor alleles can be mathematically deconvoluted from the shape and distribution spread of PCR-amplified fragments (allele(s) and corresponding stutter peaks), the highest sensitivity and specificity of such detection is achieved when novel tumor alleles are clearly separable as new local maxima in the peak distribution of the tumor sample.
[0166] To demonstrate the applicability of this new technology for PCR-based detection of MSI status, normal and tumor-derived DNA samples were amplified with either standard Taq or a chimera of thioredoxin and Taq-TBD linked at the N-terminus (TRX-Taq-TBD) using primers specific for the MONO-27 locus. When amplified with standard Taq, assignment of MSI status required mathematical comparison of the shape and spread of the peak distribution. In contrast, amplification with TRX-Taq-TBD visualized a new local maximum in the peak distribution (Figure 5, indicated by the red arrow), likely due to a reduced contribution of stutter artifacts. Detection of this new local maximum significantly improved the sensitivity and specificity of frequent MSI diagnosis.
[0167] Example 4 Testing chimeric polymerases that reduce stutter This example describes the investigation of the stutter-reducing properties of several variations of the Taq-TBD or TRX-Taq-TBD polymerases described herein. The variants investigated included an exonuclease domain deletion variant (FIG. 6A), cysteine mutants, including a cysteine null mutation variant (FIG. 6B), a variant with a point mutation (I677T) converting the T3 TBD sequence to a T7 sequence (FIG. 6C), variants with linkers of various lengths used to connect the TRX to the Taq-TBD construct at the N-terminus (FIG. 6D), variants with linkers of various lengths used to connect the TRX to the Taq-TBD construct at the C-terminus (FIG. 6E), variants with repeated TRXs connected in tandem at the N-terminus (FIG. 6F), and variants with tandemly repeated TBDs inserted within Taq (FIG. 6G). The stuttering properties of the polymerases were investigated by amplifying them using Promega's PowerPlex® Fusion multiplex system and analyzing the amplification products by capillary electrophoresis.
[0168] Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only at loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap).
[0169] The data demonstrate that the investigated variants significantly reduced stutter compared to the matched control condition (i.e., multiplex amplified with standard Taq). A "#" indicates that a dropout (i.e., no allele peak) was observed at the indicated locus for that condition.
[0170] Example 5 The Taq-TBD and thioredoxin chimera reduces stutter in STR multiplexes when analyzed by next-generation sequencing (NGS). In addition to capillary electrophoresis (CE) workflows, next-generation sequencing (NGS) can be used to analyze STR data. Such workflows will often include a target amplification step in which the regions to be sequenced (in this example, autosomal and sex-linked STRs) are amplified using PCR. The resulting amplicons are then processed (e.g., by various library preparation chemistries) and sequenced (e.g., by sequencing-by-synthesis). The sequencing data can then be analyzed to determine which alleles are present in the sample for the loci amplified in the target amplification multiplex PCR reaction.
[0171] Here, we tested whether the Taq-TBD construct, in the presence of excess TRX, was suitable for use in the target amplification step of an NGS workflow. Reactions used primers from the PowerSeq® 46GY System, library preparation was performed using the Illumina® TruSeq® DNA PCR-Free library preparation workflow, and sequencing was performed on an Illumina MiSeq™ instrument. Control amplifications using standard Taq were performed in parallel. Similar to what was observed with the CE-based workflow, Taq-TBD-mediated amplifications showed significantly reduced stutter compared to standard Taq-mediated amplifications across multiplexes (Figure 7).
[0172] Example 6 Insertion test of TBD-interacting sequences This example describes the investigation of the amplification and stutter reduction properties of a variant TRX-Taq-TIS-TBD construct (SEQ ID NO: 48), in which a putative TBD-interacting sequence (SEQ ID NO: 20) has been substituted into the 5' exonuclease domain of the TRX-Taq-TBD construct. The stutter properties of the polymerase were investigated by amplifying using Promega's PowerPlex® Fusion multiplex system and analyzing the amplified products by capillary electrophoresis.
[0173] Stutter frequency was determined by comparing the peak height of the stutter allele with the peak height of the corresponding allele. Stutter percentages were determined only at loci where the allele and stutter peaks could be clearly separated (e.g., the allele and stutter peaks did not overlap).
[0174] The data demonstrate that while overall amplification efficiency appears similar (Figure 8A), amplification at select loci appears to be enhanced by the inclusion of a TIS (Figure 8B). Notably, at these select loci, amplification with TRX-Taq-TBD appeared to result in reduced peak heights compared to Taq, which was partially restored when amplified with the TRX-Taq-TIS-TBD construct. Importantly, amplification with the TRX-Taq-TIS-TBD and TRX-Taq-TBD constructs similarly reduced stutter artifacts compared to control Taq reactions.
[0175] Example 7 Development of a medium-throughput screen to assess stuttering. This example describes a medium-throughput screen developed to quantify the incidence of stutter artifacts during PCR amplification using candidate polymerases. Two key process improvements made this medium-throughput screen possible: 1) the ability to use clarified lysate as the polymerase source, and 2) the use of primer duplexes to amplify a target STR locus pair (DYS481 and D22S1045) that has a naturally high rate of stutter.
[0176] TRX-Taq-TBD (SEQ ID NO: 35) and Taq-TBD containing the point mutation H914F were expressed in E. coli using the KRX autoinduction system (Promega catalog #L3002). After expression, the culture was centrifuged, the cell pellet was resuspended in lysis buffer to a fraction of the original culture volume, and FastBreak™ Cell Lysis Reagent (Promega catalog #V8571) was added to promote cell lysis. The lysate was then exposed to heat at 65°C and subsequently clarified by centrifugation. Heat-treated clarified cell lysate was used as a polymerase source for PCR reactions amplified with primers targeting the DYS481 STR locus (Promega PowerPlex® Y23 System, Catalog #: DC2305), the D22S1045 STR locus (Promega PowerPlex® Fusion System, Catalog #: DC2402), or both the DYS481 and D22S1045 loci simultaneously.
[0177] For the polymerase variants examined here, amplification of both STR loci was observed in independent monoplex and duplex reactions. Figure 9A shows an exemplary electropherogram of amplification using TRX-Taq-TBD. The observed peaks correspond to the expected alleles in the 2800M control DNA (DYS481 allele 22 and homozygous allele 16 at D22S1045). Duplex reactions were chosen for stutter quantification in this study and for future screening. Stutter was calculated as a percentage of the peak height of the associated allele (amplitude of the stutter peak divided by the amplitude of the allele peak, Figure 9B). Figure 9B demonstrates that the H914F mutation further reduced stutter compared to the base TRX-Taq-TBD (SEQ ID NO: 35) construct.
[0178] This medium-throughput screen is also useful for assessing stutter artifacts when Taq-TBD and TRX are used as separate proteins. Taq-TBD can be present as a clarified lysate and purified TRX can be included (Figures 9C and 9D), or TRX can be present as a clarified lysate and purified Taq-TBD can be included (Figures 9E and 9F). In both cases, Taq-TBD in the presence of TRX significantly reduced stutter compared to the Taq control.
[0179] Example 8 Site saturation at A913, H914, and R915 Example 7 described the development of a medium-throughput cell lysate screen and the observation that the H914F mutation reduces the formation of stutter artifacts. Here, we further probed position H914 and two adjacent residues, A913 and R915, by generating site-saturation libraries for these positions. These libraries were examined using the cell lysate screen for assessing stutter, as described in Example 7.
[0180] Position H914 was probed against the background of the TRX-Taq-TBD sequence with a 90-aa linker (SEQ ID NO: 54). A library of constructs was generated in which every available amino acid was substituted at this position. Table 1 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. Taq and unmutated TRX-Taq-TBD were assayed in parallel as control lysates. [Table 1-1] [Table 1-2]
[0181] Similarly, adjacent residues A913 (Table 2) and R915 (Table 3) were investigated using site saturation libraries. For these constructs, TRX-Taq-TBD (SEQ ID NO: 35) with a 60 aa linker was used as the background construct against which substitutions were made. [Table 2-1] [Table 2-2] [Table 3-1] [Table 3-2]
[0182] Mutations to A913 and H914 were well tolerated, as any amino acid substituted at this position showed reduced stutter compared to the Taq control. One particular substitution to R915 was not tolerated, as it failed to amplify at least one of the allele peaks. However, mutations to this residue generally increased stutter above the rate observed with TRX-Taq-TBD, but all of the mutations that successfully amplified duplexes showed less stutter than the Taq control.
[0183] Example 9 Mutations throughout the Taq backbone The cell lysate screen for assessing stutter, described in Example 7, was used to assay the effects of point mutations throughout the Taq backbone. To aid in the selection of mutations to evaluate, the ESM-1b and ESM-1v protein language models, trained on millions of naturally occurring protein sequences, were applied across the complete Taq protein sequence (Hie, BL, et al., Nature Biotechnology volume 42, pages 275-283 (2024) (incorporated by reference in its entirety)). These algorithms incorporate contextual information for each residue in the protein, taking into account its interactions with other residues and its evolutionary relationships. Mutations are scored against the amino acid at a given position in a reference sequence (typically the wild-type protein). The probability assigned to the mutant amino acid is compared to the probability assigned to the wild-type amino acid. The scored mutations at each position were then used to generate a list of mutations based on the aggregate rankings from the individual algorithms. A subset of these was selected for evaluation based on two criteria: The first are residues that are in positions that can contact DNA during polymerization, and the second are mutations that have a high tolerance score across the entire Taq sequence.
[0184] Table 4 summarizes the tendency of this library to form backstutter products, expressed as a percentage of the allele peak height. [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5]
[0185] Example 10 Insertion site for thioredoxin-binding domain Comparison of the crystal structures of Taq DNA polymerase large fragment (PBD:3KTQ) and T7 DNA polymerase (PDB:1T7P) clearly shows that the TBD deviates from the otherwise relatively well-aligned core components of the two polymerases. In this structural comparison, the T7 TBD extends 76 residues from the thumb domain before the structural alignment between the two enzymes is restored. Davidson et al. (incorporated by reference in its entirety) replaced six amino acid residues (HPFNLN) in the thumb domain of Taq with these 76 residues, and this substitution was used in TRX-Taq-TBD (SEQ ID NO:35). The last four residues in the TBD insert differ from the corresponding TAQ deletion sequence by only two residues (FNPS vs. FNLN, respectively). To determine whether the exact site of TBD insertion and the flanking sequences could be altered, we created constructs with alternative TBD insertion sites. These sequences differ only within the TBD insertion site, as exemplified below (Taq-derived sequence highlighted in bold): [ka]
[0186] These alternative TBD insertion site constructs were evaluated for stutter in the cell lysate duplex screen described in Example 7.
[0187] Table 5 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 5]
[0188] Example 11 Point mutations throughout the TBD domain To assess the impact of mutations within the TBD insert on stutter, a library of constructs was generated by applying the ESM-1b and ESM-1v protein language models to the entire TBD sequence, as described in Example 9 (Hie, BL, et al., Nature Biotechnology volume 42, pages 275-283 (2024) (incorporated by reference in its entirety)). The impact of these changes on the formation of stutter artifacts was then assayed using the cell lysate screen described in Example 7.
[0189] Table 6 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5]
[0190] The following mutations in the TBD were examined, but none of the allele peaks could be amplified: W482P, Y483P, Q484P, P485E, K486E, G488E, G488L, K508P, I509P, P510E, K511E, G513E, G513K, I515G, F516P, K517P, P519L, L533E, D534L, V539E, The following residues were mutated: A542L, Y544P, T545P, P546E, V547E, V547P, E548P, H549G, Y483K, K486P, P510G, D534V, V539P, G541P, A542K, Y544G, T545D, P546G, V547G, E548G, V550P, F552P, N553P, Q524L, N521E, and P554G. A second library was constructed by independently mutating all residues within the TBD to lysine residues unless the protein language model already predicted the change to lysine or a lysine was naturally present at the desired position. This library was also subjected to cell lysate assays. [Table 7-1] [Table 7-2] [Table 7-3]
[0191] Example 12 TBD site saturation The TBD point mutation screen described above identified several sites of interest for further stutter reduction. Six of these (T489, R506, T535, E537, E548, and S555) were selected for further evaluation using site-saturation mutagenesis at each indicated position within the context of SEQ ID NO: 55. The resulting constructs were evaluated for stutter using the cell lysate screen described in Example 7.
[0192] Table 8 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 8-1] [Table 8-2] [Table 8-3] [Table 8-4] [Table 8-5]
[0193] Example 13 TRX point mutation screening In this example, a library of point mutations within the TRX sequence of TRX-Taq-TBD (SEQ ID NO: 35) was generated using the ESM-1b and ESM-1v protein language models described in Example 9. This library was examined using the cell lysate screen described in Example 7 to assess whether these point mutations were tolerated and, if so, how they affected stutter propensity.
[0194] Table 9 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 9-1] [Table 9-2] [Table 9-3] [Table 9-4] [Table 9-5] [Table 9-6] [Table 9-7]
[0195] Some of these point mutations were also examined in a binary context where TRX and Taq-TBD are separate proteins: various TRX constructs were expressed as clarified lysates and examined in combination with purified Taq-TBD (SEQ ID NO: 50). [Table 10]
[0196] Example 14 Site saturation at E31 The TRX point mutation screen identified several sites of interest for further stutter reduction. One of these (E31) was selected for further evaluation using site-saturation mutagenesis in the pATG7620 (SEQ ID NO: 56) background. The resulting construct was evaluated for stutter using the cell lysate screen described in Example 7. [Table 11-1] [Table 11-2]
[0197] Example 15 Alternative linker sequences Previous examples have described flexible linkers (e.g., pATG7346, SEQ ID NO: 35) that fuse Taq-TBD and thioredoxin into a single polypeptide. As previously demonstrated in Figures 6D and 6E, the nature of the linker influences stutter frequency. In this example, we further varied the length and composition of the linker region to determine the effect on stutter product formation.
[0198] TRX-Taq-TBD (pATG7346, SEQ ID NO: 35) was designed with a 60-aa linker composed of repeats of glycine and serine residues. This base construct was modified to allow linker lengths ranging from 0 to 200 aa while maintaining a flexible glycine / serine repeat composition. The resulting constructs were expressed in E. coli for analysis in the cell lysate screen described in Example 7 to assess how the length of the N-terminal linker affects stutter propensity. [Table 12-1] [Table 12-2]
[0199] Linker configurations were also explored by inserting various linker motifs into a TRX-Taq-TBD-based construct (pATG7860, SEQ ID NO: 57) with an 82 aa linker containing several mutations throughout the protein. These motifs included a more rigid linker (pATG8264, SEQ ID NO: 74), linkers with additional functionality (e.g., a protease (e.g., TEV) cleavage site: SLEPTTEDLYFQSDND, pATG8297, SEQ ID NO: 76), a linker with an alternative flexible sequence (pATG8285, SEQ ID NO: 77), and other linker sequences such as: SEQ ID NO: 78, pATG8298:GS(24)-CASSIDYKRISRMPSKIMDAVIDTLNICKLANCE-GS(24) SEQ ID NO: 79, pATG8286:GS(24)-CASSIDYKRISRMPAVLADAVIDTLNICKLANCE-GS(24)
[0200] These constructs were expressed in E. coli for analysis in the cell lysate screen described in Example 7 to assess how the composition of the N-terminal linker affects stutter propensity.
[0201] Table 13 summarizes the tendency of each linker construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 13]
[0202] Figure 6E demonstrates that TRX can be fused to the C-terminus of Taq-TBD (Taq-TBD-TRX). In this example, C-terminal linkers of various lengths were further investigated. These C-terminal linkers were modified to lengths ranging from 25 to 55 amino acids while maintaining a flexible glycine / serine repeat composition. The resulting constructs were expressed in E. coli for analysis in the cell lysate screen described in Example 7 to assess how the length of the C-terminal linker affects stutter propensity.
[0203] Table 14 summarizes the tendency of each construct with a shorter C-terminal linker length to form backstutter products, expressed as a percentage of the allele peak height. [Table 14]
[0204] Example 16
[0205] Combination mutations across TRX and TBD Mutations of interest identified in the TBD and TrX in the above examples were aligned to generate various combinations of monomer constructs, and the resulting variants were evaluated for stuttering using the cell lysate screen described in Example 7. [Table 15-1] [Table 15-2] [Table 15-3] [Table 15-4] [Table 15-5] [Table 15-6] [Table 15-7] [Table 15-8]
[0206] Example 17 TBD orthologues reduce stuttering The TBD sequences used in experiments conducted during the development of the embodiments herein were derived from either the T3 or T7 bacteriophage DNA polymerase, which differ by a single amino acid residue (the threonine residue at position 30 of the T3 TBD is replaced by an isoleucine residue). The T3 and T7 bacteriophages are specific to E. coli. However, TBD orthologs have been found in DNA polymerases of phages that infect other species of bacteria. To determine whether these alternative TBD sequences could also function to reduce stutter in the context of Taq DNA polymerase, the E. coli phage TBD sequence found in TRX-Taq-TBD (SEQ ID NO: 35) was independently replaced with phage TBD sequences from Klebsiella pneumoniae (SEQ ID NO: 103), Salmonella enterica (SEQ ID NO: 101), and Aeromonas hydrophila (SEQ ID NO: 102). These TBD orthologs share 85%, 72%, and 49% sequence identity with the T3 and T7 phage TBDs, respectively. The resulting monomeric constructs were then evaluated for stutter in a cell lysate duplex screen as described in Example 7.
[0207] Table 16 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 16]
[0208] These alternative TBD sequences were also examined in a binary system where TRX and Taq-TBD are separate proteins. As above, the E. coli phage TBD sequence found in Taq-TBD (SEQ ID NO: 15) was independently replaced with phage TBD sequences from Klebsiella pneumoniae (SEQ ID NO: 103), Salmonella enterica (SEQ ID NO: 101), and Aeromonas hydrophila (SEQ ID NO: 102). These various Taq-TBD constructs were expressed as clarified lysates and examined in combination with purified TRX (SEQ ID NO: 16).
[0209] Table 17 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 17]
[0210] Example 18 TRX orthologues reduce stuttering Although the thioredoxin sequence (e.g., SEQ ID NO: 16) found in many experiments conducted during the development of the embodiments herein was derived from E. coli, thioredoxin is ubiquitously expressed in all organisms. To determine whether alternative TRX sequences could also function to reduce stutter in the context of the described stutter-reduced PCR system, the E. coli TRX sequence found in TRX-Taq-TBD (SEQ ID NO: 35) was independently replaced with TRX sequences from Alishewanella jeotgali (SEQ ID NO: 94) and Thiococcus pfennigii (SEQ ID NO: 93). The resulting constructs were then evaluated for stutter in a cell lysate duplex screen as described in Example 7.
[0211] Table 18 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 18]
[0212] Additionally, these thioredoxin orthologs were expressed as independent proteins (Alishewanella jeotgali (SEQ ID NO: 94), Thiococcus pfennigii (SEQ ID NO: 93)) and used in combination with purified Taq-TBD (SEQ ID NO: 50). These constructs were evaluated for stuttering in a binary cell lysate duplex screen as described in Example 7.
[0213] Table 19 summarizes the tendency of each construct to form backstutter products in the binary system, expressed as a percentage of the allele peak height. [Table 19]
[0214] Thioredoxin orthologs from Alishewanella jeotgali and Thiococcus pfennigii reduced the formation of stutter artifacts compared to the Taq control when substituted into TRX-Taq-TBD fusion constructs and used in Taq-TBD binary systems. Notably, these orthologs share 77% and 70% sequence identity, respectively, to E. coli TRX.
[0215] Example 19 Polymerase orthologues reduce stutter Based on its structure and function, Taq is classified as a Family A, DNA polymerase I. Polymerases within this family share a high degree of structural similarity. The high degree of structural similarity among Family A DNA polymerases suggests that insertion of a TBD may be achieved in any member of this family to achieve reduced stutter under appropriate conditions. Here, we demonstrate this concept by inserting a TBD domain into Tne polymerase (SEQ ID NO: 105). The resulting Tne-TBD construct was tested in a dual cell lysate assay using either TRX (SEQ ID NO: 16) or the TRX mutant construct E31P (relative to SEQ ID NO: 16) (SEQ ID NO: 107).
[0216] Table 20 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 20]
[0217] When used in combination with either TRX, the Tne-TBD construct exhibited reduced stuttering propensity both compared to Taq and compared to Tne lacking the TBD domain, which is particularly noteworthy given that Tne and Taq share only 42.8% sequence identity, and even when the T3 TBD is inserted within each thumb domain, the sequence identity only rises to 48.0%.
[0218] Experiments were conducted during the development of the embodiments herein to determine whether a TBD domain could be inserted into a Tfl-Taq chimera using a sequence similar to that described previously (Villbrandt et al., Protein Engineering, Design and Selection, Volume 13, Issue 9, September 2000, Pages 645-654, incorporated by reference in its entirety). In the Tfl-Taq chimera examined here, the intervening domain of Taq was replaced with the respective sequence from Tfl. The intervening domain is a 150-residue sequence that exhibits 3'-5' exonuclease activity in other pol I-like DNA polymerases, but this activity is absent in both Taq and Tfl. The resulting protein sequences are 96.8% identical, although within the exchanged domain (residues 462-611), they share only 77.3% sequence identity. This chimera was expressed in E. coli and tested in the monomeric cell lysate assay described in Example 7.
[0219] Table 21 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 21]
[0220] Example 20 Combination of TBD, TRX, and DNA polymerase orthologs Previous examples have demonstrated that TBD, TRX, and core polymerase sequences can be replaced with orthologous sequences and retain polymerase function to reduce stutter. In this example, we further explore this concept by replacing multiple domains within the same construct. Examples of such combinations include pairing TBD domains from Klebsiella pneumoniae, Salmonella enterica, and Aeromonas hydrophila bacteriophages with thioredoxins from Alishewanella jeotgali and Thiococcus pfennigii in the context of TRX-Taq-TBD in all possible combinations. The resulting constructs were evaluated for stuttering in a cell lysate duplex screen as described in Example 7.
[0221] Table 22 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 22]
[0222] Table 23 summarizes the percent sequence identity (Madeira et al. Search and sequence analysis tools services from EMBL-EBI in 2022. Nucleic Acids Research, April 12, 2022 (incorporated by reference in its entirety)) of the various constructs with TRX-Taq-TBD construct 7346 (SEQ ID NO: 35) for the indicated domain(s). [Table 23]
[0223] Experiments were performed during development of embodiments herein to determine whether a TBD ortholog could be substituted into Tne polymerase and used in a binary system to amplify STR loci with reduced stutter, similar to the approach described in Example 19. The resulting constructs were evaluated for stutter in a cell lysate duplex screen as described in Example 7.
[0224] Table 24 summarizes the tendency of each construct to form backstutter products, expressed as a percentage of the allele peak height. [Table 24-1] [Table 24-2]
[0225] Example 21 AI-engineered thioredoxin homologues The previous examples demonstrated that thioredoxin orthologs from different bacteria function to reduce stutter in the described system, even when paired with various TBD orthologs. Given this compatibility, experiments were conducted during development of the embodiments herein to determine whether proteins that (1) have a similar overall structure to TRX and (2) interact sufficiently with TBD could function as substitutes for TRX in the described stutter-reducing PCR system. To test this, an artificial intelligence (AI) model trained to generate protein structures and sequences was employed to replace the biogenic thioredoxin sequence in the stutter-reducing system.
[0226] A diffusion probability model was configured to hold selected residues in the putative TBD-TRX binding interface fixed (Figure 10A), and the model was applied to generate candidate protein structures. A second AI model was applied to perform reverse folding with the same residues fixed to generate sequences that would fold into the candidate structures identified by the diffusion probability model. The three-dimensional structures of the generated sequences were predicted by a third AI model, without any other information about the sequence's origin. Finally, the predicted structures were compared with the diffusion model structures (Figure 10B), and candidate sequences were selected for further investigation.
[0227] Selected engineered TRX sequences were independently fused to the N-terminus of Taq-TBD with a 60-aa linker (SEQ ID NOs: 52, 53, and 54) and tested in the cell lysate duplex screen described in Example 7. Three of the sequences with engineered TRX amplified the correct DYS481 allele with reduced stutter compared to Taq control amplification (Figures 10C, D). Under these conditions, two of these sequences failed to amplify the larger D22S1045 locus, whereas one of the engineered sequences amplified the correct allele with reduced stutter compared to the Taq control (Figures 10C, D). Notably, these three engineered sequences share only approximately 50% sequence identity with each other and with the E. coli thioredoxin sequence used as design input (Table 25). This demonstrates that using structure as a key determinant in in silico-generated thioredoxins can generate sequences that are low in identity to E. coli TRX and function to reduce stutter. [Table 25]
[0228] AI-based protein sequence generation and candidate selection methods Residues potentially interacting with the TBD were identified in the PDB model 6N7W of TRX (residues 29-37, 60-77, and 89-98, 37 of 108, or 34.3%). RFDiffusion v1.1.0 was configured to fix these residues and maintain the overall size of TRX. It was also configured to be aware of the TBD but vary only the non-fixed residues of TRX. The model generated by RFDiffusion was then input into ProteinMPNN 1.0.1, fixing the same set of residues. The sequence output from ProteinMPNN was input into ESMFold v1 for 3D structure prediction. The ESMFold 3D model was compared back to the corresponding RFDiffusion model using TMAlign v20170708, and the TM score was calculated. The isoelectric point and instability index were also calculated. Candidate sequences were selected for laboratory testing based on the TM score, instability index, isoelectric point, and ProteinMPNN score. The structural predictions of the final candidates were reviewed by manual inspection in 3D visualization software.
[0229] Tools used ESMFold (Zeming Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023) (incorporated by reference in its entirety)), isoelectric point and instability index (Cock, et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics, Bioinformatics, Volume 25, Issue 11, June 2009, Pages 1422-1423, which is incorporated by reference in its entirety); ProteinMPNN (J. Dauparas et al., Robust deep learning-based protein sequence design using ProteinMPNN. Science 378, 49-56 (2022) (incorporated by reference in its entirety)), RFDiffusion (Watson, et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089-1100 (2023) (incorporated by reference in its entirety)); TMAlign (Zhang and Skolnick, TM-align: A protein structure alignment algorithm based on TM-score, Nucleic Acids Research, 33:2302-2309 (2005) (incorporated by reference in its entirety)).
[0230] Example 22 Stutter-reducing mixture of binary and monomeric polymerases Previous examples of binary systems require exogenous addition of TRX, either as a cell lysate or purified component. Supplemented TRX can also be added as a fusion with another protein or polypeptide modifier. In this example, mixtures of purified Taq-TBD and TRX-Taq-TBD were created and evaluated for stuttering using the lysate screening assay described in Example 7. In this context, exogenous TRX is present on the TRX-Taq-TBD construct but not on the Taq-TBD variant. As shown in the examples below, homogenous mixtures of Taq-TBD exhibit low peak heights and high stuttering, while homogenous mixtures of TRX-Taq-TBD exhibit high peak heights and low stuttering (Figures 11A and 11B). Decreasing the ratio of TRX-Taq-TBD to Taq-TBD to 9:1, 3:1, and 1:1 has little effect on peak height and stuttering (Figures 11A and 11B). This provides evidence that the TRX on TRX-Taq-TBD interacts intermolecularly with Taq-TBD in a binary fashion, since otherwise, peak heights would be expected to decrease and stutter to increase as the ratio of TRX-Taq-TBD to Taq-TBD decreased. This observation is also consistent with Example 1H, which demonstrates that stutter can be reduced even at a 0.6X ratio of TRX to Taq-TBD.
[0231] array Below are the sequences provided herein, identified by SEQ ID NO: In addition to modifications described and permitted within the embodiments herein, any of the sequences herein may further be provided with or without sequences intended for purification or other related uses (such as a His tag or other purification tag as understood in the art), which may or may not be present in the sequences provided herein.
[0232] The "X" residue present in the sequences below can be any amino acid residue or the absence of an amino acid residue. Sequences that include any residue at the X position below (or sequences with no residue at the X position) are within the scope of the present specification.
[0233] In the sequences below and throughout this specification, when a position number is given without reference to a specific base sequence, the assumed reference sequences are SEQ ID NO: 1 for the DNA polymerase, SEQ ID NO: 15 for the TBD, SEQ ID NO: 16 for the TRX, and SEQ ID NO: 35 for the TRX-Taq-TBD construct. For example, H914 refers to the histidine at position 914 of the TRX-Taq-TBD with a 60-amino acid linker. For TRX-Taq-TBD constructs with linkers of different lengths, when H914 is mentioned without other context, it refers to the histidine at the position in that construct that corresponds to position H914 in SEQ ID NO: 35. Similarly, mutations made in the TBD (T489, R506, T535, E537, E548, and S555) are assigned based on their positions in SEQ ID NO: 50 (pATG6979), which encodes Taq-TBD without fused thioredoxin. For the TRX-Taq-TBD construct, when these residues are mentioned without other context, it refers to their positions in SEQ ID NO:50.
[0234] SEQ ID NO:1 - Taq DNA polymerase RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDL LGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLE EWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEAGERAALSERLFANLW GRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGHPFNLNSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGD ENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRY VPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD SEQ ID NO:2 - Taq DNA polymerase N-terminal segment RGMLPLFEPKGRVLLVDGHHLAYRTFHALKG SEQ ID NO:3 - Taq DNA polymerase insertion site A LTTSRGEPVQAVYGFAKSLLKALKE SEQ ID NO:4—Taq DNA polymerase internal segment 1 DGDAVIVVFDA SEQ ID NO:5 - Taq DNA polymerase insertion site B KAPSFRHEAYGGYKAG SEQ ID NO:6—Taq DNA polymerase internal segment 2 RAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERL SEQ ID NO:7 - Taq DNA polymerase insertion site C RAFLER
[0235] SEQ ID NO:8—Taq DNA polymerase internal segment 3 LEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNT SEQ ID NO:9 - Taq DNA polymerase insertion site D TPEGVARRYGGEWTEE SEQ ID NO:10—Taq DNA polymerase internal segment 4 AGERAALSERLFANWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAG SEQ ID NO: 11 - Taq DNA polymerase thumb insertion site HPFNLN SEQ ID NO:12 - Taq DNA polymerase C-terminal segment SRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMR RAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD SEQ ID NO: 13 - Taq DNA polymerase exonuclease domain RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEA DDVLASLAKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAH SEQ ID NO: 14—Taq DNA polymerase, large N-terminal portion RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADD VLASLAKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLK LSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVL ALREGLLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAG
[0236] SEQ ID NO: 15 - Thioredoxin binding domain (TBD) GSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPS SEQ ID NO: 16 - Thioredoxin (TRX) SDKIIHLTDDSFTDDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA SEQ ID NO: 17 - Thioredoxin C36S (TRX-C36S) SDKIIHLTDDSFTDDVLKADGAILVDFWAEWCGPSKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA SEQ ID NO: 18-T7 TBD / TRX interacting sequence (TIS) DTDMGLLRSGKLPGKRFGSHALEAWGYRLGEMKGEYKDDFKRMLEEQGEEYVDGMEWWNFNEEMMDY SEQ ID NO: 19-T7 TIS truncation A AWGYRLGEMKGEYKDDFKRMLEEQGEEYVDG SEQ ID NO: 20-T7 TIS truncation B GEMKGEYKDDFKRMLEEQGEEYVDG SEQ ID NO: 21-T7 TIS truncation C EYKDDFKRMLEEQGEEYV
[0237] SEQ ID NO:22 - Exemplary DNA polymerase / TBD / TRX chimera (N-terminal TRX) X1 X10SDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLAX20 X300RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLAR LEVPGYEADDVLASLAKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRL KPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPE PYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVA LDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVE TLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKX101 X110
[0238] SEQ ID NO:23 - Exemplary DNA polymerase / TBD / TRX chimera (N-terminal TRX) 60 residue linker
[0239] SEQ ID NO:24 - Exemplary DNA polymerase / TBD / TRX chimera (C-terminal TRX) X1 X10RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRL KPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGPGPDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRYGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKX11 X100SDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLAX101 X110
[0240] SEQ ID NO:25 - Exemplary DNA polymerase / TBD / TRX chimera (C-terminal TRX) 40 residue linker
[0241] SEQ ID NO:26 - Exemplary DNA polymerase / TBD / TRX chimera (dual TRX) X1 X10SDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLAX11 X100RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAAEKEGYEVRILTADKDLYQLLSDRHIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRL KPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGPGPDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKX101 X200DKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLAX201 X220
[0242] SEQ ID NO:27 - Exemplary DNA polymerase / TBD / TRX chimera (dual TRX)
[0243] SEQ ID NO:28 - Exemplary DNA polymerase / TBD chimera X1X10RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLA RLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDR LKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPE PYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVLF DELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVAL DYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETL FGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKX11X20
[0244] SEQ ID NO:29 - Exemplary DNA polymerase / TBD chimera MKHHHHHHMRGMPLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDL LGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLK NLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVH RAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEAT GVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLE RVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWL LVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGY VETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0245] SEQ ID NO: 30 - Exemplary Taq-TBD (pATG6370) MKHHHHHHMRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0246] SEQ ID NO:31 - Exemplary Taq-TBD Δ235Exo (pATG7221) MDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEA GERAALSERLFANWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVLFDELG LPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTIN FGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0247] SEQ ID NO:32 - Exemplary Taq-TBD(C493S), pATG7354 MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFSHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0248] SEQ ID NO: 33 - Exemplary Taq - TBD (C531S), pATG7355 MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPSELDTREYVAGAPYTPVEHVVFNPSSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0249] SEQ ID NO: 34 - Exemplary Taq - TBD (C493S / C531S), pATG#7356 MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLE VPGYEADDVLASLAKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLK PAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFSHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPSELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0250] SEQ ID NO:35 - Exemplary TRX-Taq-TBD, N-terminal TRX with 60 residue linker, (pATG7346)
[0251] SEQ ID NO:36 - Exemplary TRX-Taq-TBD, N-terminal TRX with 60 residue linker, I677T (pATG7435)
[0252] SEQ ID NO:37 - Exemplary TRX-Taq-TBD, N-terminal TRX with 55 residue linker, (pATG7360)
[0253] SEQ ID NO:38 - Exemplary TRX-Taq-TBD, N-terminal TRX with 50 residue linker, (pATG7376)
[0254] SEQ ID NO:39 - Exemplary Taq-TBD-TRX, C-terminal TRX with 45 residue linker, (pATG7358)
[0255] SEQ ID NO:40 - Exemplary Taq-TBD-TRX, C-terminal TRX with 40 residue linker, (pATG7347)
[0256] SEQ ID NO:41 - Exemplary Taq-TBD-TRX, C-terminal TRX with 35 residue linker, (pATG7357)
[0257] SEQ ID NO:42 - Exemplary TRX-Taq-TBD-TRX, a doubly fused TRX with 60 and 40 residue linkers, (pATG7350)
[0258] SEQ ID NO:43 - Exemplary 2xTRX-Taq-TBD (pATG7427)
[0259] SEQ ID NO:44 - Exemplary 3xTRX-Taq-TBD (pATG7426)
[0260] TRX-Taq-TBD with SEQ ID NO: 45-60 residue linker and T7 TIS sequence at insertion site D
[0261] Exonuclease-deficient TRX-Taq-TBD with residue linker SEQ ID NO: 46-60 and T7 TIS sequence at insertion site D MSDKIIHLTDDSFTDDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLAGS SGGGGSGGGSGSSGGSGSSGGSSGGGGSGGGSGSSGGSGSSGSGSSGSSGGGGSGGSSMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLL ESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTDTDMGLLRSGKLPG KRFGSHALEAWGYRLGEMKGEYKDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRAL SLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVLFDELGL PAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDY SQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVET LFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0262] TRX-Taq-TBD with SEQ ID NO: 47-60 residue linker and T7 TIS truncation A at insertion site B
[0263] TRX-Taq-TBD with SEQ ID NO: 48-60 residue linker and T7 TIS truncation B at insertion site C
[0264] TRX-Taq-TBD with SEQ ID NO: 49-60 residue linker and T7 TIS truncation C at insertion site A
[0265] sequence number 50-Taq-TBD、pATG6979 MRGMLPLFEPKGRVLVDGHHLAYRTPHALKGLTTSRGEPVQAVYGFAKSLLKALKEDDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNGVLPKGIGEKTARKLLEEWGSLEALLKNLDRLK PAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERRLRAFLERLEFGSLLHEFGLLESPKALEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGPGPDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRYGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD SEQ ID NO:51 - AI-generated TRX sequence no. 1, pATG8280 MKKIKTVKTDEELEKILEESNKYTIVLFGAEWCGPCKMFLQENGETLEKISKEKGFEVLLVNIDQNPGTAPKYGIRGIPTVVVFKNGKKIATKVGALSKGQLLEFADAV SEQ ID NO:52 - AI-generated TRX sequence no. 2, pATG8279 MVKKVTEEELEKLIEEAKKKGEKIMVIFSAEWCGPCKMLLKELEEIEDELKALGIKEVLELNIDQNPGTAPKYGIRGIPTIMFIDANGIVYTKVGALSKGQILELAKAV SEQ ID NO:53 - AI-generated TRX sequence no. 3, pATG8281 MIKEVNGEELDKIIKEESPKRKIIIDFGAEWCGPCKMLKAELEKIAKELEEKYGYDIYLLNIDQNPGTAPKYGIRGIPTLIILTPSGKKLTKVGALSKGQILQLVKAA
[0266] SEQ ID NO:54 - TRX-90GS linker-Taq-TBD, pATG7580
[0267] SEQ ID NO:55-TRX(E31P)-82GS linker-Taq-TBD H914F, pATG7646
[0268] SEQ ID NO:56-TRX-82GS linker-Taq-TBD H914F, pATG7620
[0269] SEQ ID NO:57-TRX(E31P)-82GS linker-Taq-TBD E537V / S555N(H914F), pATG7860
[0270] SEQ ID NO:58 - TRX-TaqTBD without linker, pATG8388
[0271] SEQ ID NO:59 - TRX-5GS linker-TaqTBD, pATG8387 MSDKIIHLTDDSFTDTVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLASGGSS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0272] SEQ ID NO:60-TRX-10GS linker-TaqTBD, pATG8386 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA SGGGGSGGSSRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDL LGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLK NLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVH RAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEAT GVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLE RVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWL LVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGY VETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0273] SEQ ID NO:61 - TRX-15GS linker-TaqTBD, pATG8385
[0274] SEQ ID NO:62-TRX-20GS linker-TaqTBD, pATG8384 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSGSSGSSGGGGSGGSSRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQL ALIKELVDLLGLARLEVPGYEADDVLASLAKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEW GSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAA ARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLA HMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSR DQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEE GWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRR GYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0275] SEQ ID NO:63-TRX-25GS linker-TaqTBD, pATG8383 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSGSSGSGSGSSGGGGSGGSS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0276] SEQ ID NO:64 - TRX-30GS linker-TaqTBD, pATG8382 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSSGGSSGGGGSGSGSGSSSGGSSGSSGS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0277] SEQ ID NO:65-TRX-35GS linker-TaqTBD, pATG8381 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSSGGSSGGGGSGSGSGSSSGGSSGSSGSSGGSS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0278] SEQ ID NO:66 - TRX-40GS linker-TaqTBD, pATG8380 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSSGGSSGGGGSGSGSGSSSGGSGSSGSGSGSGSSGGSS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0279] SEQ ID NO:67 - TRX-73GS linker-TaqTBD, pATG7584 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSSGSSGGGSGGGSGSSGGSGSSGGSSGGGGSGSGSGSSSGGSGSSGSGSSSGSSGGGGSGGSSGSSGGSGS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0280] SEQ ID NO:68 - TRX-100GS linker-TaqTBD, pATG7585 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA
[0281] SEQ ID NO:69-TRX-110GS linker-TaqTBD, pATG8345 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSSGSSGGGSGGGSGSSGSSGSSGGGSGGGGSSGSGGSGGGSGSSGSSGSSGGGSGGGSGSGGGSGSGGGSGGGSGSGSSGGSGSSGSGSSGSSGGGGSGGSS RGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEV PGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKP AIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEP YKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSRDQLERVL FDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLV ALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYV ETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0282] SEQ ID NO:70-TRX-130GS linker-TaqTBD, pATG8351 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA
[0283] SEQ ID NO:71 - TRX-150GS linker-TaqTBD, pATG8346 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA
[0284] SEQ ID NO:72 - TRX-170GS linker-TaqTBD, pATG8347 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA
[0285] SEQ ID NO:73 - TRX-200GS linker-TaqTBD, pATG8348 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA
[0286] SEQ ID NO:74 - TRX-18GS(EAAAK46)18GS linker-TaqTBD[537V 555N](H914F), pATG8264
[0287] SEQ ID NO:75 - TRX-18GS (EAAAK repeats) 18GS linker-TaqTBD [537V 555N] (H914A), pATG8292
[0288] SEQ ID NO:76 - TRX-HaloTag linker (GS50) HaloTag linker-TaqTBD [537V 555N] (H914F), pATG8297 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA SLEPTTEDLYFQSDNDSGSSGSSGSGGSGGGSGGSGGSGGSGGSGGSGGSGSGSGSGSGSGSGSGSGSSGLPPDLPFLEPKGRVLLVDGHHLAYRTFHALKLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERALSER LFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLVEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0289] SEQ ID NO: 77-TRX-GSAT linker (82)-TaqTBD [537V 555N] (H914F), pATG8285 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSTAAGAAATAAGSTGAAGAATAAGGSAGGTGSGSATGSSGASGTGTAGGTGAGSGTGSGAAGAATAAGTSGAAGAATAATSGRGMLPLFEPKGRVLLVDGHHLAYRTFHALKLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSER LFANLWGRLEGEERLLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGGSWYQPCGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREREPCELDTREYVAGAPYPTPVEHVFNPSSRDQLERVLFDELGLPAIGKTEKTRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVR TPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEEVGIGEDWLSAKGD
[0290] SEQ ID NO: 78-TRX(E31P)-24GS linker (new sequence 1) 24GS linker-TaqSTBD[537V 555N]H914F, pATG8298 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSGSSGGGSGGGSGGSGSGSGGCASSIDYKRISRMPSKIMDAVIDTLNICKLANCESGGGSGSGSGSGSSGGGSGGSSRGMSRGMLPLFEPKGRVLLVDGHHLAYRTFHALKLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERALSER LFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLVEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0291] SEQ ID NO: 79-TRX(E31P)-24GS linker (new sequence 2) 24GS linker-TaqSTBD[537V 555N]H914F, pATG 8286 MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA GSSGSGSSGGGSGGGSGGSGSGSGGCASSIDYKRISRMPAVLADAVIDTLNICKLANCESGGGSGSGSGSGSGSSGGGSGGSSRGMSRGMLPLFPEKGRVLLVDGHHLAYRTFHALKLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERALSER LFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLVEVAEEIARLEAEVFRLAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0292] SEQ ID NO:80 - TaqTBD-55GS linker-TRX, pATG7589
[0293] SEQ ID NO:81 - TaqTBD-55GS linker-TRX, pATG7588
[0294] SEQ ID NO:82 - TRX-45GS linker-TaqTBD, pATG7375
[0295] SEQ ID NO:83 - TaqSTBD-30GS linker-TRX, pATG7389
[0296] SEQ ID NO:84 - TaqTBD-25GS linker-TRX, pATG7428
[0297] SEQ ID NO:85 - Alishewanella jeotgali TRX-60GS linker-Taq-TBD, pATG8151
[0298] SEQ ID NO:86 - Alishewanella jeotgali TRX-82GS linker-Taq-Klebsiella phage Kp_GWPB35 TBD, pATG8258
[0299] SEQ ID NO:87 - Alishewanella jeotgali TRX-82GS linker-Taq- Salmonella enterica phage TBD, pATG8261
[0300] SEQ ID NO:88 - Alishewanella jeotgali TRX-82GS linker-Taq-Aeromonas phage PZL-Ah152 TBD, pATG8262
[0301] SEQ ID NO:89 - Thiococcus pfennigii TRX-60GS-Taq-TBD, pATG8156
[0302] SEQ ID NO:90 - Thiococcus pfennigii TRX-65GS-Taq-Klebsiella phage Kp_GWPB35 TBD, pATG8259
[0303] SEQ ID NO:91 - Thiococcus pfennigii TRX-65GS-Taq-Salmonella enterica phage TBD, pATG8260
[0304] SEQ ID NO: 92 - Thiococcus pfennigii TRX-82GS-Taq-Aeromonas phage PZL-Ah152 TBD, pATG8263
[0305] SEQ ID NO:93 - Thiococcus pfennigii TRX, pATG8291 MSDSIVHVTDDSFERDVLQSSEPVLVDYWADWCGPCKMIAPVLDEIATEYAGRIRVAKLNIDENPNTPPRYGIRGIPTLMLFKDGEVEATKVGAVSKSQLTAFIDSNL SEQ ID NO:94 - Alishewanella jeotgali TRX, pATG8290 MSEHILQVSDDSFETDVLKAEAPVLVDFWAEWCGPCKMIAPILDDVAAEYAGKVTVAKVNIDQNPNTPPKFGIRGIPTLLLFKNGQVAATKVGALSKTQLKQFLDSNI
[0306] SEQ ID NO:95-TRX-60GS linker-Taq-Salmonella enterica phage TBD, pATG7606
[0307] SEQ ID NO:96 - TRX-60GS linker-Taq-Aeromonas phage PZL-Ah152 TBD, pATG7610
[0308] SEQ ID NO:97 - TRX-60GS linker-Taq-Klebsiella phage Kp_GWPB35 TBD, pATG7602
[0309] sequence number 98-Taq-Salmonella entericaファージTBD、pATG8303 MRGMLPLFEPKGRVLVDGHHLAYRTPHALKGLTTSRGEPVQAVYGFAKSLLKALKEDDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNGVLPKGIGEKTARKLLEEWGSLEALLKNLDRLK PAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERRLRAFLERLEFGSLLHEFGLLESPKALEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGPGPDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGSWYAAPKGKEFFRHPRTGKDLPKYPRVVYPKVGGIFKKPKNKAQRLGLEPCERDTRDTMEGAPFTPITYVEFNPGSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0310] Array number 99 - Taq - Aeromonas phage PZL - Ah152 TBD H914H, pATG8167 MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKGLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRLDVAYLRALSLEVAEEIARLEAEVFRLAGSSWYRPKGGKAFFRHPVTGKDLTNYPRVIYPKAGEIYTKGGKLAKTLYCKDRPFTPIEYTVFNPGSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRGYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD
[0311] Sequence number 100-Taq-Klebsiella phage Kp_GWPB35 TBD, pATG8304 MRGMLPLFEPKGRVLLVDGHHLAYRTFHALKLTTSRGEPVQAVYGFAKSLLKALKEDGDAVIVVFDAKAPSFRHEAYGGYKAGRAPTPEDFPRQLALIKELVDLLGLARLEVPGYEADDVLASLAKKAEKEGYEVRILTADKDLYQLLSDRIHVLHPEGYLITPAWLWEKYGLRPDQWADYRALTGDESDNLPGVKGIGEKTARKLLEEWGSLEALLKNLDRLKPAIREKILAHMDDLKLSWDLAKVRTDLPLEVDFAKRREPDRERLRAFLERLEFGSLLHEFGLLESPKALEEAPWPPPEGAFVGFVLSRKEPMWADLLALAAARGGRVHRAPEPYKALRDLKEARGLLAKDLSVLALREGLGLPPGDDPMLLAYLLDPSNTTPEGVARRYGGEWTEEAGERAALSERLFANLWGRLEGEERLLWLYREVERPLSAVLAHMEATGVRL DVAYLRALSLEVAEEIARLEAEVFRLAGGTWYQPKGGTELFLHPRTGKPLGKYPRVKYPKQGGIYKKPKNKAQREGREPCDLDTRDYVEGAPYTPVEHVVFNPSRDQLERVLFDELGLPAIGKTEKTGKRSTSAAVLEALREAHPIVEKILQYRELTKLKSTYIDPLPDLIHPRTGRLHTRFNQTATATGRLSSDPNLQNIPVRTPLGQRIRRAFIAEEGWLLVALDYSQIELRVLAHLSGDENLIRVFQEGRDIHTETASWMFGVPREAVDPLMRRAAKTINFGVLYGMSAHRLSQELAIPYEEAQAFIERYFQSFPKVRAWIEKTLEEGRRRYVETLFGRRRYVPDLEARVKSVREAAERMAFNMPVQGTAADLMKLAMVKLFPRLEEMGARMLLQVHDELVLEAPKERAEAVARLAKEVMEGVYPLAVPLEVEVGIGEDWLSAKGD SEQ ID NO: 101 - Salmonella enterica phage TBD sequence GSWYAPKGGKEFFRHPRTGKDLPKYPRVVYPKVGGIFKKPKNKAQRLGLEPCERDTTRDTMEGAPFTPITYVEFNPG SEQ ID NO: 102 - Aeromonas phage PZL-Ah152 TBD sequence SSWYRPKGGKAFFRHPVTGKDLTNYPRVIYPKAGEIYTKGGKLAKTLYCKDRPFFTIEYTVFNPG SEQ ID NO:103 - Klebsiella phage Kp_GWPB35 TBD sequence GTWYQPKGGTELFLHPRTGKPLGKYPRVKYPKQGGIYKKPKNKAQREGREPCDLDTRDYVEGAPYTPVEHVVFNPS
[0312] SEQ ID NO:104 - Tne polymerase (v2), pATG8293 MKELQLYEEAEPTGYEIVKDHKTFEDLIEKLKEVPSFALALETSSLDPFNCEIVGISVSFKPKTAYYIPLHHRNAQNLDETLVLSKLKEILEDPSSKIVGQNLKYAYKVLMVKGISPVYPHFDTMIAAYLLEPNEKKFNLEDLSLKFLGYKMTSYQELMSFSSPLFGFSFADVPVDKAANYSCEDADITYRLYKILSMKLHEAELENVFYRIEMPLVNVLARMELNGVYVDTEFLKKLSEEYGKKLEELAEKIYQIAGEPFNINSPKQVSKILFEKLGIKPRGKTTKTGAYSTRIEVLEEIANEHEIVPLILEYRKIQKLKSTYIDTLPKLVNPKTGRIHASFHQTGTATGRLSSSDPNLQNLPTKSEEGKEIRKAIVPQDPDWWIVSADYSQIELRILAHLSGDENLVKAFEEGIDVHTLTASRIYNVKPEEVNEEMRRVGKMVNYSIIYGVTPYGLSVRLGIPVKEAEKMIISYFTLYPKVRSYIQQVVAEAKEKGYVRTLFGRKRDIPQLMARDKNTQSEGERIAINTPIQGTAADIIKLAMIDIDEELRKRNMKSRMIIQVHDELVFEVPDEEKEELVDLVKNKMTNVVKLSVPLEVDISIGKSWS Array No. 105 - Tne polymerase (v2) - TBD chimera, pATG8294 MKELQLYEEAEPTGYEIVKDHKTFEDLIEKLKEVPSFALALETSSLDPFNCEIVGISVSFKPKTAYYIPLHHRNAQNLDETLVLSKLKEILEDPSSKIVGQNLKYAYKVLMVKGISPVYPHFDTMIAAYLLEPNEKKFNLEDLSLKFLGYKMTSYQELMSFSSPLFGFSFADVPVDKAANYSCEDADITYRLYKILSMKLHEAELENVFYRIEMPLVNVLARMELNGVYVDTEFLKKLSEEYGKKLEELAEKIYQIAGGSWYQPKGGTEMFCHPRTGKPLPKYPRIKIPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVFNPSSPKQVSKILFEKLGIKPRGKTTKTGAYSTRIEVLEEIANEHEIVPLILEYRKIQKLKSTYIDTLPKLVNPKTGRIHASFHQTGTATGRLSSSDPNLQNLPTKSEEGKEIRKAIVPQDPDWWIVSADYSQIELRILAHLSGDENLVKAFEEGIDVHTLTASRIYNVKPEEVNEEMRRVGKMVNYSIIYGVTPYGLSVRLGIPVKEAEKMIISYFTLYPKVRSYIQQVVAEAKEKGYVRTLFGRKRDIPQLMARDKNTQSEGERIAINTPIQGTAADIIKLAMIDIDEELRKRNMKSRMIIQVHDELVFEVPDEEKEELVDLVKNKMTNVVKLSVPLEVDISIGKSWS
[0313] Array No. 106 - TRX - 60GS linker - Tfl - Taq chimera - TBD, pATG8266 SEQ ID NO:107 - TRX E31P, pATG8143 MSDKIIHLTDDSFDTDVLKADGAILVDFWAPWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDANLA
[0314] SEQ ID NO: 108-TRX E31P-82GS linker-Taq-TBD[E537V / S555N]H914A, pATG8100
[0315] SEQ ID NO:109 - TRX V56I / A88V-60GS linker-Taq-TBD, pATG7633
[0316] SEQ ID NO: 110-TRX S2M E31P-82GS linker-TaqTBD[537V, 555N]+H914F, pATG7932
[0317] SEQ ID NO: 111-TRXA D27E E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG7933
[0318] SEQ ID NO: 112-TRXA S2M D27E E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG7934
[0319] SEQ ID NO: 113-TRX E31P-82GS linker-Taq-TBD[T535K, E537V]+H914F, pATG7840
[0320] SEQ ID NO: 114-TRX E31P-82GS linker-Taq-TBD[T535N, S555N]+H914F, pATG7864
[0321] SEQ ID NO: 115-TRX E31P-82GS linker-Taq-TBD [T535N, E537V, S555N] + H914F, pATG7865
[0322] SEQ ID NO: 116 - TRX S2M E31P-82GS linker-Taq-TBD+H914F, pATG7866
[0323] SEQ ID NO: 117 - TRX S2M E31P D48E-82GS linker-Taq-TBD+H914F, pATG7867
[0324] SEQ ID NO: 118 - TRX S2E E31P D48E-82GS linker-Taq-TBD+H914F, pATG7868
[0325] SEQ ID NO: 119-TRX S2M E31P-82GS linker-Taq-TBD[E537V]+H914F, pATG7869
[0326] SEQ ID NO: 120-TRX S2M E31P D48E-82GS linker-Taq-TBD[E537V]+H914F, pATG7870
[0327] SEQ ID NO: 121-TRX S2E E31P D48E-82GS linker-Taq-TBD[E537V]+H914F, pATG7871
[0328] SEQ ID NO: 122-TRX[D27E E31P]-82GS linker-Taq-TBD+H914F, pATG7895
[0329] SEQ ID NO: 123-TRX D27E E31P-82GS linker-Taq-TBD[537V]+H914F, pATG7896
[0330] SEQ ID NO: 124 - TRX S2E E31P-82GS linker-Taq-TBD+H914F, pATG7898
[0331] SEQ ID NO: 125-TRX S2E E31P-82GS linker-Taq-TBD[537V]+H914F, pATG7899
[0332] SEQ ID NO: 126-TRX[E31P D48E]-82GS linker-Taq-TBD+H914F, pATG7900
[0333] SEQ ID NO: 127-TRX E31P D48E-82GS linker-Taq-TBD[537V]+H914F, pATG7901
[0334] SEQ ID NO: 128 - TRX E31P P65Q-82GS linker-Taq-TBD+H914F, pATG7902
[0335] SEQ ID NO: 129-TRX E31P-82GS linker-Taq-TBD [E537T, S555N] + H914F, pATG7903
[0336] SEQ ID NO: 130-TRX E31P-82GS linker-Taq-TBD[E537S, S555R]+H914F, pATG7904
[0337] SEQ ID NO: 131-TRX E31P-82GS linker-Taq-TBD[E537S, S555N]+H914F, pATG7905
[0338] SEQ ID NO: 132-TRX E31P-82GS linker-Taq-TBD [E537A, S555N] + H914F, pATG7906
[0339] SEQ ID NO: 133-TRX E31P-82GS linker-Taq-TBD [T535N, E537S, S555N] + H914F, pATG7907
[0340] SEQ ID NO: 134-TRX E31P-82GS linker-Taq-TBD [E537T, S555R] + H914F, pATG7908
[0341] SEQ ID NO: 135-TRX E31P-82GS linker-Taq-TBD[E537A, S555R]+H914F, pATG7909
[0342] SEQ ID NO: 136-TRX E31P-82GS linker-Taq-TBD [T535N, E537S, S555R] + H914F, pATG7910
[0343] SEQ ID NO: 137-TRX E31P-82GS linker-Taq-TBD [T535N, E537A, S555R] + H914F, pATG7911
[0344] SEQ ID NO: 138-TRX E31P-82GS linker-Taq-TBD[535N, 537A, 555N]+H914F, pATG7937
[0345] SEQ ID NO: 139-TRX E31P-82GS linker-Taq-TBD[535N, 537T, 555N]+H914F, pATG7938
[0346] SEQ ID NO: 140-TRX E31P-82GS linker-Taq-TBD[535N, 537T, 555R]+H914F, pATG7939
[0347] SEQ ID NO: 141-TRX E31P-82GS linker-Taq-TBD[535N, 537V, 555N]+H914F, pATG7940
[0348] SEQ ID NO: 142-TRX E31P-82GS linker-Taq-TBD[535N, 537V, 555R]+H914F, pATG7941
[0349] SEQ ID NO: 143-TRX[D48E, P65Q]E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG8041
[0350] SEQ ID NO: 144-TRX[S2M, D48E, P65Q]E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG8042
[0351] SEQ ID NO: 145-TRX[D27E, D48E, P65Q]E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG8043
[0352] SEQ ID NO: 146-TRX[S2M, D27E, D48E, P65Q]E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG8044
[0353] SEQ ID NO: 147-TRX[S2E, D48E, P65Q]E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG8045
[0354] SEQ ID NO: 148-TRX[S2E, D27E, D48E, P65Q]E31P-82GS linker-Taq-TBD[537V, 555N]+H914F, pATG8046
[0355] SEQ ID NO: 149-TRX E31P-82GS linker-Taq-TBD[537V / S555N], pATG8168
[0356] SEQ ID NO: 150-TRX-82GS linker-Taq-TBD[537V / S555N]+H914F, pATG8170
[0357] SEQ ID NO:151-TRX[E31P]-60GS linker-Taq-TBD, pATG7626
[0358] SEQ ID NO:152-TRX[E31P]-82GS linker-Taq-TBD, pATG7642
[0359] SEQ ID NO: 153-TRX[E31P]-82GS linker-Taq-TBD[S555K]+H914F, pATG7697
[0360] SEQ ID NO: 154-TRX[E31P]-82GS linker-Taq-TBD[S555N]+H914F, pATG7698
[0361] SEQ ID NO: 155-TRX[E31P]-82GS linker-Taq-TBD[E537K]+H914F, pATG7714
[0362] SEQ ID NO: 156-TRX[E31P]-82GS linker-Taq-TBD[E537L]+H914F, pATG7699
[0363] SEQ ID NO: 157-TRX[E31P]-82GS linker-Taq-TBD[T489K]+H914F, pATG7712
[0364] SEQ ID NO: 158-TRX[E31P]-82GS linker-Taq-TBD[E537V]+H914F, pATG7713
[0365] SEQ ID NO: 159-TRX[E31P]-82GS linker-Taq-TBD[E548K]+H914F, pATG7717
[0366] SEQ ID NO: 160-TRX[E31P]-82GS linker-Taq-TBD[T535K]+H914F, pATG7718
[0367] SEQ ID NO: 161-TRX[E31P]-82GS linker-Taq-TBD[R506K]+H914F, pATG7743
[0368] SEQ ID NO: 162-TRX[E31P]-82GS linker-Taq-TBD[S555R]+H914F, pATG7853
[0369] SEQ ID NO: 163-TRX L104I-82GS linker-Taq-TBD, pATG7643
[0370] SEQ ID NO: 164-TRX[L104I]-73GS linker-Taq-TBD, pATG7644
[0371] SEQ ID NO: 165-TRX[E31P]-73GS linker-Taq-TBD, pATG7645
[0372] SEQ ID NO: 166-TRX E31P-82GS linker-Taq-TBD[537V / S555N]+H914F, T7 clamp, pATG8130
[0373] SEQ ID NO: 167-TRX-60GS linker-Taq-TBD+H914F, pATG7569
[0374] SEQ ID NO:168-TRX-77GS linker-Taq-TBD, pATG7640
[0375] SEQ ID NO:169-TRX-90GS linker-Taq-TBD H914F, pATG7622
[0376] SEQ ID NO: 170-TRX E31P-60GS linker-Taq-TBD[537V / S555N]+H914F, pATG8169
[0377] SEQ ID NO: 171-TRX E31P-82GS linker-Taq-TBD[T535K, S555N]+H914F, pATG7863
[0378] SEQ ID NO: 172-TRX E31P-82GS linker-Taq-TBD [T535K, E537V, S555N] + H914F, pATG7841
[0379] SEQ ID NO:173 - TRX-60GS linker-TaqS alternative TBD#1, pATG7567
[0380] SEQ ID NO: 174-TRX-60GS linker-Taq-Alt_TBD_5(GGSWY_TBD_FNLN), pATG7682
[0381] SEQ ID NO: 175-TRX-60GS linker-Taq-Alt_TBD_6(GSWY_TBD_FNPS), pATG7683
[0382] SEQ ID NO:176-TRX-82GS linker-Taq-TBD, pATG7359
[0383] SEQ ID NO: 177-TRX-60GS linker-Taq-TBD+D816N E980A, pATG8332
[0384] SEQ ID NO: 178-TRX-60GS linker-Taq-TBD+D816N V869A, pATG8333
[0385] SEQ ID NO: 179-TRX-60GS linker-Taq-TBD+E980A V869A, pATG8334
[0386] SEQ ID NO: 180-TRX-60GS linker-Taq-TBD+D816N H914A, pATG8357
[0387] SEQ ID NO: 181-TRX-60GS linker-Taq-TBD+E980A H914A, pATG8358
[0388] SEQ ID NO: 182-TRX-60GS linker-Taq-TBD+V869A H914A, pATG8359
[0389] SEQ ID NO: 183-TRX-60GS linker-Taq-TBD+D816N V869A H914A, pATG8360
[0390] SEQ ID NO: 184-TRX-60GS linker-Taq-TBD+D816N E980A H914A, pATG8361
[0391] SEQ ID NO: 185-TRX-60GS linker-Taq-TBD+D816N E980A V869A H914A, pATG8362
[0392] SEQ ID NO: 186-TRX-60GS linker-Taq-TBD+E980A V869A H914A, pATG8363
[0393] SEQ ID NO: 187 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H936A V891A, pATG8320
[0394] SEQ ID NO: 188 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H936A D838N, pATG8321
[0395] SEQ ID NO: 189 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H914A E1002A, pATG8322
[0396] SEQ ID NO: 190 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H936A D838N E1002A, pATG8329
[0397] SEQ ID NO: 191 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H936A D838N V891A, pATG8330
[0398] SEQ ID NO: 192 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H936A E1002A V891A, pATG8331
[0399] SEQ ID NO: 193 - TRX E31P-18GS (EAAAK repeats) 18GS linker-Taq TBD [537V 555N] + H936A D838N E1002A V891A, pATG8339
[0400] SEQ ID NO:194 - Tne polymerase (v1) - Aeromonas TBD, pATG8302 MARLFLFDGTALAYRAYYALDRSLSTSTGIPTNAVYGVARMLVKFIKEHIIPEKDYAAVAFDKKAATFRHKLLEAYKAQRPKTPDLLVQQLPYIKRLIEALGFKVLELEGYEADDIIATLAVKGCTFFDEIFIITGDKDMLQLVNEKIKVWRIVKGISDLELYDSKKVKERYGVEPHQIPDLLALTGDEIDNIPGVTGIGEKTAVQLLGKYRNLEDILEHARELPQRVRKALLRDREVAILSKKLATLVTNAPVEVDWEEMKYRGYDKRKLLPILKELEFASIMKELQLYEEAEPTGYEIVKDHKTFEDLIEKLKEVPSFALDLETSSLDPFNCEIVGISVSFKPKTAYYIPLHHRNAQNLDETLVLSKLKEILEDPSSKIVGQNLKYDYKVLMVKGISPVYPHFDTMIAAYLLEPNEKKFNLEDLSLKFLGYKMTSYQELMSFSSPLFGFSFADVPVDKAANYSCEDADITYRLY KILSMKLHEAELENVFYRIEMPLVNVLARMELNGVYVDTEFLKKLSEEYGKKLEELAEKIYQIAGSSWYRPKGGKAFFRHPVTGKDLTNYPRVIYPKAGEIYTKGGKLAKTLYCKDRPFTPPIEYTVFNPGSPKQVSKILFEKLGIKPRGKTTKTGEYSTRIEVLEEIANEHEIVPLILEYRKIQKLKSTYIDTLPKLVNPKTGRIHASFHQTGTATGRLSSSDPNQNLPTKSEEGKEIRKAIVPQDPDWWIVSADYSQIELRILAHLSGDENLVKAFEEGIDVHTLTASRIYNVKPEEVNEEMRRVGKMVNFSIIYGVTPYGLSVRLGIPVKEAEKMIISYFTLYPKVRSYIQQVVAEAKEKGYVRTLFGRKRDPQLMARDKNTQSEGERIAINTPIQGTAADIIKLAMIDIDEELRKRNMKSRMIIQVHDELVFEVPDEEKEEKELVDLVKNKMTNVVKLSVPLEVDISIGKSWS
[0401] SEQ ID NO:195 - Tne polymerase (v1) - Klebsiella TBD, pATG8353 。
[0402] SEQ ID NO:196 - Tne polymerase (v1) - Salmonella TBD, pATG8355 。
[0403] SEQ ID NO:197-Tneポリメラズ(v2)-Klebsiella TBD、pATG8352 。 SEQ ID NO:198-Tneポリメラズ(v2)-Salmonella TBD、pATG8354 。
Claims
1. (a) a DNA polymerase domain; (b) a thioredoxin binding domain (TBD), and (c) A DNA polymerase system containing a thioredoxin (TRX) domain.
2. 2. The composition of claim 1, wherein the DNA polymerase system is capable of synthesizing a DNA product from a deoxynucleoside triphosphate in the presence of a template DNA under appropriate reaction conditions.
3. 2. The composition of claim 1, wherein the DNA polymerase system exhibits a reduced tendency to stutter compared to a DNA polymerase comprising the DNA polymerase domain in the absence of the TBD and / or TRX.
4. 4. The composition of claim 3, wherein the DNA polymerase system exhibits at least a two-fold reduction in stuttering tendency compared to a DNA polymerase comprising the DNA polymerase domain in the absence of the TBD and / or TRX.
5. The composition of claim 1 , wherein the DNA polymerase system comprises the DNA polymerase domain conjugated to the TBD and / or TRX.
6. The composition of claim 1 , wherein the DNA polymerase system comprises the DNA polymerase domain genetically fused to the TBD and / or TRX.
7. The composition of claim 6 , wherein the DNA polymerase system comprises a genetic fusion of the DNA polymerase domain, a TBD, and a TRX.
8. 10. The composition of claim 1, wherein one or more of the DNA polymerase domain, TBD, and TRX are not conjugated to other components of the system.
9. 9. The composition of claim 8, comprising free TRX and a DNA polymerase domain conjugated or genetically fused to a TBD.
10. 9. The composition of claim 8, comprising a free TBD and a DNA polymerase domain conjugated or genetically fused to TRX.
11. 9. The composition of claim 8, comprising a free DNA polymerase domain and a TRX conjugated or genetically fused to a TBD.
12. 9. The composition of claim 8, comprising a DNA polymerase domain, a TRX, and a TBD conjugated or genetically fused to each other.
13. 1. A composition comprising a chimeric DNA polymerase with reduced stuttering tendency, said chimeric DNA polymerase comprising: (a) a DNA polymerase domain; (b) a thioredoxin binding domain (TBD), and (c) the composition comprising a genetic fusion of a thioredoxin (TRX) domain.
14. The composition of claim 13 , wherein the DNA polymerase domain is thermophilic.
15. The composition of claim 14 , wherein the DNA polymerase domain is derived from a naturally occurring thermophilic DNA polymerase.
16. 16. The composition of claim 15, wherein the naturally occurring thermophilic DNA polymerase is selected from the group consisting of Thermus aquaticus DNA polymerase, Thermus thermophilus DNA polymerase, Thermus flavus DNA polymerase, Thermotoga neapolitana polymerase, and Geobacillus stearothermophilus DNA polymerase.
17. The composition of claim 13 , wherein the DNA polymerase domain is derived from a Family A DNA polymerase.
18. 14. The composition of claim 13, wherein the DNA polymerase domain comprises at least 40% sequence identity with SEQ ID NO:
1.
19. The composition of claim 1 , wherein the DNA polymerase domain further comprises an internal amino acid sequence insertion.
20. 20. The composition of claim 19, wherein the DNA polymerase domain comprises an N-terminal portion having 40% sequence identity to SEQ ID NO: 14 and a C-terminal portion having 40% sequence identity to SEQ ID NO: 12, wherein the N-terminal portion and the C-terminal portion are separated by the internal amino acid sequence insertion.
21. 21. The composition of claim 20, wherein the internal amino acid sequence insertion comprises the TBD.
22. 22. The composition of claim 21, wherein the TBD is derived from the thioredoxin binding domain of a T3 or T7 bacteriophage DNA polymerase.
23. 23. The composition of claim 22, wherein the TBD comprises at least 50% sequence identity with SEQ ID NO:
15.
24. 22. The composition of claim 21, wherein the TBD is derived from the thioredoxin binding domain of a Klebsiella pneumoniae, Salmonella enterica, or Aeromonas hydrophila phage DNA polymerase.
25. 25. The composition of claim 24, wherein the TBD comprises at least 50% sequence identity with one of SEQ ID NOs: 101-103.
26. 14. The composition of claim 13, wherein the TBD sequence is internal to the DNA polymerase domain sequence.
27. 27. The composition of claim 26, wherein the TBD sequence is within and / or replaces all or part of the thumb domain of the DNA polymerase domain sequence.
28. 14. The composition of claim 13, wherein the TRX domain is derived from Escherichia coli thioredoxin.
29. 29. The composition of claim 28, wherein the TRX domain comprises at least 50% sequence identity with SEQ ID NO: 16, 17, or 107.
30. 14. The composition of claim 13, wherein the TRX domain is derived from Alishewanella jeotgalli or Thiococcus pfennigii thioredoxin.
31. 31. The composition of claim 30, wherein the TRX domain comprises at least 50% sequence identity with one of SEQ ID NOs: 93 or 94.
32. The composition of claim 13, wherein the TRX domain comprises at least 50% sequence identity with one of SEQ ID NOs: 51-53.
33. The composition of claim 13 , wherein the TRX sequence is fused to the N- or C-terminus of the DNA polymerase domain.
34. 34. The composition of claim 33, wherein the TRX sequence is fused to the DNA polymerase domain by a linker peptide or polypeptide.
35. 35. The composition of claim 34, wherein the linker peptide or polypeptide is 1-300 amino acids in length.
36. The composition of claim 35, wherein the linker peptide or polypeptide is 30-70 amino acids in length.
37. 35. The composition of claim 34, wherein the linker is a flexible linker.
38. 38. The composition of claim 37, wherein 50-100% of the flexible linker is glycine and serine residues.
39. The composition of claim 34 , wherein the linker comprises a rigid segment.
40. 40. The composition of claim 39, wherein the rigid segment comprises one or more EAAAK peptide segments.
41. 14. The composition of claim 13, comprising a sequence having at least 60% sequence identity to one of SEQ ID NOs: 22-27 and 35-44.
42. A composition comprising a DNA polymerase domain conjugated to a thioredoxin binding domain (TBD).
43. 43. The composition of claim 42, wherein the DNA polymerase domain is genetically fused to the thioredoxin binding domain (TBD).
44. 43. The composition of claim 42, wherein the DNA polymerase domain is derived from a Family A DNA polymerase.
45. 43. The composition of claim 42, wherein the DNA polymerase domain is thermophilic.
46. 46. The composition of claim 45, wherein the DNA polymerase domain is derived from a naturally occurring thermophilic DNA polymerase.
47. 43. The composition of claim 42, wherein the naturally occurring thermophilic DNA polymerase is selected from the group consisting of Thermus aquaticus DNA polymerase, Thermus thermophilus DNA polymerase, Thermus flavus DNA polymerase, Thermotoga neapolitana polymerase, and Geobacillus stearothermophilus DNA polymerase.
48. 43. The composition of claim 42, wherein the DNA polymerase domain comprises at least 40% sequence identity with SEQ ID NO:
1.
49. 43. The composition of claim 42, wherein the TBD sequence is internal to the DNA polymerase domain sequence.
50. 50. The composition of claim 49, wherein the TBD sequence is within and / or replaces all or part of the thumb domain of the DNA polymerase domain sequence.
51. 50. The composition of claim 49, wherein the DNA polymerase domain comprises an N-terminal portion having 40% sequence identity to SEQ ID NO: 14 and a C-terminal portion having 40% sequence identity to SEQ ID NO: 12, wherein the N-terminal portion and the C-terminal portion are separated by the TBD.
52. 52. The composition of claim 51, wherein the TBD is derived from the thioredoxin binding domain of T3 or T7 bacteriophage DNA polymerase.
53. 53. The composition of claim 52, wherein the TBD comprises at least 50% sequence identity with SEQ ID NO:
15.
54. 52. The composition of claim 51, wherein the TBD is derived from the thioredoxin binding domain of a Klebsiella pneumoniae, Salmonella enterica, or Aeromonas hydrophila phage DNA polymerase.
55. 43. The composition of claim 42, wherein the TBD comprises at least 50% sequence identity with one of SEQ ID NOs: 101-103.
56. 43. The composition of claim 42, further comprising thioredoxin (TRX), wherein the thioredoxin is present in the composition in an amount not greater than 800 molar excess relative to the TBD.
57. 57. The composition of claim 56, wherein the TRX is derived from Escherichia coli thioredoxin.
58. 58. The composition of claim 57, wherein the TRX comprises at least 50% sequence identity with SEQ ID NO: 16, 17, or 107.
59. 57. The composition of claim 56, wherein the TRX is derived from Alishewanella jeotgalli or Thiococcus pfennigii thioredoxin.
60. 60. The composition of claim 59, wherein the TRX comprises at least 50% sequence identity with one of SEQ ID NOs: 93 or 94.
61. 57. The composition of claim 56, wherein the TRX comprises at least 50% sequence identity with SEQ ID NOs: 51-53.
62. 57. The composition of claim 56, wherein the TRX is a fusion with an additional polypeptide sequence.
63. 63. The composition of claim 62, wherein the additional polypeptide sequence is a DNA binding protein, an amino acid sequence capable of binding to DNA, a protein associated with a DNA replication site, a TBD, and / or a DNA polymerase.
64. 63. The composition of claim 62, wherein the additional polypeptide sequence is fused to the TRX by a linker peptide or polypeptide.
65. 65. The composition of claim 64, wherein the linker peptide or polypeptide is 1-300 amino acids in length.
66. 66. The composition of claim 65, wherein the linker peptide or polypeptide is 30-70 amino acids in length.
67. 65. The composition of claim 64, wherein 50-100% of the linker peptide or polypeptide are glycine and serine residues.
68. 57. The composition of claim 56, wherein the TRX is present in the composition in less than a 50 molar excess relative to the TBD.
69. 63. The composition of claim 62, wherein the fusion protein comprises a sequence having at least 50% sequence identity to one of SEQ ID NOs: 28-34.
70. comprising a DNA polymerase domain corresponding to SEQ ID NO: 1; and (a) a segment having at least 40% sequence identity to one of SEQ ID NOs: 2, 4, 6, 8, 10, and 12; (b) A DNA polymerase comprising (i) a segment having at least 40% sequence identity to SEQ ID NOs: 3, 5, 7, 9, and 11, or (ii) a segment in which all or a portion of the sequence in SEQ ID NO: 1 corresponding to one or more of SEQ ID NOs: 3, 5, 7, 9, and 11 is replaced with a heterologous sequence selected from a TBD, a TRX, and a TBD / TRX interacting sequence (TIS).
71. 71. The DNA polymerase of claim 70, comprising a TBD having at least 50% sequence identity to SEQ ID NO:
15.
72. 72. The DNA polymerase of claim 71, wherein the TBD is inserted into and / or replaces all or part of one of SEQ ID NOs: 3, 5, 7, 9, and 11 located at the C-terminus, N-terminus, or SEQ ID NO: 3, 5, 7, 9, and 11.
73. 71. The DNA polymerase of claim 70, comprising a TRX having at least 50% sequence identity to SEQ ID NO: 16, 17, or 107.
74. 74. The DNA polymerase of claim 73, wherein the TRX is inserted into one of SEQ ID NOs: 3, 5, 7, 9, and 11 located at the C-terminus, N-terminus, and / or is substituted for all or part of one of SEQ ID NOs: 3, 5, 7, 9, and 11.
75. 71. The DNA polymerase of claim 70, comprising a TIS having at least 40% sequence identity to one of SEQ ID NOs: 18-21.
76. 76. The DNA polymerase of claim 75, wherein the TIS is inserted into one of SEQ ID NOs: 3, 5, 7, 9, and 11 located at the C-terminus, N-terminus, and / or is substituted for all or part of one of SEQ ID NOs: 3, 5, 7, 9, and 11.
77. 71. The DNA polymerase of claim 70, wherein all or part of the exonuclease domain of SEQ ID NO: 13 is removed from the sequence corresponding to SEQ ID NO:
1.
78. below: (a) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(SEQ ID NO:7)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12), (b) (SEQ ID NO: 2) - (one of SEQ ID NOs: 18-21) - (SEQ ID NO: 4) - (SEQ ID NO: 5) - (SEQ ID NO: 6) - (SEQ ID NO: 7) - (SEQ ID NO: 8) - (SEQ ID NO: 9) - (SEQ ID NO: 10) - (SEQ ID NO: 15) - (SEQ ID NO: 12); (c) (SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (d) (SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (e) (SEQ ID NO: 2) - (SEQ ID NO: 3) - (SEQ ID NO: 4) - (SEQ ID NO: 5) - (SEQ ID NO: 6) - (SEQ ID NO: 7) - (SEQ ID NO: 8) - (one of SEQ ID NOs: 18-21) - (SEQ ID NO: 10) - (SEQ ID NO: 15) - (SEQ ID NO: 12); (f) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(SEQ ID NO:7)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107), (g) (SEQ ID NO:2) - (one of SEQ ID NOs:18-21) - (SEQ ID NO:4) - (SEQ ID NO:5) - (SEQ ID NO:6) - (SEQ ID NO:7) - (SEQ ID NO:8) - (SEQ ID NO:9) - (SEQ ID NO:10) - (SEQ ID NO:15) - (SEQ ID NO:12) - (SEQ ID NO:16, 17, or 107); (h) (SEQ ID NO: 2) - (SEQ ID NO: 3) - (SEQ ID NO: 4) - (one of SEQ ID NOs: 18-21) - (SEQ ID NO: 6) - (SEQ ID NO: 7) - (SEQ ID NO: 8) - (SEQ ID NO: 9) - (SEQ ID NO: 10) - (SEQ ID NO: 15) - (SEQ ID NO: 12) - (SEQ ID NO: 16, 17, or 107); (i) (SEQ ID NO:2)-(SEQ ID NO:3)-(SEQ ID NO:4)-(SEQ ID NO:5)-(SEQ ID NO:6)-(one of SEQ ID NOs:18-21)-(SEQ ID NO:8)-(SEQ ID NO:9)-(SEQ ID NO:10)-(SEQ ID NO:15)-(SEQ ID NO:12)-(SEQ ID NO:16, 17, or 107); (j) (SEQ ID NO: 2) - (SEQ ID NO: 3) - (SEQ ID NO: 4) - (SEQ ID NO: 5) - (SEQ ID NO: 6) - (SEQ ID NO: 7) - (SEQ ID NO: 8) - (one of SEQ ID NOs: 18-21) - (SEQ ID NO: 10) - (SEQ ID NO: 15) - (SEQ ID NO: 12) - (SEQ ID NO: 16, 17, or 107); (k) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12), (l) (SEQ ID NO: 16 or 17) - (SEQ ID NO: 2) - (one of SEQ ID NOs: 18-21) - (SEQ ID NO: 4) - (SEQ ID NO: 5) - (SEQ ID NO: 6) - (SEQ ID NO: 7) - (SEQ ID NO: 8) - (SEQ ID NO: 9) - (SEQ ID NO: 10) - (SEQ ID NO: 15) - (SEQ ID NO: 12); (m) (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(SEQ ID NO: 9)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12); (n) (SEQ ID NO: 16 or 17) - (SEQ ID NO: 2) - (SEQ ID NO: 3) - (SEQ ID NO: 4) - (SEQ ID NO: 5) - (SEQ ID NO: 6) - (one of SEQ ID NOs: 18-21) - (SEQ ID NO: 8) - (SEQ ID NO: 9) - (SEQ ID NO: 10) - (SEQ ID NO: 15) - (SEQ ID NO: 12), and (o) The DNA polymerase of claim 70, comprising a sequence having at least 60% sequence identity to one of (SEQ ID NO: 16 or 17)-(SEQ ID NO: 2)-(SEQ ID NO: 3)-(SEQ ID NO: 4)-(SEQ ID NO: 5)-(SEQ ID NO: 6)-(SEQ ID NO: 7)-(SEQ ID NO: 8)-(one of SEQ ID NOs: 18-21)-(SEQ ID NO: 10)-(SEQ ID NO: 15)-(SEQ ID NO: 12).
79. 71. The DNA polymerase of claim 70, comprising a sequence having at least 60% sequence identity to one of SEQ ID NOs: 22-49.
80. below, (a) one DNA polymerase domain, one TBD, and one TRX; (b) one DNA polymerase domain, one TBD, and two or more TRXs; (c) one DNA polymerase domain, two or more TBDs, and one TRX; (d) one exonuclease-deficient DNA polymerase domain, one TBD, and one TRX; (e) one exonuclease-deficient DNA polymerase domain, one TBD, and two or more TRXs; (f) an exonuclease-deficient DNA polymerase domain, two or more TBDs, and one TRX; (g) one DNA polymerase domain, one TBD, one TRX, and one TIS; (h) one DNA polymerase domain, one TBD, two or more TRXs, and one TIS; (i) one DNA polymerase domain, two or more TBDs, one TRX, and one TIS; (j) one exonuclease-deficient DNA polymerase domain, one TBD, one TRX, and one TIS; (k) one exonuclease-deficient DNA polymerase domain, one TBD, two or more TRXs, and one TIS; or (l) A DNA polymerase comprising one exonuclease-deficient DNA polymerase domain, two or more TBDs, one TRX, and one TIS.
81. 81. The DNA polymerase of claim 80, wherein the DNA polymerase domain has at least 40% sequence identity to SEQ ID NO: 1 or an ordered combination of eight or more of SEQ ID NOs: 2-12.
82. 81. The DNA polymerase of claim 80, wherein the TBD has at least 50% sequence identity to SEQ ID NO:
15.
83. 81. The DNA polymerase of claim 80, wherein the TRX has at least 50% sequence identity to SEQ ID NO: 16, 17, or 107.
84. 81. The DNA polymerase of claim 80, wherein the TIS has at least 40% sequence identity to one of SEQ ID NOs: 18-21.
85. 81. The DNA polymerase of claim 80, wherein the exonuclease-deficient DNA polymerase domain lacks all or a portion of SEQ ID NO:
13.
86. 86. A reaction mixture comprising the composition or DNA polymerase of any one of claims 1-85 and amplification reagents sufficient to amplify a DNA target sequence.
87. 87. The reaction mixture of claim 86, wherein the amplification reagents comprise one or more of oligonucleotide primers, deoxynucleotide triphosphates, magnesium chloride, a buffer, water, and a template DNA comprising the DNA target sequence.
88. 87. The reaction mixture of claim 86, wherein the DNA target sequence comprises one or more short tandem repeats (STRs).
89. 87. The reaction mixture of claim 86, wherein the DNA target sequence comprises one or more mononucleotide repeats.
90. 89. The reaction mixture of claim 88, wherein the STR comprises repeating units of 1-50 nucleotides ranging in length from 10-500 nucleotides.
91. 87. The reaction mixture of claim 86, further comprising a reducing agent.
92. 92. The reaction mixture of claim 91, wherein the reducing agent is a thiol reducing agent or a non-thiol reducing agent.
93. 93. The reaction mixture of claim 92, wherein the reducing agent is dithiothreitol (DTT) or tris(2-carboxyethyl)phosphine (TCEP).
94. A method for amplifying a DNA target sequence comprising exposing the reaction mixture of any one of claims 86-93 to polymerase chain reaction temperature cycling conditions.
95. A thioredoxin (TRX) polypeptide that comprises 100% sequence similarity to positions 29-37, 60-77, and 89-98 of SEQ ID NO: 16, and is capable of binding to a TRX binding domain (TBD) having the amino acid sequence of SEQ ID NO:
15.
96. 96. The TRX polypeptide of claim 95, wherein the TRX polypeptide has at least 90% sequence identity to positions 29-37, 60-77, and 89-98 of SEQ ID NO:
16.
97. 97. The TRX polypeptide of claim 96, wherein the TRX polypeptide has 100% sequence identity to positions 29-37, 60-77, and 89-98 of SEQ ID NO:
16.
98. 96. The TRX polypeptide of claim 95, wherein the TRX polypeptide comprises at least 40% sequence identity with SEQ ID NO:
16.
99. 99. The TRX polypeptide of claim 98, wherein the TRX polypeptide comprises 40-60% sequence identity with SEQ ID NO:
16.
100. 96. The TRX polypeptide of claim 95, wherein the length of the TRX polypeptide is 100-120 amino acids.
101. 96. The TRX polypeptide of claim 95, wherein the TRX polypeptide has a 3D fold threshold of 0.8 or greater for TRX in protein database model 6N7W.
102. 96. The TRX polypeptide of claim 95, wherein the TRX polypeptide has an instability score of less than 40.
103. A thioredoxin (TRX) polypeptide capable of binding to a TRX binding domain (TBD) and having (i) a 3D fold threshold of 0.8 or greater relative to TRX in protein database model 6N7W, and / or (ii) an instability score of less than 40.
104. The TRX polypeptide of claim 103, wherein the TRX polypeptide comprises 100% sequence similarity to positions 29-37, 60-77, and 89-98 of SEQ ID NO:
16.
105. 104. The TRX polypeptide of claim 103, wherein the length of the TRX polypeptide is 100-120 amino acids.
106. A thioredoxin (TRX) polypeptide capable of binding to a TRX binding domain (TBD), wherein the root mean square deviation (RMSD) calculated for the alpha carbons of at least 70% of the amino acid residues corresponding to amino acids 29-37, 60-77, and 89-98 of SEQ ID NO: 16 in a 3D molecular structure of the TRX polypeptide is 3.0 Å or less relative to TRX in protein database model 6N7W.
107. 107. The TRX polypeptide of claim 106, wherein the alpha carbon RMSD of the TBD-interacting residues of the TRX is 3.0 Å or less relative to protein database model 6N7W.
108. 107. The TRX polypeptide of claim 106, wherein the TBD-interacting residues of the TRX have at least 70% sequence similarity to SEQ ID NO:
16.
109. 109. The TRX polypeptide of claim 108, wherein the TBD-interacting residues of the TRX have at least 70% sequence identity to SEQ ID NO:
16.
110. below, (a) a DNA polymerase domain that shares at least 40% sequence identity with a Family A DNA polymerase; (b) a thioredoxin binding domain (TBD) having at least 50% sequence identity to a native phage-derived TBD; and (c) a DNA polymerase system comprising a thioredoxin (TRX) domain according to any one of claims 95 to 109.
111. below, (a) (i) a DNA polymerase domain; (ii) a thioredoxin binding domain (TBD), and (iii) a first polypeptide comprising a thioredoxin (TRX) domain; and (b) (i) a DNA polymerase domain, and (ii) a second polypeptide comprising a TBD.