Engineered DNA ligase mutants
Engineered DNA ligases with optimized amino acid sequences improve ligation efficiency and temperature tolerance, overcoming the limitations of traditional ligases for DNA manipulation.
Patent Information
- Application Number
- JP2025542157
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-23
- Filing Date
- 2024-01-23
- Publication Date
- 2026-02-10
AI Technical Summary
Existing DNA ligases, such as T4 DNA ligase, are limited by temperature sensitivity, buffer additive inhibition, and sequence bias, making them less efficient for ligation of DNA substrates in molecular biology and diagnostic applications.
Engineered DNA ligase polypeptides with specific amino acid sequences and substitutions, providing improved ligation efficiency and tolerance to higher temperatures and buffer additives.
The engineered DNA ligases offer enhanced ligation efficiency and broader substrate compatibility, addressing the limitations of traditional ligases and facilitating more efficient DNA manipulation.
Smart Images

Figure 2026504945000001 
Figure 2026504945000002 
Figure 2026504945000003
Abstract
Description
Related Applications
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 481,158, filed January 23, 2023, which is incorporated herein by reference in its entirety.
[0002] Reference to a sequence listing, table or computer program The Sequence Listing submitted concurrently herewith under file name CX9-235WO2_ST26.xml, created on January 22, 2024, with a file size of 2,686,547 bytes, is a part of the present specification and is incorporated herein by reference. [Technical Field]
[0003] The present disclosure provides engineered DNA ligase polypeptides and compositions thereof, polynucleotides encoding the engineered DNA ligase polypeptides, and methods of using the engineered DNA ligases for polynucleotide synthesis, molecular biology tools, and diagnostic applications, etc. [Background technology]
[0004] DNA ligases are a family of enzymes that catalyze the covalent joining of DNA molecules by catalyzing the formation of a phosphodiester bond between the 3' hydroxyl end of one DNA substrate and the 5' phosphorylated end of another. DNA ligases are involved in maintaining genome integrity by repairing single-strand breaks in double-stranded DNA during replication, repair, and recombination. Some ligases, such as T4 phage ligase and eukaryotic DNA ligase, use ATP, while others, such as E. coli ligase, use NAD as a cofactor. DNA ligases can join dsDNA fragments with fully base-paired blunt ends or ends with complementary single-stranded overhangs. DNA ligases have found great utility in molecular biology and diagnostic applications, including restriction enzyme cloning, adapter ligation for cloning / sequencing, SNP or sequence analysis, and assembling DNA fragments from multiple smaller fragments.
[0005] The prototype DNA ligase is derived from bacteriophage T4, the most commonly used ligase in molecular biology and diagnostic applications. T4 DNA ligase can ligate the cohesive or "sticky" ends of DNA, oligonucleotides, and some RNA and RNA-DNA hybrids. T4 DNA ligase can also ligate blunt-ended DNA with high efficiency. T4 DNA ligase uses ATP as a cofactor. T4 DNA ligase is typically active between 4°C and 37°C but loses activity at higher temperatures. T4 DNA ligase is also sensitive to buffer additives, such as monovalent salts, which inhibit its activity, particularly its end-joining activity. Furthermore, T4 DNA ligase is biased toward the sequences of the final and penultimate bases of the ligation site. Therefore, while T4 DNA ligase has become an important tool for ligating DNA molecules in research and diagnostic applications, a ligase that provides easier and more efficient ligation of DNA substrates is desirable. Summary of the Invention
[0006] The present disclosure provides engineered DNA ligase polypeptides and compositions thereof, as well as polynucleotides encoding the engineered DNA ligase polypeptides. The present disclosure also provides methods of using the engineered DNA ligase polypeptides and compositions thereof to ligate polynucleotides.
[0007] In one aspect, the disclosure provides a nucleic acid sequence that is at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more similar to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2 and even-numbered SEQ ID NOs: 40-1184, or a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2 and even-numbered SEQ ID NOs: 40-1184. and a functional fragment thereof, comprising an amino acid sequence having at least one sequence identity to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, 62, 138, 318, 722, or 938, or one or more substitutions relative to a reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722, or 938.
[0008] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or a reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or relative to the reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938.
[0009] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or to a reference sequence corresponding to SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0010] In some embodiments, the engineered DNA ligase comprises a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or to a reference sequence corresponding to SEQ ID NO: 2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2.
[0011] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or to a reference sequence corresponding to an even-numbered SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0012] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises the amino acid sequences at amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 148, 149, 150, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 288, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, and / or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or a reference sequence corresponding to SEQ ID NO:2.
[0013] Its specific ingredient is the specificity of the DNA fragment It also has a range of 11D, 12A / I and 13G / R、14G / S / T / V、18D / N / S、30C / H / S、31R、33M / R / V、34L / R、36T / Y、37G / L / N / S 44S, 50G / I / S / T, 56P, 59E, 60Y, 61T / V, 63F / R, 67R, 68A / M / S / V / Y, 69T, 71G / L / P / R、73C / K / P / T / V / W、74S、76F / G / H / L / N / R、77D、82R、88V、95A / L / R / V、9 6A / G / T / V、97G、99G / I、100V、101R、102G / K / L / S、103V、104K、105K / S / T、106 L / S / V, 110R, 112M, 113A / T, 117G / S / V / Y, 125R / T, 128C, 130T, 132R, 138L / R 139T, 148P, 149P, 150C / F / T, 155R, 156C, 159Q, 161R / V, 162W, 164A / R K, 177G, 186A / C / E / H / L / M / R / T / V, 188A, 189C / T, 190R, 191T, 195R, 196E / V 197R、198A / D / K / L / N / R / V / W、201L / S、205E / G / K、207L、208D / F / H、212F / G / M / S / W、220V、226D / E / Q / S / V、228E / I / M / S、230L / M、231P、232R、233G / T / W、23 5W、237L / M / R / S / V / Y、239M / N / P / Q / S / T / V / W、240E / G / K / Q / R / S / Y、242P / Q / T 251L, 254G / S, 258L / S / V, 263G / L / Q / T, 264A / C, 266M / T, 267D / W / Y, 269L 71A / G / N / S、273A / G / S、277Q / R、278E、282G / L / M / T / V / Y、283A / G / K / L / M / R / S / V、284D、286F / L / S、288I、289A / L / S / V、290L、294L、295K、297W、300G / T、30 1F / L、305K、306I / K / S / V、308K / L / S、309G / R、317Q、323S、328R、334L / R、337 G / L / M / P / R / S 339Y 349E 355S 356A / V / W 357H / K / P / R / S / V 358C 359N / R360H / M / P, 362G, 363R, 364R, 367C / L, 370C / G, 372N / Q, 374A / S, 375W, 378T, 379A / G / P, 380T, 381K / R, 382V, 38 4C / V, 386F, 387G, 388K / Y, 389K / L / Q / R, 390E, 392C / I / K / L / R / S, 396C / H, 397K / L / M, 404S, 405I, 408C / V, 414A / L / Q / R / T / V, 415A / C / E / H / I / K / L / V, 416K, 417D / G / L, 418A / G / I / L / M / P / S / T, 419G, 421R, 422N, 423R / T, or 428F / R / S, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or relative to a reference sequence corresponding to SEQ ID NO:2.
[0014] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions at amino acid position(s) 233, 317, 191, 288, 207, 149, 251, 205, 269, 164, 36, 428, 105 / 132, or 105, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0015] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0016] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to the reference sequence corresponding to SEQ ID NO:2.
[0017] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0018] In some embodiments, the engineered DNA ligase comprises a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0019] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:40-1184.
[0020] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938.
[0021] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, 138, 318, 722, or 938, or relative to the reference sequence corresponding to SEQ ID NO:62, 138, 318, 722, or 938.
[0022] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 19 8, 139, 148, 149, 150, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 288, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, 384, 386, 387, 388, 389, 390, 392, 393 6, 397, 404, 405, 408, 414, 415, 416, 417, 418, 419, 421, 422, 423, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 62, 138, 318, 722, or 938, or relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0023] Its specific ingredient is the specificity of the DNA fragment It also has a range of 11D, 12A / I and 13G / R、14G / S / T / V、18D / N / S、30C / H / S、31R、33M / R / V、34L / R、36T / Y、37G / L / N / S 44S, 50G / I / S / T, 56P, 59E, 60Y, 61T / V, 63F / R, 67R, 68A / M / S / V / Y, 69T, 71G / L / P / R、73C / K / P / T / V / W、74S、76F / G / H / L / N / R、77D、82R、88V、95A / L / R / V、96 A / G / T / V、97G、99G / I、100V、101R、102G / K / L / S、103V、104K、105K / S / T、106L / S / V、110R、112M、113A / T、117G / S / V / Y、125R / T、128C、130T、132R、138L / R、1 39T, 148P, 149P, 150C / F / T, 155R, 156C, 159Q, 161R / V, 162W, 164A / R, 165K 177G、186A / C / E / H / L / M / R / T / V、188A、189C / T、190R、191T、195R、196E / V、197 R、198A / D / K / L / N / R / V / W、201L / S、205E / G / K、207L、208D / F / H、212F / G / M / S / W、220V、226D / E / Q / S / V、228E / I / M / S、230L / M、231P、232R、233G / T / W、235W、2 37L / M / R / S / V / Y、239M / N / P / Q / S / T / V / W、240E / G / K / Q / R / S / Y、242P / Q / T / V、2 51L, 254G / S, 258L / S / V, 263G / L / Q / T, 264A / C, 266M / T, 267D / W / Y, 269L, 271A / G / N / S、273A / G / S、277Q / R、278E、282G / L / M / T / V / Y、283A / F / G / K / L / M / R / S / V、284D、286F / L / S / W、288I、289A / L / S / V、290L、294L、295K、297W、300G / T、30 1F / L、305K、306I / K / S / V、308K / L / S、309G、317T / Q、323S、328R、334L / R、337 G / L / M / P / R / S 339Y 349E 355S 356A / V / W 357H / K / P / R / S / V 358C 359N / R360H / M / P, 362G, 364R, 367C / L, 370C / G, 372N / Q, 374A / S, 375W, 378T, 379A / G / P, 380T, 381K / R, 382V, 384C / V, 386F, 387G, 3 88K / Y, 389K / L / Q / R, 390E, 392C / I / K / L / R / S, 396C / H, 397K / L / M, 404S, 405I, 408C / V, 414A / L / Q / R / S / T / V, 415A / C / E / H / I / K 421R, 422N, 423R / T, or 428E / F / R / S, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0024] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0025] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises a substitution at least at amino acid residue 63R, 242Q, 283L, 286S, 317Q, 414Q, 418S, or 428R, or a combination thereof, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0026] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to SEQ ID NO:62.
[0027] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to an even-numbered SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62.
[0028] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid position(s) 196, 242, 337, 33, 277, 30, 359, 283, 415, 387, 379, 205, 186, 389, 102, 164, 301, 375, 267, 380, 254, 317, 77 / 139 / 317 / 417, 105 / 317 / 417, 317 / 349 / 362 / 386, 105 / 317, 139 / 317 / 362 , 233 / 317 / 405, 139 / 317, 162, 286, 414, 417, 226, 61, 105, 230, 418, 370, 297, 237, 428, 362, 233, 235, 148, 100, 97, 382, or 358, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:62, or are relative to a reference sequence corresponding to SEQ ID NO:62.
[0029] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or to a reference sequence corresponding to SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138 or to the reference sequence corresponding to SEQ ID NO: 138.
[0030] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138.
[0031] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid at position(s) 242 / 283 / 286 / 359 / 418, 283 / 286, 283 / 286 / 418, 186 / 242 / 283 / 286 / 418, 205 / 286 / 359, 283, 242 / 286 / 418, 186 / 205 / 242, 283 / 286 / 359, 277 / 286 / 359 / 418, 186 / 205 / 283 / 286 / 359, 242 / 277 / 418, 242 / 283 / 286 / 418, 112 / 19 6 / 389, 286, 283 / 286 / 359 / 418, 242 / 359 / 418, 186 / 283, 186 / 359 / 418, 205 / 242 / 418, 186 / 205 / 242 / 283 / 286 / 359 / 418, 205 / 359 / 418, 186 / 205 / 283 / 286 / 418, 186 / 283 / 359, 186 / 242, 186 / 242 / 359, 283 / 359 / 418, 277 / 418, 186 / 188 / 283, 186 / 286 / 418, 186 / 242 / 286 / 359 / 418, 418, 1 86 / 242 / 283 / 286 / 359 / 418, 186 / 277 / 359 / 418, 242 / 283 / 286, 205 / 418, 30 / 297, 205 / 242 / 286 / 359 / 418, 186 / 205 / 359 / 418, 359 / 418, 186, 230, 33 / 297, 186 / 205, 186 / 283 / 359 / 418, 186 / 418, 205 / 242 / 283 / 359 / 418, 33 / 375 / 389, 33 / 230, 196 / 242 / 283 / 286 / 359 / 418, 186 / 242 / 283 / and at least a substitution or set of substitutions in 359 / 418, 186 / 359, 33 / 196, 186 / 277 / 418, 242, 33 / 196 / 297 / 301, 205 / 237 / 242 / 283 / 286 / 359, 186 / 205 / 283 / 359 / 418, 33 / 389, or 186 / 196 / 242 / 283 / 286 / 359 / 418, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 138, or are relative to a reference sequence corresponding to SEQ ID NO: 138.
[0032] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or to a reference sequence corresponding to SEQ ID NO:318, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318 or to the reference sequence corresponding to SEQ ID NO:318.
[0033] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or to a reference sequence corresponding to an even-numbered SEQ ID NO:460-936, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318.
[0034] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid at position(s) 363, 63, 389, 381, 197, 359, 102, 165, 388, 414, 337, 164, 416, 101, 415, 423, 364, 73, 50, 71, 388 / 419, 357, 396, 68, 76, 14, 271, 360, 266, 208, 74, 263, 264, 13, 378, 372, 300, 294, 290, 397, 95, 258, 161, 212, 198 , 138, 18, 404, 273, 117, 240, 69, 278, 289, 82, 328, 61 / 186 / 417, 186 / 370 / 417, 267, 61 / 370, 186 / 267 / 370 / 417, 61 / 186, 370 / 417, 417, 61 / 370 / 382, 61 / 186 / 267 / 370 / 417, 61, 267 / 370 / 417, 267 / 370, 61 / 417, 61 / 186 / 237 / 267 / 370, 61 / 237 / 370 / 417, 370, 186 / 370, 61 / 18 6 / 267 / 417, 61 / 186 / 370 / 382, 370 / 382 / 417, 61 / 186 / 370, 237 / 267 / 370 / 417, 61 / 186 / 382, 61 / 267, 61 / 267 / 417, 61 / 237 / 267 / 382, 186 / 370 / 382, 237 / 267 / 370, 61 / 186 / 267 / 370, 186 / 237 / 267 / 370, 61 / 186 / 267, 61 / 186 / 237, 186, 186 / 267, 237 / 370 / 417, 242 / 414, 162 / 4 or 390, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318 or to a reference sequence corresponding to SEQ ID NO:318.
[0035] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:722, or to a reference sequence corresponding to SEQ ID NO:722, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722.
[0036] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 938-1098, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 938-1098, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 722, or to the reference sequence corresponding to SEQ ID NO: 722.
[0037] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 63, 63 / 96 / 370, 389, 13 / 267 / 363 / 389, 13 / 186 / 389, 50 / 267 / 363 / 370 / 389, 363 / 370, 96 / 370, 61 / 63 / 212, 11 / 305, 11, 242 / 283 / 286 / 317 / 414 / 418, 323, 334, 339, 356, 384, 408, 67, 392, 104, 355, 1 59, 155, 367, 31, 231, 36, 150, 239, 103, 125, 228, 37, 189, 177, 422, 128, 220, 130, 56, 190, 156, 232, 423, 34, 99, 59, 60, 421, or 195, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:722, or are relative to a reference sequence corresponding to SEQ ID NO:722.
[0038] In some embodiments, the engineered DNA ligase comprises a reference sequence corresponding to residues 12-437 of SEQ ID NO:938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to SEQ ID NO:938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938 or to the reference sequence corresponding to SEQ ID NO:938.
[0039] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 1100-1184, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 1100-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 938.
[0040] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 308 / 357 / 390, 74 / 76 / 201 / 308 / 357, 61 / 74 / 76 / 186 / 201 / 308 / 309 / 357 / 390, 14 / 201 / 240 / 289 / 357, 308 / 415, 76 / 357 / 396, 263 / 308 / 396, 61 / 76 / 96 / 240 / 308 / 309, 14 / 306 / 415, 14 / 73 / 106 / 415, 12 / 14 / 258 / 263 / 289 / 308 / 309 / 396, 74 / 76 / 117 / 309 / 357, 14 / 258 / 263 / 357 / 396, 14 / 96 / 106 / 306, 14 / 106, 12 / 14 / 308 / 309, 14 / 357 / 390, 14 / 117 / 258 / 309 / 3 57, 390, 240 / 273 / 357 / 390, 61 / 76 / 186 / 201 / 308 / 309, 14 / 396, 309, 14, 106 / 306 / 308, 14 / 240 / 306 / 308, 12 / 14 / 186 / 357, 309 / 390, 14 / 306, 14 / 76 / 308, 117 / 208 / 258 / 263 / 289 / 308 / 309, 14 / 73 / 106, 76 / and at least a substitution or set of substitutions at 208 / 263, 357, 14 / 308, 263, 76, 14 / 300 / 308 / 415, 240, 33 / 357 / 390, 14 / 76 / 273, or 74, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 938, or are relative to a reference sequence corresponding to SEQ ID NO: 938.
[0041] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0042] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0043] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising at least a substitution or set of substitutions provided in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0044] In some embodiments, the engineered DNA ligase comprises an amino acid sequence comprising residues 12-437 of an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, or a sequence comprising an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2.
[0045] In some embodiments, the engineered DNA ligase comprises an amino acid sequence comprising residues 12-437 of an even-numbered SEQ ID NO:40-1184, or an amino acid sequence comprising an even-numbered SEQ ID NO:40-1184, optionally, the amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions.
[0046] In some embodiments, the engineered DNA ligase comprises an amino acid sequence comprising residues 12-437 of SEQ ID NO: 62, 138, 318, 722, 938, or 1108, or an amino acid sequence comprising SEQ ID NO: 62, 138, 318, 722, 938, or 1108, optionally, the amino acid sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions.
[0047] In some embodiments, the engineered DNA ligase has DNA ligase activity. In some embodiments, the engineered DNA ligase has DNA ligase activity and is characterized by at least one improved property compared to a reference DNA ligase. In some embodiments, the improved property of the engineered DNA ligase is selected from: i) increased activity, ii) increased stability, iii) increased thermostability, iv) increased product yield, v) increased solubility, vi) reduced sequence bias, and vii) insensitivity or reduced sensitivity to input DNA concentration, or any combination of i), ii), iii), iv), v), vi), and vii), compared to the reference DNA ligase. In some embodiments, the improved property of the engineered DNA ligase is compared to a reference DNA ligase having a sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722, or 938, or a sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722, or 938. In some embodiments, the improved properties of the engineered DNA ligase are compared to a reference DNA ligase having a sequence corresponding to residues 12-437 of SEQ ID NO: 2, or a sequence corresponding to SEQ ID NO: 2. In some embodiments, the reference DNA ligase is wild-type T4 DNA ligase.
[0048] In some further embodiments, the engineered DNA ligase is purified, hi some embodiments, the engineered DNA ligase is provided in solution, provided as a lyophilizate, or immobilized on a substrate, such as a solid substrate, a porous substrate, a membrane, or a particle.
[0049] In another aspect, the present disclosure provides a recombinant polynucleotide comprising a polynucleotide sequence encoding an engineered DNA ligase disclosed herein.
[0050] In some embodiments, the recombinant polynucleotide comprises a reference polynucleotide sequence corresponding to nucleotide residues 34 to 1311 of SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a polynucleotide sequence having at least 70%, 75%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to SEQ ID NO: 1, 61, 137, 317, 721, or 937, wherein the recombinant polynucleotide encodes an engineered DNA ligase.
[0051] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to nucleotide residues 34-1311 of an odd-numbered SEQ ID NO:39-1183, or to a reference polynucleotide sequence corresponding to an odd-numbered SEQ ID NO:39-1183, wherein the recombinant polynucleotide encodes an engineered DNA ligase.
[0052] In some embodiments, the polynucleotide sequence of the recombinant polynucleotide encoding the engineered DNA ligase is codon-optimized for expression in an organism or cell type thereof, such as a bacterial cell, a fungal cell, an insect cell, or a mammalian cell.
[0053] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence comprising nucleotide residues 34 to 1311 of SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a polynucleotide sequence comprising SEQ ID NO: 1, 61, 137, 317, 721, or 937.
[0054] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence comprising nucleotide residues 34 to 1311 of an odd-numbered SEQ ID NO: 39-1183, or a polynucleotide sequence comprising an odd-numbered SEQ ID NO: 39-1183.
[0055] In a further aspect, the present disclosure provides an expression vector comprising a recombinant polynucleotide provided herein encoding an engineered DNA ligase. In some embodiments, the recombinant polynucleotide of the expression vector is operably linked to a regulatory sequence. In some embodiments, the regulatory sequence comprises a promoter, particularly a heterologous promoter.
[0056] In another aspect, the present disclosure also provides a host cell comprising a recombinant polynucleotide or expression vector provided herein. In some embodiments, the host cell is a prokaryotic or eukaryotic cell. In some embodiments, the host cell is a bacterial cell, a fungal cell, an insect cell, or a mammalian cell.
[0057] In a further aspect, the disclosure provides methods for producing an engineered DNA ligase polypeptide, the method comprising culturing a host cell described herein under suitable culture conditions such that at least one engineered DNA ligase is produced. In some embodiments, the method further comprises recovering or isolating the engineered DNA ligase from the medium and / or the host cell. In some embodiments, the method further comprises purifying the engineered DNA ligase.
[0058] In another aspect, the present disclosure provides compositions comprising at least one engineered DNA ligase disclosed herein. In some embodiments, the compositions comprise at least a buffer. In some embodiments, the compositions further comprise a nucleotide substrate (e.g., ATP) and / or one or more DNA ligase substrates. In some embodiments, the DNA ligase substrate comprises an adaptor or linker.
[0059] In a further aspect, the present disclosure provides a method of ligating at least a first DNA strand and a second DNA strand, comprising contacting the first DNA strand and the second DNA strand with an engineered DNA ligase described herein in the presence of a nucleotide substrate under conditions suitable for ligating the first DNA strand to the second DNA strand, wherein the first DNA strand comprises a ligatable 5'-end and the second DNA strand comprises a ligatable 3'-end to the 5'-end of the first DNA strand. In some embodiments, the 3'-end of the second DNA strand is a 3'-hydroxyl and the 5'-end of the first DNA strand is a 5'-phosphate.
[0060] In some embodiments, the method further includes a third DNA or polynucleotide strand, wherein the first DNA strand and the second DNA strand hybridize adjacent to each other on the third DNA or polynucleotide strand, positioning the 5' end of the first DNA strand adjacent to the 3' end of the second DNA strand. In some embodiments, the third DNA or polynucleotide strand is contiguous with the first DNA strand or the second DNA strand. In some embodiments, the third DNA strand is contiguous with the first DNA strand and the second DNA strand to form a single, contiguous DNA ligase substrate.
[0061] In some embodiments of the method, the first DNA strand hybridizes to the third DNA strand to form a first dsDNA substrate, and the second DNA strand hybridizes to the fourth DNA strand to form a second dsDNA substrate. In some embodiments, the first dsDNA substrate comprises a blunt-ended 5'-end of the first DNA strand, and the second dsDNA substrate comprises a blunt-ended 3'-end of the second DNA strand. In some embodiments of the method, the first dsDNA substrate comprises an overhang on at least one end of the first dsDNA substrate, and the second dsDNA substrate comprises an overhang on at least one end of the second dsDNA substrate, wherein the overhangs on the first dsDNA substrate and the overhangs on the second dsDNA substrate are complementary and can hybridize to each other, forming one or more nicks that can be ligated.
[0062] In a further aspect, the present disclosure also provides kits comprising at least one engineered DNA ligase disclosed herein, hi some embodiments, the kits further comprise one or more of a buffer, a nucleotide substrate, a reducing agent, one or more DNA ligase substrates, and / or a ligation enhancer. DETAILED DESCRIPTION OF THE INVENTION
[0063] The present disclosure provides engineered DNA ligase polypeptides and compositions thereof, as well as polynucleotides encoding the engineered DNA ligase polypeptides. The present disclosure also provides methods for using the engineered DNA ligase polypeptides and compositions thereof for molecular biological, diagnostic, and other purposes. In some embodiments, the engineered DNA ligase polypeptides exhibit, among other things, increased activity, increased stability, increased thermostability, increased solubility, and / or reduced sequence bias.
[0064] Abbreviations and Definitions Unless otherwise defined, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Generally, the nomenclature used herein and the laboratory procedures of cell culture, molecular genetics, microbiology, organic chemistry, analytical chemistry, and nucleic acid chemistry described below are those well known and commonly used in the art.
[0065] Although any suitable methods and materials similar or equivalent to those described herein can be used to practice the present invention, exemplary methods and materials are described herein. It is understood that the present invention is not limited to the specific methodology, protocols, and reagents described, which may vary depending on the context in which they are used by those skilled in the art. Accordingly, the terms defined below are more fully described by reference to the entire application.
[0066] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.
[0067] As used herein, the term "comprising" and its cognates are used in an inclusive sense (i.e., equivalent to the term "including" and its corresponding cognates).
[0068] It should also be understood that where the description of an embodiment uses the term "comprising" and its cognates, the embodiment can also be described using the terms "consisting essentially of" or "consisting of."
[0069] Moreover, numerical ranges are inclusive of the numbers defining the range. Accordingly, every numerical range disclosed herein is intended to include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly written herein. Also, every maximum (or minimum) numerical limitation disclosed herein is intended to include every lower (or higher) numerical limitation, as if such lower (or higher) numerical limitations were all expressly written herein.
[0070] As used herein, the term "about" refers to an acceptable degree of error for a particular value. In some cases, "about" means within 0.05%, 0.5%, 1.0%, or 2.0% of a given value's range. In other cases, "about" means within 1, 2, 3, or 4 standard deviations of a given value.
[0071] Furthermore, the headings provided herein are not limitations of the various aspects or embodiments of the invention that may be had by reference to this application as a whole. Accordingly, the terms defined immediately below are more fully defined by reference to this application as a whole.
[0072] The "EC" number refers to the Enzyme Nomenclature of the International Union of Biochemistry and Molecular Biology (NC-IUBMB) Committee on Nomenclature. The IUBMB biochemical classification is a numerical classification system for enzymes based on the chemical reaction they catalyze.
[0073] "ATCC" refers to the American Type Culture Collection, whose biorepository collection includes genes and strains.
[0074] "NCBI" refers to the National Center for Biological Information and the sequence databases provided therein.
[0075] "Protein," "polypeptide," and "peptide" are used interchangeably to refer to a polymer of at least two amino acids covalently joined by amide bonds, regardless of length or post-translational modification (e.g., glycosylation or phosphorylation).
[0076] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Abbreviations used for genetically encoded amino acids are conventional and are alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine (Cys or C), glutamic acid (Glu or E), glycine (Gly or G), glutamine (Gln or Q), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V). When three-letter abbreviations are used, amino acids may be in either the L- or D-configuration about the α-carbon (Cα) unless specifically preceded by "L" or "D" or unless otherwise apparent from the context in which the abbreviation is used. For example, "Ala" indicates alanine without specifying the configuration about the α-carbon, whereas "D-Ala" and "L-Ala" indicate D-alanine and L-alanine, respectively. When single-letter abbreviations are used, uppercase letters indicate amino acids in the L-configuration about the α-carbon, and lowercase letters indicate amino acids in the D-configuration about the α-carbon. For example, "A" indicates L-alanine and "a" indicates D-alanine. When polypeptide sequences are presented as a series of single-letter or three-letter abbreviations (or mixtures thereof), the sequences are presented in the amino (N) to carboxy (C) direction, according to common convention.
[0077] "Fusion protein" and "chimeric protein" and "chimera" refer to a hybrid protein created by the joining of two or more polynucleotides that originally encoded separate proteins. In some embodiments, fusion proteins are created by recombinant techniques.
[0078] "DNA ligase" refers to an enzyme that covalently joins the 5'-phosphoryl end ("donor") and 3'-hydroxyl end ("acceptor") of DNA to one another. DNA ligases can be classified into two families based on cofactor requirements: ATP-dependent ligases and NAD+-dependent ligases. Eukaryotic and archaeal DNA ligases are generally ATP-dependent. DNA ligases of eubacterial origin are generally NAD+-dependent. DNA ligases include enzymes within the general class of EC 6.5.1.
[0079] The terms "polynucleotide," "nucleic acid," or "oligonucleotide" are used herein to refer to a polymer containing at least two nucleotides, where the nucleotides are either deoxyribonucleotides or ribonucleotides, or a mixture of deoxyribonucleotides and ribonucleotides. In some embodiments, the abbreviations used to genetically encode nucleosides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless specifically indicated, the abbreviated nucleoside may be either a ribonucleoside or a 2'-deoxyribonucleoside. Nucleosides may be identified individually or collectively as either a ribonucleoside or a 2'-deoxyribonucleoside. When a polynucleotide, nucleic acid, or oligonucleotide sequence is presented as a series of single-letter abbreviations, the sequence is presented in the 5' to 3' direction according to common convention, and phosphates are not indicated. The term "DNA" refers to deoxyribonucleic acid. The term "RNA" refers to ribonucleic acid. A polynucleotide or nucleic acid can be single-stranded or double-stranded, or can contain both single-stranded and double-stranded regions.
[0080] "Duplex" and "ds" refer to a double-stranded nucleic acid (e.g., DNA or RNA) molecule composed of two single-stranded polynucleotides whose sequences are complementary (A pairs with T or U, and C pairs with G), arranged in an antiparallel 5' to 3' orientation, and held together by hydrogen bonds between the nucleobases (i.e., adenine [A], guanine [G], cytosine [C], thymine [T], uridine [U]).
[0081] " Complementary " is used herein to describe the structural relationship between nucleotide bases that can form base pairs with each other.For example, the purine nucleotide bases present in polynucleotides that are complementary to pyrimidine nucleotide bases on polynucleotides can form base pairs by forming hydrogen bonds with each other.Complementary nucleotide bases can be base paired through Watson-Crick base pairing or in any other manner to form a stable duplex or other nucleic acid structure.
[0082] "Watson / Crick base pairing" refers to specific pairing patterns of nucleobases and analogs that bind together through sequence-specific hydrogen bonds, e.g., A pairs with T or U, G pairs with C, etc.
[0083] "Annealing" or "hybridization" refers to the base-pairing interaction between one nucleobase polymer (e.g., polynucleotides and oligonucleotides) and another nucleobase polymer, resulting in the formation of a duplex, triplex, or quadruplex structure. Annealing or hybridization can occur via Watson-Crick base pairing interactions, but can also be mediated by other hydrogen-bonding interactions, such as Hoogsteen base pairing. In some embodiments, the nucleobase polymer that anneals or hybridizes to another is a single nucleobase polymer, while in other embodiments, the nucleobase polymers are separate nucleobase polymers.
[0084] When used with respect to a cell, polynucleotide, or polypeptide, "engineered," "recombinant," "non-naturally occurring," and "mutant" refer to material that is not otherwise found in nature or is identical to it, but that has been modified in a manner produced or derived from synthetic material and / or by manipulation using recombinant techniques, or material that corresponds to the natural or native form of that material.
[0085] "Wild-type" and "naturally occurring" refer to forms found in nature. For example, a wild-type polypeptide or polynucleotide sequence is one that can be isolated from a natural source and is present in an organism that has not been intentionally modified by human manipulation.
[0086] "Coding sequence" refers to that portion of a nucleic acid (eg, a gene) that codes for the amino acid sequence of a protein.
[0087] "Percent (%) sequence identity" refers to a comparison between a polynucleotide and a polypeptide and is determined by comparing two optimally aligned sequences over a comparison window. The portion of the polynucleotide or polypeptide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence due to optimal alignment of the two sequences. The percentage can be calculated by determining the number of positions where the identical nucleic acid base or amino acid residue appears in both sequences to produce the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to produce the percentage of sequence identity. Alternatively, the percentage can be calculated by determining the number of positions where the identical nucleic acid base or amino acid residue occurs in both sequences, or where the nucleic acid base or amino acid residue aligns with a gap to produce the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Those skilled in the art will appreciate that there are many established algorithms available for aligning two sequences. Optimal alignment of sequences for comparison can be achieved, for example, by the local homology algorithm of Smith and Waterman (Smith and Waterman, Adv. Appl. Math., 1981, 2:482), by the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch, J. Mol. Biol., 1970, 48:443), by the similarity search method of Pearson and Lipman (Pearson and Lipman, Proc. Natl. Acad. Sci. USA, 1988, 85:2444), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection as known in the art.Examples of algorithms suitable for determining percent sequence identity and percent sequence similarity include, but are not limited to, the BLAST and BLAST 2.0 algorithms (e.g., Altschul et al., J. Mol. Biol., 1990, 215:403-410; and Altschul et al., Nucleic Acids Res., 1977, 3389-3402). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length "W" in the query sequence that match or meet some positive threshold score "T" when aligned with words of the same length in database sequences. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits serve as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. For nucleotide sequences, cumulative scores are calculated using the parameters "M" (reward score for a pair of matching residues; always greater than 0) and "N" (penalty score for mismatching residues; always less than 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction is halted if the cumulative alignment score falls by an amount "X" from its maximum achieved value, if the cumulative score becomes zero or less than zero due to the accumulation of one or more negative-scoring residue alignments, or if the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands.For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see, e.g., Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA, 1989, 89:10915). Exemplary sequence alignments and determinations of percent sequence identity can use the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison, Wis.) using the default parameters provided.
[0088] A "reference sequence" refers to a defined sequence used as the basis for sequence comparison. A reference sequence can be a subset of a larger sequence, such as a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, at least 100 residues in length, or the entire length of a nucleic acid or polypeptide. Because two polynucleotides or polypeptides can each contain (1) similar sequences (i.e., portions of the complete sequence) between the two sequences and (2) additional sequences that differ between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is typically performed by comparing the sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, which can have one or more changes in the primary sequence. For example, the phrase "a reference sequence corresponding to SEQ ID NO: 2 having an aspartic acid at the residue corresponding to X11" (or "a reference sequence corresponding to SEQ ID NO: 2 having an aspartic acid at the residue corresponding to position 11") refers to a reference sequence in which the residue corresponding to position X11 (e.g., glycine) in SEQ ID NO: 2 has been changed to aspartic acid.
[0089] A "comparison window" refers to a conceptual segment of contiguous nucleotide positions or amino acid residues over which a sequence can be compared to a reference sequence. In some embodiments, the comparison window is at least 15-20 contiguous nucleotides or amino acids, and the portion of the sequence in the comparison window may contain no more than 20 percent additions or deletions (i.e., gaps) compared to the reference sequence (no additions or deletions) for optimal alignment of the two sequences. In some embodiments, the comparison window may be longer than 15-20 contiguous residues, and may optionally include a window of 30, 40, 50, 100, or more.
[0090] "Corresponding," "referring to," and "relative to," when used in the context of numbering a given amino acid or polynucleotide sequence, refer to the numbering of residues in a particular reference sequence when comparing the given amino acid or polynucleotide sequence to the reference sequence. In other words, residue numbers or residue positions in a given polymer are specified with respect to the reference sequence, not by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as the amino acid sequence of an engineered DNA ligase, can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, despite the presence of gaps, the numbering of residues in a given amino acid or polynucleotide sequence is done with respect to the reference sequence to which it is aligned.
[0091] "Mutation" refers to a change in a nucleic acid sequence. In some embodiments, a mutation results in a change in the encoded polypeptide sequence (i.e., compared to the original sequence without the mutation). In some embodiments, a mutation comprises a substitution, such that a different amino acid is produced. In some alternative embodiments, a mutation comprises an addition, such that an amino acid is added to the original polypeptide sequence (e.g., an insertion). In some further embodiments, a mutation comprises a deletion, such that an amino acid is deleted from the original polypeptide sequence. There can be any number of mutations in a given sequence. In some embodiments, a "substitution" comprises an amino acid deletion, which, if present, can be represented by a "-" symbol.
[0092] "Amino acid difference" and "residue difference" refer to the difference in an amino acid residue at a position in a polypeptide sequence relative to the amino acid residue at the corresponding position in a reference sequence. The position of an amino acid difference is generally referred to herein as "Xn," where n refers to the corresponding position in the reference sequence on which the residue difference is based. For example, "residue difference at position X14 compared to SEQ ID NO:2" (or "residue difference at position 14 compared to SEQ ID NO:2") refers to the difference in the amino acid residue at the polypeptide position corresponding to position 14 of SEQ ID NO:2. Thus, if a reference polypeptide of SEQ ID NO:2 has a lysine at position 14, then "residue difference at position X14 compared to SEQ ID NO:2" refers to an amino acid substitution of any residue other than lysine at the polypeptide position corresponding to position 14 of SEQ ID NO:2. In some cases herein, a specific amino acid residue difference at a position is designated as "XnY," where "Xn" identifies the corresponding residue and position in the reference polypeptide (as described above), and "Y" is the single-letter identifier of the amino acid found in the engineered polypeptide (i.e., the residue that differs from the reference polypeptide). In some cases (e.g., in the Tables of Examples), the disclosure also provides specific amino acid differences designated by the conventional designation "AnB," where A is the single-letter identifier of the residue in the reference sequence, "n" is the number of the residue position in the reference sequence, and B is the single-letter identifier of the residue substitution in the sequence of the engineered polypeptide. In some embodiments, an amino acid difference, e.g., a substitution, is designated by the abbreviation "nB," without the identifier of the residue in the reference sequence. In some embodiments, the phrase "amino acid residue nB" indicates the presence of an amino acid residue in the engineered polypeptide, which may or may not be a substitution relative to the reference polypeptide or amino acid sequence.
[0093] In some cases, the polypeptides of the present disclosure may contain one or more amino acid residue differences relative to a reference sequence, as indicated by a list of the specific positions where the residue difference exists relative to the reference sequence. In some embodiments, when more than one amino acid can be used at a particular residue position in a polypeptide, the various amino acid residues that can be used are separated by a " / " (e.g., X12A / X12I, X12A / I, or 129A / I).
[0094] "Amino acid substitution set" and "substitution set" refer to a group of amino acid substitutions within a polypeptide sequence. In some embodiments, a substitution set includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more amino acid substitutions. In some embodiments, a substitution set refers to a set of amino acid substitutions present in any of the mutant DNA ligase polypeptides listed in any of the tables of the Examples. In these substitution sets, individual substitutions are separated by a semicolon (";"; e.g., L105K;T317Q) or a slash (" / "; e.g., L105K / T317Q or 105K / 317Q).
[0095] " Conservative amino acid substitution " refers to the substitution of a residue with a different residue that has a similar side chain, and therefore typically includes the substitution of an amino acid in a polypeptide with the amino acid of the same or similar defined amino acid class.By way of example and not limitation, the amino acid with an aliphatic side chain can be substituted with another aliphatic amino acid (for example, alanine, valine, leucine and isoleucine); the amino acid with a hydroxyl side chain can be substituted with another amino acid with a hydroxyl side chain (for example, serine and threonine); the amino acid with an aromatic side chain can be substituted with another amino acid with an aromatic side chain (for example, phenylalanine, tyrosine, tryptophan and histidine); the amino acid with a basic side chain can be substituted with another amino acid with a basic side chain (for example, lysine and arginine); the amino acid with an acidic side chain can be substituted with another amino acid with an acidic side chain (for example, aspartic acid or glutamic acid); and the hydrophobic or hydrophilic amino acid can be substituted with another hydrophobic or hydrophilic amino acid, respectively.
[0096] "Non-conservative substitution" refers to the substitution of an amino acid in a polypeptide with an amino acid having significantly different side chain properties. Non-conservative substitutions may use amino acids between defined groups rather than within them, and may affect: (a) the structure of the peptide backbone in the area of the substitution (e.g., proline for glycine); (b) charge or hydrophobicity; and / or (c) the solid portion of the side chain. By way of example and not limitation, exemplary non-conservative substitutions include an acidic amino acid substituted with a basic or aliphatic amino acid; an aromatic amino acid substituted with a small amino acid; and a hydrophilic amino acid substituted with a hydrophobic amino acid.
[0097] "Deletion" refers to modification of a polypeptide by removal of one or more amino acids from a reference polypeptide. Deletions can include removal of one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to 10% of the total number of amino acids, or up to 20% of the total number of amino acids comprising the reference enzyme, while retaining the enzymatic activity and / or improving the properties of the engineered DNA ligase. Deletions can be directed to internal and / or terminal portions of the polypeptide. In various embodiments, deletions can include contiguous segments or can be discontinuous. In some embodiments, deletions are indicated by "-" and can be present in a substitution set.
[0098] "Insertion" refers to the modification of a polypeptide by the addition of one or more amino acids from a reference polypeptide.The insertion can be internal to the polypeptide, or at the carboxy or amino terminus.As used herein, the term "insertion" includes fusion proteins known in the art.The insertion can be a continuous segment of amino acids, or can be separated by one or more amino acids in a naturally occurring polypeptide.
[0099] "Functional fragment" and "biologically active fragment" are used interchangeably herein to refer to a polypeptide that has amino- and / or carboxy-terminal deletion(s) and / or internal deletions, but whose remaining amino acid sequence is identical to the corresponding positions in the sequence to which it is being compared (e.g., a full-length engineered DNA ligase of the present disclosure), and that retains substantially all of the activity of the full-length polypeptide.
[0100] An "isolated polypeptide" refers to a polypeptide that has been substantially separated from other contaminants (e.g., proteins, lipids, and polynucleotides) that naturally accompany it. The term encompasses polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., a host cell or in vitro synthesis). The engineered DNA ligase polypeptide may be present within a cell, in cell culture medium, or prepared in various forms, such as a lysate or isolated preparation.
[0101] A "substantially pure polypeptide" refers to a composition in which the polypeptide species is the predominant species present (i.e., more abundant than any other individual macromolecular species in the composition, on a molar or weight basis); generally, a composition is substantially purified when the subject species constitutes at least about 50 percent of the macromolecular species present, on a molar or weight percent basis. Generally, a substantially pure DNA ligase composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of all macromolecular species present in the composition, on a molar or weight percent basis. In some embodiments, the subject species is purified to essential homogeneity (i.e., contaminating species cannot be detected in the composition by conventional detection methods), such that the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 Daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated recombinant DNA ligase polypeptide is a substantially pure polypeptide composition.
[0102] "Improved enzymatic properties" refers to an engineered DNA ligase polypeptide that exhibits an improved enzymatic property compared to a reference DNA ligase polypeptide, such as a wild-type DNA ligase polypeptide or another engineered DNA ligase polypeptide. Improved properties may include, but are not limited to, properties such as increased protein expression, increased thermal activity, increased thermostability, increased stability, increased enzymatic activity, increased substrate specificity and / or affinity, increased substrate range, increased specific activity, increased resistance to substrate and / or end-product inhibition, increased chemical stability, improved solvent stability, increased solubility, and increased inhibitor resistance / tolerance.
[0103] "Increased enzymatic activity" and "enhanced catalytic activity" refer to improved properties of an engineered DNA ligase polypeptide and can be expressed as an increase in specific activity (e.g., product produced / time / weight protein) and / or an increase in the rate of substrate-to-product conversion (e.g., the rate of conversion of a starting amount of substrate to product in a specified period of time using a specified amount of DNA ligase) compared to a reference DNA ligase enzyme (e.g., a wild-type DNA ligase and / or another engineered DNA ligase). Exemplary methods for determining enzymatic activity are provided in the Examples. m , V max or k cat Any property associated with enzymatic activity can be affected, including classical enzymatic properties of, and the change can result in increased enzymatic activity. Improved enzymatic activity can range from about 1.1-fold the enzymatic activity of the corresponding wild-type enzyme to about 1.5-fold, 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold or more enzymatic activity than a naturally occurring DNA ligase or another engineered DNA ligase from which the DNA ligase polypeptide is derived.
[0104] "Hybridization stringency" refers to hybridization conditions, such as washing conditions, in nucleic acid hybridization. Generally, hybridization reactions are carried out under conditions of lower stringency, followed by washing under various but higher stringency conditions (e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York, 2001; Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, 2003). The term "moderately stringent hybridization" refers to conditions that allow target DNA to bind to complementary nucleic acids that have about 90% or more identity with target polynucleotides, and about 60% identity, preferably about 75% identity, about 85% identity with target DNA. Exemplary moderately stringent conditions are those equivalent to hybridization in 50% formamide, 5x Denhardt's solution, 5x SSPE, 0.2% SDS at 42°C, followed by a wash in 0.2x SSPE, 0.2% SDS at 42°C. "High stringency hybridization" generally refers to hybridization at a temperature higher than the thermal melting temperature, T, determined under solution conditions for a defined polynucleotide sequence. mTo about 10°C or less. In some embodiments, high stringency conditions refer to conditions that allow hybridization of only nucleic acid sequences that form stable hybrids in 0.018M NaCl at 65°C (i.e., if a hybrid is not stable in 0.018M NaCl at 65°C, it is not stable under high stringency conditions as contemplated herein). High stringency conditions can be provided, for example, by hybridization in conditions equivalent to 50% formamide, 5x Denhardt's solution, 5x SSPE, 0.2% SDS at 42°C, followed by washing in 0.1x SSPE and 0.1% SDS at 65°C. Another high stringency condition includes hybridizing in conditions equivalent to hybridizing in 5x SSC containing 0.1% (w:v) SDS at 65°C and washing in 0.1x SSC containing 0.1% SDS at 65°C. Other high stringency hybridization conditions, as well as moderately stringent conditions, are described in the references cited above.
[0105] "Optimized codons" refers to the alteration of codons in a polynucleotide encoding a protein relative to those preferentially used in a particular organism, so that the encoded protein is more efficiently expressed in that organism. Although the genetic code is degenerate, it is well known that codon usage by a particular organism is non-random and biased toward certain codon triplets, in that most amino acids are represented by several codons called "synonyms" or "synonymous" codons. This codon usage bias can be higher for a given gene, for genes with a common function or ancestral origin, for highly expressed proteins versus low copy number proteins, and for aggregated protein-coding regions of an organism's genome. In some embodiments, a polynucleotide encoding a DNA ligase enzyme is codon-optimized for optimal production from the host organism selected for expression.
[0106] The term "control sequences," as used herein, refers to all components necessary or advantageous for expression of the polynucleotides and / or polypeptides of the present disclosure. Each control sequence may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, a leader, polyadenylation sequence, propeptide sequence, promoter sequence, signal peptide sequence, initiation sequence, and transcription terminator. At a minimum, control sequences include a promoter, and transcriptional and translational stop signals. In some embodiments, linkers are provided in the control sequences for the purpose of introducing specific restriction sites that facilitate ligation of the control sequences with the coding region of the nucleic acid sequence encoding the polypeptide.
[0107] "Operably linked" or "operably linked" refers to a configuration in which a control sequence is placed in an appropriate position (i.e., in a functional relationship) with respect to a polynucleotide of interest such that the control sequence directs or regulates expression of the polynucleotide and, in some embodiments, expression of the encoded polypeptide of interest.
[0108] A "promoter" or "promoter sequence" refers to a nucleic acid sequence recognized by a host cell for expression of a polynucleotide of interest, such as a coding sequence. A promoter sequence comprises transcriptional control sequences that mediate expression of the polynucleotide of interest. The promoter may be any nucleic acid sequence (including mutant promoters, truncated promoters, and hybrid promoters) that exhibits transcriptional activity in the host cell of choice, and may be derived from genes encoding extracellular or intracellular polypeptides that are homologous or heterologous to the host cell.
[0109] "Suitable reaction conditions" or "suitable conditions" refer to conditions in an enzyme conversion reaction solution (e.g., ranges of enzyme load, substrate load, temperature, pH, buffer, cosolvent, cofactor, etc.) under which a DNA ligase polypeptide of the disclosure can convert polynucleotide substrate(s) into a desired ligated product polynucleotide. Exemplary "suitable reaction conditions" are provided herein (see Examples).
[0110] "Product," in the context of an enzymatic conversion process, refers to a compound or molecule that results from the action of a DNA ligase polypeptide on a substrate.
[0111] "Culturing" refers to the growth of a population of cells under appropriate conditions using any suitable medium (eg, liquid, gel, or solid).
[0112] "Vector" refers to a recombinant construct for introducing a polynucleotide of interest into a cell. In some embodiments, the vector is an expression vector operably linked to a suitable control sequence capable of effecting expression of the polynucleotide or a polypeptide encoded by the polynucleotide in a suitable host. In some embodiments, an "expression vector" has a promoter sequence operably linked to a polynucleotide (e.g., a transgene) to drive expression in a host cell, and in some embodiments, also includes a transcription terminator sequence.
[0113] "Expression" includes any step involved in producing a polypeptide of interest, including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from the cell.
[0114] "Produce" refers to the production of proteins and / or other compounds by a cell. The term is intended to encompass any step involved in producing a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from the cell.
[0115] "Heterologous" or "recombinant" refers to the relationship between two or more nucleic acid or polypeptide sequences (e.g., promoter sequences, signal peptides, terminator sequences, etc.) that are derived from different sources and are not related in nature.
[0116] "Host cell" and "host strain" refer to suitable hosts for expression vectors containing polynucleotides provided herein (e.g., polynucleotide sequences encoding at least one DNA ligase mutant). In some embodiments, host cells are prokaryotic or eukaryotic cells transformed or transfected with vectors constructed using recombinant DNA techniques known in the art.
[0117] Engineered DNA ligase polypeptides In one aspect, the present disclosure provides DNA ligases, including engineered DNA ligase polypeptide variants, that have DNA ligase activity and are characterized by improved properties compared to naturally occurring wild-type DNA ligases. In some embodiments, DNA ligases and engineered DNA ligase polypeptide variants are useful for ligating polynucleotide substrates, particularly DNA substrates. In some embodiments, engineered DNA ligases can be prepared and used as non-fusion or fusion polypeptides.
[0118] In some embodiments, the engineered DNA ligase, or functional fragment thereof, comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 and even-numbered SEQ ID NOs:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:2 and even-numbered SEQ ID NOs:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, 62, 138, 318, 722, or 938, or a reference sequence corresponding to SEQ ID NO:2, 62, 138, 318, 722, or 938.
[0119] In some embodiments, the engineered DNA ligase, or functional fragment thereof, comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or to a reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938.
[0120] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or to a reference sequence corresponding to SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0121] In some embodiments, the engineered DNA ligase comprises a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or to a reference sequence corresponding to SEQ ID NO: 2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2.
[0122] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or to a reference sequence corresponding to an even-numbered SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0123] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises the amino acid sequences at amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 148, 149, 150, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 288, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, and / or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or a reference sequence corresponding to SEQ ID NO:2.
[0124] Its specific ingredient is the specificity of the DNA fragment It also has a range of 11D, 12A / I and 13G / R、14G / S / T / V、18D / N / S、30C / H / S、31R、33M / R / V、34L / R、36T / Y、37G / L / N / S 44S, 50G / I / S / T, 56P, 59E, 60Y, 61T / V, 63F / R, 67R, 68A / M / S / V / Y, 69T, 71G / L / P / R、73C / K / P / T / V / W、74S、76F / G / H / L / N / R、77D、82R、88V、95A / L / R / V、9 6A / G / T / V、97G、99G / I、100V、101R、102G / K / L / S、103V、104K、105K / S / T、106 L / S / V, 110R, 112M, 113A / T, 117G / S / V / Y, 125R / T, 128C, 130T, 132R, 138L / R 139T, 148P, 149P, 150C / F / T, 155R, 156C, 159Q, 161R / V, 162W, 164A / R K, 177G, 186A / C / E / H / L / M / R / T / V, 188A, 189C / T, 190R, 191T, 195R, 196E / V 197R、198A / D / K / L / N / R / V / W、201L / S、205E / G / K、207L、208D / F / H、212F / G / M / S / W、220V、226D / E / Q / S / V、228E / I / M / S、230L / M、231P、232R、233G / T / W、23 5W、237L / M / R / S / V / Y、239M / N / P / Q / S / T / V / W、240E / G / K / Q / R / S / Y、242P / Q / T 251L, 254G / S, 258L / S / V, 263G / L / Q / T, 264A / C, 266M / T, 267D / W / Y, 269L 71A / G / N / S、273A / G / S、277Q / R、278E、282G / L / M / T / V / Y、283A / G / K / L / M / R / S / V、284D、286F / L / S、288I、289A / L / S / V、290L、294L、295K、297W、300G / T、30 1F / L、305K、306I / K / S / V、308K / L / S、309G / R、317Q、323S、328R、334L / R、337 G / L / M / P / R / S 339Y 349E 355S 356A / V / W 357H / K / P / R / S / V 358C 359N / R360H / M / P, 362G, 363R, 364R, 367C / L, 370C / G, 372N / Q, 374A / S, 375W, 378T, 379A / G / P, 380T, 381K / R, 382V, 38 4C / V, 386F, 387G, 388K / Y, 389K / L / Q / R, 390E, 392C / I / K / L / R / S, 396C / H, 397K / L / M, 404S, 405I, 408C / V, 414A / L / Q / R / T / V, 415A / C / E / H / I / K / L / V, 416K, 417D / G / L, 418A / G / I / L / M / P / S / T, 419G, 421R, 422N, 423R / T, or 428F / R / S, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or relative to a reference sequence corresponding to SEQ ID NO:2.
[0125] Its specific ingredient is the specificity of the DNA fragment It also includes G11D, M12A / I, T13G / R and K14 G / S / T / V, Q18D / N / S, T30C / H / S, S31R, T33M / R / V, A34L / R, E36T / Y, D37G / L / N / S、D44S、E50G / I / S / T、Y56P、D59E、L60Y、I61T / V、G63F / R、K67R、I68A / M / S / V / Y、K69T、K71G / L / P / R、L73C / K / P / T / V / W、P74S、D76F / G / H / L / N / R、Y77D、D8 2R、L88V、Y95A / L / R / V、N96A / G / T / V、R97G、L99G / I、T100V、G101R、N102G / K / L / S, A103V, A104K, L105K / S / T, D106L / S / V, V110R, L112M, S113A / T, Q117G / S / V / Y, K125R / T, Q128C, D130T, K132R, K138L / R, S139T, I148P, A149P, E150 C / F / T, L155R, A156C, L159Q, K161R / V, Y162W, K164A / R, R165K, C177G, K186 A / C / E / H / L / M / R / T / V、V188A、L189C / T、K190R、S191T、K195R、I196E / V、I197 R、T198A / D / K / L / N / R / V / W、T201L / S、Q205E / G / K、I207L、A208D / F / H、V212F / G / M / S / W、I220V、K226D / E / Q / S / V、A228E / I / M / S、V230L / M、Q231P、K232R、S2 33G / T / W、F235W、K237L / M / R / S / V / Y、D239M / N / P / Q / S / T / V / W、D240E / G / K / Q / R / S / Y、V242P / Q / T、V251L、H254G / S、A258L / S / V、R263G / L / Q / T、I264A / C、E2 66M / T、Q267D / W / Y、I269L、F271A / G / N / S、R273A / G / S、E277Q / R、Q278E、E282 G / L / M / T / V / Y F283A / G / K / L / M / R / S / V P284D W286F / L / S L288I E289A / L / S / V W290L D294L G295K F297W S300G / T E301F / L Q305K E306I / K / S / VF308K / L / S, H309G / R, T317Q, M323S, D328R, K334L / R, F337G / L / M / P / R / S, I339 Y, D349E, F355S, E356A / V / W, E357H / K / P / R / S / V, G358C, K359N / R, E360H / M / P, T 362G, K363R, N364R, V367C / L, A370C / G, V372N / Q, E374A / S, Y375W, N378T, E37 9A / G / P, V380T, S381K / R, I382V, G384C / V, Y386F, T387G, D388K / Y, E389K / L / Q / R, M390E, V392C / I / K / L / R / S, A396C / H, R397K / L / M, K404S, V405I, I408C / V, S414A / L / Q / R / T / V, T415A / C / E / H / I / K / L / V, S416K, S417D / G / L, K418A / G / I / L / M / P / S / T, T419G, K421R, K422N, S423R / T, or E428F / R / S or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0126] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2 or relative to a reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution 63F / R, 242P / Q / T, 283A / G / K / L / M / R / S / V, 286F / L / S, 317Q, 414A / L / Q / R / T / V, 418A / G / I / L / M / P / S / T, or 428F / R / S, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2 or relative to a reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least substitutions 63R, 242Q, 283L, 286S, 317Q, 414Q, 418S, or 428R, or a combination thereof, and the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0127] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 428, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution 428R, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution E428R, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:2.
[0128] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 317, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution 317Q, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution T317Q, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2.
[0129] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 283, 286, or 418, or a combination thereof, where the amino acid position is relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution 283L, 286S, or 418S, or a combination thereof, where the amino acid position is relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution F283L, W286S, or K418S, or a combination thereof, where the amino acid position is relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2.
[0130] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 242 or 414, or a combination thereof, where the amino acid position is relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution 242Q or 414Q, or a combination thereof, where the amino acid position is relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution V242Q or S414Q, or a combination thereof, where the amino acid position is relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or relative to the reference sequence corresponding to SEQ ID NO: 2.
[0131] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 63, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution 63R, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:2. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution G63R, the amino acid position relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:2.
[0132] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions at amino acid position(s) 233, 317, 191, 288, 207, 149, 251, 205, 269, 164, 36, 428, 105 / 132, or 105, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0133] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions 233T, 317Q, 191T, 288I, 207L, 149P, 251L, 205E, 269L, 164A, 36T, 428R, 105K / 132R or 105K, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0134] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions S233T, T317Q, S191T, L288I, I207L, A149P, V251L, Q205E, I269L, K164A, E36T, E428R, L105K / K132R, or L105K, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0135] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises the amino acid sequence at amino acid position(s) 196 / 428, 242 / 428, 337 / 428, 33 / 428, 277 / 428, 30 / 428, 359 / 428, 283 / 428, 415 / 428, 387 / 428, 379 / 428, 205 / 428, 186 / 428, 389 / 428, 102 / 428, 164 / 428, 301 / 428, 375 / 428, 267 / 428, 380 / 428, 254 / 428, 317 / 428, 77 / 139 / 317 / 417 / 428, 105 / 317 / 417 / 428, 317 / 349 / 362 / 386 / 428, 105 / 317 / 428, 139 / 317 428, 362 / 428, 233 / 317 / 405 / 428, 139 / 317 / 428, 162 / 428, 286 / 428, 414 / 428, 417 / 428, 226 / 428, 61 / 428, 105 / 428, 230 / 428, 418 / 428, 370 / 428, 297 / 428, 237 / 428, 428, 362 / 428, 233 / 428, 235 / 428, 148 / 428, 100 / 428, 97 / 428, 382 / 428, or 358 / 428, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0136] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 196E / 428R, 242P / 428R, 337L / 428R, 33R / 428R, 33V / 428R, 277Q / 428R, 337P / 428R, 30C / 428R, 359N / 428R, 337G / 428R, 283S / 428R, 242Q / 428R, 415A / 428R, 387G / 428R, 415H / 428R, 379A / 428R, 205G / 428R, 186M / 428R, 337M / 428R, 389Q / 428R, 102L ... / 428R, 389L / 428R, 359R / 428R, 205K / 428R, 164A / 428R, 415E / 428R, 301F / 4 28R, 30H / 428R, 379G / 428R, 375W / 428R, 267W / 428R, 380T / 428R, 415C / 428R , 254G / 428R, 186T / 428R, 317Q / 428R, 77D / 139T / 317Q / 417G / 428R, 105K / 31 7Q / 417G / 428R, 317Q / 349E / 362G / 386F / 428R, 105K / 317Q / 428R, 139T / 317Q / 362G / 428R, 233T / 317Q / 405I / 428R, 139T / 317Q / 428R, 162W / 428R, 283V / 42 8R, 286S / 428R, 414L / 428R, 417L / 428R, 226S / 428R, 61T / 428R, 226Q / 428R, 105S / 428R, 186E / 428R, 230M / 428R, 61V / 428R, 186V / 428R, 186L / 428R, 379 P / 428R, 415V / 428R, 415I / 428R, 186A / 428R, 283R / 428R, 226V / 428R, 418L / 4 28R, 418S / 428R, 370G / 428R, 283A / 428R, 337S / 428R, 297W / 428R, 237S / 428 R, 428S, 186C / 428R, 362G / 428R, 283G / 428R, 196V / 428R, 237L / 428R, 415L / 428R, 233W / 428R, 242T / 428R, 277R / 428R, 235W / 428R, 148P / 428R, 254S / 42 8R, 102G / 428R, 105T / 428R, 301L / 428R, 286F / 428R, 100V / 428R, 237Y / 428R,30S / 428R, 428F, 414Q / 428R, 414V / 428R, 283L / 428R, 414A / 428R, 414R / 428R, 370C / 428R, 283K / 428R, 286L / 428R, 186H / 428 R, 417D / 428R, 226D / 428R, 233G / 428R, 267Y / 428R, 97G / 428R, 226E / 428R, 418T / 428R, 230L / 428R, 414T / 428R, 418A / 428R, 2 37V / 428R, 418G / 428R, 382V / 428R, 267D / 428R, 283M / 428R, 418P / 428R, 237R / 428R, 237M / 428R, 418I / 428R, 358C / 428R, 186R / 428R, 418M / 428R, or 102S / 428R, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0137] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 11.2 and 12.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0138] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid position(s) 242 / 283 / 286 / 317 / 359 / 418 / 428, 283 / 286 / 317 / 428, 283 / 286 / 317 / 418 / 428, 186 / 242 / 283 / 286 / 317 / 418 / 428, 205 / 286 / 317 / 359 / 428, 283 / 317 / 428, 242 / 286 / 317 / 418 / 428, 186 / 205 / 242 / 317 / 428, 283 / 286 / 317 / 359 / 428, 277 / 286 / 317 / 3 59 / 418 / 428, 186 / 205 / 283 / 286 / 317 / 359 / 428, 242 / 277 / 317 / 418 / 428, 242 / 283 / 286 / 317 / 418 / 428, 112 / 196 / 317 / 389 / 428, 286 / 317 / 428, 283 / 286 / 317 / 359 / 418 / 428, 242 / 317 / 359 / 418 / 428, 186 / 283 / 317 / 428, 186 / 317 / 359 / 418 / 428, 205 / 242 / 317 / 418 / 428, 186 / 205 / 242 / 283 / 286 / 317 / 359 / 418 / 428, 205 / 317 / 359 / 418 / 428, 186 / 205 / 283 / 286 / 317 / 418 / 428, 186 / 283 / 317 / 359 / 428, 186 / 242 / 317 / 428, 186 / 242 / 317 / 359 / 428, 283 / 317 / 359 / 418 / 428, 277 / 317 / 418 / 428, 186 / 188 / 283 / 317 / 428, 186 / 286 / 317 / 418 / 428, 186 / 242 / 286 / 317 / 359 / 418 / 428, 317 / 418 / 428, 186 / 242 / 283 / 286 / 317 / 359 / 418 / 428, 186 / 277 / 317 / 359 / 418 / 428, 242 / 283 / 286 / 317 / 428, 205 / 317 / 418 / 428, 30 / 297 / 317 / 428, 205 / 242 / 286 / 317 / 359 / 418 / 428, 186 / 205 / 317 / 359 / 418 / 428, 317 / 359 / 418 / 428, 186 / 317 / 428, 230 / 317 / 428, 33 / 297 / 317 / 428, 186 / 205 / 317 / 428, 186 / 283 / 317 / 359 / 418 / 428, 186 / 317 / 418 / 428,205 / 242 / 283 / 317 / 359 / 418 / 428, 33 / 317 / 375 / 389 / 428, 33 / 230 / 317 / 428, 196 / 242 / 283 / 286 / 317 / 359 / 418 / 428, 186 / 242 / 283 / 317 / 359 / 418 / 428, 186 / 317 / 359 / 428, 33 / 196 / 317 / 428, 186 / 277 / 317 / 418 / 428, 242 / 317 / 428, 33 / 196 / 297 / 301 / 317 / 4 At least one substitution or set of substitutions at 28, 205 / 237 / 242 / 283 / 286 / 317 / 359 / 428, 186 / 205 / 283 / 317 / 359 / 418 / 428, 33 / 317 / 389 / 428, or 186 / 196 / 242 / 283 / 286 / 317 / 359 / 418 / 428, amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, or relative to a reference sequence corresponding to SEQ ID NO: 2.
[0139] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 242Q / 283V / 286S / 317Q / 359R / 418L / 428R, 283L / 286S / 317Q / 428R, 283L / 286S / 317Q / 418S / 428R, 186E / 242Q / 283V / 286S / 317Q / 418L / 428R, 205K / 286S / 317Q / 359R / 428R, 283V / 317Q / 428R, 242P / 286S / 317Q / 418T / 428R, 186E / 205K / 242T / 3 17Q / 428R, 283L / 286S / 317Q / 359R / 428R, 277Q / 286S / 317Q / 359N / 418T / 42 8R, 186A / 205K / 283V / 286S / 317Q / 359R / 428R, 242Q / 277Q / 317Q / 418L / 428 R, 242Q / 283S / 286S / 317Q / 418T / 428R, 112M / 196V / 317Q / 389Q / 428R, 286S / 317Q / 428R, 242P / 283R / 286S / 317Q / 359R / 418L / 428R, 283L / 286S / 317Q / 3 59R / 418S / 428R, 242Q / 317Q / 359N / 418S / 428R, 186E / 283L / 317Q / 428R, 18 6E / 317Q / 359R / 418S / 428R, 205K / 242T / 317Q / 418S / 428R, 186E / 205K / 242 Q / 283S / 286S / 317Q / 359R / 418S / 428R, 205K / 317Q / 359N / 418L / 428R, 186A / 205K / 283L / 286S / 317Q / 418L / 428R, 242T / 283V / 286S / 317Q / 359R / 418S / 428R, 186E / 283L / 317Q / 359R / 428R, 186E / 242P / 317Q / 428R, 283V / 286S / 3 17Q / 418S / 428R, 186E / 242T / 317Q / 359R / 428R, 283L / 317Q / 359R / 418L / 42 8R, 277Q / 317Q / 418S / 428R, 186E / 188A / 283V / 317Q / 428R, 186E / 286S / 317 Q / 418S / 428R, 186A / 242T / 286S / 317Q / 359N / 418L / 428R, 317Q / 418L / 428R,186A / 242P / 283V / 286S / 317Q / 359R / 418S / 428R、186A / 277Q / 317Q / 359R / 418S / 428R、242Q / 283R / 286S / 317Q / 428R、205K / 317Q / 418S / 428R、30C / 297W / 317Q / 428R、283L / 317Q / 359R / 418S / 428R、186E / 286S / 317Q / 418 L / 428R、317Q / 418T / 428R、205K / 242T / 286S / 317Q / 359R / 418S / 428R、18 6E / 205K / 317Q / 359N / 418T / 428R、317Q / 418S / 428R、186E / 242T / 283S / 2 86S / 317Q / 418S / 428R、317Q / 359R / 418S / 428R、242T / 286S / 317Q / 418S / 428R、186E / 317Q / 428R、230L / 317Q / 428R、33V / 297W / 317Q / 428R、186E / 205K / 317Q / 428R、186A / 283L / 317Q / 359R / 418S / 428R、186E / 317Q / 418L / 428R、186E / 277Q / 317Q / 359R / 418S / 428R、205K / 242T / 283R / 317Q / 359 R / 418L / 428R、186E / 317Q / 418T / 428R、33V / 317Q / 375W / 389Q / 428R、33R / 230L / 317Q / 428R、196E / 242T / 283V / 286S / 317Q / 359R / 418T / 428R、186 E / 242T / 283L / 317Q / 359R / 418S / 428R、186E / 317Q / 359R / 428R、33V / 196 V / 317Q / 428R、33R / 317Q / 375W / 389Q / 428R、186A / 277Q / 317Q / 418S / 428 R、186A / 242T / 283L / 286S / 317Q / 418L / 428R、242T / 317Q / 428R、33V / 196 V / 297W / 301F / 317Q / 428R、205K / 237Y / 242P / 283L / 286S / 317Q / 359R / 42 8R、186A / 205K / 283L / 317Q / 359R / 418T / 428R、33V / 317Q / 389Q / 428R、or 186A / 196E / 242T / 283L / 286S / 317Q / 359N / 418T / 428RIt is relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, or it is relative to a reference sequence corresponding to SEQ ID NO: 2.
[0140] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant shown in Table 13.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0141] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 283 / 286 / 317 / 363 / 418 / 428, 63 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 389 / 418 / 428, 283 / 286 / 317 / 381 / 418 / 428, 197 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 359 / 418 / 428, 102 / 283 / 286 / 317 / 418 / 428, 165 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 38 8 / 418 / 428, 283 / 286 / 317 / 414 / 418 / 428, 283 / 286 / 317 / 337 / 418 / 428, 164 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 416 / 418 / 428, 101 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 415 / 418 / 428, 283 / 286 / 317 / 418 / 423 / 428, 283 / 286 / 317 / 364 / 418 / 428, 73 / 283 / 286 / 317 / 418 / 428, 50 / 283 / 286 / 317 / 418 / 428, 71 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 388 / 418 / 419 / 428, 283 / 286 / 317 / 357 / 418 / 428, 283 / 286 / 317 / 396 / 418 / 428, 68 / 283 / 286 / 317 / 418 / 428, 76 / 28 3 / 286 / 317 / 418 / 428, 14 / 283 / 286 / 317 / 418 / 428, 271 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 360 / 418 / 428, 266 / 283 / 286 / 317 / 418 / 428, 208 / 283 / 286 / 317 / 418 / 428, 74 / 283 / 286 / 317 / 418 / 428, 263 / 283 / 286 / 317 / 418 / 428, 264 / 283 / 286 / 317 / 418 / 428, 13 / 283 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 378 / 418 / 428, 283 / 286 / 317 / 372 / 418 / 428, 283 / 286 / 300 / 317 / 418 / 428, 283 / 286 / 294 / 317 / 418 / 428, 283 / 286 / 290 / 317 / 418 / 428, 283 / 286 / 317 / 397 / 418 / 428,95 / 283 / 286 / 317 / 418 / 428、258 / 283 / 286 / 317 / 418 / 428、161 / 283 / 286 / 317 / 418 / 428、212 / 283 / 286 / 317 / 418 / 428、198 / 283 / 286 / 317 / 418 / 428、138 / 283 / 286 / 317 / 418 / 428、18 / 283 / 286 / 317 / 418 / 428、283 / 286 / 317 / 404 / 418 / 428、273 / 283 / 286 / 317 / 418 / 428、117 / 283 / 286 / 317 / 418 / 428、240 / 283 / 286 / 317 / 418 / 428、69 / 283 / 286 / 317 / 418 / 428、278 / 283 / 286 / 317 / 418 / 428、283 / 286 / 289 / 317 / 418 / 428、82 / 283 / 286 / 317 / 418 / 428、283 / 286 / 317 / 328 / 418 / 428、61 / 186 / 283 / 286 / 317 / 417 / 418 / 428、186 / 283 / 286 / 317 / 370 / 417 / 418 / 428、267 / 283 / 286 / 317 / 418 / 428、61 / 283 / 286 / 317 / 370 / 418 / 428、186 / 267 / 283 / 286 / 317 / 370 / 417 / 418 / 428、61 / 186 / 283 / 286 / 317 / 418 / 428、283 / 286 / 317 / 370 / 417 / 418 / 428、283 / 286 / 317 / 417 / 418 / 428、61 / 283 / 286 / 317 / 370 / 382 / 418 / 428、61 / 186 / 267 / 283 / 286 / 317 / 370 / 417 / 418 / 428、61 / 283 / 286 / 317 / 418 / 428、267 / 283 / 286 / 317 / 370 / 417 / 418 / 428、267 / 283 / 286 / 317 / 370 / 418 / 428、61 / 283 / 286 / 317 / 417 / 418 / 428、61 / 186 / 237 / 267 / 283 / 286 / 317 / 370 / 418 / 428、61 / 237 / 283 / 286 / 317 / 370 / 417 / 418 / 428、283 / 286 / 317 / 370 / 418 / 428、186 / 283 / 286 / 317 / 370 / 418 / 428、61 / 186 / 267 / 283 / 286 / 317 / 417 / 418 / 428、61 / 186 / 283 / 286 / 317 / 370 / 382 / 418 / 428、283 / 286 / 317 / 370 / 382 / 417 / 418 / 428、61 / 186 / 283 / 286 / 317 / 370 / 418 / 428、237 / 267 / 283 / 286 / 317 / 370 / 417 / 418 / 428、61 / 186 / 283 / 286 / 317 / 382 / 418 / 428、61 / 267 / 283 / 286 / 317 / 418 / 428、61 / 267 / 283 / 286 / 317 / 417 / 418 / 428、61 / 237 / 267 / 283 / 286 / 317 / 382 / 418 / 428、186 / 283 / 286 / 317 / 370 / 382 / 418 / 428、237 / 267 / 283 / 286 / 317 / 370 / 418 / 428、61 / 186 / 267 / 283 / 286 / 317 / 370 / 418 / 428、186 / 237 / 267 / 283 / 286 / 317 / 370 / 418 / 428、61 / 186 / 267 / 283 / 286 / 317 / 418 / 428、61 / 186 / 237 / 283 / 286 / 317 / 418 / 428、186 / 283 / 286 / 317 / 418 / 428、186 / 267 / 283 / 286 / 317 / 418 / 428、237 / 283 / 286 / 317 / 370 / 417 / 418 / 428、242 / 283 / 286 / 317 / 414 / 418 / 428、162 / 283 / 286 / 317 / 414 / 418 / 428、267 / 283 / 286 / 317 / 414 / 418 / 428、105 / 283 / 286 / 317 / 414 / 418 / 428、162 / 242 / 283 / 286 / 317 / 414 / 418 / 428、97 / 162 / 283 / 286 / 317 / 414 / 418 / 428、105 / 162 / 267 / 283 / 286 / 317 / 414 / 418 / 428、283 / 286 / 317 / 356 / 418 / 428、283 / 286 / 317 / 392 / 418 / 428、106 / 283 / 286 / 317 / 418 / 428、283 / 286 / 308 / 317 / 418 / 428、283 / 286 / 306 / 317 / 418 / 428、96 / 283 / 286 / 317 / 418 / 428、282 / 283 / 286 / 317 / 418 / 428、113 / 283 / 286 / 317 / 418 / 428、283 / 286 / 309 / 317 / 418 / 428、110 / 283 / 286 / 317 / 418 / 428、37 / 283 / 286 / 317 / 418 / 428、201 / 283 / 286 / 317 / 418 / 428、The amino acid sequences include at least a substitution or set of substitutions at 283 / 284 / 286 / 317 / 418 / 428, 283 / 286 / 317 / 374 / 418 / 428, 283 / 286 / 295 / 317 / 418 / 428, 88 / 283 / 286 / 317 / 418 / 428, 44 / 283 / 286 / 317 / 418 / 428, 12 / 283 / 286 / 317 / 418 / 428, or 283 / 286 / 317 / 390 / 418 / 428, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0142] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 283L / 286S / 317Q / 363R / 418S / 428R, 63R / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 389K / 418S / 428R, 283L / 286S / 317Q / 389K / 418S / 428R, 283L / 286S / 317Q / 389K / 418S / 428R, 197R / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 389K / 418S / 428R, 283L / 286S / 317Q / 389K / 418S / 428R, 197R / 283L / 286S / 317Q / 363R / 418S / 428R, ... 59R / 418S / 428R, 283L / 286S / 317Q / 381K / 418S / 428R, 102K / 283L / 286S / 31 7Q / 418S / 428R, 165K / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 388 K / 418S / 428R, 283L / 286S / 317Q / 414R / 418S / 428R, 283L / 286S / 317Q / 337R / 418S / 428R, 164R / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 416K / 418S / 428R, 101R / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 415K / 418S / 428R, 283L / 286S / 317Q / 418S / 423R / 428R, 283L / 286S / 317Q / 364R / 4 18S / 428R, 73P / 283L / 286S / 317Q / 418S / 428R, 50I / 283L / 286S / 317Q / 418S / 428R, 71L / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 388Y / 418S / 4 19G / 428R, 50G / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 357K / 418 S / 428R, 283L / 286S / 317Q / 396H / 418S / 428R, 68Y / 283L / 286S / 317Q / 418S / 428R, 76F / 283L / 286S / 317Q / 418S / 428R, 71P / 283L / 286S / 317Q / 418S / 428 R, 14V / 283L / 286S / 317Q / 418S / 428R, 271G / 283L / 286S / 317Q / 418S / 428R,68A / 283L / 286S / 317Q / 418S / 428R、68M / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 360P / 418S / 428R、266M / 283L / 286S / 317Q / 418S / 428R、208D / 283L / 286S / 317Q / 418S / 428R、74S / 283L / 286S / 317Q / 418S / 428R、263T / 283L / 286S / 317Q / 418S / 428R、264C / 283L / 286S / 317Q / 418S / 428R、13R / 283L / 286S / 317Q / 418S / 428R、263L / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 360M / 418S / 428R、283L / 286S / 317Q / 378T / 418S / 428R、76H / 283L / 286S / 317Q / 418S / 428R、13G / 283L / 286S / 317Q / 418S / 428R、76R / 283L / 286S / 317Q / 418S / 428R、50T / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 372Q / 418S / 428R、50S / 283L / 286S / 317Q / 418S / 428R、283L / 286S / S300G / 317Q / 418S / 428R、73T / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 294L / 317Q / 418S / 428R、283L / 286S / 317Q / 372N / 418S / 428R、76N / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 290L / 317Q / 418S / 428R、283L / 286S / 317Q / 397L / 418S / 428R、76G / 283L / 286S / 317Q / 418S / 428R、95L / 283L / 286S / 317Q / 418S / 428R、258L / 283L / 286S / 317Q / 418S / 428R、161V / 283L / 286S / 317Q / 418S / 428R、212M / 283L / 286S / 317Q / 418S / 428R、198V / 283L / 286S / 317Q / 418S / 428R、138L / 283L / 286S / 317Q / 418S / 428R、18D / 283L / 286S / 317Q / 418S / 428R、18S / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 404S / 418S / 428R、273G / 283L / 286S / 317Q / 418S / 428R、117G / 283L / 286S / 317Q / 418S / 428R、240E / 283L / 286S / 317Q / 418S / 428R、68S / 283L / 286S / 317Q / 418S / 428R、73C / 283L / 286S / 317Q / 418S / 428R、69T / 283L / 286S / 317Q / 418S / 428R、278E / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 360H / 418S / 428R、198D / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 300T / 317Q / 418S / 428R、161R / 283L / 286S / 317Q / 418S / 428R、73K / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 289V / 317Q / 418S / 428R、82R / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 328R / 418S / 428R、61T / 186A / 283L / 286S / 317Q / 417D / 418S / 428R、186A / 283L / 286S / 317Q / 370C / 417L / 418S / 428R、267Y / 283L / 286S / 317Q / 418S / 428R、61T / 283L / 286S / 317Q / 370C / 418S / 428R、186A / 267Y / 283L / 286S / 317Q / 370C / 417L / 418S / 428R、61T / 186C / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 370C / 417D / 418S / 428R、283L / 286S / 317Q / 417D / 418S / 428R、61T / 283L / 286S / 317Q / 370C / 382V / 418S / 428R、61T / 186A / 267Y / 283L / 286S / 317Q / 370C / 417D / 418S / 428R、61T / 283L / 286S / 317Q / 418S / 428R、267Y / 283L / 286S / 317Q / 370C / 417D / 418S / 428R、267Y / 283L / 286S / 317Q / 370C / 418S / 428R、61T / 283L / 286S / 317Q / 417L / 418S / 428R、61T / 186H / 237R / 267Y / 283L / 286S / 317Q / 370C / 418S / 428R、61T / 237R / 283L / 286S / 317Q / 370C / 417L / 418S / 428R、283L / 286S / 317Q / 370C / 417L / 418S / 428R、283L / 286S / 317Q / 370C / 418S / 428R、186E / 283L / 286S / 317Q / 370C / 418S / 428R、61T / 186E / 267Y / 283L / 286S / 317Q / 417D / 418S / 428R、61T / 186C / 283L / 286S / 317Q / 370C / 382V / 418S / 428R、283L / 286S / 317Q / 370C / 382V / 417D / 418S / 428R、61T / 186C / 267Y / 283L / 286S / 317Q / 417D / 418S / 428R、61T / 186C / 283L / 286S / 317Q / 370C / 418S / 428R、237R / 267Y / 283L / 286S / 317Q / 370C / 417D / 418S / 428R、61T / 186E / 283L / 286S / 317Q / 417D / 418S / 428R、61T / 186H / 283L / 286S / 317Q / 418S / 428R、61T / 186C / 283L / 286S / 317Q / 382V / 418S / 428R、61T / 186H / 283L / 286S / 317Q / 370C / 418S / 428R、61T / 267Y / 283L / 286S / 317Q / 418S / 428R、61T / 267Y / 283L / 286S / 317Q / 417L / 418S / 428R、61T / 237R / 267Y / 283L / 286S / 317Q / 382V / 418S / 428R、186H / 283L / 286S / 317Q / 370C / 382V / 418S / 428R、237R / 267Y / 283L / 286S / 317Q / 370C / 418S / 428R、186H / 283L / 286S / 317Q / 370C / 417D / 418S / 428R、61T / 186V / 267Y / 283L / 286S / 317Q / 370C / 418S / 428R、267Y / 283L / 286S / 317Q / 370C / 417L / 418S / 428R、61T / 186V / 283L / 286S / 317Q / 370C / 418S / 428R、186H / 237R / 267Y / 283L / 286S / 317Q / 370C / 418S / 428R、61T / 186E / 267Y / 283L / 286S / 317Q / 418S / 428R、61T / 186C / 237R / 283L / 286S / 317Q / 418S / 428R、61T / 186H / 237R / 283L / 286S / 317Q / 418S / 428R、61T / 186C / 283L / 286S / 317Q / 417L / 418S / 428R、186H / 283L / 286S / 317Q / 418S / 428R、186A / 267Y / 283L / 286S / 317Q / 418S / 428R、186C / 283L / 286S / 317Q / 370C / 418S / 428R、61T / 186C / 237R / 267Y / 283L / 286S / 317Q / 370C / 418S / 428R、186E / 267Y / 283L / 286S / 317Q / 418S / 428R、237R / 283L / 286S / 317Q / 370C / 417L / 418S / 428R、186H / 283L / 286S / 317Q / 370C / 418S / 428R、242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、283L / 286S / 317Q / 414T / 418S / 428R、283L / 286S / 317Q / 414Q / 418S / 428R、162W / 283L / 286S / 317Q / 414Q / 418S / 428R、162W / 283L / 286S / 317Q / 414V / 418S / 428R、267D / 283L / 286S / 317Q / 414T / 418S / 428R、283L / 286S / 317Q / 414V / 418S / 428R、267D / 283L / 286S / 317Q / 414Q / 418S / 428R、283L / 286S / 317Q / 414A / 418S / 428R、162W / 283L / 286S / 317Q / 414L / 418S / 428R、105S / 283L / 286S / 317Q / 414T / 418S / 428R、162W / 242Q / 283L / 286S / 317Q / 414T / 418S / 428R、97G / 162W / 283L / 286S / 317Q / 414Q / 418S / 428R、105S / 162W / 267D / 283L / 286S / 317Q / 414V / 418S / 428R、283L / 286S / 317Q / 356V / 418S / 428R、273A / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 357P / 418S / 428R、14G / 283L / 286S / 317Q / 418S / 428R、14S / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 396C / 418S / 428R、240R / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 392K / 418S / 428R、273S、 / 283L / 286S / 317Q / 418S / 428R、106V / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 308S / 317Q / 418S / 428R、283L / 286S / 308L / 317Q / 418S / 428R、283L / 286S / 306I / 317Q / 418S / 428R、96A / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 397K / 418S / 428R、263Q / 283L / 286S / 317Q / 418S / 428R、282T / 283L / 286S / 317Q / 418S / 428R、138R / 283L / 286S / 317Q / 418S / 428R、258S / 283L / 286S / 317Q / 418S / 428R、76L / 283L / 286S / 317Q / 418S / 428R、14T / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 289S / 317Q / 418S / 428R、240G / 283L / 286S / 317Q / 418S / 428R、106S / 283L / 286S / 317Q / 418S / 428R、117S / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 357S / 418S / 428R、96G / 283L / 286S / 317Q / 418S / 428R、113T / 283L / 286S / 317Q / 418S / 428R、71G / 283L / 286S / 317Q / 418S / 428R、282V / 283L / 286S / 317Q / 418S / 428R、117V / 283L / 286S / 317Q / 418S / 428R、282M / 283L / 286S / 317Q / 418S / 428R、258V / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 309G / 317Q / 418S / 428R、283L / 286S / 306K / 317Q / 418S / 428R、283L / 286S / 308K / 317Q / 418S / 428R、283L / 286S / 317Q / 357R / 418S / 428R、212S / 283L / 286S / 317Q / 418S / 428R、18N / 283L / 286S / 317Q / 418S / 428R、113A / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 306S / 317Q / 418S / 428R、212F / 283L / 286S / 317Q / 418S / 428R、198N / 283L / 286S / 317Q / 418S / 428R、110R / 283L / 286S / 317Q / 418S / 428R、240K / 283L / 286S / 317Q / 418S / 428R、198L / 283L / 286S / 317Q / 418S / 428R、71R / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 357V / 418S / 428R、198R / 283L / 286S / 317Q / 418S / 428R、271A / 283L / 286S / 317Q / 418S / 428R、282Y / 283L / 286S / 317Q / 418S / 428R、212G / 283L / 286S / 317Q / 418S / 428R、95V / 283L / 286S / 317Q / 418S / 428R、37S / 283L / 286S / 317Q / 418S / 428R、68V / 283L / 286S / 317Q / 418S / 428R、240S / 283L / 286S / 317Q / 418S / 428R、117Y / 283L / 286S / 317Q / 418S / 428R、201S / 283L / 286S / 317Q / 418S / 428R、271N / 283L / 286S / 317Q / 418S / 428R、240Q / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 289L / 317Q / 418S / 428R、212W / 283L / 286S / 317Q / 418S / 428R、198W / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 357H / 418S / 428R、208F / 283L / 286S / 317Q / 418S / 428R、282L / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 306V / 317Q / 418S / 428R、283L / 284D / 286S / 317Q / 418S / 428R、96T / 283L / 286S / 317Q / 418S / 428R、266T / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / 374S / 418S / 428R、95R / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 309R / 317Q / 418S / 428R、73W / 283L / 286S / 317Q / 418S / 428R、283L / 286S / 317Q / R397M / 418S / 428R, 283L / 286S / 289A / 317Q / 418S / 428R, 283L / 286S / 295K / 317Q / 41 8S / 428R, 106L / 283L / 286S / 317Q / 418S / 428R, 95A / 283L / 286S / 317Q / 418S / 428R, 283L / 286S / 317Q / 37 4A / 418S / 428R, 88V / 283L / 286S / 317Q / 418S / 428R, 96V / 283L / 286S / 317Q / 418S / 428R, 271S / 283L / 286 S / 317Q / 418S / 428R, 44S / 283L / 286S / 317Q / 418S / 428R, 208H / 283L / 286S / 317Q / 418S / 428R, 263G / 283 L / 286S / 317Q / 418S / 428R, 12A / 283L / 286S / 317Q / 418S / 428R, 201L / 283L / 286S / 317Q / 418S / 428R, 19 8A / 283L / 286S / 317Q / 418S / 428R, 198K / 283L / 286S / 317Q / 418S / 428R, 282G / 283L / 286S / 317Q / 418S / 4 28R, 283L / 286S / 317Q / 390E / 418S / 428R, 264A / 283L / 286S / 317Q / 418S / 428R, or 73V / 283L / 286S / 317Q / 418S / 428R, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or relative to a reference sequence corresponding to SEQ ID NO:2.
[0143] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or substitutions set forth in Tables 14.2, 15.2, and 16.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0144] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 63 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 63 / 96 / 242 / 283 / 286 / 317 / 370 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 389 / 414 / 418 / 428, 13 / 242 / 267 / 283 / 286 / 317 / 363 / 389 / 414 / 418 / 428, 13 / 186 / 242 / 283 / 286 / 317 / 389 / 414 / 418 / 428, 50 / 242 / 267 / 283 / 286 / 31 7 / 363 / 370 / 389 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 363 / 370 / 414 / 418 / 428, 96 / 242 / 283 / 286 / 317 / 370 / 414 / 418 / 428, 61 / 63 / 212 / 242 / 283 / 286 / 317 / 4 14 / 418 / 428, 11 / 242 / 283 / 286 / 305 / 317 / 414 / 418 / 428, 11 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 428, 242 / 283 / 286 / 317 / 323 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 334 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 339 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 356 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 384 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 408 / 414 / 418 / 428, 67 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 392 / 414 / 418 / 428, 104 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 242 / 2 83 / 286 / 317 / 355 / 414 / 418 / 428, 159 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 155 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 367 / 414 / 418 / 428, 3 1 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 231 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 36 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 150 / 242 / 283 / 286 / 317 / 414 / 418 / 428,239 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 103 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 125 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 228 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 37 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 189 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 177 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 414 / 418 / 422 / 428, 128 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 220 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 130 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 56 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 190 / 242 / 283 / 28 6 / 317 / 414 / 418 / 428, 156 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 232 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 414 / 418 / 423 / 428, 34 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 99 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 59 / 242 / 283 / 286 / 317 / 414 242 / 283 / 286 / 317 / 414 / 418 / 421 / 428, 60 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 242 / 283 / 286 / 317 / 414 / 418 / 421 / 428, or 195 / 242 / 283 / 286 / 317 / 414 / 418 / 428, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or relative to a reference sequence corresponding to SEQ ID NO:2.
[0145] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 63R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 63R / 96T / 242Q / 283L / 286S / 317Q / 370C / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 389K / 414Q / 418S / 428R, 13R / 242Q / 267Y / 283L / 286S / 317Q / 363R / 389K / 414Q / 418S / 428R, 13R / 186H ... 2Q / 283L / 286S / 317Q / 389K / 414Q / 418S / 428R, 50T / 242Q / 267Y / 283L / 286 S / 317Q / 363R / 370C / 389K / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 363 R / 370C / 414Q / 418S / 428R, 96T / 242Q / 283L / 286S / 317Q / 370C / 414Q / 418S / 428R, 61T / 63R / 212W / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 11D / 242 Q / 283L / 286S / Q305K / 317Q / 414Q / 418S / 428R, 11D / 242Q / 283L / 286S / 317 Q / 414Q / 418S / 428R, 428R, 242Q / 283L / 286S / 317Q / 323S / 414Q / 418S / 428 R, 242Q / 283L / 286S / 317Q / 334L / 414Q / 418S / 428R, 242Q / 283L / 286S / 317 Q / 339Y / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 356V / 414Q / 418S / 428 R, 242Q / 283L / 286S / 317Q / 384V / 414Q / 418S / 428R, 242Q / 283L / 286S / 317 Q / 408C / 414Q / 418S / 428R, 67R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R , 242Q / 283L / 286S / 317Q / 392R / 414Q / 418S / 428R, 104K / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 355S / 414Q / 418S / 428R,159Q / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、155R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 367L / 414Q / 418S / 428R、31R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、231P / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、36Y / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、150T / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239V / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、103V / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 356A / 414Q / 418S / 428R、63F / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、125T / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 367C / 414Q / 418S / 428R、228I / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、37S / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、37G / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239M / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239W / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239Q / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、189C / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239S / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239T / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、177G / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239P / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、37N / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 414Q / 418S / 422N / 428R、228S / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、125R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、128C / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、189T / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 356W / 414Q / 418S / 428R、220V / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 408V / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 392S / 414Q / 418S / 428R、130T / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、228M / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、56P / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、228E / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、190R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 392I / 414Q / 418S / 428R、156C / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、232R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、150F / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、150C / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、239N / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 392L / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 414Q / 418S / 423T / 428R、34L / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、99I / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、242Q / 283L / 286S / 317Q / 384C / 414Q / 418S / 428R、59E / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 334R / 414Q / 418S / 428R, 60Y / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 99G / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 3 4R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 37L / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 392C / 414Q / 418S / 428R, 242Q / 283L / 286S / 317Q / 414Q / 418S / 421R / 428R, or 195R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0146] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant shown in Table 17.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0147] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 63 / 242 / 283 / 286 / 308 / 317 / 357 / 390 / 414 / 418 / 428, 63 / 74 / 76 / 201 / 242 / 283 / 286 / 308 / 317 / 357 / 414 / 418 / 428, 61 / 63 / 74 / 76 / 186 / 201 / 242 / 283 / 286 / 308 / 309 / 317 / 357 / 390 / 414 / 418 / 428, 14 / 63 / 201 / 240 / 242 / 283 / 286 / 289 / 317 / 357 / 414 / 4 18 / 428, 63 / 242 / 283 / 286 / 308 / 317 / 414 / 415 / 418 / 428, 63 / 76 / 242 / 283 / 286 / 317 / 357 / 396 / 414 / 418 / 428, 63 / 242 / 263 / 283 / 286 / 308 / 317 / 396 / 414 / 418 / 428, 61 / 63 / 76 / 96 / 240 / 242 / 283 / 286 / 308 / 309 / 317 / 414 / 418 / 428, 14 / 63 / 242 / 283 / 286 / 306 / 317 / 414 / 415 / 418 / 428, 14 / 63 / 73 / 106 / 242 / 28 3 / 286 / 317 / 414 / 415 / 418 / 428, 12 / 14 / 63 / 242 / 258 / 263 / 283 / 286 / 289 / 308 / 309 / 317 / 396 / 414 / 418 / 428, 63 / 74 / 76 / 117 / 242 / 283 / 286 / 309 / 317 / 357 / 414 / 418 / 428, 14 / 63 / 242 / 258 / 263 / 283 / 286 / 317 / 357 / 396 / 414 / 418 / 428, 14 / 63 / 96 / 106 / 242 / 283 / 286 / 306 / 317 / 414 / 418 / 428, 14 / 63 / 106 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 12 / 14 / 63 / 242 / 283 / 286 / 308 / 309 / 317 / 414 / 418 / 428, 14 / 63 / 242 / 283 / 286 / 317 / 357 / 390 / 414 / 418 / 428, 14 / 63 / 117 / 242 / 258 / 283 / 286 / 309 / 317 / 357 / 414 / 418 / 428, 63 / 242 / 283 / 286 / 317 / 390 / 414 / 418 / 428, 63 / 240 / 242 / 273 / 283 / 286 / 317 / 357 / 390 / 414 / 418 / 428,61 / 63 / 76 / 186 / 201 / 242 / 283 / 286 / 308 / 309 / 317 / 414 / 418 / 428, 14 / 63 / 242 / 283 / 286 / 317 / 396 / 414 / 418 / 428, 63 / 242 / 283 / 286 / 309 / 317 / 414 / 418 / 4 28, 14 / 63 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 63 / 106 / 242 / 283 / 286 / 306 / 308 / 317 / 414 / 418 / 428, 14 / 63 / 240 / 242 / 283 / 286 / 306 / 308 / 317 / 414 / 418 / 42 8, 12 / 14 / 63 / 186 / 242 / 283 / 286 / 317 / 357 / 414 / 418 / 428, 63 / 242 / 283 / 286 / 309 / 317 / 390 / 414 / 418 / 428, 14 / 63 / 242 / 283 / 286 / 306 / 317 / 414 / 418 / 428 , 14 / 63 / 76 / 242 / 283 / 286 / 308 / 317 / 414 / 418 / 428, 63 / 117 / 208 / 242 / 258 / 263 / 283 / 286 / 289 / 308 / 309 / 317 / 414 / 418 / 428, 14 / 63 / 73 / 106 / 242 / 283 / 28 6 / 317 / 414 / 418 / 428, 63 / 76 / 208 / 242 / 263 / 283 / 286 / 317 / 414 / 418 / 428, 63 / 242 / 283 / 286 / 317 / 357 / 414 / 418 / 428, 14 / 63 / 242 / 283 / 286 / 308 / 317 / 41 4 / 418 / 428, 63 / 242 / 263 / 283 / 286 / 317 / 414 / 418 / 428, 63 / 76 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 14 / 63 / 242 / 283 / 286 / 300 / 308 / 317 / 414 / 415 / 418 / 428 , 63 / 240 / 242 / 283 / 286 / 317 / 414 / 418 / 428, 33 / 63 / 242 / 283 / 286 / 317 / 357 / 390 / 414 / 418 / 428, 14 / 63 / 76 / 242 / 273 / 283 / 286 / 317 / 414 / 418 / 428, or 63 / 74 / 242 / 283 / 286 / 317 / 414 / 418 / 428, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0148] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 63R / 242Q / 283L / 286S / 308L / 317Q / 357S / 390E / 414Q / 418S / 428R, 63R / 74S / 76L / 201S / 242Q / 283L / 286S / 308L / 317Q / 357P / 414Q / 418S / 428R, 61T / 63R / 74S / 76L / 186A / 201S / 242Q / 283L / 286S / 308L / 309R / 317Q / 357P / 390E / 414Q / 418S / 428R, 63R / 74S / 76L / 186A / 201S / 242Q / 283L / 286S / 308L / 309R / 317Q / 357P / 390E / 414Q / 418S / 428R, 8S / 428R, 14G / 63R / 201S / 240R / 242Q / 283L / 286S / 289S / 317Q / 357P / 414 Q / 418S / 428R, 63R / 242Q / 283L / 286S / 308S / 317Q / 414Q / 415I / 418S / 428 R, 63R / 76L / 242Q / 283L / 286S / 317Q / 357S / 396C / 414Q / 418S / 428R, 63R / 242Q / 263L / 283L / 286S / 308L / 317Q / 396C / 414Q / 418S / 428R, 61T / 63R / 76 L / 96A / 240R / 242Q / 283L / 286S / 308L / 309R / 317Q / 414Q / 418S / 428R, 14S / 63R / 242Q / 283L / 286S / 306I / 317Q / 414Q / 415I / 418S / 428R, 14S / 63R / 7 3W / 106S / 242Q / 283L / 286S / 317Q / 414Q / 415I / 418S / 428R, 12A / 14G / 63R / 242Q / 258S / 263Q / 283L / 286S / 289S / 308L / 309R / 317Q / 396C / 414Q / 418S / 428R, 63R / 74S / 76L / Q117V / 242Q / 283L / 286S / 309R / 317Q / 357S / 414Q / 418S / 428R, 14T / 63R / 242Q / 258S / 263Q / 283L / 286S / 317Q / 357S / 396C / 4 14Q / 418S / 428R, 14S / 63R / 96G / 106S / 242Q / 283L / 286S / 306I / 317Q / 414 Q / 418S / 428R, 14S / 63R / 106S / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R,12A / 14G / 63R / 242Q / 283L / 286S / 308L / 309R / 317Q / 414Q / 418S / 428R、14G / 63R / 242Q / 283L / 286S / 317Q / 357S / 390E / 414Q / 418S / 428R、14G / 63R / 117V / 242Q / 258S / 283L / 286S / 309R / 317Q / 357S / 414Q / 418S / 428R、63R / 242Q / 283L / 286S / 317Q / 390E / 414Q / 418S / 428R、63R / 240R / 242Q / 273S / 283L / 286S / 317Q / 357P / 390E / 414Q / 418S / 428R、61T / 63R / 76L / 186V / 201S / 242Q / 283L / 286S / 308L / 309R / 317Q / 414Q / 418S / 428R、14G / 63R / 242Q / 283L / 286S / 317Q / 396C / 414Q / 418S / 428R、63R / 242Q / 283L / 286S / 309R / 317Q / 414Q / 418S / 428R、14S / 63R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、63R / 106V / 242Q / 283L / 286S / 306I / 308S / 317Q / 414Q / 418S / 428R、14S / 63R / 240S / 242Q / 283L / 286S / 306I / 308S / 317Q / 414Q / 418S / 428R、12I / 14G / 63R / 186V / 242Q / 283L / 286S / 317Q / 357S / 414Q / 418S / 428R、63R / 242Q / 283L / 286S / 309R / 317Q / 390E / 414Q / 418S / 428R、14S / 63R / 242Q / 283L / 286S / 306I / 317Q / 414Q / 418S / 428R、14T / 63R / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、14S / 63R / 76R / 242Q / 283L / 286S / 308S / 317Q / 414Q / 418S / 428R、63R / 117V / 208D / 242Q / 258S / 263Q / 283L / 286S / 289S / 308L / 309R / 317Q / 414Q / 418S / 428R、14S / 63R / 73W / 106S / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R、63R / 76L / 208D / 242Q / 263Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 63R / 242Q / 283L / 28 6S / 317Q / 357S / 414Q / 418S / 428R, 14S / 63R / 242Q / 283L / 286S / 308S / 317Q / 414Q / 418S / 428R, 63R / 242Q / 263L / 283L / 286S / 317Q / 414Q / 418S / 428R, 63R / 76L / 242Q / 283L / 2 86S / 317Q / 414Q / 418S / 428R, 14S / 63R / 242Q / 283L / 286S / 300G / 308S / 317Q / 414Q / 415 I / 418S / 428R, 63R / 240Y / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, 33M / 63R / 242Q / 283L / 286S / 317Q / 357S / 390E / 414Q / 418S / 428R, 14G / 63R / 76L / 242Q / 273S / 283L / 286S / 317Q / 414Q / 418S / 428R, or 63R / 74S / 242Q / 283L / 286S / 317Q / 414Q / 418S / 428R, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0149] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Table 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0150] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises substitutions at least at the amino acid positions set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0151] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0152] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions at an amino acid position(s) set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0153] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0154] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to the reference sequence corresponding to SEQ ID NO:2.
[0155] In some embodiments, the engineered DNA ligase comprises a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0156] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:40-1184.
[0157] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938.
[0158] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, 138, 318, 722, or 938, or relative to the reference sequence corresponding to SEQ ID NO:62, 138, 318, 722, or 938.
[0159] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 19 8, 139, 148, 149, 150, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 288, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, 384, 386, 387, 388, 389, 390, 392, 393 6, 397, 404, 405, 408, 414, 415, 416, 417, 418, 419, 421, 422, 423, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 62, 138, 318, 722, or 938, or relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0160] Its specific ingredient is the specificity of the DNA fragment It also has a range of 11D, 12A / I and 13G / R、14G / S / T / V、18D / N / S、30C / H / S、31R、33M / R / V、34L / R、36T / Y、37G / L / N / S 44S, 50G / I / S / T, 56P, 59E, 60Y, 61T / V, 63F / R, 67R, 68A / M / S / V / Y, 69T, 71G / L / P / R、73C / K / P / T / V / W、74S、76F / G / H / L / N / R、77D、82R、88V、95A / L / R / V、96 A / G / T / V、97G、99G / I、100V、101R、102G / K / L / S、103V、104K、105K / S / T、106L / S / V、110R、112M、113A / T、117G / S / V / Y、125R / T、128C、130T、132R、138L / R、1 39T, 148P, 149P, 150C / F / T, 155R, 156C, 159Q, 161R / V, 162W, 164A / R, 165K 177G、186A / C / E / H / L / M / R / T / V、188A、189C / T、190R、191T、195R、196E / V、197 R、198A / D / K / L / N / R / V / W、201L / S、205E / G / K、207L、208D / F / H、212F / G / M / S / W、220V、226D / E / Q / S / V、228E / I / M / S、230L / M、231P、232R、233G / T / W、235W、2 37L / M / R / S / V / Y、239M / N / P / Q / S / T / V / W、240E / G / K / Q / R / S / Y、242P / Q / T / V、2 51L, 254G / S, 258L / S / V, 263G / L / Q / T, 264A / C, 266M / T, 267D / W / Y, 269L, 271A / G / N / S、273A / G / S、277Q / R、278E、282G / L / M / T / V / Y、283A / F / G / K / L / M / R / S / V、284D、286F / L / S / W、288I、289A / L / S / V、290L、294L、295K、297W、300G / T、30 1F / L、305K、306I / K / S / V、308K / L / S、309G、317T / Q、323S、328R、334L / R、337 G / L / M / P / R / S 339Y 349E 355S 356A / V / W 357H / K / P / R / S / V 358C 359N / R360H / M / P, 362G, 364R, 367C / L, 370C / G, 372N / Q, 374A / S, 375W, 378T, 379A / G / P, 380T, 381K / R, 382V, 384C / V, 386F, 387G, 3 88K / Y, 389K / L / Q / R, 390E, 392C / I / K / L / R / S, 396C / H, 397K / L / M, 404S, 405I, 408C / V, 414A / L / Q / R / S / T / V, 415A / C / E / H / I / K 421R, 422N, 423R / T, or 428E / F / R / S, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0161] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises substitutions or amino acid residues 63F / R, 242P / Q / T / V, 283A / F / G / K / L / M / R / S / V, 286F / L / S / W, 317T / Q, 414A / L / Q / R / S / T / V, 418A / G / I / K / L / M / P / S / T, or 428E / F / R / S, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938. In some embodiments, the amino acid sequence of the engineered DNA ligase comprises a substitution at least at amino acid residue 63R, 242Q, 283L, 286S, 317Q, 414Q, 418S, or 428R, or a combination thereof, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0162] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to SEQ ID NO:62.
[0163] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to an even-numbered SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62.
[0164] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid position(s) 196, 242, 337, 33, 277, 30, 359, 283, 415, 387, 379, 205, 186, 389, 102, 164, 301, 375, 267, 380, 254, 317, 77 / 139 / 317 / 417, 105 / 317 / 417, 317 / 349 / 362 / 386, 105 / 317, 139 / 317 / 362 , 233 / 317 / 405, 139 / 317, 162, 286, 414, 417, 226, 61, 105, 230, 418, 370, 297, 237, 428, 362, 233, 235, 148, 100, 97, 382, or 358, and the amino acid sequence includes one or more substitutions relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 62 or relative to a reference sequence corresponding to SEQ ID NO: 62.
[0165] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions: 196E, 242P, 337L, 33R, 33V, 277Q, 337P, 30C, 359N, 337G, 283S, 242Q, 415A, 387G, 415H, 379A, 205G, 186M, 337M, 389Q, 102L, 389L, 359R, 205K, 164A, 415E, 301F, 30H, 379G, 375W, 267W, 379G, 375H ... 80T, 415C, 254G, 186T, 317Q, 77D / 139T / 317Q / 417G, 105K / 317Q / 417G, 317Q / 349E / 362G / 386F, 105K / 317Q, 139T / 317Q / 36 2G, 233T / 317Q / 405I, 139T / 317Q, 162W, 283V, 286S, 414L, 417L, 226S, 61T, 226Q, 105S, 186E, 230M, 61V, 186V, 186L, 379P , 415V, 415I, 186A, 283R, 226V, 418L, 418S, 370G, 283A, 337S, 297W, 237S, 428S, 186C, 362G, 283G, 196V, 237L, 415L, 233W , 242T, 277R, 235W, 148P, 254S, 102G, 105T, 301L, 286F, 100V, 237Y, 30S, 428F, 414Q, 414V, 283L, 414A, 414R, 370C, 283K, 286L, 186H, 417D, 226D, 233G, 267Y, 97G, 226E, 418T, 230L, 414T, 418A, 237V, 418G, 382V, 267D, 283M, 418P, 237R, 237M, 418I, 358C, 186R, 418M, or 102S, and the amino acid sequence contains one or more substitutions relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:62 or a reference sequence corresponding to SEQ ID NO:62.
[0166] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions I196E, V242P, F337L, T33R, T33V, E277Q, F337P, T30C, K359N, F337G, F283S, V242Q, T415A, T387G, T415H, E379A, Q205G, K186M, F337M, E389Q, N102L, E389L, K359R, Q205K, K164A, T415E, E301F, T30H, E379G, Y375W, Q267W, V380 T, T415C, H254G, K186T, T317Q, Y77D / S139T / T317Q / S417G, L105K / T31 7Q / S417G, T317Q / D349E / T362G / Y386F, L105K / T317Q, S139T / T317Q / T3 62G, S233T / T317Q / V405I, S139T / T317Q, Y162W, F283V, W286S, S414L, S417L, K226S, I61T, K226Q, L105S, K186E, V230M, I61V, K186V, K186L, E 379P, T415V, T415I, K186A, F283R, K226V, K418L, K418S, A370G, F283A , F337S, F297W, K237S, R428S, K186C, T362G, F283G, I196V, K237L, T415 L, S233W, V242T, E277R, F235W, I148P, H254S, N102G, L105T, E301L, W2 86F, T100V, K237Y, T30S, R428F, S414Q, S414V, F283L, S414A, S414R, A3 70C, F283K, W286L, K186H, S417D, K226D, S233G, Q267Y, R97G, K226E, K418T, V230L, S414T, K418A, K237V, K418G, I382V, Q267D, F283M, K418P, K237R, K237M, K418I, G358C, K186R, K418M, or N102S, and the amino acid sequence comprises one or more substitutions relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62 or relative to a reference sequence corresponding to SEQ ID NO:62.
[0167] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or to a reference sequence corresponding to SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138 or to the reference sequence corresponding to SEQ ID NO: 138.
[0168] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138.
[0169] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid at position(s) 242 / 283 / 286 / 359 / 418, 283 / 286, 283 / 286 / 418, 186 / 242 / 283 / 286 / 418, 205 / 286 / 359, 283, 242 / 286 / 418, 186 / 205 / 242, 283 / 286 / 359, 277 / 286 / 359 / 418, 186 / 205 / 283 / 286 / 359, 242 / 277 / 418, 242 / 283 / 286 / 418, 112 / 19 6 / 389, 286, 283 / 286 / 359 / 418, 242 / 359 / 418, 186 / 283, 186 / 359 / 418, 205 / 242 / 418, 186 / 205 / 242 / 283 / 286 / 359 / 418, 205 / 359 / 418, 186 / 205 / 283 / 286 / 418, 186 / 283 / 359, 186 / 242, 186 / 242 / 359, 283 / 359 / 418, 277 / 418, 186 / 188 / 283, 186 / 286 / 418, 186 / 242 / 286 / 359 / 418, 418, 1 86 / 242 / 283 / 286 / 359 / 418, 186 / 277 / 359 / 418, 242 / 283 / 286, 205 / 418, 30 / 297, 205 / 242 / 286 / 359 / 418, 186 / 205 / 359 / 418, 359 / 418, 186, 230, 33 / 297, 186 / 205, 186 / 283 / 359 / 418, 186 / 418, 205 / 242 / 283 / 359 / 418, 33 / 375 / 389, 33 / 230, 196 / 242 / 283 / 286 / 359 / 418, 186 / 242 / 283 / and at least a substitution or set of substitutions in 359 / 418, 186 / 359, 33 / 196, 186 / 277 / 418, 242, 33 / 196 / 297 / 301, 205 / 237 / 242 / 283 / 286 / 359, 186 / 205 / 283 / 359 / 418, 33 / 389, or 186 / 196 / 242 / 283 / 286 / 359 / 418, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 138, or are relative to a reference sequence corresponding to SEQ ID NO: 138.
[0170] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 242Q / 283V / 286S / 359R / 418L, 283L / 286S, 283L / 286S / 418S, 186E / 242Q / 283V / 286S / 418L, 205K / 286S / 359R, 283V, 242P / 286S / 418T, 186E / 205K / 242T, 283L / 286S / 359R, 277Q / 286S / 359N / 418T, 186A / 205K / 283V / 286S / 359R, 242Q / 277Q / 418 L, 242Q / 283S / 286S / 418T, 112M / 196V / 389Q, 286S, 242P / 283R / 286S / 359R / 418L, 283L / 286S / 359R / 418S, 242Q / 359N / 418S, 186E / 283L, 186E / 359R / 4 18S, 205K / 242T / 418S, 186E / 205K / 242Q / 283S / 286S / 359R / 418S, 205K / 35 9N / 418L, 186A / 205K / 283L / 286S / 418L, 242T / 283V / 286S / 359R / 418S, 186E / 283L / 359R, 186E / 242P, 283V / 286S / 418S, 186E / 242T / 359R, 283L / 359R / 418L, 277Q / 418S, 186E / 188A / 283V, 186E / 286S / 418S, 186A / 242T / 286S / 35 9N / 418L, 418L, 186A / 242P / 283V / 286S / 359R / 418S, 186A / 277Q / 359R / 418 S, 242Q / 283R / 286S, 205K / 418S, 30C / 297W, 283L / 359R / 418S, 186E / 286S / 4 18L, 418T, 205K / 242T / 286S / 359R / 418S, 186E / 205K / 359N / 418T, 418S, 18 6E / 242T / 283S / 286S / 418S, 359R / 418S, 242T / 286S / 418S, 186E, 230L, 33V / 297W, 186E / 205K, 186A / 283L / 359R / 418S, 186E / 418L, 186E / 277Q / 359R / 4 18S, 205K / 242T / 283R / 359R / 418L, 186E / 418T, 33V / 375W / 389Q, 33R / 230L,196E / 242T / 283V / 286S / 359R / 418T, 186E / 242T / 283L / 359R / 418S, 186E / 359R, 33V / 196V, 33R / 375W / 389Q, 186A / 277Q / 418S, 186A / 242T / 283L / 286S / 418L, 242T, 33V / 196V / 297W / 301F, 205K / 237Y / 242 P / 283L / 286S / 359R, 186A / 205K / 283L / 359R / 418T, 33V / 389Q, or 186A / 196E / 242T / 283L / 286S / 359N / 418T, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 138, or are relative to a reference sequence corresponding to SEQ ID NO: 138.
[0171] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions V242Q / F283V / W286S / K359R / K418L, F283L / W286S, F283L / W286S / K418S, K186E / V242Q / F283V / W286S / K418L, Q205K / W286S / K359R, F283V, V242P / W286S / K418T, K186E / Q205K / V242T, F283L / W286S / K359R, E277Q / W286S / K359N / K418T, K186A / Q 205K / F283V / W286S / K359R, V242Q / E277Q / K418L, V242Q / F283S / W286S / K4 18T, L112M / I196V / E389Q, W286S, V242P / F283R / W286S / K359R / K418L, F283 L / W286S / K359R / K418S, V242Q / K359N / K418S, K186E / F283L, K186E / K359R / K418S, Q205K / V242T / K418S, K186E / Q205K / V242Q / F283S / W286S / K359R / K4 18S, Q205K / K359N / K418L, K186A / Q205K / F283L / W286S / K418L, V242T / F28 3V / W286S / K359R / K418S, K186E / F283L / K359R, K186E / V242P, F283V / W286S / K418S, K186E / V242T / K359R, F283L / K359R / K418L, E277Q / K418S, K186E / V 188A / F283V, K186E / W286S / K418S, K186A / V242T / W286S / K359N / K418L, K41 8L, K186A / V242P / F283V / W286S / K359R / K418S, K186A / E277Q / K359R / K418S , V242Q / F283R / W286S, Q205K / K418S, T30C / F297W, F283L / K359R / K418S, K1 86E / W286S / K418L, K418T, Q205K / V242T / W286S / K359R / K418S, K186E / Q205 K / K359N / K418T, K418S, K186E / V242T / F283S / W286S / K418S, K359R / K418S,V242T / W286S / K418S, K186E, V230L, T33V / F297W, K186E / Q205K, K186A / F283L / K359R / K418S, K186E / K418L, K186E / E277Q / K359R / K418S, Q205K / V242T / F283R / K359R / K418 L, K186E / K418T, T33V / Y375W / E389Q, T33R / V230L, I196E / V242T / F283V / W286S / K359 R / K418T, K186E / V242T / F283L / K359R / K418S, K186E / K359R, T33V / I196V, T33R / Y375W / E389Q, K186A / E277Q / K418S, K186A / V242T / F283L / W286S / K418L, V242T, T33V / I196V / F297W / E301F, Q205K / K237Y / V242P / F283L / W286S / K359R, K186A / Q205K / F283L / K359R / K418T, T33V / E389Q, or K186A / I196E / V242T / F283L / W286S / K359N / K418T, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 138, or are relative to a reference sequence corresponding to SEQ ID NO: 138.
[0172] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or to a reference sequence corresponding to SEQ ID NO:318, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318 or to the reference sequence corresponding to SEQ ID NO:318.
[0173] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or to a reference sequence corresponding to an even-numbered SEQ ID NO:460-936, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318.
[0174] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one amino acid position(s) 363, 63, 389, 381, 197, 359, 102, 165, 388, 414, 337, 164, 416, 101, 415, 423, 364, 73, 50, 71, 388 / 419, 357, 396, 68, 76, 14, 271, 360, 266, 208, 74, 263, 264, 13, 378, 372, 300, 294, 290, 397, 95, 258, 161, 212, 198, 138, 18, 404, 273, 117, 240, 69, 278, 289, 82, 328, 61 / 186 / 417, 186 / 370 / 417, 267, 61 / 370, 186 / 267 / 370 / 417, 61 / 186, 370 / 417, 417, 61 / 370 / 382, 61 / 186 / 267 / 370 / 417, 61, 267 / 370 / 417, 267 / 370, 61 / 417, 61 / 186 / 237 / 267 / 370, 61 / 237 / 370 / 417, 370, 186 / 370, 61 / 186 / 267 / 417, 61 / 186 / 370 / 382, 370 / 382 / 417, 61 / 186 / 370, 237 / 267 / 370 / 417, 61 / 186 / 382, 61 / 267, 61 / 267 / 417, 61 / 237 / 267 / 382, 186 / 370 / 382, 237 / 267 / 370, 61 / 186 / 267 / 370, 186 / 237 / 267 / 370, 61 / 186 / 267, 61 / 186 / 237, 186, 186 / 267, 237 / 370 / 417, 242 / 414, 162 / 414, or 390, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or relative to a reference sequence corresponding to SEQ ID NO:318.
[0175] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 363R, 63R, 389K, 381R, 389R, 197R, 359R, 381K, 102K, 165K, 388K, 414R, 337R, 164R, 416K, 101R, 415K, 423R, 364R, 73P, 50I, 71L, 388Y / 419G, 50G, 357K, 396H, 68Y, 76F, 71P, 14V, 271G, 68A, 68M, 360P, 266M, 208D, 74S, 263T, 264C, 13R, 263L, 360M, 378T, 76H, 13G, 76R, 50T, 372Q, 50S, 300G, 73T, 294L, 372N, 76N, 290L, 397L , 76G, 95L, 258L, 161V, 212M, 198V, 138L, 18D, 18S, 404S, 273G, 117G, 240E, 6 8S, 73C, 69T, 278E, 360H, 198D, 300T, 161R, 73K, 289V, 82R, 328R, 61T / 186A / 417D, 186A / 370C / 417L, 267Y, 61T / 370C, 186A / 267Y / 370C / 417L, 61T / 186C , 370C / 417D, 417D, 61T / 370C / 382V, 61T / 186A / 267Y / 370C / 417D, 61T, 267Y / 370C / 417D, 267Y / 370C, 61T / 417L, 61T / 186H / 237R / 267Y / 370C, 61T / 237R / 370C / 417L, 370C / 417L, 370C, 186E / 370C, 61T / 186E / 267Y / 417D, 61T / 186C / 370C / 382V, 370C / 382V / 417D, 61T / 186C / 267Y / 417D, 61T / 186C / 370C, 237R / 267Y / 370C / 417D, 61T / 186E / 417D, 61T / 186H, 61T / 186C / 382V, 61T / 186H / 370C, 61T / 267Y, 61T / 267Y / 417L, 61T / 237R / 267Y / 382V, 186H / 370C / 382V, 2 37R / 267Y / 370C, 186H / 370C / 417D, 61T / 186V / 267Y / 370C, 267Y / 370C / 417L , 61T / 186V / 370C, 186H / 237R / 267Y / 370C, 61T / 186E / 267Y, 61T / 186C / 237R,61T / 186H / 237R, 61T / 186C / 417L, 186H, 186A / 267Y, 186C / 370C, 61T / 186C / 237R / 267Y / 370C, 186E / 267Y, 237R / 370C / 417L, 186H / 370C, 242Q / 414Q, 414T, 414Q, 162W / 414Q, 162W / 414V, 267D / 414T, 414V, 267D / 414Q, 414A, 162W / 414L, 10 5S / 414T, 162W / 242Q / 414T, 97G / 162W / 414Q, 105S / 162W / 267D / 414V, 356V, 273A, 357P, 14G, 14S, 396C, 240R, 392K, 27 3S, 106V, 308S, 308L, 306I, 96A, 397K, 263Q, 282T, 138R, 258S, 76L, 14T, 289S, 240G, 106S, 117S, 357S, 96G, 113T, 71G, 282V, 117V, 282M, 258V, 309G, 306K, 308K, 357R, 212S, 18N, 113A, 306S, 212F, 198N, 110R, 240K, 198L, 71R, 357V, 198R , 271A, 282Y, 212G, 95V, 37S, 68V, 240S, 117Y, 201S, 271N, 240Q, 289L, 212W, 198W, 357H, 208F, 282L, 306V, 284D, 96T, 2 66T, 374S, 95R, 309R, 73W, 397M, 289A, 295K, 106L, 95A, 374A, 88V, 96V, 271S, 44S, 208H, 263G, 12A, 201L, 198A, 198K, 282G, 390E, 264A, or 73V, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 318, or are relative to a reference sequence corresponding to SEQ ID NO: 318.
[0176] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions: K363R, G63R, E389K, S381R, E389R, I197R, K359R, S381K, N102K, R165K, D388K, S414R, F337R, K164R, S416K, G101R, T415K, S423R, N364R, L73P, E50I, K71L, D388Y / T419G, E50G, E357K, A396H, I68Y, D76F, K71P, K14V, F271G, I68A, I68M, E360P, E266M, A208D, P74S, R263T, I264C, T13R, R263L, E360M, N378T, D76H, T13G, D76R, E50T, V372Q, E50S, S300G, L73T, D294L, V372N, D76N, W290L, R397L, D 76G, Y95L, A258L, K161V, V212M, T198V, K138L, Q18D, Q18S, K404S, R273G, Q 117G, D240E, I68S, L73C, K69T, Q278E, E360H, T198D, S300T, K161R, L73K, E 289V, D82R, D328R, I61T / K186A / S417D, K186A / A370C / S417L, Q267Y, I61T / A370C, K186A / Q267Y / A370C / S417L, I61T / K186C, A370C / S417D, S417D, I61 T / A370C / I382V, I61T / K186A / Q267Y / A370C / S417D, I61T, Q267Y / A370C / S4 17D, Q267Y / A370C, I61T / S417L, I61T / K186H / K237R / Q267Y / A370C, I61T / K 237R / A370C / S417L, A370C / S417L, A370C, K186E / A370C, I61T / K186E / Q267 Y / S417D, I61T / K186C / A370C / I382V, A370C / I382V / S417D, I61T / K186C / Q2 67Y / S417D, I61T / K186C / A370C, K237R / Q267Y / A370C / S417D, I61T / K186E / S417D, I61T / K186H, I61T / K186C / I382V, I61T / K186H / A370C, I61T / Q267Y,I61T / Q267Y / S417L、I61T / K237R / Q267Y / I382V、K186H / A370C / I382V、K237R / Q267Y / A370C、K186H / A370C / S417D、I61T / K186V / Q267Y / A370C、Q267Y / A370C / S417L、I61T / K186V / A370C、K186H / K237R / Q267Y / A370C、I61T / K186E / Q267Y、I61T / K186C / K237R、I61T / K186H / K237R、I61T / K186C / S417L、K186H、K186A / Q267Y、K186C / A370C、I61T / K186C / K237R / Q267Y / A370C、K186E / Q267Y、K237R / A370C / S417L、K186H / A370C、V242Q / S414Q、S414T、S414Q、Y162W / S414Q、Y162W / S414V、Q267D / S414T、S414V、Q267D / S414Q、S414A、Y162W / S414L、L105S / S414T、Y162W / V242Q / S414T、R97G / Y162W / S414Q、L105S / Y162W / Q267D / S414V、E356V、R273A、E357P、K14G、K14S、A396C、D240R、V392K、R273S、D106V、F308S、F308L、E306I、N96A、R397K、R263Q、E282T、K138R、A258S、D76L、K14T、E289S、D240G、D106S、Q117S、E357S、N96G、S113T、K71G、E282V、Q117V、E282M、A258V、H309G、E306K、F308K、E357R、V212S、Q18N、S113A、E306S、V212F、T198N、V110R、D240K、T198L、K71R、E357V、T198R、F271A、E282Y、V212G、Y95V、D37S、I68V、D240S、Q117Y、T201S、F271N、D240Q、E289L、V212W、T198W、E357H、A208F、E282L、E306V、P284D、N96T、E266T、E374S、Y95R、H309R、L73W、R397M、E289A、G295K、D106L、Y95A、E374A、L88V、N96V、F271S、D44S、A208H, R263G, M12A, T201L, T198A, T198K, E282G, M390E, I264A, or L73V, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 318, or relative to a reference sequence corresponding to SEQ ID NO: 318.
[0177] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:722, or to a reference sequence corresponding to SEQ ID NO:722, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722.
[0178] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 938-1098, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 938-1098, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 722, or to the reference sequence corresponding to SEQ ID NO: 722.
[0179] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 63, 63 / 96 / 370, 389, 13 / 267 / 363 / 389, 13 / 186 / 389, 50 / 267 / 363 / 370 / 389, 363 / 370, 96 / 370, 61 / 63 / 212, 11 / 305, 11, 242 / 283 / 286 / 317 / 414 / 418, 323, 334, 339, 356, 384, 408, 67, 392, 104, 355, 1 59, 155, 367, 31, 231, 36, 150, 239, 103, 125, 228, 37, 189, 177, 422, 128, 220, 130, 56, 190, 156, 232, 423, 34, 99, 59, 60, 421, or 195, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:722, or are relative to a reference sequence corresponding to SEQ ID NO:722.
[0180] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 63R, 63R / 96T / 370C, 389K, 13R / 267Y / 363R / 389K, 13R / 186H / 389K, 50T / 267Y / 363R / 370C / 389K, 363R / 370C, 96T / 370C, 61T / 63 R / 212W, 11D / 305K, 11D, 242V / 283F / 286W / 317T / 414S / 418K, 323S, 334L, 339Y, 356V, 384V, 4 08C, 67R, 392R, 104K, 355S, 159Q, 155R, 367L, 31R, 231P, 36Y, 150T, 239V, 103V, 356A, 63F, 12 5T, 367C, 228I, 37S, 37G, 239M, 239W, 239Q, 189C, 239S, 239T, 177G, 239P, 37N, 422N, 228S, 1 25R, 128C, 189T, 356W, 220V, 408V, 392S, 130T, 228M, 56P, 228E, 190R, 392I, 156C, 232R, 150F , 150C, 239N, 392L, 423T, 34L, 99I, 384C, 59E, 334R, 60Y, 99G, 34R, 37L, 392C, 421R, or 195R, and the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 722, or are relative to a reference sequence corresponding to SEQ ID NO: 722.
[0181] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions G63R, G63R / N96T / A370C, E389K, T13R / Q267Y / K363R / E389K, T13R / K186H / E389K, E50T / Q267Y / K363R / A370C / E389K, K363R / A370C, N96T / A370C, I61T / G63R / V 212W, G11D / Q305K, G11D, Q242V / L283F / S286W / Q317T / Q414S / S418K, M323S, K334L, I339Y, E356V, G384V, I4 08C, K67R, V392R, A104K, F355S, L159Q, L155R, V367L, S31R, Q231P, E36Y, E150T, D239V, A103V, E356A, G63F , K125T, V367C, A228I, D37S, D37G, D239M, D239W, D239Q, L189C, D239S, D239T, C177G, D239P, D37N, K422N, A 228S, K125R, Q128C, L189T, E356W, I220V, I408V, V392S, D130T, A228M, Y56P, A228E, K190R, V392I, A156C, K 232R, E150F, E150C, D239N, V392L, S423T, A34L, L99I, G384C, D59E, K334R, L60Y, L99G, A34R, D37L, V392C, K421R, or K195R, and the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 722, or are relative to a reference sequence corresponding to SEQ ID NO: 722.
[0182] In some embodiments, the engineered DNA ligase comprises a reference sequence corresponding to residues 12-437 of SEQ ID NO:938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to SEQ ID NO:938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938 or to the reference sequence corresponding to SEQ ID NO:938.
[0183] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 1100-1184, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 1100-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 938.
[0184] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises amino acid position(s) 308 / 357 / 390, 74 / 76 / 201 / 308 / 357, 61 / 74 / 76 / 186 / 201 / 308 / 309 / 357 / 390, 14 / 201 / 240 / 289 / 357, 308 / 415, 76 / 357 / 396, 263 / 308 / 396, 61 / 76 / 96 / 240 / 308 / 309, 14 / 306 / 415, 14 / 73 / 106 / 415, 12 / 14 / 258 / 263 / 289 / 308 / 309 / 396, 74 / 76 / 117 / 309 / 357, 14 / 258 / 263 / 357 / 396, 14 / 96 / 106 / 306, 14 / 106, 12 / 14 / 308 / 309, 14 / 357 / 390, 14 / 117 / 258 / 309 / 3 57, 390, 240 / 273 / 357 / 390, 61 / 76 / 186 / 201 / 308 / 309, 14 / 396, 309, 14, 106 / 306 / 308, 14 / 240 / 306 / 308, 12 / 14 / 186 / 357, 309 / 390, 14 / 306, 14 / 76 / 308, 117 / 208 / 258 / 263 / 289 / 308 / 309, 14 / 73 / 106, 76 / and at least a substitution or set of substitutions at 208 / 263, 357, 14 / 308, 263, 76, 14 / 300 / 308 / 415, 240, 33 / 357 / 390, 14 / 76 / 273, or 74, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 938, or are relative to a reference sequence corresponding to SEQ ID NO: 938.
[0185] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 308L / 357S / 390E, 74S / 76L / 201S / 308L / 357P, 61T / 74S / 76L / 186A / 201S / 308L / 309R / 357P / 390E, 14G / 201S / 240R / 289S / 357P, 308S / 415I, 76L / 357S / 396C, 263L / 308L / 396C, 61T / 76L / 96A / 240R / 308L / 309R, 14S / 306I / 415I, 14S / 73W / 106S / 415I, 12A / 14G / 258S / 263Q / 289S / 308L / 309R / 396C, 74S / 76L / 11 7V / 309R / 357S, 14T / 258S / 263Q / 357S / 396C, 14S / 96G / 106S / 306I, 14S / 106S, 12A / 14G / 308L / 309R, 14G / 357S / 390E, 14G / 117V / 258S / 309R / 357S, 390E, 240R / 273S / 357P / 390E, 61T / 76L / 186V / 201S / 308L / 309R, 14G / 396C, 309R, 14S, 106V / 306 I / 308S, 14S / 240S / 306I / 308S, 12I / 14G / 186V / 357S, 309R / 390E, 14S / 306I, 14T, 14S / 76R / 308S, 117V / 208D / 258S / 263Q / 289S / 308L / 309R, 14S / 73W / 106S, 76L / 208D / 263Q, 357S, 14S / 308S, 263L, 76L, 14S / 300G / 308S / 415I, 240Y, 33M / 357S / 390E, 14G / 76L / 273S, or 74S, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 938 or are relative to a reference sequence corresponding to SEQ ID NO: 938.
[0186] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions F308L / E357S / M390E, P74S / D76L / T201S / F308L / E357P, I61T / P74S / D76L / K186A / T201S / F308L / H309R / E357P / M390E, K14G / T201S / D240R / E289S / E357P, F308S / T415I, D76L / E357S / A396C, R263L / F308L / A396C, I61T / D76L / N96 A / D240R / F308L / H309R, K14S / E306I / T415I, K14S / L73W / D106S / T415I , M12A / K14G / A258S / R263Q / E289S / F308L / H309R / A396C, P74S / D76L / Q1 17V / H309R / E357S, K14T / A258S / R263Q / E357S / A396C, K14S / N96G / D106 S / E306I, K14S / D106S, M12A / K14G / F308L / H309R, K14G / E357S / M390E, K 14G / Q117V / A258S / H309R / E357S, M390E, D240R / R273S / E357P / M390E, I61T / D76L / K186V / T201S / F308L / H309R, K14G / A396C, H309R, K14S, D10 6V / E306I / F308S, K14S / D240S / E306I / F308S, M12I / K14G / K186V / E357S , H309R / M390E, K14S / E306I, K14T, K14S / D76R / F308S, Q117V / A208D / A2 58S / R263Q / E289S / F308L / H309R, K14S / L73W / D106S, D76L / A208D / R263Q, E357S, K14S / F308S, R263L, D76L, K14S / S300G / F308S / T415I, D240Y, T33M / E357S / M390E, K14G / D76L / R273S, or P74S, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 938 or are relative to a reference sequence corresponding to SEQ ID NO: 938.
[0187] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution at an amino acid position set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0188] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0189] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions at an amino acid position(s) set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0190] In some embodiments, the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0191] In some embodiments, the engineered DNA ligase comprises an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, where the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0192] In some embodiments, the engineered DNA ligase comprises a sequence corresponding to residues 12-437 of an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, or a sequence corresponding to residues 12-437 of an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2. and 18.2.
[0193] In some embodiments, the engineered DNA ligase is selected from the group consisting of SEQ ID NOs: 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 2 4, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226 , 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 351 52, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 41 4, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476 , 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538,540、542、544、546、548、550、552、554、556、558、560、562、564、568、570、572、574、576、578、580、582、584、586、588、590、592、594、596、598、600、602、604、606、608、610、612、614、616、618、620、622、624、626、628、630、632、634、636、638、640、642、644、646、648、650、652、654、656、658、660、662、664、666、668、670、672、674、676、678、680、682、684、686、688、700、702、704、706、708、710、712、714、716、718、720、722、724、728、730、732、734、736、738、740、742、744、746、748、750、752、754、756、758、760、762、764、766、768、770、772、774、776、778、780、782、784、786、788、790、792、794、796、798、800、802、804、806、808、810、812、814、816、818、820、822、824、828、830、832、834、836、838、840、842、844、846、848、850、852、854、856、858、860、862、864、866、868、870、872、874、876、878、880、882、884、886、888、890、892、894、896、898、900、902、904、906、908、910、912、914、916、918、920、922、924、928、930、932、934、936、938、940、942、944、946、948、950、952、954、956、958、960、962、964、966、968、970、972、974、976、978、980、982、984、986、988、990、992、994、996、998、1000、1002、1004、1006、1008、1010、1012、1014、1016、1018、1020、1022、1024、1028、1030、1032、1034、1036、1038、1040、1042、1044、1046、1048, 1050, 1052, 1054, 1056, 1058, 1060, 1062, 1064, 1066, 1068, 1070, 1072, 1074, 1076, 1078, 1080, 1082, 1084, 1086, 1088, 1090, 1092, 1094, 1096, 1098, 1100, 1102, 1104, 1106, 1108, 1110, 1112, 1114, 1116, 1118, 1120, 1122, 1124, 1128, 1130, 1132, 1134, 1136, 1138, 1140, 1142, 1144, 1 1146, 1148, 1150, 1152, 1154, 1156, 1158, 1160, 1162, 1164, 1166, 1168, 1170, 1172, 1174, 1176, 1178, 1180, 1182, or 1184.
[0194] In some embodiments, the engineered DNA ligase is selected from the group consisting of SEQ ID NOs: 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 2 4, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226 , 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 351 52, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 41 4, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476 , 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538,540、542、544、546、548、550、552、554、556、558、560、562、564、568、570、572、574、576、578、580、582、584、586、588、590、592、594、596、598、600、602、604、606、608、610、612、614、616、618、620、622、624、626、628、630、632、634、636、638、640、642、644、646、648、650、652、654、656、658、660、662、664、666、668、670、672、674、676、678、680、682、684、686、688、700、702、704、706、708、710、712、714、716、718、720、722、724、728、730、732、734、736、738、740、742、744、746、748、750、752、754、756、758、760、762、764、766、768、770、772、774、776、778、780、782、784、786、788、790、792、794、796、798、800、802、804、806、808、810、812、814、816、818、820、822、824、828、830、832、834、836、838、840、842、844、846、848、850、852、854、856、858、860、862、864、866、868、870、872、874、876、878、880、882、884、886、888、890、892、894、896、898、900、902、904、906、908、910、912、914、916、918、920、922、924、928、930、932、934、936、938、940、942、944、946、948、950、952、954、956、958、960、962、964、966、968、970、972、974、976、978、980、982、984、986、988、990、992、994、996、998、1000、1002、1004、1006、1008、1010、1012、1014、1016、1018、1020、1022、1024、1028、1030、1032、1034、1036、1038、1040、1042、1044、1046、1048, 1050, 1052, 1054, 1056, 1058, 1060, 1062, 1064, 1066, 1068, 1070, 1072, 1074, 1076, 1078, 1080, 1082, 1084, 1086, 1088, 1090, 1092, 1094, 1096, 1098, 1100, 1102, 1104, 1106, 1108, 1110, 1112, 1114, 1116, 1118, 1120, 1122, 1124, 1128, 1130, 1132, 1134, 1136, 1138, 1140, 1142, 1 1144, 1146, 1148, 1150, 1152, 1154, 1156, 1158, 1160, 1162, 1164, 1166, 1168, 1170, 1172, 1174, 1176, 1178, 1180, 1182, or 1184.
[0195] In some embodiments, the engineered DNA ligase comprises an amino acid sequence comprising residues 12-437 of an even-numbered SEQ ID NO:40-1184, or an amino acid sequence comprising an even-numbered SEQ ID NO:40-1184. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions.
[0196] In some embodiments, the engineered DNA ligase is selected from the group consisting of SEQ ID NOs: 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 2 4, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226 , 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 351 52, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 41 4, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476 , 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538,540、542、544、546、548、550、552、554、556、558、560、562、564、568、570、572、574、576、578、580、582、584、586、588、590、592、594、596、598、600、602、604、606、608、610、612、614、616、618、620、622、624、626、628、630、632、634、636、638、640、642、644、646、648、650、652、654、656、658、660、662、664、666、668、670、672、674、676、678、680、682、684、686、688、700、702、704、706、708、710、712、714、716、718、720、722、724、728、730、732、734、736、738、740、742、744、746、748、750、752、754、756、758、760、762、764、766、768、770、772、774、776、778、780、782、784、786、788、790、792、794、796、798、800、802、804、806、808、810、812、814、816、818、820、822、824、828、830、832、834、836、838、840、842、844、846、848、850、852、854、856、858、860、862、864、866、868、870、872、874、876、878、880、882、884、886、888、890、892、894、896、898、900、902、904、906、908、910、912、914、916、918、920、922、924、928、930、932、934、936、938、940、942、944、946、948、950、952、954、956、958、960、962、964、966、968、970、972、974、976、978、980、982、984、986、988、990、992、994、996、998、1000、1002、1004、1006、1008、1010、1012、1014、1016、1018、1020、1022、1024、1028、1030、1032、1034、1036、1038、1040、1042、1044、1046、1048, 1050, 1052, 1054, 1056, 1058, 1060, 1062, 1064, 1066, 1068, 1070, 1072, 1074, 1076, 1078, 1080, 1082, 1084, 1086, 1088, 1090, 1092, 1094, 1096, 1098, 1100, 1102, 1104, 1106, 1108, 1110, 1112, 1114, 1116, 1118, 1119 1150, 1152, 1154, 1156, 1158, 1160, 1162, 1164, 1166, 1168, 1170, 1172, 1174, 1176, 1178, 1180, 1182, or 1184. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions.
[0197] In some embodiments, the engineered DNA ligase is selected from the group consisting of SEQ ID NOs: 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 2 4, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226 , 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 351 52, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 41 4, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476 , 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538,540、542、544、546、548、550、552、554、556、558、560、562、564、568、570、572、574、576、578、580、582、584、586、588、590、592、594、596、598、600、602、604、606、608、610、612、614、616、618、620、622、624、626、628、630、632、634、636、638、640、642、644、646、648、650、652、654、656、658、660、662、664、666、668、670、672、674、676、678、680、682、684、686、688、700、702、704、706、708、710、712、714、716、718、720、722、724、728、730、732、734、736、738、740、742、744、746、748、750、752、754、756、758、760、762、764、766、768、770、772、774、776、778、780、782、784、786、788、790、792、794、796、798、800、802、804、806、808、810、812、814、816、818、820、822、824、828、830、832、834、836、838、840、842、844、846、848、850、852、854、856、858、860、862、864、866、868、870、872、874、876、878、880、882、884、886、888、890、892、894、896、898、900、902、904、906、908、910、912、914、916、918、920、922、924、928、930、932、934、936、938、940、942、944、946、948、950、952、954、956、958、960、962、964、966、968、970、972、974、976、978、980、982、984、986、988、990、992、994、996、998、1000、1002、1004、1006、1008、1010、1012、1014、1016、1018、1020、1022、1024、1028、1030、1032、1034、1036、1038、1040、1042、1044、1046、1048, 1050, 1052, 1054, 1056, 1058, 1060, 1062, 1064, 1066, 1068, 1070, 1072, 1074, 1076, 1078, 1080, 1082, 1084, 1086, 1088, 1090, 1092, 1094, 1096, 1098, 1100, 1102, 1104, 1106, 1108, 1110, 1112, 1114, 1116, 1117 1160, 1162, 1164, 1166, 1168, 1170, 1172, 1174, 1176, 1178, 1180, 1182, or 1184. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions.
[0198] In some embodiments, the engineered DNA ligase comprises an amino acid sequence comprising residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or an amino acid sequence comprising SEQ ID NO: 62, 138, 318, 722, or 938. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions, insertions, and / or deletions. In some embodiments, the amino acid sequence of the engineered DNA ligase optionally comprises 1, 2, 3, 4, or 5 substitutions.
[0199] In some of the foregoing embodiments, the engineered DNA ligase polypeptide has 1, 2, 3, 4, or up to 5 substitutions in the amino acid sequence. In some embodiments, the engineered DNA ligase polypeptide has 1, 2, 3, or 4 substitutions in the amino acid sequence. In some embodiments, the substitutions comprise non-conservative or conservative substitutions. In some embodiments, the substitutions comprise conservative substitutions. In some embodiments, the substitutions comprise non-conservative substitutions. In some embodiments, guidance regarding non-conservative and conservative substitutions is provided by the variants disclosed herein.
[0200] In some embodiments, the engineered DNA ligases of the present disclosure have DNA ligase activity. In some embodiments, the engineered DNA ligases have DNA ligase activity and are characterized by or exhibit one or more improved or enhanced properties described herein compared to a reference DNA ligase.
[0201] In some embodiments, the engineered DNA ligase has increased activity compared to a reference DNA ligase. In some embodiments, the increased activity is about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 2-fold, 2.5-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold or more compared to the reference DNA ligase. Exemplary activity improvements are provided in the Examples.
[0202] In some embodiments, the engineered DNA ligase has increased stability compared to a reference DNA ligase. In some embodiments, the engineered DNA ligase has increased thermostability compared to a reference DNA ligase. In some embodiments, the increased thermostability is at temperatures between about 25°C and 55°C, between about 30°C and about 45°C, between 35°C and about 40°C, particularly between about 40°C and about 50°C. In some embodiments, the increased thermostability is 25°C, 30°C, 35°C, 40°C, 41°C, 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, 49°C, or 50°C. In some embodiments, the increased thermostability is achieved by treatment at a particular temperature for 15 minutes, 30 minutes, 45 minutes, or 1 hour.
[0203] In some embodiments, the engineered DNA ligase has an increased product yield compared to a reference DNA ligase, in some embodiments, the increased product yield is under the substrates and reaction conditions provided in the examples.
[0204] In some embodiments, the engineered DNA ligase has a higher solubility than the reference DNA ligase. In some embodiments, the engineered DNA ligase has a reduced sequence bias than the reference DNA ligase. In some embodiments, the engineered DNA ligase is insensitive or has a reduced sensitivity to the input DNA substrate concentration.
[0205] In some embodiments, the reference DNA ligase has a sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, 938, or 1108, or a sequence corresponding to SEQ ID NO: 62, 138, 318, 722, 938, or 1108. In some embodiments, the reference DNA ligase has a sequence corresponding to residues 12-437 of SEQ ID NO: 2, or a sequence corresponding to SEQ ID NO: 2. In some embodiments, the reference DNA ligase is wild-type T4 DNA ligase.
[0206] In some embodiments, the engineered DNA ligase has improved properties compared to a reference DNA ligase selected from: i) increased activity, ii) increased stability, iii) increased thermostability, iv) increased product yield, v) increased solubility, vi) reduced sequence bias, and vii) insensitivity or reduced sensitivity to input DNA substrate concentration, or any combination of i), ii), iii), iv), v), vi), and vii). In some embodiments, the reference DNA ligase has a sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, 62, 138, 318, 722, or 938, or a sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722, or 938. In some embodiments, the reference DNA ligase has a sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, or a sequence corresponding to SEQ ID NO: 2. In some embodiments, the reference DNA ligase is wild-type T4 DNA ligase.
[0207] In some embodiments, the present disclosure provides: (a) a sequence corresponding to residues 12 to 437 of SEQ ID NO:2; a sequence corresponding to residues 12 to 613 of SEQ ID NO:4; residues 12 to 614 of SEQ ID NO:6; residues 12 to 610 of SEQ ID NO:8; residues 12 to 606 of SEQ ID NO:10; residues 12 to 615 of SEQ ID NO:12; residues 12 to 594 of SEQ ID NO:14; residues 12 to 620 of SEQ ID NO:16; residues 12 to 608 of SEQ ID NO:18; residues 12 to 611 of SEQ ID NO:20; residues 12 to 614 of SEQ ID NO:22, residues 12 to 611 of SEQ ID NO:24, residues 12 to 609 of SEQ ID NO:26, residues 12 to 422 of SEQ ID NO:28, residues 12 to 518 of SEQ ID NO:30, residues 12 to 438 of SEQ ID NO:32, residues 12 to 381 of SEQ ID NO:34, residues 12 to 424 of SEQ ID NO:36, or residues 12 to 390 of SEQ ID NO:38, or (b) an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the sequence corresponding to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38.
[0208] In some embodiments, the engineered DNA ligase (a) comprising residues 12 to 437 of SEQ ID NO:2; residues 12 to 613 of SEQ ID NO:4; residues 12 to 614 of SEQ ID NO:6; residues 12 to 610 of SEQ ID NO:8; residues 12 to 606 of SEQ ID NO:10; residues 12 to 615 of SEQ ID NO:12; residues 12 to 594 of SEQ ID NO:14; residues 12 to 620 of SEQ ID NO:16; residues 12 to 608 of SEQ ID NO:18; residues 12 to 611 of SEQ ID NO:20; residues 12 to 614 of SEQ ID NO:22, residues 12 to 611 of SEQ ID NO:24, residues 12 to 609 of SEQ ID NO:26, residues 12 to 422 of SEQ ID NO:28, residues 12 to 518 of SEQ ID NO:30, residues 12 to 438 of SEQ ID NO:32, residues 12 to 381 of SEQ ID NO:34, residues 12 to 424 of SEQ ID NO:36, or residues 12 to 390 of SEQ ID NO:38; or (b) comprises an amino acid sequence comprising SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38.
[0209] In some embodiments, the engineered DNA ligase is expressed as a fusion protein. In some embodiments, the engineered DNA ligase described herein can be fused to various polypeptide sequences, such as, by way of example and not limitation, polypeptide tags that can be used for detection and / or purification. In some embodiments, the engineered DNA ligase fusion protein comprises a glycine-histidine or histidine tag (His tag). In some embodiments, the engineered DNA ligase fusion protein comprises an epitope tag, such as c-myc, FLAG, V5, or hemagglutinin (HA). In some embodiments, the engineered DNA ligase fusion protein comprises a GST, SUMO, Strep, MBP, or GFP tag. In some embodiments, the fusion is to the amino (N-) terminus of the engineered DNA ligase polypeptide. In some embodiments, the fusion is to the carboxy (C-) terminus of the engineered DNA ligase polypeptide.
[0210] In some embodiments, the engineered DNA ligase polypeptides described herein are isolated compositions, hi some embodiments, the engineered DNA ligase polypeptides are purified or are purified preparations, as further discussed herein.
[0211] In some embodiments, the present disclosure further provides functional or biologically active fragments of the engineered DNA ligase polypeptides described herein. Thus, for each and every embodiment of an engineered DNA ligase herein, a functional or biologically active fragment of the engineered DNA ligase is provided herein. In some embodiments, a functional or biologically active fragment of an engineered DNA ligase comprises at least about 90%, 95%, 96%, 97%, 98%, 99% or more of the activity of the DNA ligase polypeptide (i.e., the parent DNA ligase) from which it is derived. In some embodiments, a functional or biologically active fragment comprises at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the parent sequence of the DNA ligase. In some embodiments, functional fragments will be truncated by fewer than 5, fewer than 10, fewer than 15, fewer than 10, fewer than 25, fewer than 30, fewer than 35, fewer than 40, fewer than 45, fewer than 50, fewer than 55, fewer than 60, fewer than 65, or fewer than 70 amino acids.
[0212] In some embodiments, a functional fragment of an engineered DNA ligase herein comprises at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the parent sequence of the engineered DNA ligase. In some embodiments, the functional fragment will be truncated by fewer than 5, fewer than 10, fewer than 15, fewer than 10, fewer than 25, fewer than 30, fewer than 35, fewer than 40, fewer than 45, fewer than 50, fewer than 55, fewer than 60, fewer than 65, or fewer than 70 amino acids.
[0213] In some embodiments, functional or biologically active fragments of the engineered DNA ligase polypeptides described herein comprise at least a mutation or set of mutations in the amino acid sequence of an engineered DNA ligase described herein. Thus, in some embodiments, functional or biologically active fragments of the engineered DNA ligase exhibit enhanced or improved properties relative to the mutation or set of mutations in the parent DNA ligase.
[0214] Polynucleotides encoding engineered polypeptides, expression vectors and host cells In another aspect, the present disclosure provides recombinant polynucleotides encoding the engineered DNA ligases described herein. In some embodiments, the recombinant polynucleotides are operably linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide construct capable of expressing the engineered DNA ligase. In some embodiments, an expression construct comprising at least one heterologous polynucleotide encoding an engineered DNA ligase polypeptide(s) is introduced into a suitable host cell to express the corresponding DNA ligase polypeptide(s).
[0215] As will be apparent to those skilled in the art, the availability of protein sequences and knowledge of the codons corresponding to various amino acids provides a description of all polynucleotides capable of encoding the subject polypeptides. The degeneracy of the genetic code, in which the same amino acid is encoded by alternative or synonymous codons, allows for the creation of an extremely large number of nucleic acids, all of which encode the engineered DNA ligases of the present disclosure. Thus, the present disclosure provides methods and compositions for generating any and all possible variations of polynucleotides that can be made that encode the engineered DNA ligase polypeptides described herein by selecting combinations based on possible codon choices, and all such polynucleotide sequence variations shall be considered to be specifically disclosed for any polypeptide described herein, including the amino acid sequences presented in the Examples (e.g., Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2) and Sequence Listing.
[0216] In some embodiments, codons are preferably optimized for utilization by the host cell selected for protein production. In some embodiments, bacterially preferred codons are used for expression in bacteria. In some embodiments, fungal cell preferred codons are used for expression in fungal cells. In some embodiments, mammalian cell preferred codons are used for expression in mammalian cells. In some embodiments, insect cell preferred codons are used for expression in insect cells. In some embodiments, a codon-optimized polynucleotide encoding an engineered DNA ligase polypeptide described herein comprises preferred codons at about 40%, 50%, 60%, 70%, 80%, 90%, or more than 90% of the codon positions in the full-length coding region.
[0217] Thus, in some embodiments, a recombinant polynucleotide of the present disclosure encodes an engineered DNA ligase polypeptide described herein. In some embodiments, the polynucleotide sequence of the recombinant polynucleotide is codon-optimized.
[0218] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 and even-numbered SEQ ID NOs:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:2 and even-numbered SEQ ID NOs:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, 62, 138, 318, 722, or 938, or the reference sequence corresponding to SEQ ID NO:2, 62, 138, 318, 722, or 938.
[0219] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or to a reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938.
[0220] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or to a reference sequence corresponding to SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0221] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or to a reference sequence corresponding to SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0222] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or a reference sequence corresponding to an even-numbered SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2.
[0223] In some embodiments, the recombinant polynucleotide comprises a sequence identical to that at amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 148, 149, 150, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 288, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, 384, 386, 387, 388, 389, 390, 392, 396, and a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution in SEQ ID NO: 397, 404, 405, 408, 414, 415, 416, 417, 418, 419, 421, 422, 423, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, or are relative to a reference sequence corresponding to SEQ ID NO: 2.
[0224] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence including at least a substitution at positions 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0225] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions at amino acid position(s) 233, 317, 191, 288, 207, 149, 251, 205, 269, 164, 36, 428, 105 / 132, or 105, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0226] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising substitutions at least at the amino acid positions set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0227] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0228] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions at the amino acid position(s) set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to a reference sequence corresponding to SEQ ID NO:2.
[0229] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to SEQ ID NO:2.
[0230] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2 or are relative to the reference sequence corresponding to SEQ ID NO:2.
[0231] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0232] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or to a reference sequence corresponding to an even-numbered SEQ ID NO:40-1184.
[0233] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938.
[0234] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or a reference sequence corresponding to an even-numbered SEQ ID NO:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, 138, 318, 722, or 938, or the reference sequence corresponding to SEQ ID NO:62, 138, 318, 722, or 938.
[0235] In some embodiments, the recombinant polynucleotide comprises a sequence identical to that at amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 199, 100, 101, 102, , 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 288, 2 89, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, 384, 386, 387, 388, 389, 390, 392, 396, 397, 404, 405, 408, 414, 415, 416, 417 , 418, 419, 421, 422, 423, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 62, 138, 318, 722, or 938, or relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0236] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising a substitution at least at amino acid positions 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0237] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to a reference sequence corresponding to SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62.
[0238] In some embodiments, the recombinant polynucleotide comprises a sequence identical to that of at least amino acid position(s) 196, 242, 337, 33, 277, 30, 359, 283, 415, 387, 379, 205, 186, 389, 102, 164, 301, 375, 267, 380, 254, 317, 77 / 139 / 317 / 417, 105 / 317 / 417, 317 / 349 / 362 / 386, 105 / 317, 139 / 317 / 362, 233 / 317 / 405, 139 / 317, 162, The present invention also includes a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence including at least a substitution or set of substitutions at 286, 414, 417, 226, 61, 105, 230, 418, 370, 297, 237, 428, 362, 233, 235, 148, 100, 97, 382, or 358, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:62, or are relative to a reference sequence corresponding to SEQ ID NO:62.
[0239] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or a reference sequence corresponding to an even-numbered SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62.
[0240] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or to a reference sequence corresponding to SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138 or to the reference sequence corresponding to SEQ ID NO: 138.
[0241] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or a reference sequence corresponding to an even-numbered SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138.
[0242] In some embodiments, the recombinant polynucleotide comprises a sequence encoding an amino acid sequence at amino acid position(s) 242 / 283 / 286 / 359 / 418, 283 / 286, 283 / 286 / 418, 186 / 242 / 283 / 286 / 418, 205 / 286 / 359, 283, 242 / 286 / 418, 186 / 205 / 242, 283 / 286 / 359, 277 / 286 / 359 / 418, 186 / 205 / 283 / 286 / 359, 242 / 277 / 418, 242 / 283 / 286 / 418, 112 / 196 / 389, 286, 283 / 286 / 359 59 / 418, 242 / 359 / 418, 186 / 283, 186 / 359 / 418, 205 / 242 / 418, 186 / 205 / 242 / 283 / 286 / 359 / 418, 205 / 359 / 418, 186 / 205 / 283 / 286 / 418, 186 / 283 / 359, 186 / 242, 186 / 242 / 359, 283 / 359 / 418, 277 / 418, 186 / 188 / 283, 186 / 286 / 418, 186 / 242 / 286 / 359 / 418, 418, 186 / 242 / 283 / 286 / 359 / 418, 186 / 277 / 359 / 418, 242 / 283 / 286, 205 / 418, 30 / 297, 205 / 242 / 286 / 359 / 418, 186 / 205 / 359 / 418, 359 / 418, 186, 230, 33 / 297, 186 / 205, 186 / 283 / 359 / 418, 186 / 418, 205 / 242 / 283 / 359 / 418, 33 / 375 / 389, 33 / 230, 196 / 242 / 283 / 286 / 359 / 418, 186 / 242 / 283 / 359 / 418, 186 / 359, 33 / 196, 186 / 277 / 41 and a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence including at least a substitution or set of substitutions in SEQ ID NO: 8, 242, 33 / 196 / 297 / 301, 205 / 237 / 242 / 283 / 286 / 359, 186 / 205 / 283 / 359 / 418, 33 / 389, or 186 / 196 / 242 / 283 / 286 / 359 / 418, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 138, or are relative to a reference sequence corresponding to SEQ ID NO: 138.
[0243] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or to a reference sequence corresponding to SEQ ID NO:318, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318.
[0244] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or a reference sequence corresponding to an even-numbered SEQ ID NO:460-936, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318.
[0245] In some embodiments, the recombinant polynucleotide comprises a sequence identical to that of at least amino acid position(s) 363, 63, 389, 381, 197, 359, 102, 165, 388, 414, 337, 164, 416, 101, 415, 423, 364, 73, 50, 71, 388 / 419, 357, 396, 68, 76, 14, 271, 360, 266, 208, 74, 263, 264, 13, 378, 372, 300, 294, 290, 397, 95, 258, 161, 212, 198, 138, 18, 404, 273, 117, 240, 69, 278, 289, 82, 328, 61 / 186 / 417, 186 / 370 / 417, 267, 61 / 370, 186 / 267 / 370 / 417, 61 / 186, 370 / 417, 417, 61 / 370 / 382, 61 / 186 / 267 / 370 / 417, 61, 267 / 370 / 417, 267 / 370, 61 / 417, 61 / 186 / 237 / 267 / 370, 61 / 237 / 370 / 417, 370, 186 / 370, 61 / 186 / 267 / 417, 61 / 186 / 370 / 382, 370 / 382 / 417, 61 / 186 / 370, 237 / 267 / 370 / 417, 61 / 186 / 382, 61 / 267, 61 / 267 / 417, 61 / 237 / 267 / 382, 186 / 370 / 382, 237 / 267 / 370, 61 / 186 / 267 / 370, 186 / 237 / 267 / 370, 61 / 186 / 267, 61 / 186 / 237, 186, 186 / 267, 237 / 370 / 417, 242 / 414, 162 / 414, 267 / 414, 105 / 414, 162 / 242 / 414, 97 / 162 / 414, 105 The present invention also includes a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions at residues 162, 267, 414, 356, 392, 106, 308, 306, 96, 282, 113, 309, 110, 37, 201, 284, 374, 295, 88, 44, 12, or 390, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO: 318, or are relative to a reference sequence corresponding to SEQ ID NO: 318.
[0246] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:722, or to a reference sequence corresponding to SEQ ID NO:722, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722.
[0247] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 938-1098, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 938-1098, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 722, or to the reference sequence corresponding to SEQ ID NO: 722.
[0248] In some embodiments, the recombinant polynucleotide comprises a sequence encoding an amino acid sequence at amino acid position(s) 63, 63 / 96 / 370, 389, 13 / 267 / 363 / 389, 13 / 186 / 389, 50 / 267 / 363 / 370 / 389, 363 / 370, 96 / 370, 61 / 63 / 212, 11 / 305, 11, 242 / 283 / 286 / 317 / 414 / 418, 323, 334, 339, 356, 384, 408, 67, 392, 104, 355, 159, 155, 367, 31, 231, 36, 150 , 239, 103, 125, 228, 37, 189, 177, 422, 128, 220, 130, 56, 190, 156, 232, 423, 34, 99, 59, 60, 421, or 195, where the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:722, or are relative to a reference sequence corresponding to SEQ ID NO:722.
[0249] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:938, or to a reference sequence corresponding to SEQ ID NO:938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938 or to the reference sequence corresponding to SEQ ID NO:938.
[0250] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 1100-1184, or to a reference sequence corresponding to an even-numbered SEQ ID NO: 1100-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 938, or to the reference sequence corresponding to SEQ ID NO: 938.
[0251] In some embodiments, the recombinant polynucleotide comprises a sequence encoding an amino acid sequence at amino acid position(s) 308 / 357 / 390, 74 / 76 / 201 / 308 / 357, 61 / 74 / 76 / 186 / 201 / 308 / 309 / 357 / 390, 14 / 201 / 240 / 289 / 357, 308 / 415, 76 / 357 / 396, 263 / 308 / 396, 61 / 76 / 96 / 240 / 308 / 309, 14 / 306 / 415, 14 / 73 / 106 / 415, 12 / 14 / 258 / 263 / 289 / 308 / 309 / 396, 74 / 76 / 117 / 309 / 357, 14 / 258 / 263 / 357 / 396, 14 / 96 / 106 / 306, 14 / 106, 12 / 14 / 308 / 309, 14 / 357 / 390, 14 / 117 / 258 / 309 / 357, 390, 240 / 273 / 357 / 390, 61 / 76 / 18 6 / 201 / 308 / 309, 14 / 396, 309, 14, 106 / 306 / 308, 14 / 240 / 306 / 308, 12 / 14 / 186 / 357, 309 / 390, 14 / 306, 14 / 76 / 308, 117 / 208 / 258 / 263 / 289 / 308 / 309, 14 / 73 / 106, 76 / 208 / 263, 357, 14 / 308, 263, 76, 14 / 300 / 308 / 415, 2 The polynucleotide sequence includes a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions at 40, 33 / 357 / 390, 14 / 76 / 273, or 74, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12 to 437 of SEQ ID NO:938, or are relative to a reference sequence corresponding to SEQ ID NO:938.
[0252] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising substitutions at least at the amino acid positions set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0253] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0254] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions at amino acid position(s) set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0255] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to a reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to a reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0256] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising at least a substitution or set of substitutions provided in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
[0257] In some embodiments, the recombinant polynucleotide comprises a sequence corresponding to residues 12-437 of an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, or a sequence corresponding to residues 12-437 of an engineered DNA ligase variant shown in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2. The polynucleotide sequences encoding engineered DNA ligases include those comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a sequence corresponding to a DNA ligase variant.
[0258] In some embodiments, the recombinant polynucleotide comprises a polynucleotide encoding an engineered DNA ligase comprising an amino acid sequence comprising residues 12-437 of an even-numbered SEQ ID NO:40-1184, or an amino acid sequence comprising an even-numbered SEQ ID NO:40-1184, optionally, the amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions.
[0259] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising residues 12-437 of SEQ ID NO: 62, 138, 318, 722, 938, or 1108, or an amino acid sequence comprising SEQ ID NO: 62, 138, 318, 722, 938, or 1108, optionally, the amino acid sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions.
[0260] In some embodiments, the recombinant polynucleotide comprises a reference polynucleotide sequence corresponding to nucleotide residues 34 to 1311 of SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a polynucleotide sequence having at least 70%, 75%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to SEQ ID NO: 1, 61, 137, 317, 721, or 937, wherein the recombinant polynucleotide encodes an engineered DNA ligase.
[0261] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to nucleotide residues 34-1311 of an odd-numbered SEQ ID NO:39-1183, or to a reference polynucleotide sequence corresponding to an odd-numbered SEQ ID NO:39-1183, wherein said recombinant polynucleotide encodes an engineered DNA ligase.
[0262] In some embodiments, the recombinant polynucleotide is selected from the group consisting of SEQ ID NOs: 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163 , 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 2 89, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 3 51, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 41 3, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475 , 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537,539、541、543、545、547、549、551、553、555、557、559、561、563、565、567、569、571、573、575、577、579、581、583、585、587、589、591、593、595、597、599、601、603、605、607、609、611、613、615、617、619、621、623、625、627、629、631、633、635、637、639、641、643、645、647、649、651、653、655、657、659、661、663、665、667、669、671、673、675、677、679、681、683、685、687、689、691、693、695、697、699、701、703、705、707、709、711、713、715、717、719、721、723、725、727、729、731、733、735、737、739、741、743、745、747、749、751、753、755、757、759、761、763、765、767、769、771、773、775、777、779、781、783、785、787、789、791、793、795、797、799、801、803、805、807、809、811、813、815、817、819、821、823、825、827、829、831、833、835、837、839、841、843、845、847、849、851、853、855、857、859、861、863、865、867、869、871、873、875、877、879、881、883、885、887、889、891、893、895、897、899、901、903、905、907、909、911、913、915、917、919、921、923、927、929、931、933、935、937、939、941、943、945、947、949、951、953、955、957、959、961、963、965、967、969、971、973、975、977、979、981、983、985、987、989、991、993、995、997、999、1001、1003、1005、1007、1009、1011、1013、1015、1017、1019、1021、1023、1025、1027、1029、1031、1033, 1035, 1037, 1039, 1041, 1043, 1045, 1047, 1049, 1051, 1053, 1055, 1057, 1059, 1061, 1063, 1065, 1067, 1069, 1071, 1073, 1075, 1077, 1079, 1081, 1083, 1085, 1087, 1089, 1091, 1093, 1095, 1097, 1099, 1101, 1103, 1105, 1107, 1109, 1111, 1113, 1115, 1117, 1119, 1121, 1123, 1125, 1127, 1129, 1131, 1133, 1135, 1137, 1139, 1141, 1143, and a polynucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the sequence corresponding to nucleotide residues 34 to 1311 of 1145, 1147, 1149, 1151, 1153, 1155, 1157, 1159, 1161, 1163, 1165, 1167, 1169, 1171, 1173, 1175, 1177, 1179, 1181, or 1183, wherein the recombinant polynucleotide encodes a DNA ligase.
[0263] In some embodiments, the recombinant polynucleotide is selected from the group consisting of SEQ ID NOs: 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163 , 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 2 89, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 3 51, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 41 3, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475 , 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537,539、541、543、545、547、549、551、553、555、557、559、561、563、565、567、569、571、573、575、577、579、581、583、585、587、589、591、593、595、597、599、601、603、605、607、609、611、613、615、617、619、621、623、625、627、629、631、633、635、637、639、641、643、645、647、649、651、653、655、657、659、661、663、665、667、669、671、673、675、677、679、681、683、685、687、689、691、693、695、697、699、701、703、705、707、709、711、713、715、717、719、721、723、725、727、729、731、733、735、737、739、741、743、745、747、749、751、753、755、757、759、761、763、765、767、769、771、773、775、777、779、781、783、785、787、789、791、793、795、797、799、801、803、805、807、809、811、813、815、817、819、821、823、825、827、829、831、833、835、837、839、841、843、845、847、849、851、853、855、857、859、861、863、865、867、869、871、873、875、877、879、881、883、885、887、889、891、893、895、897、899、901、903、905、907、909、911、913、915、917、919、921、923、927、929、931、933、935、937、939、941、943、945、947、949、951、953、955、957、959、961、963、965、967、969、971、973、975、977、979、981、983、985、987、989、991、993、995、997、999、1001、1003、1005、1007、1009、1011、1013、1015、1017、1019、1021、1023、1025、1027、1029、1031、1033, 1035, 1037, 1039, 1041, 1043, 1045, 1047, 1049, 1051, 1053, 1055, 1057, 1059, 1061, 1063, 1065, 1067, 1069, 1071, 1073, 1075, 1077, 1079, 1081, 1083, 1085, 1087, 1089, 1091, 1093, 1095, 1097, 1099, 1101, 1103, 1105, 1107, 1109, 1111, 1113, 1115, 1117, 1119, 1121, 1123, 1125, 1127, 1129, 1131, 1133, 1135, 1137, 1139, 1 141, 1143, 1145, 1147, 1149, 1151, 1153, 1155, 1157, 1159, 1161, 1163, 1165, 1167, 1169, 1171, 1173, 1175, 1177, 1179, 1181, or 1183, wherein the recombinant polynucleotide encodes a DNA ligase.
[0264] In some embodiments, as described above, the recombinant polynucleotide comprises a polynucleotide sequence that is codon-optimized for expression of the encoded engineered DNA ligase. In some embodiments, the polynucleotide sequence is codon-optimized for expression in a prokaryotic cell. In some embodiments, the polynucleotide sequence is codon-optimized for expression in a eukaryotic cell. In some embodiments, the polynucleotide sequence is codon-optimized for expression in a bacterial cell.
[0265] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence comprising nucleotide residues 34 to 1311 of an odd-numbered SEQ ID NO: 39-1183, or a polynucleotide sequence comprising an odd-numbered SEQ ID NO: 39-1183.
[0266] In some embodiments, the recombinant polynucleotide is selected from the group consisting of SEQ ID NOs: 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163 , 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 2 89, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 3 51, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 41 3, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475 , 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537,539、541、543、545、547、549、551、553、555、557、559、561、563、565、567、569、571、573、575、577、579、581、583、585、587、589、591、593、595、597、599、601、603、605、607、609、611、613、615、617、619、621、623、625、627、629、631、633、635、637、639、641、643、645、647、649、651、653、655、657、659、661、663、665、667、669、671、673、675、677、679、681、683、685、687、689、691、693、695、697、699、701、703、705、707、709、711、713、715、717、719、721、723、725、727、729、731、733、735、737、739、741、743、745、747、749、751、753、755、757、759、761、763、765、767、769、771、773、775、777、779、781、783、785、787、789、791、793、795、797、799、801、803、805、807、809、811、813、815、817、819、821、823、825、827、829、831、833、835、837、839、841、843、845、847、849、851、853、855、857、859、861、863、865、867、869、871、873、875、877、879、881、883、885、887、889、891、893、895、897、899、901、903、905、907、909、911、913、915、917、919、921、923、927、929、931、933、935、937、939、941、943、945、947、949、951、953、955、957、959、961、963、965、967、969、971、973、975、977、979、981、983、985、987、989、991、993、995、997、999、1001、1003、1005、1007、1009、1011、1013、1015、1017、1019、1021、1023、1025、1027、1029、1031、1033, 1035, 1037, 1039, 1041, 1043, 1045, 1047, 1049, 1051, 1053, 1055, 1057, 1059, 1061, 1063, 1065, 1067, 1069, 1071, 1073, 1075, 1077, 1079, 1081, 1083, 1085, 1087, 1089, 1091, 1093, 1095, 1097, 1099, 1101, 1103, 1105, 1107, 1109, 1111, 1113, 1114 1151, 1153, 1155, 1157, 1159, 1161, 1163, 1165, 1167, 1169, 1171, 1173, 1175, 1177, 1179, 1181, or 1183.
[0267] In some embodiments, the recombinant polynucleotide is selected from the group consisting of SEQ ID NOs: 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163 , 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 2 89, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 3 51, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 41 3, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475 , 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537,539、541、543、545、547、549、551、553、555、557、559、561、563、565、567、569、571、573、575、577、579、581、583、585、587、589、591、593、595、597、599、601、603、605、607、609、611、613、615、617、619、621、623、625、627、629、631、633、635、637、639、641、643、645、647、649、651、653、655、657、659、661、663、665、667、669、671、673、675、677、679、681、683、685、687、689、691、693、695、697、699、701、703、705、707、709、711、713、715、717、719、721、723、725、727、729、731、733、735、737、739、741、743、745、747、749、751、753、755、757、759、761、763、765、767、769、771、773、775、777、779、781、783、785、787、789、791、793、795、797、799、801、803、805、807、809、811、813、815、817、819、821、823、825、827、829、831、833、835、837、839、841、843、845、847、849、851、853、855、857、859、861、863、865、867、869、871、873、875、877、879、881、883、885、887、889、891、893、895、897、899、901、903、905、907、909、911、913、915、917、919、921、923、927、929、931、933、935、937、939、941、943、945、947、949、951、953、955、957、959、961、963、965、967、969、971、973、975、977、979、981、983、985、987、989、991、993、995、997、999、1001、1003、1005、1007、1009、1011、1013、1015、1017、1019、1021、1023、1025、1027、1029、1031、1033, 1035, 1037, 1039, 1041, 1043, 1045, 1047, 1049, 1051, 1053, 1055, 1057, 1059, 1061, 1063, 1065, 1067, 1069, 1071, 1073, 1075, 1077, 1079, 1081, 1083, 1085, 1087, 1089, 1091, 1093, 1095, 1097, 1099, 1101, 1103, 1105, 1107, 1109, 1111 , 1113, 1115, 1117, 1119, 1121, 1123, 1125, 1127, 1129, 1131, 1133, 1135, 1137, 1139, 1141, 1143, 1145, 1147, 1149, 1151, 1153, 1155, 1157, 1159, 1161, 1163, 1165, 1167, 1169, 1171, 1173, 1175, 1177, 1179, 1181, or 1183.
[0268] In some embodiments, the recombinant polynucleotide comprises a polynucleotide sequence comprising nucleotide residues 34 to 1311 of SEQ ID NO: 1, 61, 137, 317, 721, 937, or 1107, or a polynucleotide sequence comprising SEQ ID NO: 1, 61, 137, 317, 721, 937, or 1107.
[0269] In some embodiments, the recombinant polynucleotide hybridizes under highly stringent conditions to a reference polynucleotide sequence described herein encoding an engineered DNA ligase, e.g., a recombinant polynucleotide provided in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, or a reverse complement thereof. In some embodiments, the reference polynucleotide sequence corresponds to a polynucleotide sequence encoding nucleotide residues 34 to 1311 of SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a sequence corresponding to SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a reverse complement thereof, or any of the other engineered DNA ligases provided herein. In some embodiments, the recombinant polynucleotide encodes a DNA ligase and hybridizes under highly stringent conditions to the reverse complement of a reference polynucleotide sequence corresponding to nucleotide residues 34-1311 of an odd-numbered SEQ ID NO:39-1183, or to a reference polynucleotide sequence corresponding to an odd-numbered SEQ ID NO:39-1183.
[0270] In some embodiments, the recombinant polynucleotide hybridizes under highly stringent conditions to the reverse complement of a reference polynucleotide sequence encoding an engineered DNA ligase, wherein the engineered DNA ligase comprises an amino acid sequence having one or more amino acid differences compared to SEQ ID NOs: 2, 62, 138, 318, 722, or 938 at a residue position selected from any of the positions set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2. In some embodiments, polynucleotides that hybridize under highly stringent conditions comprise a reference polynucleotide sequence corresponding to nucleotide residues 39 to 1183 of SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a polynucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to SEQ ID NO: 1, 61, 137, 317, 721, or 937. In some further embodiments, polynucleotides that hybridize under highly stringent conditions comprise nucleotide residues 39 to 1183 of a polynucleotide sequence set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, or a polynucleotide sequence having at least 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to a polynucleotide sequence set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2.
[0271] In some embodiments, the recombinant polynucleotide comprises: (a) a sequence corresponding to residues 12 to 437 of SEQ ID NO:2; a sequence corresponding to residues 12 to 613 of SEQ ID NO:4; residues 12 to 614 of SEQ ID NO:6; residues 12 to 610 of SEQ ID NO:8; residues 12 to 606 of SEQ ID NO:10; residues 12 to 615 of SEQ ID NO:12; residues 12 to 594 of SEQ ID NO:14; residues 12 to 620 of SEQ ID NO:16; residues 12 to 608 of SEQ ID NO:18; residues 12 to 611 of SEQ ID NO:20; residues 12 to 614 of SEQ ID NO:22, residues 12 to 611 of SEQ ID NO:24, residues 12 to 609 of SEQ ID NO:26, residues 12 to 422 of SEQ ID NO:28, residues 12 to 518 of SEQ ID NO:30, residues 12 to 438 of SEQ ID NO:32, residues 12 to 381 of SEQ ID NO:34, residues 12 to 424 of SEQ ID NO:36, or residues 12 to 390 of SEQ ID NO:38, or (b) a polynucleotide sequence encoding an engineered DNA ligase, comprising a polynucleotide sequence comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the sequence corresponding to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38.
[0272] In some embodiments, the recombinant polynucleotide comprises: (a) comprising residues 12 to 437 of SEQ ID NO:2; residues 12 to 613 of SEQ ID NO:4; residues 12 to 614 of SEQ ID NO:6; residues 12 to 610 of SEQ ID NO:8; residues 12 to 606 of SEQ ID NO:10; residues 12 to 615 of SEQ ID NO:12; residues 12 to 594 of SEQ ID NO:14; residues 12 to 620 of SEQ ID NO:16; residues 12 to 608 of SEQ ID NO:18; residues 12 to 611 of SEQ ID NO:20; residues 12 to 614 of SEQ ID NO:22, residues 12 to 611 of SEQ ID NO:24, residues 12 to 609 of SEQ ID NO:26, residues 12 to 422 of SEQ ID NO:28, residues 12 to 518 of SEQ ID NO:30, residues 12 to 438 of SEQ ID NO:32, residues 12 to 381 of SEQ ID NO:34, residues 12 to 424 of SEQ ID NO:36, or residues 12 to 390 of SEQ ID NO:38; or (b) comprises a polynucleotide sequence encoding an engineered DNA ligase comprising an amino acid sequence comprising SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38.
[0273] In some embodiments, the recombinant polynucleotide comprises: (a) Residues 34 to 1311 of SEQ ID NO: 1, residues 34 to 1839 of SEQ ID NO: 3, residues 34 to 1842 of SEQ ID NO: 5, residues 34 to 1830 of SEQ ID NO: 7, residues 34 to 1818 of SEQ ID NO: 9, residues 34 to 1845 of SEQ ID NO: 11, residues 34 to 1782 of SEQ ID NO: 13, residues 34 to 1860 of SEQ ID NO: 15, residues 34 to 1824 of SEQ ID NO: 17, residues 34 to 1833 of SEQ ID NO: 19 ...860 of SEQ ID NO: 15, residues 34 to 1824 of SEQ ID NO: 17, residues 34 to 1833 of SEQ ID NO: 19, residues 34 to 1818 of SEQ ID NO: 9, residues 34 to 1845 of SEQ ID NO: 11, residues 34 to 1860 of SEQ ID NO: 15, residues 34 to 182 a polynucleotide sequence corresponding to residues 34 to 1842 of SEQ ID NO: 21, residues 34 to 1833 of SEQ ID NO: 23, residues 34 to 1827 of SEQ ID NO: 25, residues 34 to 1266 of SEQ ID NO: 27, residues 34 to 1554 of SEQ ID NO: 29, residues 34 to 1314 of SEQ ID NO: 31, residues 34 to 1143 of SEQ ID NO: 33, residues 34 to 1272 of SEQ ID NO: 35, or residues 34 to 1170 of SEQ ID NO: 37; or (b) a polynucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a polynucleotide sequence corresponding to SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, or 37, wherein the polynucleotide sequence encodes a DNA ligase.
[0274] In some embodiments, the recombinant polynucleotide comprises (a) residues 34 to 1311 of SEQ ID NO: 1, residues 34 to 1839 of SEQ ID NO: 3, residues 34 to 1842 of SEQ ID NO: 5, residues 34 to 1830 of SEQ ID NO: 7, residues 34 to 1818 of SEQ ID NO: 9, residues 34 to 1845 of SEQ ID NO: 11, residues 34 to 1782 of SEQ ID NO: 13, residues 34 to 1860 of SEQ ID NO: 15, residues 34 to 1824 of SEQ ID NO: 17, residues 34 to 1839 of SEQ ID NO: 18, residues 34 to 1842 of SEQ ID NO: 19, residues 34 to 1818 of SEQ ID NO: 20, residues 34 to 1845 of SEQ ID NO: 21, residues 34 to 1860 of SEQ ID NO: 22, residues 34 to 1824 of SEQ ID NO: 23, residues 34 to 1839 of SEQ ID NO: 24, residues 34 to 1842 of SEQ ID NO: 25, residues 34 to 1824 of SEQ ID NO: 26, residues 34 to 1839 of SEQ ID NO: 27, residues 34 to 1845 of SEQ ID NO: 28, residues 34 to 1845 of SEQ ID NO: 29, residues 34 to 1860 of SEQ ID NO: 30, residues 34 to 1824 of SEQ ID NO: 31, residues 34 to 1839 of SEQ ID NO: 32, residues 34 to 1833 of SEQ ID NO: 19, residues 34 to 1842 of SEQ ID NO: 21, residues 34 to 1833 of SEQ ID NO: 23, residues 34 to 1827 of SEQ ID NO: 25, residues 34 to 1266 of SEQ ID NO: 27, residues 34 to 1554 of SEQ ID NO: 29, residues 34 to 1314 of SEQ ID NO: 31, residues 34 to 1143 of SEQ ID NO: 33, residues 34 to 1272 of SEQ ID NO: 35, or residues 34 to 1170 of SEQ ID NO: 37; or (b) comprises a polynucleotide sequence comprising SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, or 37.
[0275] In some embodiments, a recombinant polynucleotide encoding any of the DNA ligases herein is manipulated in various ways to facilitate expression of the DNA ligase polypeptide. In some embodiments, the recombinant polynucleotide encoding the DNA ligase comprises an expression vector in which one or more control sequences are present to regulate expression of the DNA ligase polynucleotide and / or polypeptide. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. In some embodiments, control sequences include, among others, promoters, leader sequences, polyadenylation sequences, propeptide sequences, signal peptide sequences, and transcription terminators.
[0276] In some embodiments, an appropriate promoter is selected based on host cell selection. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include the promoters for the E. coli lac operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis alpha-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens alpha-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus Examples of suitable promoters include, but are not limited to, promoters obtained from the (A. subtilis) xylA and xylB genes, and prokaryotic beta-lactamase genes (see, e.g., Villa-Kamaroff et al., Proc. Natl. Acad. Sci. USA, 1978, 75:3727-3731), and the tac promoter (see, e.g., DeBoer et al., Proc. Natl. Acad. Sci. USA, 1983, 80:21-25).Exemplary promoters for filamentous fungal host cells include those encoding Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral alpha-amylase, Aspergillus niger acid-stable alpha-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans Examples of promoters that may be used include, but are not limited to, promoters obtained from the genes for Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (see, e.g., WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the genes for Aspergillus niger neutral alpha-amylase and Aspergillus oryzae triose phosphate isomerase), and mutant, truncated, and hybrid promoters thereof. Exemplary yeast cell promoters can be derived from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase.Other useful promoters for yeast host cells are known in the art (see, e.g., Romanos et al., Yeast, 1992, 8:423-488). Exemplary promoters for use in insect cells include, but are not limited to, the polyhedrin, p10, ELT, OpIE2, and hr5 / ie1 promoters. Exemplary promoters for use in mammalian cells include, but are not limited to, promoters derived from cytomegalovirus (CMV), chicken β-actin promoter fused with a CMV enhancer, promoters derived from simian vacuolating virus 40 (SV40), promoters derived from Homo sapiens phosphoglycerate kinase, promoters derived from beta-actin, promoters derived from elongation factor-1a or glyceraldehyde-3-phosphate dehydrogenase, and promoters derived from Gallus β-actin.
[0277] In some embodiments, the control sequence is a suitable transcription terminator sequence (i.e., a sequence recognized by a host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the DNA ligase polypeptide. Any suitable terminator that is functional in the selected host cell can be used in the present invention. In the case of bacterial expression, the transcription terminator can be a Rho-dependent terminator that depends on Rho transcription factors, or a Rho-independent or endogenous terminator that does not require transcription factors. For exemplary bacterial transcription terminators, see Peters et al., J Mol Biol., 2011, 412(5):793-813. Exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known in the art (see, e.g., Romanos et al., supra).Exemplary terminators for insect and mammalian cells include, but are not limited to, those from cytomegalovirus (CMV), simian virus 40 (SV40), homo sapiens growth hormone (hGH), bovine growth hormone (BGH), and human or rabbit beta globulin.
[0278] In some embodiments, the control sequence is a suitable leader sequence, a non-translated region of an mRNA that is important for translation by the host cell. In some embodiments, the leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the DNA ligase polypeptide. Any suitable leader sequence that is functional in the selected host cell finds use in the present invention. Exemplary leaders for filamentous fungal host cells are obtained from the genes for Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Suitable leaders for yeast host cells are obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP). Suitable leaders for mammalian host cells include, but are not limited to, the 5'-UTR elements present in orthopoxvirus mRNAs.
[0279] In some embodiments, the regulatory sequence is a polyadenylation sequence (i.e., a sequence operably linked to the 3' end of a nucleic acid sequence that, upon transcription, is recognized by a host cell as a signal for adding polyadenosine residues to the transcribed mRNA). Any suitable polyadenylation sequence that is functional in the host cell of choice finds use in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase. Useful polyadenylation sequences for yeast host cells are known (see, e.g., Guo and Sherman, Mol. Cell. Biol., 1995, 15:5983-5990). Useful polyadenylation and 3' UTR sequences for insect and mammalian host cells include, but are not limited to, the OpIE2 polyA sequence, the D. melanogaster metallothionein (Mt) polyA signal sequence, the D. melanogaster alcohol dehydrogenase (adh), the SV40 polyA signal sequence, and the 3'-UTRs of alpha- and beta-globin mRNAs containing sequence elements that increase mRNA stability and translation.
[0280] In some embodiments, the control sequence is also a signal peptide (i.e., a coding region encoding an amino acid sequence linked to the amino terminus of a polypeptide that directs the encoded polypeptide into the secretory pathway of a cell). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence inherently contains a signal peptide coding region naturally linked in translation reading frame with the segment of the coding region that encodes the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a host cell of choice is used to express the engineered polypeptide(s). Effective signal peptide coding regions for bacterial host cells include, but are not limited to, those obtained from the genes encoding Bacillus NClB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Additional signal peptides are known in the art (see, e.g., Simonen and Palva, Microbiol. Rev., 1993, 57:109-137).In some embodiments, useful signal peptide coding regions for filamentous fungal host cells include, but are not limited to, the signal peptide coding regions obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Useful signal peptides for yeast host cells include, but are not limited to, those derived from the genes for Saccharomyces cerevisiae alpha-factor and Saccharomyces cerevisiae invertase. Useful signal peptides for insect and mammalian host cells include, but are not limited to, those from the immunoglobulin gamma (IgG) gene and signal peptides in human secreted proteins such as the human β-galactosidase polypeptide.
[0281] In some embodiments, the control sequence is a propeptide coding region that encodes an amino acid sequence positioned at the amino terminus of a polypeptide. The resulting polypeptide is referred to as a "proenzyme," "propolypeptide," or "zymogen." A propolypeptide can be converted to a mature, active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. The propeptide coding region can be obtained from any suitable source, including, but not limited to, genes for Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae alpha-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (see, e.g., WO 95 / 33836). When both the signal peptide region and the propeptide region are present at the amino terminus of a polypeptide, the propeptide region is positioned adjacent to the amino terminus of the polypeptide and the signal peptide region is positioned adjacent to the amino terminus of the propeptide region.
[0282] In some embodiments, regulatory sequences are also utilized. These sequences facilitate regulation of polypeptide expression relative to host cell growth. Examples of regulatory systems are those that turn gene expression on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter.
[0283] In another aspect, the present disclosure provides recombinant expression vectors comprising a polynucleotide encoding an engineered DNA ligase polypeptide and, depending on the type of host into which it will be introduced, one or more expression control regions, such as a promoter and terminator, an origin of replication, etc. In some embodiments, the various nucleic acids and control sequences described herein are joined together (i.e., operably linked) to create a recombinant expression vector. Alternatively, in some embodiments, the nucleic acid sequences of the present disclosure are expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression.
[0284] The recombinant expression vector can be any suitable vector (e.g., a plasmid or a virus) that can be conveniently subjected to recombinant DNA procedures and can result in the expression of the DNA ligase polynucleotide sequence. The selection of a vector typically depends on the compatibility of the vector with the host cell into which the vector is introduced. The vector can be a linear or closed circular plasmid.
[0285] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extrachromosomal entity whose replication is independent of chromosomal replication, e.g., a plasmid, extrachromosomal element, minichromosome, or artificial chromosome). The vector may include any means for ensuring self-replication. In some alternative embodiments, the vector is one that, upon introduction into a host cell, is integrated into the genome and replicates along with the chromosome(s) into which it has been integrated. Furthermore, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids, and / or transposons, that together comprise the total DNA to be introduced into the genome of the host cell are utilized.
[0286] In some embodiments, the recombinant polynucleotide can be provided on a non-replicating expression vector or plasmid. In some embodiments, the non-replicating expression vector or plasmid can be based on a replication-deficient viral vector (see, for example, Travieso et al., npj Vaccines, 2022, Vol. 7, Article 75).
[0287] In some embodiments, the expression vector contains one or more selectable markers that allow for easy selection of transformed cells. A "selectable marker" is a gene whose product provides biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, etc. Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance, such as ampicillin, kanamycin, chloramphenicol, or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase; e.g., from A. nidulans or A. orzyae), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase; e.g., from S. hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase; e.g., from A. nidulans or A. orzyae), sC (sulfate adenyltransferase), and trpC (anthranilate synthase), and equivalents thereof.
[0288] In another aspect, the disclosure provides host cells comprising a polynucleotide encoding at least one engineered DNA ligase polypeptide of the disclosure, wherein the polynucleotide(s) is / are operably linked to one or more control sequences for expression of the engineered DNA ligase enzyme(s) in the host cell. In some embodiments, the host cell comprises an expression vector comprising a polynucleotide encoding an engineered DNA ligase polypeptide described herein, wherein the polynucleotide is operably linked to one or more control sequences. Suitable host cells for use in expressing the polypeptides encoded by the expression vectors of the invention are known in the art and include, but are not limited to, bacterial cells, such as E. coli, B. subtilis, Vibrio fluvialis, Streptomyces, and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)); insect cells, such as Drosophila S2 and Spodoptera Sf9 cells; animal cells, such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Exemplary host cells also include various Escherichia coli strains (eg, W3110(ΔfhuA) and BL21).
[0289] In another aspect, the disclosure provides a method of producing an engineered DNA ligase polypeptide, the method comprising culturing a host cell capable of expressing a polynucleotide encoding the engineered DNA ligase polypeptide under conditions suitable for expression of the polypeptide, such that the engineered DNA ligase is produced. In some embodiments, the method further comprises the step(s) of isolating and / or purifying the DNA ligase polypeptide described herein.
[0290] Suitable culture medium and growth conditions for host cells are known in the art.Any suitable method for introducing polynucleotide for expressing DNA ligase polypeptide into cell is contemplated to be used in the present invention.Suitable techniques include but are not limited to electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection and protoplast fusion.
[0291] In some embodiments, recombinant polypeptides (e.g., DNA ligase variants) can be produced using any suitable method known in the art. For example, there are a wide variety of different mutagenesis techniques well known to those of skill in the art. Additionally, mutagenesis kits are also available from many commercial molecular biology suppliers. Methods are available for making specific substitutions at defined amino acids (site-directed), specific or random mutations in local regions of a gene (region-directed), or random mutagenesis throughout a gene (e.g., saturation mutagenesis). Numerous suitable methods for generating enzyme variants are known to those of skill in the art, including, but not limited to, site-directed mutagenesis of single- or double-stranded DNA using PCR, cassette mutagenesis, gene synthesis, error-prone PCR, shuffling, and chemical saturation mutagenesis, or any other suitable method known in the art. Non-limiting examples of methods used in DNA and protein manipulation are provided in the following patents: U.S. Patent No. 6,117,679; U.S. Patent No. 6,420,175; U.S. Patent No. 6,376,246; U.S. Patent No. 6,586,182; U.S. Patent No. 7,747,391; U.S. Patent No. 7,747,393; U.S. Patent No. 7,783,428; and U.S. Patent No. 8,383,346. After variants are generated, they can be screened for any desired property (e.g., high or increased activity, or low or decreased activity, increased thermal activity, increased stability, increased substrate range, increased inhibitor resistance / tolerance, increased solubility, pH stability, etc.).
[0292] In some embodiments, engineered DNA ligase polypeptides having the properties disclosed herein can be obtained by subjecting a polynucleotide encoding a naturally occurring or engineered DNA ligase polypeptide to suitable mutagenesis and / or directed evolution methods known in the art, for example, as described herein. Exemplary directed evolution techniques are mutagenesis and / or DNA shuffling (e.g., Stemmer, Proc. Natl. Acad. Sci. USA, 1994, 91:10747-10751; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767 and U.S. Patent No. 6,537,746). Other directed evolution procedures that can be used include, inter alia, the staggered extension process (StEP), in vitro recombination (e.g., Zhao et al., Nat. Biotechnol., 1998, 16:258-261), mutagenic PCR (e.g., Caldwell et al., PCR Methods Appl., 1994, 3:S136-S140), and cassette mutagenesis (e.g., Black et al., Proc. Natl. Acad. Sci. USA, 1996, 93:3525-3529).
[0293] Mutagenesis and directed evolution methods can be applied to polynucleotides encoding DNA ligases to generate libraries of variants that can be expressed, screened, and assayed. Any suitable mutagenesis and directed evolution method can be used in the present disclosure and is known in the art (e.g., U.S. Patent Nos. 5,605,793, 5,811,238, 5,830,721, 5,834,252, 5,837,458, 5,928,905, 6,096,548, 6,117,679, 6,132,970, 6,165,793, 6,180,406, 6,251,674, 6,265,201, 6,277,638, 6,287, No. 861, No. 6,287,862, No. 6,291,242, No. 6,297,053, No. 6,303,344, No. 6,309,883, No. 6,319,713, No. 6,319,714, No. 6,323,030, No. 6,326,204, No. No. 6,335,160, No. 6,335,198, No. 6,344,356, No. 6,352,859, No. 6,355,484, No. 6,358,740, No. 6,358,742, No. 6,365,377, No. 6,365,408, No. 6,368, No. 861, No. 6,372,497, No. 6,337,186, No. 6,376,246, No. 6,379,964, No. 6,387,702, No. 6,391,552, No. 6,391,640, No. 6,395,547, No. 6,406,855, No. No. 6,406,910, No. 6,413,745, No. 6,413,774, No. 6,420,175, No. 6,423,542, No. 6,426,224, No. 6,436,675, No. 6,444,468, No. 6,455,253, No. 6,479, No. 652, No. 6,482,647, No. 6,483,011, No. 6,484,105, No. 6,489,146, No. 6,500,617, No. 6,500,639, No. 6,506,602, No. 6,506,603, No. 6,518,065, No. No. 6,519,065, No. 6,521,453, No. 6,528,311, No. 6,537,746, No. 6,573,098, No. 6,576,467, No. 6,579,678, No. 6,586,182, No. 6,602,986, No. 6,605,No. 430, No. 6,613,514, No. 6,653,072, No. 6,686,515, No. 6,703,240, No. 6,716,631, No. 6,825,001, No. 6,902,922, No. 6,917,882, No. 6,946,296, No. 6,961,664, No. 6,995,017, No. 7,024,312, No. 7,058,515, No. 7,105,297, No. 7,148,0 No. 54, No. 7,220,566, No. 7,288,375, No. 7,384,387, No. 7,421,347, No. 7,430,477, No. 7,462,469, No. 7,534,564, No. 7,620,500, No. 7,620,502, No. 7,629,170, No. 7,702,464, No. 7,747,391, No. 7,747,393, No. 7,751,986, No. 7,776,598 Numbers: 7,783,428, 7,795,030, 7,853,410, 7,868,138, 7,783,428, 7,873,477, 7,873,499, 7,904,249, 7,957,912, 7,981,614, 8,014,961, 8,029,988, 8,048,674, 8,058,001, 8,076,138 No. 8,108,150, No. 8,170,806, No. 8,224,580, No. 8,377,681, No. 8,383,346, No. 8,457,903, No. 8,504,498, No. 8,589, No. 085, No. 8,762,066, No. 8,768,871, No. 9,593,326, No. 9,665,694, No. 9,684,771, PCT Non-US Corporation; Ling et al., Anal. Biochem., 1997, 254(2): 157-78; Dale et al., Meth. Mol. Biol., 1996, 57: 369-74; Smith, Ann. Rev. Genet., 1985, 19: 423-462; Botstein et al. al.,Science,1985,229:1193-1201;Carter,Biochem.J.,1986,237:1-7;Kramer et al.,Cell,1984,38:879-887;Wells et al.,Gene,1985,34:315-323;Minshull et al.,Curr.Op.Chem.Biol.,1999,3:284-290;Christians et al.,Nat.Biotechnol.,1999,17:259-264;Crameri et al.,Nature,1998,391:288-291;Crameri, et al. al.,Nat.Biotechnol.,1997,15:436-438;Zhang et al.,Proc.Nat.Acad.Sci.USA,1997,94:4504-4509;Crameri et al. al., Nat. Biotechnol., 1996, 14:315-319; Stemmer, Nature, 1994, 366:389-391; Stemmer, Proc. Nat. Acad. Sci. USA, 1994, 91:10747-10751; European Patent No. 3 049 973; International Publication No. WO 95 / 22625; International Publication No. WO 97 / 0078; International Publication No. WO 97 / 35966; International Publication No. WO 98 / 27230; International Publication No. WO 00 / 42651; International Publication No. WO 01 / 75767; International Publication No. WO 2009 / 152336; and International Publication No. WO 2015 / 048573 (all of which are incorporated herein by reference).
[0294] In some embodiments, clones obtained after mutagenesis treatment are screened by subjecting the enzyme preparation to defined treatment or assay conditions (e.g., temperature, pH conditions, type of ligase substrate, input substrate concentration, etc.) and measuring DNA ligase activity or other appropriate assay conditions after treatment. Clones containing polynucleotides encoding the polypeptides of interest are then isolated from the gene, sequenced to identify nucleotide sequence changes (if any), and used to express the enzyme in host cells. Measurement of enzyme activity from expression libraries can be performed using any suitable method known in the art and described in the Examples.
[0295] For engineered polypeptides of known sequence, polynucleotides encoding the polypeptides can be prepared by standard solid-phase synthesis methods according to known synthesis methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then joined (e.g., by enzymatic or chemical ligation, or polymerase-mediated methods) to form any desired contiguous sequence (e.g., Hughes et al., Cold Spring Harb Perspect Biol. 2017 Jan;9(1):a023812). For example, the polynucleotides and oligonucleotides disclosed herein can be prepared by chemical synthesis using the phosphoramidite method (e.g., Beaucage et al., Tet. Lett., 1981, 22:1859-69; and Matthes et al., EMBO J., 1984, 3:801-05), as typically performed in automated synthesis methods.
[0296] In some embodiments, a method for preparing an engineered DNA ligase polypeptide can include (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence selected from the amino acid sequences of any of the variants described herein, and (b) expressing the DNA ligase polypeptide encoded by the polynucleotide. In some embodiments of the method, the amino acid sequence encoded by the polynucleotide can have one or several (e.g., up to 3, 4, 5, or up to 10) amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1 to 2, 1 to 3, 1 to 4, 1 to 5, 1 to 6, 1 to 7, 1 to 8, 1 to 9, 1 to 10, 1 to 15, 1 to 20, 1 to 21, 1 to 22, 1 to 23, 1 to 24, 1 to 25, 1 to 30, 1 to 35, 1 to 40, 1 to 45, or 1 to 50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has deletions, insertions, and / or substitutions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residues. In some embodiments, the amino acid sequence optionally has deletions, insertions, and / or substitutions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residues. In some embodiments, the substitutions are conservative or non-conservative.
[0297] In some embodiments, any of the engineered DNA ligase polypeptides expressed in the host cells are recovered and / or purified from the cells and / or culture medium using any one or more of known techniques for protein purification, including lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography, among others.
[0298] Chromatographic techniques for isolating / purifying DNA ligase polypeptides include, among others, reverse-phase chromatography, high-performance liquid chromatography, ion-exchange chromatography, hydrophobic interaction chromatography, size-exclusion chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying DNA ligases depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, and the like, and will be apparent to those skilled in the art. In some embodiments, affinity techniques can be used to isolate DNA ligases. For affinity chromatography purification, antibodies that specifically bind to DNA ligase polypeptides can be used. In some embodiments, an affinity tag, such as a His tag, can be introduced into the DNA ligase polypeptide for isolation / purification purposes.
[0299] composition In a further aspect, the present disclosure provides compositions of the DNA ligases disclosed herein. In some embodiments, the compositions comprise at least one engineered DNA ligase polypeptide described herein. In some embodiments, the engineered DNA ligase polypeptide in the composition is isolated or purified. In some embodiments, the DNA ligase is combined with other components and compounds to provide compositions and formulations comprising the engineered DNA ligase polypeptide suitable for different applications and uses.
[0300] In some embodiments, the composition comprises at least one engineered DNA ligase described herein, e.g., an engineered DNA ligase provided in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the composition comprises an engineered DNA ligase comprising an amino acid sequence comprising residues 12-437 of an even-numbered SEQ ID NO:40-1184, or an amino acid sequence comprising an even-numbered SEQ ID NO:40-1184.
[0301] In some embodiments, the composition comprises a DNA ligase provided in Table 8.1. In some embodiments, the composition comprises: (a) comprising residues 12 to 437 of SEQ ID NO:2; residues 12 to 613 of SEQ ID NO:4; residues 12 to 614 of SEQ ID NO:6; residues 12 to 610 of SEQ ID NO:8; residues 12 to 606 of SEQ ID NO:10; residues 12 to 615 of SEQ ID NO:12; residues 12 to 594 of SEQ ID NO:14; residues 12 to 620 of SEQ ID NO:16; residues 12 to 608 of SEQ ID NO:18; residues 12 to 611 of SEQ ID NO:20; residues 12 to 614 of SEQ ID NO:22, residues 12 to 611 of SEQ ID NO:24, residues 12 to 609 of SEQ ID NO:26, residues 12 to 422 of SEQ ID NO:28, residues 12 to 518 of SEQ ID NO:30, residues 12 to 438 of SEQ ID NO:32, residues 12 to 381 of SEQ ID NO:34, residues 12 to 424 of SEQ ID NO:36, or residues 12 to 390 of SEQ ID NO:38; or (b) comprises an amino acid sequence comprising SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38.
[0302] In some embodiments, the composition further comprises a buffer, a nucleotide substrate (e.g., ATP), and / or at least one or more DNA ligase substrates, such as one or more of a synthetic oligonucleotide or a recombinant polynucleotide. In some embodiments, the DNA ligase substrate comprises a DNA adaptor or linker. In some embodiments, the composition further comprises a divalent cation (e.g., Mg +2 or Mn +2 ), a reducing agent (e.g., DTT), and / or a salt (e.g., KCl and / or NaCl).
[0303] In some embodiments, the composition further comprises an additive or ligation enhancer, including one or more of DMSO, betaine, polyethylene glycol (e.g., PEG 6000, PEG 8000, etc.), bovine serum albumin, Ficoll, and dextran (e.g., dextran 6000), among others. In some embodiments, the composition comprises 1% to 40% v / v DMSO. In some embodiments, the composition comprises 0.1 M to 3 M betaine. In some embodiments, the composition comprises 0.5% to 20% w / v PEG (e.g., PEG 6000 or PEG 8000).
[0304] In some embodiments, the composition further comprises a second nucleic acid-modifying enzyme, including, among others, a DNA polymerase (e.g., a thermal DNA polymerase, an isothermal polymerase, a reverse transcriptase, an RNA polymerase, etc.), a restriction enzyme, a polynucleotide kinase, or a terminal transferase.
[0305] In some embodiments, the engineered DNA ligase described herein is provided as a solution, lyophilizate, or immobilized on a substrate. In some embodiments, the substrate is a solid substrate, a porous substrate, a membrane, or a particle. The enzyme can be encapsulated in a matrix or membrane. In some embodiments, the matrix includes polymeric materials such as calcium alginate, agar, k-carrageenan, polyacrylamide, agarose or its derivatives (e.g., cross-linked agarose), and collagen, or solid matrices such as activated carbon, porous ceramic, and diatomaceous earth. In some embodiments, the matrix is a particle, membrane, or fiber. Types of membranes include nylon, cellulose, polysulfone, or polyacrylate, among others.
[0306] In some embodiments, the enzyme is immobilized on the surface of a support material. In some embodiments, the enzyme is adsorbed to the support material. In some embodiments, the enzyme is immobilized on the support material by a covalent bond. Support materials include inorganic materials such as alumina, silica, porous glass, ceramic, diatomaceous earth, clay, and bentonite, among others, or organic materials such as cellulose (CMC, DEAE-cellulose), starch, activated carbon, polyacrylamide, polystyrene, and ion exchange resins such as Amberlite, Sephadex, and Dowex.
[0307] Uses of Engineered DNA Ligase Polypeptides and Kits In another aspect, the disclosure provides for the use of engineered DNA ligases for ligating or joining DNA molecules, such as for polynucleotide synthesis, DNA repair, or other molecular biological and diagnostic uses.
[0308] In some embodiments, engineered DNA ligases are used to ligate polynucleotides or oligonucleotides. In some embodiments, engineered DNA ligases are used to synthesize polynucleotides from shorter oligonucleotides or polynucleotides. In some embodiments, a method for ligating at least a first polynucleotide strand and a second polynucleotide strand comprises contacting the first polynucleotide strand and the second polynucleotide strand with an engineered DNA ligase described herein in the presence of a nucleotide substrate under conditions suitable for ligating the first polynucleotide strand to the second polynucleotide strand, wherein the first polynucleotide strand comprises a ligatable 5' end and the second polynucleotide strand comprises a ligatable 3' end to the 5' end of the first polynucleotide strand. In some embodiments, the 5' end of the first polynucleotide strand has a 5'-phosphate (e.g., -PO4) and the 3' end of the second polynucleotide strand has a 3' hydroxyl (OH). In some embodiments, the first polynucleotide strand and the second polynucleotide strand comprise DNA or a mixture of DNA and RNA.
[0309] In some embodiments, the method further includes a third polynucleotide strand, wherein the first polynucleotide strand and the second polynucleotide strand hybridize adjacent to each other on the third polynucleotide strand, positioning the 5' end of the first polynucleotide strand adjacent to the 3' end of the second polynucleotide strand. In some embodiments of the method, the third polynucleotide strand is contiguous with the first polynucleotide strand or the second polynucleotide strand. In some embodiments of the method, the third polynucleotide strand is contiguous with the first polynucleotide strand and the second polynucleotide strand to form a single, continuous polynucleotide ligase substrate. In some embodiments, the first polynucleotide strand, the second polynucleotide strand, and the third polynucleotide strand comprise separate polynucleotides. In some embodiments, the third polynucleotide strand is RNA, DNA, or a mixture of RNA and DNA, preferably DNA.
[0310] In some embodiments of the method, the third polynucleotide comprises a splint or bridge polynucleotide, wherein the 5' terminal sequence of the first polynucleotide strand and the 3' terminal sequence of the second polynucleotide strand hybridize adjacent to each other on the splint or bridge polynucleotide, positioning the 5' end of the first polynucleotide strand adjacent to the 3' end of the second polynucleotide strand. In some embodiments, the splint or bridge polynucleotide is RNA, DNA, or a mixture of RNA and DNA, preferably DNA.
[0311] In some embodiments of the method, the first polynucleotide strand hybridizes to a third polynucleotide strand to form a first double-stranded polynucleotide substrate, and the second polynucleotide strand hybridizes to a fourth polynucleotide strand to form a second double-stranded polynucleotide substrate.
[0312] In some embodiments, the first double-stranded polynucleotide substrate comprises a blunt-ended 5' end of a first polynucleotide strand, and the second double-stranded polynucleotide substrate comprises a blunt-ended 3' end of a second polynucleotide strand.
[0313] In some embodiments, the first double-stranded polynucleotide substrate comprises an overhang on at least one end of the first double-stranded polynucleotide substrate, and the second double-stranded polynucleotide substrate comprises an overhang on at least one end of the second double-stranded polynucleotide substrate, wherein the overhangs on the first double-stranded polynucleotide substrate and the overhangs on the second double-stranded polynucleotide substrate are complementary and thus can hybridize to each other, forming one or more ligatable nicks. In some embodiments, the nicks are between the 5' end of the first polynucleotide strand and the 3' end of the second polynucleotide strand.
[0314] In some embodiments, the first double-stranded polynucleotide substrate and the second double-stranded polynucleotide substrate comprise DNA, for example, a first dsDNA substrate and a second dsDNA substrate.
[0315] In some embodiments, the DNA ligase substrate comprises an adaptor or linker. In some embodiments, the adaptor or linker has a blunt end or a cohesive (sticky) end. In some embodiments, the linker substrate has two blunt ends. In some embodiments, the adaptor has one blunt end and one cohesive or sticky end.
[0316] In some embodiments, the method includes a nucleotide substrate used by the DNA ligase to catalyze the joining reaction. In some embodiments, the nucleotide substrate is ATP. In some embodiments of the method, the reaction conditions for the ligation reaction include the addition of a divalent metal (e.g., Mg +2 ), buffer, reducing agent (e.g., DTT) and / or salt (KCl or NaCl).
[0317] In some embodiments, the reaction conditions also include ligation enhancing reagents, including, by way of example and not limitation, DMSO, betaine, polyethylene glycol (e.g., PEG 6000 and PEG 8000) or other molecular crowding reagents, bovine serum albumin (BSA), dextran, 1,2-propanediol, and Ficoll, among others.
[0318] In some embodiments, the ligation reaction is carried out at an appropriate temperature and reaction time. In some embodiments, the ligation reaction temperature is about 2°C to about 50°C. In some embodiments, the ligation reaction temperature is 2°C to 45°C, 4°C to 40°C, 4°C to 35°C, 4°C to 30°C, or 10°C to 25°C. In some embodiments, the ligation reaction temperature is 2°C, 3°C, 4°C, 5°C, 10°C, 15°C, 20°C, 25°C, 30°C, 37°C, 40°C, 45°C, or 50°C. In some embodiments, the ligation reaction temperature can be different, for example, a temperature at which stable hybrids form between the polynucleotide substrate(s), followed by a higher temperature to promote completion of the ligation reaction.
[0319] In some embodiments, the ligation reaction time can be sufficient for ligation of the polynucleotide substrate(s). In some embodiments, the ligation reaction time is 0.5 to 72 hours or more. In some embodiments, the ligation reaction time is 1 to 72 hours, 2 to 48 hours, or 2 to 24 hours. In some embodiments, the ligation reaction time is 0.5, 1, 2, 4, 5, 12, 24, 48, or 72 hours or more.
[0320] In some embodiments, the DNA ligases described herein are used in ligase chain reaction (LCR) or ligase detection reaction (LDR) (see, e.g., Gibriel et al., Mutat Res Rev Mutat Res. 2017, 773:66-90). In some embodiments, the DNA ligases described herein are used to generate templates for rolling circle amplification. In some embodiments, the DNA ligases are used for DNA sequencing or single nucleotide polymorphism (SNP) detection, for example, by supported oligonucleotide ligation and detection (SOLiD).
[0321] In a further aspect, the present disclosure provides kits comprising the DNA ligase described herein. In some embodiments, the kits further comprise one or more of a buffer, a nucleotide substrate (e.g., ATP), and / or one or more polynucleotide ligase substrates (particularly synthetic oligonucleotides or recombinant polynucleotides). In some embodiments, the kits further comprise an additive or ligation enhancer, including one or more of DMSO, betaine, polyethylene glycol (e.g., PEG 6000, PEG 8000, etc.), bovine serum albumin, Ficoll, and dextran (e.g., Dextran 6000). [Example]
[0322] The following examples (including the experiments and results achieved) are provided for illustrative purposes only and should not be construed as limiting the invention. Many of the reagents and equipment described below have a variety of suitable sources. It is not intended that the invention be limited to any particular source for any reagent or equipment item.
[0323] In the experimental disclosure that follows, the following abbreviations apply: M (mole); mM (millimolar), uM, and μM (micromolar); nM (nanomole); mol (mole); gm and g (gram); mg (milligram); ug and μg (microgram); L and 1 (liter); ml and mL (milliliter); ul, uL, μl, and μL (microliter); cm (centimeter); mm (millimeter); um and μιη (micrometer); sec. (second); min(s) (minute(s)); h(s) and hr(s) (hour(s)); U (unit); MW (molecular weight); rpm (revolutions per minute); psi and PSI (pounds per square inch); °C (degrees Celsius); RT and rt (room temperature); IPTG (isopropyl β-D-1-thiogalactopyranoside); LB (lysogeny broth); TB (terrific broth); SFP (shake flask powder); CDS (coding sequence); Escherichia coli (E. coli) W3110 (Coli Genetic Stock) Commonly used laboratory Escherichia coli (E. coli) strains available from the Center for Clinical Chemistry Science (CGSC), New Haven, Connecticut; HTP (high throughput); FIOPC and FIOP (fold improvement over positive control or fold improvement over parent, respectively).
[0324] Example 1 Isolation of DNA ligase gene and construction of expression vector Genes encoding N-terminal hexa-histidine tagged mutants of multiple wild-type (WT) dsDNA ligase enzymes were designed and synthesized using codon optimization for E. coli expression and subcloned into the E. coli expression vector pCK100900i (see, e.g., U.S. Patent No. 7,629,157 and U.S. Patent Application Publication No. 2016 / 0244787, both of which are incorporated herein by reference). W3110-derived E. coli strains were transformed with these plasmid constructs. A library of gene variants was generated from these plasmids using directed evolution techniques commonly known to those skilled in the art (see, e.g., U.S. Patent No. 8,383,346 and WO 2010 / 144103, both of which are incorporated herein by reference). Substitutions in the enzyme variants described herein are indicated with reference to an N-terminal 6-histidine tagged version of the referenced WT dsDNA ligase enzyme (i.e., SEQ ID NO: 2) or variants thereof, as indicated.
[0325] Example 2 DNA ligase expression and thermal lysis in high-throughput (HTP) High-throughput (HTP) amplification of dsDNA ligase enzymes and mutants Transformed E. coli cells were selected by plating on LB agar plates containing 1% glucose and 30 μg / ml chloramphenicol. After overnight incubation at 37°C, colonies were placed into wells of 96-well shallow, flat-bottom NUNC™ (Thermo-Scientific) plates filled with 180 μl / well LB medium supplemented with 1% glucose and 30 μg / ml chloramphenicol. Cultures were grown overnight for 18–20 hours in a shaker (200 rpm, 30°C, 85% relative humidity; Kuhner). An overnight growth sample (20 μL) was transferred to a Costar 96-well deep plate filled with 380 μL of Terrific Broth supplemented with 30 μg / ml chloramphenicol. The plates were incubated in a shaker (250 rpm, 30°C, 85% relative humidity; Kuhner) until OD .600 The cells were incubated for 120 min until the RI reached 0.4–0.8. The cells were then induced with 40 μL of 10 mM IPTG in sterile water and incubated overnight in a shaker (250 rpm, 30°C, 85% relative humidity; Kuhner). The cells were pelleted (4000 rpm x 20 min), the supernatant was discarded, and the cells were frozen at -80°C prior to analysis.
[0326] Thermal lysis of HTP cell pellets with lysozyme 200 μL of buffer containing 50 mM Tris-HCl buffer (pH 7.5) was added to the cell pellet in each well. The cells were completely resuspended for 15-20 minutes at room temperature while shaking on a benchtop shaker. Simultaneously, 60 μL of lysis buffer containing 50 mM Tris-HCl buffer (pH 7.5) and 0.2 g / L lysozyme was transferred to the BioRad plate. 60 μL of the completely resuspended cell pellet was then transferred to the BioRad plate. The plate was thoroughly mixed and heated at a specific temperature for 1 hour. The plate was then centrifuged at 4,000 rpm and 4°C for 15 minutes. The clear supernatant was then used in biocatalytic reactions to determine their activity levels.
[0327] Example 3 DNA ligase expression and purification in high-throughput (HTP) Transformed E. coli cells were selected by plating on LB agar plates containing 1% glucose and 30 μg / ml chloramphenicol. After overnight incubation at 37°C, colonies were placed into wells of 96-well shallow, flat-bottom NUNC™ (Thermo-Scientific) plates filled with 180 μl / well LB medium supplemented with 1% glucose and 30 μg / ml chloramphenicol. Cultures were grown overnight for 18–20 hours in a shaker (200 rpm, 30°C, 85% relative humidity; Kuhner). An overnight growth sample (20 μL) was transferred to a Costar 96-well deep plate filled with 380 μL of Terrific Broth supplemented with 30 μg / ml chloramphenicol. The plates were incubated in a shaker (250 rpm, 30°C, 85% relative humidity; Kuhner) until OD . 600 The cells were incubated for 120 min until the RI reached 0.4–0.8. The cells were then induced with 40 μL of 10 mM IPTG in sterile water and incubated overnight in a shaker (250 rpm, 30°C, 85% relative humidity; Kuhner). The cells were pelleted (4000 rpm x 20 min), the supernatant was discarded, and the cells were frozen at -80°C prior to analysis.
[0328] Lysis of HTP cell pellets for HTP purification of dsDNA ligase from crude lysates The cell pellet was resuspended in 400 μl / well of lysis mixture [50 mM Tris-HCl buffer, pH 7.5, 75% (v / v) B-Per reagent (Thermo Fisher Scientific), 300 mM NaCl, 10 mM imidazole, and 0.2% (v / v) Triton X-100]. The mixture was stirred at room temperature for 2 hours, pelleted (4000 rpm x 20 min), and the supernatant was saved for purification.
[0329] dsDNA ligase was purified from crude Escherichia coli (E. coli) extracts by metal affinity chromatography using HIS-Select® High Capacity (HC) nickel-coated plates (Sigma) according to the manufacturer's instructions. The HIS-Select plates were equilibrated with a total of 800 μl of wash buffer (50 mM Tris-HCl, pH 7.5, 500 mM NaCl, 20 mM imidazole, 0.02% v / v Triton X-100 reagent) per well. 200 μl of HTP lysate containing dsDNA ligase per well was then loaded onto the plate and centrifuged at 2115 relative centrifugal force (rcf) and 4°C for 1 minute. The plate was washed twice with 400 μl of wash buffer / well, centrifuged at 2115 rcf and 4°C for 2–3 minutes for each wash. The dsDNA ligase samples were eluted by adding 120 μl of elution buffer (50 mM Tris-HCl, pH 7.5, 500 mM NaCl, 250 mM imidazole, 0.02% v / v Triton X-100 reagent) per well by centrifugation at 2115 rcf and 4°C for 1 min.
[0330] The eluate was buffer exchanged using a Zeba™ Spin desalting plate (Thermo Fisher). Briefly, the plate was equilibrated twice with 375 μl of 2× dsDNA ligase storage buffer (80 mM Tris-HCl pH 7.5, 200 mM KCl, and 0.2 mM EDTA) per well and centrifuged at 2115 relative centrifugal force (rcf) and 4°C for 1 minute. 90 μl of HIS-Select sample eluate was loaded onto the desalting plate and centrifuged at 2115 rcf and 4°C for 2 minutes. The eluate from the desalting plate was retained and mixed with an equal volume of glycerol to a final storage buffer concentration of 40 mM Tris-HCl pH 7.5, 100 mM KCl, 0.1 mM EDTA, and 50% glycerol.
[0331] Example 4 Shake flask expression and purification of DNA ligase Shake flask expression Selected HTP cultures grown as described above were plated onto LB agar plates containing 1% glucose and 30 μg / ml chloramphenicol and grown overnight at 37°C. A single colony from each culture was transferred to 5 ml LB broth containing 1% glucose and 30 μg / ml chloramphenicol. Cultures were grown at 30°C and 250 rpm for 20 hours and then subcultured at a dilution of approximately 1:50 into 250 ml of Terrific Broth containing 30 μg / ml chloramphenicol to a final OD of approximately 0.05. 600 The culture was grown at 30°C, 250 rpm for approximately 195 minutes until it reached an OD of approximately 0.6. 600 The cells were incubated at 30°C for 60 min and then induced by the addition of IPTG to a final concentration of 1 mM. The induced culture was incubated at 30°C and 250 rpm for 20 h. After this incubation period, the culture was centrifuged at 4000 rpm for 10 min. The culture supernatant was discarded, and the pellet was resuspended in 35 ml of 20 mM Tris-HCl buffer (pH 7.5). The cell suspension was cooled in an ice bath and lysed using a Microfluidizer cell disrupter (Microfluidics M-110L). The crude lysate was pelleted by centrifugation (11,000 rpm, 4°C for 60 min), and the supernatant was then filtered through a 0.2 μm PES membrane to further clarify the lysate.
[0332] Purification of DNA ligase from shake flask lysates Additional NaCl and imidazole were added to the clarified dsDNA ligase lysate to adjust their concentrations to 500 mM NaCl and 20 mM imidazole, respectively. The lysate was then purified using an AKTA Pure purification system and a 5 ml HisTrap FF column (GE Healthcare) using the AC Step HiF setting (run parameters are listed below). The SF wash buffer contained 50 mM Tris-HCl, 500 mM NaCl, 20 mM imidazole, and 0.02% v / v Triton X-100 reagent. [Table 1]
[0333] The three most concentrated 1 ml fractions were identified by UV absorption (280 nm) and dialyzed overnight in dialysis buffer (40 mM Tris-HCl, pH 7.5, 100 mM KCl, 0.1 mM EDTA, and 50% glycerol) in a 3.5K Slide-A-Lyzer™ dialysis cassette (Thermo Fisher) for buffer exchange. The concentration of the purified and dialyzed dsDNA ligase samples was measured by absorbance at 280 nm.
[0334] Example 5 Preparation of dsDNA 50-mer inserts for HTP screening A 6-carboxyfluorescein-labeled double-stranded 50-mer DNA fragment (Integrated DNA Technologies) consisting of two single-stranded, HPLC-purified synthetic oligonucleotides was prepared by annealing these two oligonucleotides in 1x annealing buffer (10 mM Tris-HCl pH 7.5, 50 mM NaCl, 10 mM EDTA). The resulting double-stranded "50-mer fluorescently labeled insert" possesses a monobasic deoxyadenine 3' overhang and a 5' monophosphate terminus at both ends of the molecule and is internally labeled with a fluorescent dye (6-carboxyfluorescein attached to the 5-position of the thymine ring via a 6-carbon spacer arm).
[0335] Example 6 Preparation of dsDNA 20mer adapters for HTP screening A double-stranded "20-mer adapter" molecule (SEQ ID NO: 1217 and SEQ ID NO: 1218) containing two single-stranded HPLC-purified oligonucleotides (Integrated DNA Technologies) was prepared by annealing in 1× annealing buffer (10 mM Tris pH 7.5, 50 mM NaCl, 10 mM EDTA). The resulting 20-mer adapter duplex has phosphorothioate-protected 5′-deoxythymidine overhangs and 5′-phosphates at the ligation-compatible ends.
[0336] Example 7 Capillary electrophoresis (CE) analysis of oligonucleotides Sample preparation for reaction analysis using CE Reaction samples were analyzed by capillary electrophoresis using an ABI 3500xl Genetic Analyzer (Thermo Fisher). Reactions (25 μL) were quenched by adding 5 μL of 50 mM aqueous EDTA. 2 μL of this quenched solution was transferred to a new 96-well or 384-well MicroAmp Optical PCR plate containing 18 μL of Hi-Di™ formamide (Thermo Fisher) with appropriate size standards. The ABI 3500xl was configured with POP6 polymer, 50 cm capillaries, and an oven temperature of 45°C. Pre-run settings were 18 kV for 180 seconds. Injection was at 5 kV for 5 seconds, and run settings were 19.5 kV for 1500 seconds. FAM-labeled oligo substrates and products were identified by their size relative to a sizing ladder; the substrate oligo peak was approximately 49 bp, and the ligation product appeared in the approximately 85 bp region.
[0337] Example 8 E. coli shake flask expression, purification and activity verification Recombinant DNA ligase gene Twenty genes (SEQ ID NOs: 1 / 2-37 / 38) cloned for recombinant expression in E. coli as described in Example 1 were expressed and purified as described in Example 4.
[0338] Ligation reactions were performed in a 96-well, 200 μL BioRad PCR plate. The reaction contained 1 nM insert (SEQ ID NO:1185 / SEQ ID NO:1186), 200 nM adapter (SEQ ID NO:1217 / SEQ ID NO:1218), and 1× ligation buffer containing 66 mM Tris-HCl buffer at pH 7.5, 10 mM MgCl2, 1 mM DTT, and 1 mM ATP. The reactions were set up as follows: (i) all reaction components except purified dsDNA ligase were premixed in a single solution, and 80 μL of this solution was dispensed into each well of a 96-well plate. (ii) 20 μL of purified dsDNA ligase solution (final concentration: 1.28 μM) was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at 23°C for 30 minutes, then held at 4°C until the reaction was quenched.
[0339] The reaction was quenched by adding 20 μL of 90 mM EDTA, followed by 20 μL of 2.8 g / L aqueous proteinase K (final concentration: 0.2 g / L proteinase K). The quenched reaction was incubated at 50°C for 2 hours in a thermocycler and then purified using a ZR-96 DNA Clean and Concentrator-5 (catalog number D4024; the purification protocol was included in the kit). The purified mixture was analyzed on a Caliper (Perkin Elmer) Labchip GX capillary electrophoresis instrument using the DNA High Sensitivity Assay according to the manufacturer's instructions.
[0340] Conversion was calculated as the double ligation product peak area relative to the sum of the unligated insert, single ligation, and double ligation product peak areas. The results are shown in Table 8.1. The amino acid sequence identity (without His tag) between each DNA ligase tested is shown in Table 8.2. [Table 2] [Table 3]
[0341] Example 9 Improvement over SEQ ID NO: 2 in DNA ligase activity HTP screening of improved dsDNA ligase mutants SEQ ID NO:2 was selected as the parent DNA ligase enzyme. A library of engineered genes was generated from the parent gene using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). The polypeptides encoded by each gene were produced by HTP and prepared as described in Table 9.1.
[0342] Ligation reactions were performed in 96-well format 200 μL BioRad PCR plates. Reactions contained the inserts, adapters, and reaction buffers listed in Table 9.1. Reactions were set up as follows: (i) all reaction components except dsDNA ligase were premixed into a single solution, and 20 μL of this solution was dispensed into each well of a 96-well plate. (ii) 5 μL of dsDNA ligase solution was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at the indicated temperature and reaction time, then held at 4°C until the reaction was quenched. Reaction and quench details are specified in Table 9.1. Following the quenched reaction, capillary electrophoresis (CE) sample preparation was performed as described in Table 9.1. [Table 4]
[0343] Activity relative to SEQ ID NO:2 (active FIOPs) was calculated by dividing the fold improvement in conversion of the mutant by the conversion observed in the reaction with SEQ ID NO:2 (conversion can be set as the average of replicates or the best single sample as appropriate). Conversion was calculated as the double ligation product peak area relative to the sum of the unligated insert, single ligation, and double ligation product peak areas. The results are shown in Table 9.2. [Table 5]
[0344] Example 10 Improvement over SEQ ID NO: 2 in DNA ligase activity HTP screening of improved dsDNA ligase mutants SEQ ID NO:2 was selected as the parent DNA ligase enzyme. A library of engineered genes was generated from the parent gene using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). The polypeptides encoded by each gene were produced by HTP and prepared as described in Table 10.1.
[0345] Ligation reactions were performed in 96-well format 200 μL BioRad PCR plates. Reactions contained the inserts, adapters, and reaction buffers listed in Table 10.1. Reactions were set up as follows: (i) all reaction components except dsDNA ligase were premixed into a single solution, and 20 μL of this solution was dispensed into each well of a 96-well plate. (ii) 5 μL of dsDNA ligase solution was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at the indicated temperature and reaction time, then held at 4°C until the reaction was quenched. Reaction and quench details are specified in Table 10.1. Following the quenched reaction, capillary electrophoresis (CE) sample preparation was performed as described in Table 10.1. [Table 6]
[0346] Activity against SEQ ID NO:2 (active FIOP) was calculated as the fold improvement in intermediate product conversion of the mutant divided by the intermediate product conversion observed in the reaction with SEQ ID NO:2 (intermediate product conversion can be set as the average of the replicates or the highest single sample as appropriate). Intermediate product conversion was calculated as the peak area of the single ligated product relative to the sum of the peak areas of the unligated insert and the single ligated product. The results are shown in Table 10.2. [Table 7]
[0347] Example 11 Improvement over SEQ ID NO: 62 in DNA ligase activity HTP screening of improved dsDNA ligase mutants SEQ ID NO:62 was selected as the parent DNA ligase enzyme. A library of engineered genes was generated from the parent gene using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). The polypeptides encoded by each gene were produced by HTP and prepared as described in Table 11.1.
[0348] Ligation reactions were performed in 96-well format 200 μL BioRad PCR plates. Reactions contained the inserts, adapters, and reaction buffers listed in Table 11.1. Reactions were set up as follows: (i) all reaction components except dsDNA ligase were premixed into a single solution, and 20 μL of this solution was dispensed into each well of a 96-well plate. (ii) 5 μL of dsDNA ligase solution was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at the indicated temperature and reaction time, then held at 4°C until the reaction was quenched. Reaction and quench details are specified in Table 11.1. Following the quenched reaction, capillary electrophoresis (CE) sample preparation was performed as described in Table 11.1. [Table 8]
[0349] Activity relative to SEQ ID NO: 62 (active FIOPs) was calculated by dividing the fold improvement in conversion of the mutant by the conversion observed in the reaction with SEQ ID NO: 62 (conversion can be set as the average of replicates or the best single sample as appropriate). Conversion was calculated as the double ligation product peak area relative to the sum of the peak areas of the unligated insert, single ligation, and double ligation products. The results are shown in Table 11.2. [Table 9-1] [Table 9-2] [Table 9-3]
[0350] Example 12 Improvement of DNA ligase activity relative to SEQ ID NO: 62 HTP screening of improved dsDNA ligase mutants SEQ ID NO:62 was selected as the parent DNA ligase enzyme. A library of engineered genes was generated from the parent gene using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). The polypeptides encoded by each gene were produced by HTP and prepared as described in Table 12.1.
[0351] Ligation reactions were performed in 96-well format 200 μL BioRad PCR plates. Reactions contained the inserts, adapters, and reaction buffers listed in Table 12.1. Reactions were set up as follows: (i) all reaction components except dsDNA ligase were premixed into a single solution, and 20 μL of this solution was dispensed into each well of a 96-well plate. (ii) 5 μL of dsDNA ligase solution was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at the indicated temperature and reaction time, then held at 4°C until the reaction was quenched. Reaction and quench details are specified in Table 12.1. Following the quenched reaction, capillary electrophoresis (CE) sample preparation was performed as described in Table 12.1. [Table 10]
[0352] Activity relative to SEQ ID NO: 62 (active FIOP) was calculated by dividing the fold improvement in conversion of the mutant by the conversion observed in the reaction with SEQ ID NO: 62 (conversion can be set as the average of replicates or the best single sample as appropriate). Conversion was calculated as the double ligation product peak area relative to the sum of the peak areas of the unligated insert, single ligation, and double ligation products. The results are shown in Table 12.2. [Table 11]
[0353] Example 13 Improvements over SEQ ID NO: 138 in DNA ligase activity and thermostability HTP screening of improved dsDNA ligase mutants SEQ ID NO: 138 was selected as the parent DNA ligase enzyme. A library of engineered genes was generated from the parent gene using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). The polypeptides encoded by each gene were produced by HTP and prepared as described in Table 13.1.
[0354] Ligation reactions were performed in 96-well format 200 μL BioRad PCR plates. Reactions contained the inserts, adapters, and reaction buffers listed in Table 13.1. Reactions were set up as follows: (i) all reaction components except dsDNA ligase were premixed into a single solution, and 20 μL of this solution was dispensed into each well of a 96-well plate. (ii) 5 μL of dsDNA ligase solution was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at the indicated temperature and reaction time, then held at 4°C until the reaction was quenched. Reaction and quench details are specified in Table 13.1. Following the quenched reaction, capillary electrophoresis (CE) sample preparation was performed as described in Table 13.1. [Table 12]
[0355] Activity against SEQ ID NO: 138 (active FIOPs) was calculated by dividing the fold improvement in conversion of the mutant by the conversion observed in the reaction with SEQ ID NO: 138 (conversion can be set as the average of replicates or the best single sample as appropriate). Conversion was calculated as the double ligation product peak area relative to the sum of the peak areas of the unligated insert, single ligation, and double ligation products. The results are shown in Table 13.2. [Table 13-1] [Table 13-2] [Table 13-3]
[0356] Example 14 Improvement over SEQ ID NO: 318 in DNA ligase activity HTP screening of improved dsDNA ligase mutants SEQ ID NO: 318 was selected as the parent DNA ligase enzyme. A library of engineered genes was generated from the parent genes using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). The polypeptides encoded by each gene were produced by HTP and prepared as described in Table 14.1.
[0357] Ligation reactions were performed in 96-well format 200 μL BioRad PCR plates. Reactions contained the inserts, adapters, and reaction buffers listed in Table 14.1. Reactions were set up as follows: (i) all reaction components except dsDNA ligase were premixed into a single solution, and 20 μL of this solution was dispensed into each well of a 96-well plate. (ii) 5 μL of dsDNA ligase solution was then added to the well to initiate the reaction. The reaction plate was heat-sealed with a peelable aluminum seal and incubated in a thermocycler at the indicated temperature and reaction time, then held at 4°C until the reaction was quenched. Reaction and quench details are specified in Table 14.1. Following the quenched reaction, capillary electrophoresis (CE) sample preparation was performed as described in Table 14.1. [Table 14]
[0358] Activity against SEQ ID NO: 318 (active FIOPs) was calculated by dividing the fold improvement in conversion of the mutant by the conversion observed in the reaction with SEQ ID NO: 318 (conversion can be set as the average of replicates or the best single sample as appropriate). Conversion was calculated as the double ligation product peak area relative to the sum of the peak areas of the unligated insert, single ligation, and double ligation products. The results are shown in Table 14.2. [Table 15-1] [Table 15-2] [Table 15-3]
[0359] Example 15 Improvements relative to SEQ ID NO:...
Claims
1. 1. An engineered DNA ligase, or functional fragment thereof, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence corresponding to residues 12-437 of SEQ ID NO:2 and the even numbered SEQ ID NOs:40-1184, or to a reference sequence corresponding to SEQ ID NO:2 and the even numbered SEQ ID NOs:40-1184, wherein said amino acid sequence comprises one or more substitutions relative to said reference sequence corresponding to residues 12-437 of SEQ ID NO:2, 62, 138, 318, 722, or 938, or said reference sequence corresponding to SEQ ID NO:2, 62, 138, 318, 722, or 938.
2. 2. The engineered DNA ligase of Claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938.
3. 2. The engineered DNA ligase of claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or the reference sequence corresponding to SEQ ID NO:2, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:
2.
4. 2. The engineered DNA ligase of claim 1, or a functional fragment thereof, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 2, or to the reference sequence corresponding to SEQ ID NO:
2.
5. 2. The engineered DNA ligase of claim 1, or a functional fragment thereof, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or to the reference sequence corresponding to an even-numbered SEQ ID NO:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:
2.
6. 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 199, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 140, 141, 142, 143, 144, 145, 146, , 148, 149, 150, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278 8, 282, 283, 284, 286, 288, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, 384, 386, 387, 388, 389, 390, 392, 3 6. The engineered DNA ligase of any one of claims 1-5, wherein the engineered DNA ligase comprises at least a substitution at residues 96, 397, 404, 405, 408, 414, 415, 416, 417, 418, 419, 421, 422, 423, or 428, or a combination thereof, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
7. The amino acid sequence of the operated DNA ligase has at least substitutions or amino acid residues 11D, 12A / I, 13G / R, 14G / S / T / V, 18D / N / S, 30C / H / S, 31R, 33M / R / V, 34L / R, 36T / Y, 37G / L / N / S, 44S, 50G / I / S / T, 56P, 59E, 60Y, 61T / V, 63F / R, 67R, 68A / M / S / V / Y, 69T, 71G / L / P / R, 73C / K / P / T / V / W, 74S, 76F / G / H / L / N / R, 77D, 82R, 88V, 95A / L / R / V, 96A / G / T / V, 97G, 99G / I, 100V, 101R, 102G / K / L / S, 103V, 104K, 105K / S / T, 106L / S / V, 110R, 112M, 113A / T, 117G / S / V / Y, 125R / T, 128C, 130T, 132R, 138L / R, 139T, 148P, 149P, 150C / F / T, 155R, 156C, 159Q, 161R / V, 162W, 164A / R, 165K, 177G, 186A / C / E / H / L / M / R / T / V, 188A, 189C / T, 190R, 191T, 195R, 196E / V, 197R, 198A / D / K / L / N / R / V / W, 201L / S, 205E / G / K, 207L, 208D / F / H, 212F / G / M / S / W, 220V, 226D / E / Q / S / V, 228E / I / M / S, 230L / M, 231P, 232R, 233G / T / W, 235W, 237L / M / R / S / V / Y, 239M / N / P / Q / S / T / V / W, 240E / G / K / Q / R / S / Y, 242P / Q / T, 251L, 254G / S, 258L / S / V, 263G / L / Q / T, 264A / C, 266M / T, 267D / W / Y, 269L, 271A / G / N / S, 273A / G / S, 277Q / R, 278E, 282G / L / M / T / V / Y, 283A / G / K / L / M / R / S / V, 284D, 286F / L / S, 288I, 289A / L / S / V, 290L, 294L, 295K, 297W, 300G / T, 301F / L, 305K, 306I / K / S / V, 308K / L / S, 309G / R, 317Q, 323S, 328R, 334L / R, 337G / L / M / P / R / S, 339Y, 349E, 355S, 356A / V / W, 357H / K / P / R / S / V, 358C, 359N / R, 360H / M / P, 362G,363R, 364R, 367C / L, 370C / G, 372N / Q, 374A / S, 375W, 378T, 379A / G / P, 380T, 381K / R, 382V, 384C / V, 386F, 387G, 388K / Y, 389K / L / Q / R, 390E, 392C / I / K / L / R / S, 396C / H, 397K / L / M, 404S, 405I, 408C / V, 414A / L / Q / R / T / V, 415A / C / E / H / I / K 7. The engineered DNA ligase of any one of claims 1 to 6, wherein the engineered DNA ligase comprises the following amino acid residues: 416K, 417D / G / L, 418A / G / I / L / M / P / S / T, 419G, 421R, 422N, 423R / T, or 428F / R / S, or a combination thereof, and the amino acid positions are relative to the reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
8. 7. The engineered DNA ligase of any one of claims 1-6, wherein the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
9. 9. The engineered DNA ligase of any one of claims 1 to 6 and 8, wherein the amino acid sequence of the engineered DNA ligase comprises at least substitutions 63R, 242Q, 283L, 286S, 317Q, 414Q, 418S, or 428R, or a combination thereof, and wherein the amino acid positions are relative to the reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
10. 7. The engineered DNA ligase of any one of claims 1 to 6, wherein the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions at amino acid position(s) 233, 317, 191, 288, 207, 149, 251, 205, 269, 164, 36, 428, 105 / 132, or 105, wherein said amino acid positions are relative to the reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
11. 11. The engineered DNA ligase of any one of claims 1 to 6 and 10, wherein the amino acid sequence of the engineered DNA ligase comprises at least the substitution or set of substitutions 233T, 317Q, 191T, 288I, 207L, 149P, 251L, 205E, 269L, 164A, 36T, 428R, 105K / 132R or 105K, amino acid positions relative to the reference sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or relative to the reference sequence corresponding to SEQ ID NO:
2.
12. 2. The engineered DNA ligase of claim 1, wherein the amino acid sequence of the engineered DNA ligase comprises at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, and wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
13. 2. The engineered DNA ligase of claim 1, wherein the amino acid sequence of the engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to the reference sequence corresponding to SEQ ID NO:
2.
14. 10. The engineered DNA ligase of claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to said reference sequence corresponding to residues 12-437 of SEQ ID NO:2, or are relative to said reference sequence corresponding to SEQ ID NO:
2.
15. 2. The engineered DNA ligase of Claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938.
16. 2. The engineered DNA ligase of claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to said reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or to said reference sequence corresponding to an even-numbered SEQ ID NO:40-1184.
17. 2. The engineered DNA ligase of Claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938.
18. 2. The engineered DNA ligase of Claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO:40-1184, or to the reference sequence corresponding to an even-numbered SEQ ID NO:40-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, 138, 318, 722, or 938, or to the reference sequence corresponding to SEQ ID NO:62, 138, 318, 722, or 938.
19. 10. The amino acid sequence of the engineered DNA ligase is selected from the group consisting of amino acid positions 11, 12, 13, 14, 18, 30, 31, 33, 34, 36, 37, 44, 50, 56, 59, 60, 61, 63, 67, 68, 69, 71, 73, 74, 76, 77, 82, 88, 95, 96, 97, 99, 100, 101, 102, 103, 104, 105, 106, 110, 112, 113, 117, 125, 128, 130, 132, 138, 139, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 2 50, 155, 156, 159, 161, 162, 164, 165, 177, 186, 188, 189, 190, 191, 195, 196, 197, 198, 201, 205, 207, 208, 212, 220, 226, 228, 230, 231, 232, 233, 235, 237, 239, 240, 242, 251, 254, 258, 263, 264, 266, 267, 269, 271, 273, 277, 278, 282, 283, 284, 286, 2 88, 289, 290, 294, 295, 297, 300, 301, 305, 306, 308, 309, 317, 323, 328, 334, 337, 339, 349, 355, 356, 357, 358, 359, 360, 362, 364, 367, 370, 372, 374, 375, 378, 379, 380, 381, 382, 384, 386, 387, 388, 389, 390, 392, 396, 397, 404, 405, 408, 414, 415, 4 19. The engineered DNA ligase of Claim 17 or 18, wherein the engineered DNA ligase comprises at least a substitution at residues 16, 417, 418, 419, 421, 422, 423, or 428, or a combination thereof, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
20. The amino acid sequence of the operated DNA ligase has at least substitutions or amino acid residues 11D, 12A / I, 13G / R, 14G / S / T / V, 18D / N / S, 30C / H / S, 31R, 33M / R / V, 34L / R, 36T / Y, 37G / L / N / S, 44S, 50G / I / S / T, 56P, 59E, 60Y, 61T / V, 63F / R, 67R, 68A / M / S / V / Y, 69T, 71G / L / P / R, 73C / K / P / T / V / W, 74S, 76F / G / H / L / N / R, 77D, 82R, 88V, 95A / L / R / V, 96A / G / T / V, 97G, 99G / I, 100V, 101R, 102G / K / L / S, 103V, 104K, 105K / S / T, 106L / S / V, 110R, 112M, 113A / T, 117G / S / V / Y, 125R / T, 128C, 130T, 132R, 138L / R, 139T, 148P, 149P, 150C / F / T, 155R, 156C, 159Q, 161R / V, 162W, 164A / R, 165K, 177G, 186A / C / E / H / L / M / R / T / V, 188A, 189C / T, 190R, 191T, 195R, 196E / V, 197R, 198A / D / K / L / N / R / V / W, 201L / S, 205E / G / K, 207L, 208D / F / H, 212F / G / M / S / W, 220V, 2362G, 364R, 367C / L, 370C / G, 372N / Q, 374A / S, 375W, 378T, 379A / G / P, 380T, 381K / R, 382V, 384C / V, 386F, 387G, 388K / Y, 389K / L / Q / R, 390E, 392C / I / K / L / R / S, 396C / H, 397K / L / M, 404S, 405I, 408C / V, 414A / L / Q / R / S / T / V, 415A / C / E / H / I / K / L / V, 416K, 417D / G / L, 418 20. The engineered DNA ligase of any one of claims 17-19, comprising the amino acid residues A / G / I / K / L / M / P / S / T, 419G, 421R, 422N, 423R / T, or 428E / F / R / S, or a combination thereof, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or are relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
21. 20. The engineered DNA ligase of any one of claims 17-19, wherein the amino acid sequence of the engineered DNA ligase comprises at least a substitution at amino acid position 63, 242, 283, 286, 317, 414, 418, or 428, or a combination thereof, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722, or 938, or relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
22. 20. The engineered DNA ligase of any one of Claims 17-19, wherein the amino acid sequence of the engineered DNA ligase comprises at least substitutions or amino acid residues 63F / R, 242P / Q / T / V, 283A / F / G / K / L / M / R / S / V, 286F / L / S / W, 317T / Q, 414A / L / Q / R / S / T / V, 418A / G / I / K / L / M / P / S / T, or 428E / F / R / S, or a combination thereof, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 62, 138, 318, 722 or 938, or are relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722 or 938.
23. 18. The engineered DNA ligase of Claim 17, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to the reference sequence corresponding to SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62 or to the reference sequence corresponding to SEQ ID NO:
62.
24. 19. The engineered DNA ligase of Claim 18, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or to the reference sequence corresponding to an even-numbered SEQ ID NO:62, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:
62.
25. 317 / 349 / 362 / 386, 105 / 317, 139 / 317 / 362, 233 / 317 / 405, 139 / 317, 162 ...
25. The engineered DNA ligase of Claim 23 or 24, wherein the engineered DNA ligase comprises at least a substitution or set of substitutions at residues 12-437 of SEQ ID NO:62, or at least a substitution or set of substitutions at residues 12-437 of SEQ ID NO:62, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:
62.
26. The amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions: 196E, 242P, 337L, 33R, 33V, 277Q, 337P, 30C, 359N, 337G, 283S, 242Q, 415A, 387G, 415H, 379A, 205G, 186M, 337M, 389Q, 102L, 389L, 359R, 205K, 164A, 415E, 301F, 30H, 379G, 375W, 267W, 380T, 415C, 254G, 1 86T, 317Q, 77D / 139T / 317Q / 417G, 105K / 317Q / 417G, 317Q / 349E / 362G / 386F, 105K / 317Q, 139T / 317Q / 362G, 233T / 317Q / 405I, 139T / 317Q, 162W, 283V, 286S, 414L, 417L, 226S, 61T, 226Q, 105S, 186E, 230M, 61V, 186V, 186L, 379P, 415V, 415I, 186A, 283R, 2 26V, 418L, 418S, 370G, 283A, 337S, 297W, 237S, 428S, 186C, 362G, 283G, 196V, 237L, 415L, 233W, 242T, 277R, 235W, 148P, 254S , 102G, 105T, 301L, 286F, 100V, 237Y, 30S, 428F, 414Q, 414V, 283L, 414A, 414R, 370C, 283K, 286L, 186H, 417D, 226D, 233G, 267Y , 97G, 226E, 418T, 230L, 414T, 418A, 237V, 418G, 382V, 267D, 283M, 418P, 237R, 237M, 418I, 358C, 186R, 418M, or 102S, and wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:62, or are relative to the reference sequence corresponding to SEQ ID NO:
62.
27. 18. The engineered DNA ligase of Claim 17, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or the reference sequence corresponding to SEQ ID NO: 138, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138 or to the reference sequence corresponding to SEQ ID NO:
138.
28. 19. The engineered DNA ligase of Claim 18, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:314-458, or to the reference sequence corresponding to an even numbered SEQ ID NO:314-458, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:138 or to the reference sequence corresponding to SEQ ID NO:
138.
29. 286 / 286 / 359 / 418, 283 / 286, 283 / 286 / 418, 186 / 242 / 283 / 286 / 418, 205 / 286 / 359, 283, 242 / 286 / 418, 186 / 205 / 242, 283 / 286 / 359, 277 / 286 / 359 / 418, 186 / 205 / 283 / 286 / 359, 242 / 277 / 418, 242 / 283 / 286 / 418, 112 / 196 / 389, 286, 283 / 286 / 359 / 418, 242 / 359 / 418, 186 / 283, 186 / 359 / 418, 205 / 242 / 418, 186 / 205 / 242 / 283 / 286 / 359 / 418, 205 / 359 / 418, 186 / 205 / 283 / 286 / 418, 186 / 283 / 359, 186 / 242, 186 / 242 / 359, 283 / 359 / 418, 277 / 418, 186 / 188 / 283, 186 / 286 / 418, 186 / 242 / 286 / 359 / 418, 418, 186 / 242 / 283 / 286 / 359 / 418, 186 / 277 / 359 / 418, 242 / 283 / 286, 205 / 418, 30 / 297, 205 / 242 / 286 / 359 / 418, 186 / 205 / 359 / 418, 359 / 418, 186, 230, 33 / 297, 186 / 205, 186 / 283 / 359 / 418, 186 / 418, 205 / 242 / 283 / 359 / 418, 33 / 375 / 389, 33 / 230, 196 / 242 / 283 / 286 / 359 / 418, 186 / 242 / 283 / 359 / 418, 186 / 359, 33 / 196, 186 / 2 29. The engineered DNA ligase of Claim 27 or 28, comprising at least a substitution or set of substitutions in: 77 / 418, 242, 33 / 196 / 297 / 301, 205 / 237 / 242 / 283 / 286 / 359, 186 / 205 / 283 / 359 / 418, 33 / 389, or 186 / 196 / 242 / 283 / 286 / 359 / 418, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or are relative to the reference sequence corresponding to SEQ ID NO:
138.
30. [0023] The amino acid sequence of the engineered DNA ligase may contain at least one substitution or set of substitutions 242Q / 283V / 286S / 359R / 418L, 283L / 286S, 283L / 286S / 418S, 186E / 242Q / 283V / 286S / 418L, 205K / 286S / 359R, 283V, 242P / 286S / 418T, 186E / 205K / 242T, 283L / 286S / 359R, 277Q / 286S / 359N / 418T, 186A / 205K / 283V / 286S / 359R, 242Q / 277Q / 418L, 242Q / 283S / 286S / 418T, 112M / 196V / 389Q, 286S, 242P / 283R / 286S / 359R / 418L, 283L / 286S / 359R / 418S, 242Q / 359N / 418S, 186E / 283L, 186E / 359R / 418S, 2 05K / 242T / 418S, 186E / 205K / 242Q / 283S / 286S / 359R / 418S, 205K / 359N / 41 8L, 186A / 205K / 283L / 286S / 418L, 242T / 283V / 286S / 359R / 418S, 186E / 283 L / 359R, 186E / 242P, 283V / 286S / 418S, 186E / 242T / 359R, 283L / 359R / 418L , 277Q / 418S, 186E / 188A / 283V, 186E / 286S / 418S, 186A / 242T / 286S / 359N / 418L, 418L, 186A / 242P / 283V / 286S / 359R / 418S, 186A / 277Q / 359R / 418S, 2 42Q / 283R / 286S, 205K / 418S, 30C / 297W, 283L / 359R / 418S, 186E / 286S / 418 L, 418T, 205K / 242T / 286S / 359R / 418S, 186E / 205K / 359N / 418T, 418S, 186E / 242T / 283S / 286S / 418S, 359R / 418S, 242T / 286S / 418S, 186E, 230L, 33V / 2 97W, 186E / 205K, 186A / 283L / 359R / 418S, 186E / 418L, 186E / 277Q / 359R / 41 8S, 205K / 242T / 283R / 359R / 418L, 186E / 418T, 33V / 375W / 389Q, 33R / 230L,196E / 242T / 283V / 286S / 359R / 418T, 186E / 242T / 283L / 359R / 418S, 186E / 359R, 33V / 196V, 33R / 375W / 389Q, 18 6A / 277Q / 418S, 186A / 242T / 283L / 286S / 418L, 242T, 33V / 196V / 297W / 301F, 205K / 237Y / 242P / 283L / 286S / 359 29. The engineered DNA ligase of claim 27 or 28, comprising the following amino acid positions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138: 186A / 205K / 283L / 359R / 418T, 33V / 389Q, or 186A / 196E / 242T / 283L / 286S / 359N / 418T, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO: 138, or are relative to the reference sequence corresponding to SEQ ID NO:
138.
31. 18. The engineered DNA ligase of Claim 17, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or the reference sequence corresponding to SEQ ID NO:318, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318 or to the reference sequence corresponding to SEQ ID NO:
318.
32. 19. The engineered DNA ligase of Claim 18, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:460-936, or to the reference sequence corresponding to an even numbered SEQ ID NO:460-936, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:
318.
33. 271, 360, 266, 208, 74, 263, 264, 13, 378, 372, 300, 294, 290, 397, 95, 258, 161, 212, 198, 138, 18, 404, 273, 117 , 240, 69, 278, 289, 82, 328, 61 / 186 / 417, 186 / 370 / 417, 267, 61 / 370, 186 / 267 / 370 / 417, 61 / 186, 370 / 417, 417, 61 / 370 / 382, 61 / 186 / 267 / 370 / 417, 61, 267 / 370 / 417, 267 / 370, 61 / 417, 61 / 186 / 237 / 267 / 370, 61 / 237 / 370 / 417, 370, 186 / 370, 61 / 186 / 267 / 417, 61 / 186 / 370 / 382, 370 / 382 / 417, 61 / 186 / 370, 237 / 267 / 370 / 417, 61 / 186 / 382, 61 / 267, 61 / 267 / 417, 61 / 237 / 267 / 382, 186 / 370 / 382, 237 / 267 / 370, 61 / 186 / 267 / 370, 186 / 237 / 267 / 370, 61 / 186 / 267, 61 / 186 / 237, 186, 186 / 267, 237 / 370 / 417, 242 / 414, 162 / 414, 267 / 414, 105 / 414, 162 / 242 / 414, 33. The engineered DNA ligase of Claim 31 or 32, wherein the engineered DNA ligase comprises at least a substitution or set of substitutions at 97 / 162 / 414, 105 / 162 / 267 / 414, 356, 392, 106, 308, 306, 96, 282, 113, 309, 110, 37, 201, 284, 374, 295, 88, 44, 12, or 390 amino acid positions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or relative to the reference sequence corresponding to SEQ ID NO:
318.
34. 3. The amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions: 363R, 63R, 389K, 381R, 389R, 197R, 359R, 381K, 102K, 165K, 388K, 414R, 337R, 164R, 416K, 101R, 415K, 423R, 364R, 73P, 50I, 71L, 388Y / 419G, 50G, 357K, 396H, 68Y, 76F, 71P, 14V, 271G, 68A, 68M, 360P, 266M, 208D, 74S, 263T, 264C, 13R, 263L, 360M, 378T, 7 6H, 13G, 76R, 50T, 372Q, 50S, 300G, 73T, 294L, 372N, 76N, 290L, 397L, 76G, 9 5L, 258L, 161V, 212M, 198V, 138L, 18D, 18S, 404S, 273G, 117G, 240E, 68S, 73 C, 69T, 278E, 360H, 198D, 300T, 161R, 73K, 289V, 82R, 328R, 61T / 186A / 417D , 186A / 370C / 417L, 267Y, 61T / 370C, 186A / 267Y / 370C / 417L, 61T / 186C, 370 C / 417D, 417D, 61T / 370C / 382V, 61T / 186A / 267Y / 370C / 417D, 61T, 267Y / 370 C / 417D, 267Y / 370C, 61T / 417L, 61T / 186H / 237R / 267Y / 370C, 61T / 237R / 370 C / 417L, 370C / 417L, 370C, 186E / 370C, 61T / 186E / 267Y / 417D, 61T / 186C / 37 0C / 382V, 370C / 382V / 417D, 61T / 186C / 267Y / 417D, 61T / 186C / 370C, 237R / 2 67Y / 370C / 417D, 61T / 186E / 417D, 61T / 186H, 61T / 186C / 382V, 61T / 186H / 37 0C, 61T / 267Y, 61T / 267Y / 417L, 61T / 237R / 267Y / 382V, 186H / 370C / 382V, 23 7R / 267Y / 370C, 186H / 370C / 417D, 61T / 186V / 267Y / 370C, 267Y / 370C / 417L, 61T / 186V / 370C, 186H / 237R / 267Y / 370C, 61T / 186E / 267Y, 61T / 186C / 237R,61T / 186H / 237R, 61T / 186C / 417L, 186H, 186A / 267Y, 186C / 370C, 61T / 186C / 237R / 267Y / 370C, 186E / 267Y, 237R / 370C / 417L , 186H / 370C, 242Q / 414Q, 414T, 414Q, 162W / 414Q, 162W / 414V, 267D / 414T, 414V, 267D / 414Q, 414A, 162W / 414L, 105S / 414T, 162W / 242Q / 414T, 97G / 162W / 414Q, 105S / 162W / 267D / 414V, 356V, 273A, 357P, 14G, 14S, 396C, 240R, 392K, 273S, 106V, 308S , 308L, 306I, 96A, 397K, 263Q, 282T, 138R, 258S, 76L, 14T, 289S, 240G, 106S, 117S, 357S, 96G, 113T, 71G, 282V, 117V, 282M, 2 58V, 309G, 306K, 308K, 357R, 212S, 18N, 113A, 306S, 212F, 198N, 110R, 240K, 198L, 71R, 357V, 198R, 271A, 282Y, 212G, 95V, 37S, 68V, 240S, 117Y, 201S, 271N, 240Q, 289L, 212W, 198W, 357H, 208F, 282L, 306V, 284D, 96T, 266T, 374S, 95R, 309R, 73W, 3 33. The engineered DNA ligase of claim 31 or 32, comprising 97M, 289A, 295K, 106L, 95A, 374A, 88V, 96V, 271S, 44S, 208H, 263G, 12A, 201L, 198A, 198K, 282G, 390E, 264A, or 73V, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:318, or are relative to the reference sequence corresponding to SEQ ID NO:
318.
35. 18. The engineered DNA ligase of Claim 17, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722, or the reference sequence corresponding to SEQ ID NO:722, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722 or the reference sequence corresponding to SEQ ID NO:
722.
36. 19. The engineered DNA ligase of Claim 18, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938-1098, or to the reference sequence corresponding to an even numbered SEQ ID NO:938-1098, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722 or to the reference sequence corresponding to SEQ ID NO:
722.
37. The amino acid sequence of the engineered DNA ligase may be selected from the group consisting of: 63, 63 / 96 / 370, 389, 13 / 267 / 363 / 389, 13 / 186 / 389, 50 / 267 / 363 / 370 / 389, 363 / 370, 96 / 370, 61 / 63 / 212, 11 / 305, 11, 242 / 283 / 286 / 317 / 414 / 418, 323, 334, 339, 356, 384, 408, 67, 392, 104, 355, 159, 155, 367, 31, 231, 36 , 150, 239, 103, 125, 228, 37, 189, 177, 422, 128, 220, 130, 56, 190, 156, 232, 423, 34, 99, 59, 60, 421, or 195, wherein said amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722, or are relative to the reference sequence corresponding to SEQ ID NO:
722.
38. The amino acid sequence of the engineered DNA ligase comprises at least one substitution or set of substitutions 63R, 63R / 96T / 370C, 389K, 13R / 267Y / 363R / 389K, 13R / 186H / 389K, 50T / 267Y / 363R / 370C / 389K, 363R / 370C, 96T / 370C, 61T / 63R / 212W, 11D / 305K, 11D, 242V / 283F / 286W / 317T / 414S / 418K, 323S, 334L, 339Y, 356V, 384V, 408C, 67R, 392R, 104K, 3 55S, 159Q, 155R, 367L, 31R, 231P, 36Y, 150T, 239V, 103V, 356A, 63F, 125T, 367C, 228I, 37S, 37G, 239M, 239W, 239Q, 189C, 239S, 239T, 177G, 239P, 37N, 422N, 228S, 125R, 128C, 189T, 356W, 220V, 408V, 392S, 130T, 228M, 56P, 228E, 190R, 392I, 156C, 232R, 150F, 150C, 239N, 392L, 423T, 34L, 9 37. The engineered DNA ligase of Claim 35 or 36, comprising 9I, 384C, 59E, 334R, 60Y, 99G, 34R, 37L, 392C, 421R, or 195R, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:722, or are relative to the reference sequence corresponding to SEQ ID NO:
722.
39. 18. The engineered DNA ligase of Claim 17, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938, or the reference sequence corresponding to SEQ ID NO:938, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938 or to the reference sequence corresponding to SEQ ID NO:
938.
40. 19. The engineered DNA ligase of Claim 18, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the reference sequence corresponding to residues 12-437 of an even-numbered SEQ ID NO: 1100-1184, or to the reference sequence corresponding to an even-numbered SEQ ID NO: 1100-1184, wherein the amino acid sequence comprises one or more substitutions relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:
938.
41. The amino acid sequence of the engineered DNA ligase may be at amino acid position(s) 308 / 357 / 390, 74 / 76 / 201 / 308 / 357, 61 / 74 / 76 / 186 / 201 / 308 / 309 / 357 / 390, 14 / 201 / 240 / 289 / 357, 308 / 415, 76 / 357 / 396, 263 / 308 / 396, 61 / 76 / 96 / 240 / 308 / 309, 14 / 306 / 415, 14 / 73 / 106 / 415, 12 / 14 / 258 / 263 / 289 / 308 / 309 / 396, 74 / 76 / 117 / 309 / 357, 14 / 258 / 263 / 357 / 396, 14 / 96 / 106 / 306, 14 / 106, 12 / 14 / 308 / 309, 14 / 357 / 390, 14 / 117 / 258 / 309 / 357, 390, 240 / 273 / 357 / 39 0, 61 / 76 / 186 / 201 / 308 / 309, 14 / 396, 309, 14, 106 / 306 / 308, 14 / 240 / 306 / 308, 12 / 14 / 186 / 357, 309 / 390, 14 / 306, 14 / 76 / 308, 117 / 208 / 258 / 263 / 289 / 308 / 309, 14 / 73 / 106, 76 / 208 / 263, 357, 14 / 308, 263, 76, 1 41. The engineered DNA ligase of Claim 39 or 40, comprising at least a substitution or set of substitutions at 4 / 300 / 308 / 415, 240, 33 / 357 / 390, 14 / 76 / 273, or 74, wherein said amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938, or are relative to the reference sequence corresponding to SEQ ID NO:
938.
42. The amino acid sequence of the engineered DNA ligase may comprise at least one substitution or set of substitutions: 308L / 357S / 390E, 74S / 76L / 201S / 308L / 357P, 61T / 74S / 76L / 186A / 201S / 308L / 309R / 357P / 390E, 14G / 201S / 240R / 289S / 357P, 308S / 415I, 76L / 357S / 396C, 263L / 308L / 396C, 61T / 76L / 96A / 240R / 308L / 309R, 14S / 306I / 415I, 14S / 73W / 106S / 415I, 12A / 14G / 258S / 263Q / 289S / 308L / 309R / 396C, 74S / 76L / 117V / 309R / 357S, 14T / 258S / 263Q / 357S / 396C, 14S / 96G / 106S / 306I, 14S / 106S, 12A / 14G / 308L / 309R, 14G / 357S / 390E, 14G / 117V / 258S / 309R / 357S , 390E, 240R / 273S / 357P / 390E, 61T / 76L / 186V / 201S / 308L / 309R, 14G / 396C, 309R, 14S, 106V / 306I / 308S, 14S / 240S / 306I / 3 08S, 12I / 14G / 186V / 357S, 309R / 390E, 14S / 306I, 14T, 14S / 76R / 308S, 117V / 208D / 258S / 263Q / 289S / 308L / 309R, 14S / 73W / 10 41. The engineered DNA ligase of Claim 39 or 40, comprising 6S, 76L / 208D / 263Q, 357S, 14S / 308S, 263L, 76L, 14S / 300G / 308S / 415I, 240Y, 33M / 357S / 390E, 14G / 76L / 273S, or 74S, wherein the amino acid positions are relative to the reference sequence corresponding to residues 12-437 of SEQ ID NO:938, or are relative to the reference sequence corresponding to SEQ ID NO:
938.
43. 2. The engineered DNA ligase of Claim 1, wherein the amino acid sequence of the engineered DNA ligase comprises at least one substitution as set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, and wherein the amino acid position is relative to the reference sequence corresponding to SEQ ID NO: 62, 138, 318, 722, or 938.
44. 2. The engineered DNA ligase of Claim 1, wherein the amino acid sequence of said engineered DNA ligase comprises at least a substitution or set of substitutions of an engineered DNA ligase variant set forth in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions are relative to the reference sequence corresponding to SEQ ID NOs: 62, 138, 318, 722, or 938.
45. 2. The engineered DNA ligase of Claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference sequence comprising at least a substitution or set of substitutions as provided in Tables 9.2, 10.2, 11.2, 12.2, 13.2, 14.2, 15.2, 16.2, 17.2, and 18.2, wherein the amino acid positions relative to the reference sequence correspond to SEQ ID NOs: 62, 138, 318, 722, or 938.
46. 2. The engineered DNA ligase of Claim 1, wherein the amino acid sequence of the DNA ligase comprises residues 12-437 of an even-numbered SEQ ID NO:40-1184, or an even-numbered SEQ ID NO:40-1184, optionally, the amino acid sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions.
47. 2. The engineered DNA ligase of Claim 1, wherein the amino acid sequence of the DNA ligase comprises residues 12 to 437 of SEQ ID NO: 62, 138, 318, 722, 938 or 1108, or comprises SEQ ID NO: 62, 138, 318, 722, 938 or 1108, and optionally, the amino acid sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 substitutions in the amino acid sequence.
48. 48. The engineered DNA ligase of any one of claims 1 to 47, wherein the engineered DNA ligase has DNA ligase activity and improved properties compared to a reference DNA ligase having a sequence corresponding to residues 12 to 437 of SEQ ID NO: 2, 62, 138, 318, 722 or 938, or a sequence corresponding to SEQ ID NO: 2, 62, 138, 318, 722 or 938.
49. 49. The engineered DNA ligase of any one of claims 1-48, characterized by an improved property selected from: i) increased activity, ii) increased stability, iii) increased thermostability, iv) increased product yield, v) increased solubility, vi) reduced sequence bias, and vii) insensitivity or decreased sensitivity to input DNA substrate concentration, or any combination of i), ii), iii), iv), v), and vi), compared to a reference DNA ligase having a sequence corresponding to residues 12-437 of SEQ ID NO:2, 62, 138, 318, 722, or 938, or a sequence corresponding to SEQ ID NO:2, 62, 138, 318, 722, or 938.
50. 50. The engineered DNA ligase of claim 48 or 49, wherein the reference DNA ligase has a sequence corresponding to residues 12 to 437 of SEQ ID NO:2, or a sequence corresponding to SEQ ID NO:
2.
51. (a) a sequence corresponding to residues 12-437 of SEQ ID NO:2; residues 12-613 of SEQ ID NO:4; residues 12-614 of SEQ ID NO:6; residues 12-610 of SEQ ID NO:8; residues 12-606 of SEQ ID NO:10; residues 12-615 of SEQ ID NO:12; residues 12-594 of SEQ ID NO:14; residues 12-620 of SEQ ID NO:16; residues 12-608 of SEQ ID NO:18; residues 12-611 of SEQ ID NO:20; residues 12-614 of SEQ ID NO:22, residues 12-611 of SEQ ID NO:24, residues 12-609 of SEQ ID NO:26, residues 12-422 of SEQ ID NO:28, residues 12-518 of SEQ ID NO:30, residues 12-438 of SEQ ID NO:32, residues 12-381 of SEQ ID NO:34, residues 12-424 of SEQ ID NO:36, or residues 12-390 of SEQ ID NO:38; or (b) a sequence corresponding to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to
52. the amino acid sequence of the engineered DNA ligase is (a) residues 12-437 of SEQ ID NO:2; residues 12-613 of SEQ ID NO:4; residues 12-614 of SEQ ID NO:6; residues 12-610 of SEQ ID NO:8; residues 12-606 of SEQ ID NO:10; residues 12-615 of SEQ ID NO:12; residues 12-594 of SEQ ID NO:14; residues 12-620 of SEQ ID NO:16; residues 12-608 of SEQ ID NO:18; residues 12-611 of SEQ ID NO:20; residues 12-614 of SEQ ID NO:22, residues 12-611 of SEQ ID NO:24, residues 12-609 of SEQ ID NO:26, residues 12-422 of SEQ ID NO:28, residues 12-518 of SEQ ID NO:30, residues 12-438 of SEQ ID NO:32, residues 12-381 of SEQ ID NO:34, residues 12-424 of SEQ ID NO:36, or residues 12-390 of SEQ ID NO:38, or (b) SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38 52. The engineered DNA ligase of claim 51 , comprising:
53. 53. The engineered DNA ligase of any one of claims 1 to 52, wherein the engineered DNA ligase is purified.
54. 53. A recombinant polynucleotide comprising a polynucleotide sequence encoding the engineered DNA ligase of claim 51 or 52.
55. A recombinant polynucleotide comprising a polynucleotide sequence encoding the engineered DNA ligase of any one of claims 1 to 50.
56. 56. The recombinant polynucleotide of Claim 55, comprising a reference polynucleotide sequence corresponding to nucleotide residues 34 to 1311 of SEQ ID NO: 1, 61, 137, 317, 721, or 937, or a polynucleotide sequence having at least 70%, 75%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to SEQ ID NO: 1, 61, 137, 317, 721, or 937, wherein said recombinant polynucleotide encodes an engineered DNA ligase.
57. 56. The recombinant polynucleotide of Claim 55, comprising a polynucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to a reference polynucleotide sequence corresponding to nucleotide residues 34 to 1311 of an odd-numbered SEQ ID NO:39-1183, or a reference polynucleotide sequence corresponding to an odd-numbered SEQ ID NO:39-1183, wherein said recombinant polynucleotide encodes an engineered DNA ligase.
58. 58. The recombinant polynucleotide of any one of claims 55 to 57, wherein said polynucleotide sequence is codon-optimized for expression of the encoded engineered DNA ligase.
59. 56. The recombinant polynucleotide of claim 55, comprising a polynucleotide sequence comprising nucleotide residues 34 to 1311 of an odd-numbered SEQ ID NO: 39-1183, or a polynucleotide sequence comprising an odd-numbered SEQ ID NO: 39-1183.
60. 56. The recombinant polynucleotide of claim 55, comprising a polynucleotide sequence comprising nucleotide residues 34 to 1311 of, or comprising, SEQ ID NO: 1, 61, 137, 317, 721, 937 or 1107.
61. An expression vector comprising at least one recombinant polynucleotide according to any one of claims 54 to 61.
62. 62. The expression vector of claim 61, wherein the polynucleotide is operably linked to a regulatory sequence.
63. 63. The expression vector of claim 62, wherein the control sequence comprises at least a promoter.
64. A host cell comprising the expression vector of any one of claims 61 to 63.
65. 65. The host cell of claim 64, comprising a prokaryotic or eukaryotic cell.
66. 66. The host cell of claim 65, comprising a bacterial cell, a fungal cell, an insect cell, or a mammalian cell.
67. 67. A method of producing an engineered DNA ligase polypeptide in a host cell, the method comprising culturing the host cell of any one of claims 64 to 66 under suitable culture conditions such that at least one engineered DNA ligase is produced.
68. 68. The method of claim 67, further comprising recovering the engineered DNA ligase polypeptide from the culture and / or host cell.
69. 69. The method of claim 67 or 68, further comprising purifying the engineered DNA ligase.
70. 54. A composition comprising the engineered DNA ligase of any one of claims 1 to 53.
71. 71. The composition of claim 70, further comprising one or more of a buffer, ATP, a reducing agent, and / or one or more DNA ligase substrates.
72. 72. The composition of claim 70 or 71, further comprising a ligation enhancer.
73. 100. A method of ligating at least a first and a second DNA strand, comprising contacting the first and second DNA strands with the engineered DNA ligase of any one of claims 1 to 53 in the presence of a nucleotide substrate under conditions suitable for ligating the first DNA strand to the second DNA strand, wherein the first DNA strand comprises a ligatable 5' end and the second DNA strand comprises a ligatable 3' end to the 5' end of the first DNA strand.
74. 74. The method of claim 73, further comprising a third DNA strand, wherein the first DNA strand and the second DNA strand hybridize adjacent to each other on the third DNA strand, positioning the 5' end of the first DNA strand adjacent to the 3' end of the second DNA strand.
75. 75. The method of claim 74, wherein the third DNA strand is contiguous with the first DNA strand or the second DNA strand.
76. 75. The method of claim 74, wherein the third DNA strand is contiguous with the first DNA strand and the second DNA strand to form a single contiguous DNA ligase substrate.
77. 74. The method of Claim 73, wherein the first DNA strand hybridizes to a third DNA strand to form a first dsDNA substrate, and the second DNA strand hybridizes to a fourth DNA strand to form a second dsDNA substrate.
78. 78. The method of claim 77, wherein the first dsDNA substrate comprises a blunt-ended 5' end of the first DNA strand and the second dsDNA substrate comprises a blunt-ended 3' end of the second DNA strand.
79. 78. The method of Claim 77, wherein the first dsDNA substrate comprises an overhang on at least one end of the first dsDNA substrate, the second dsDNA substrate comprises an overhang on at least one end of the second dsDNA substrate, and the overhang on the first dsDNA substrate and the overhang on the second dsDNA substrate are complementary and can hybridize to each other, forming one or more nicks that can be ligated.
80. 80. The method of any one of claims 73 to 79, wherein the 3' end of the second DNA strand is a 3'-OH and the 5' end of the first DNA strand is a 5'-phosphate.
81. 81. The method of any one of claims 73 to 80, wherein the nucleotide substrate is ATP.
82. A kit comprising at least the DNA ligase according to any one of claims 1 to 53.
83. 83. The kit of claim 82, further comprising one or more of a buffer, a nucleotide substrate, a DNA substrate for the ligase, and / or a ligation enhancer.