CRISPR transposon systems and components

Modified CRISPR-Tn system proteins enhance nucleic acid integration and binding, overcoming inefficiencies in existing CRISPR/Cas systems by improving activity and utility in prokaryotic and eukaryotic cells.

JP2026507559APending Publication Date: 2026-03-04THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-14
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems face limitations in their ability to efficiently integrate and modify nucleic acids, particularly in prokaryotic and eukaryotic cells, due to variations in protein activity and nucleic acid binding capabilities.

Method used

Modified polypeptides and nucleic acids encoding transposon-associated proteins such as TnsA, TnsB, TnsC, TniQ, and Cas proteins like Cas5, Cas6, Cas7, and Cas8 are developed, enhancing nucleic acid integration activity and binding ability, suitable for CRISPR-Tn systems.

Benefits of technology

The modified proteins demonstrate improved nucleic acid integration and binding efficiency in vivo, addressing the limitations of existing CRISPR/Cas systems by increasing their effectiveness in modifying target nucleic acids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507559000047
    Figure 2026507559000047
  • Figure 2026507559000048
    Figure 2026507559000048
  • Figure 2026507559000049
    Figure 2026507559000049
Patent Text Reader

Abstract

The present disclosure provides Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CRISPR-Tn or CAST) systems, components thereof, and methods for nucleic acid modification using the systems or components. More specifically, the present disclosure provides modified Cas proteins and transposon-associated proteins for nucleic acid modification.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CRISPR-Tn or CAST) system and its components, such as Cas proteins and transposon-associated proteins.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application Nos. 63 / 484,923, filed February 14, 2023, 63 / 518,665, filed August 10, 2023, 63 / 587,916, filed October 4, 2023, and 63 / 621,894, filed January 17, 2024, the contents of each of which are incorporated herein by reference in their entirety.

[0003] Sequence Listing Description The contents of the Electronic Sequence Listing entitled COLUM-41261-601.xml (Size: 27,398 bytes, Created: February 14, 2024) are incorporated herein by reference in their entirety.

[0004] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with government support under grants HG011650, EB031935, HG009490, EB027793, EB031172, GM118062, and AI142756 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]

[0005] In bacteria and archaea, the CRISPR / Cas system provides immunity by integrating fragments of invading phage, viral, and plasmid DNA into CRISPR loci and directing degradation of homologous sequences using the corresponding CRISPR RNA ("crRNA"). Transcription of the CRISPR locus produces a "pre-crRNA," which is processed to yield a crRNA containing a spacer repeat fragment that guides an effector nuclease complex to cleave dsDNA sequences complementary to the spacer. Several different types of CRISPR systems are known (e.g., Type I, Type II, or Type III), classified based on the type of Cas protein and the use of a protospacer adjacent motif (PAM) for selection of the protospacer in the invading DNA.

[0006] Although RNA-guided targeting typically results in endonucleolytic cleavage of the bound substrate, recent studies have revealed a series of atypical pathways in which CRISPR protein-RNA effector complexes are naturally repurposed for alternative functions. For example, some Type I (Cascade) and Type II (Cas9) systems utilize truncated guide RNAs to achieve potent transcriptional repression without cleavage, and other Type I (Cascade) and Type V (Cas12) systems reside within unusual bacterial Tn7-like transposons and lack a nuclease component entirely. Summary of the Invention

[0007] Provided herein are modified polypeptides and nucleic acids encoding them that are useful in Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CRISPR-Tn or CAST) systems and methods utilizing the same. Polypeptides include transposon-associated proteins such as TnsA, TnsB, TnsC, and TniQ, as well as Cas proteins such as Cas5, Cas6, Cas7, and Cas8. The modified proteins may exhibit improved activity or utility in modifying target nucleic acids. In some embodiments, the modified proteins have increased nucleic acid integration activity compared to proteins without the disclosed modifications. In some embodiments, the modified proteins have increased or modified nucleic acid binding ability compared to proteins without the disclosed modifications. In some embodiments, the modified proteins have increased nucleic acid integration activity or efficiency in vivo (e.g., in prokaryotic or eukaryotic cells, in a subject) compared to proteins without the disclosed modifications.

[0008] In some embodiments, the polypeptide comprises one or more amino acid sequences having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any of SEQ ID NOs: 1-14, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NOs: 1-14.

[0009] In some embodiments, the polypeptide has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:1, as well as one or more amino acids at positions 2, 3, 5, 28, 57, 77, 80, 107, 110, 116, 122, 142, 155, 161, 166, 173, 177, 185, 211, 216, 227, and 230 relative to SEQ ID NO:1. amino acid substitutions; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:2, and at positions 2, 5, 22, 24, 25, 29, 75, 141, 199, 215, 319, 347, 364, 370, 383, 439, 454, 458, 485, 509, 533, 538, 565, 581, 586, 595, 596, 597, and 6 00; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:3, and one or more amino acid substitutions at positions 9, 15, 16, 18, 21, 64, 81, 86, 87, 99, 109, 142, 147, 153, 168, 180, 216, 230, 285, and 304 relative to SEQ ID NO:3; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to number 4, and at positions 4, 5, 9, 10, 12, 21, 23, 25, 26, 31, 32, 34, 35, 37, 41, 45, 47, 48, 51, 52, 55, 60, 61, 65, 67, 69, 72, 75, 79, 80, 82, 87, 88, 90, 91, 93, 94, 96, 98,one or more amino acid substitutions at 99, 100, 103, 106, 108, 113, 116, 125, 126, 128, 129, 135, 139, 143, 146, 147, 149, 153, 154, 156, 158, 159, 160, 162, 164, 166, 167, 168, 169, 170, 177, 179, 180, 182, 183, 185, 187, 188, 190, 191, 192, 193, 195, 196, 200, 204, 207, and 208; %, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 49, 52, 55, 56, 58, 60, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 119, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 110, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 1 45, 146, 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169, 173, 174, 176, 181, 182, 186, 187, 190, 195, 197, 198, 205, 208, 209, 211, 215, 21 8, 223, 226, 227, 231, 232, 235, 239, 246, 248, 250, 259, 260, 261, 262, 263, 267, 269, 273, 274, 277, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298 , 302, 303, 307, 313, 316, 317, 320, 323, 325, 331, 332, 339, 345, 348, 349, 352, 353, 354, 356, 361, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373,375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 404, 407, 408, 410, 411, 412, 413, 414, 415, 416, 421, 422, 423, 424, 425, 426, 427, 4 28, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456, 459, 460, 463, 464, 470, 472, 473, 494, 495, 498, 501, 502, 504, 505, 506, 508, 50 9, 510, 512, 513, 514, 517, 520, 521, 522, 525, 526, 527, 530, 531, 532, 533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552 , 553, 554, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 567, 568, 569, 570, 571, 574, 575, 576, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, one or more amino acid substitutions at 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:6, and at least one amino acid substitution at positions 1, 2, 3, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 9, 11, 12, 14, 21, 22, 26, 27, 31, 35, 38, 43, 44, 46, 47, 54, 59, 60, 61, 64, 65, 67, 68, 71, 72, 74, 76, 79, 80, 81, 84, 89, 95, 102, 105, 109, 110, 111, 112, 113 , 114, 116, 118, 119, 120, 123, 129, 130, 131, 132, 134, 142, 145, 146, 147, 148, 150, 154, 155, 166, 169, 178, 180, 181, 183, 184, 187, 190, 194, 197, 201,204, 207, 209, 213, 219, 221, 225, 226, 227, 229, 232, 233, 234, 236, 238, 241, 246, 251, 252, 256, 257, 261, 263, 265, 267, 269, 271, 272, 274, 280, 281, 285, 286, 288, 291, 292, 296, 299, 301, 303, 304, 306, 307, 308, 310, 313, 314, 316, 317, 318, 319, 320, 321 one or more amino acid substitutions at 3, 324, 326, 328, 330, 331, 332, 340, 341, 343, 344, 355, 412, 418, 427, 514, 1198, 1201, 1206, 1212, 1260, and 1282; ... at least 98%, or at least 99%) identity to SEQ ID NO: 7, and one or more amino acid substitutions at positions 99, 133, 189, 265, 266, 336, and 343 relative to SEQ ID NO: 7; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 8; one or more amino acid substitutions at positions 119, 134, 155, 180, 183, 274, 319, 447, 454, 458, 461, 512, 538, and 580 relative to SEQ ID NO:8; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:9; and one or more amino acid substitutions at positions 28, 82, 144, 151, 162, 182, 273, 327, and 346 relative to SEQ ID NO:9; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:10, and one or more amino acid substitutions at positions 21 and 90 relative to SEQ ID NO:10; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least and identity at positions 2, 3, 7, 9, 11, 12, 14, 16, 20, 26, 29, 32, 34, 35, 40, 43, 45, 46, 54, 61, 64, 65, 70, 77, 101, 103, 105, 106, 108, 109, 111, 119, 120, 123, 126, 127, 130, 131, 148, 149, 151, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, one or more amino acid substitutions at 9, 164, 166, 185, 194, 196, 203, 211, 217, 218, 219, 236, 242, 257, 267, 279, 283, 286, 288, 291, 293, 296, 303, 306, 313, 314, 316, 326, 331, 336, 347, 352, 361, 374, 377, 395, 396, 398, and 408;At least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 12, and at least 70% identity to SEQ ID NO: 12, and at least 70% identity to SEQ ID NO: 12, at least 70% identity to SEQ ID NO: 12, and at least 70% identity to SEQ ID NO: 12, at least 70% identity to SEQ ID NO: 12, , 77, 81, 88, 92, 93, 94, 96, 102, 105, 106, 108, 110, 121, 126, 128, 134, 138, 142, 147, 150, 151, 153, 156, 157, 160, 162, 165, 170, 171, 173, 174, 179, 181, 183, 185, 186, 187, 188, 191, 198, 201, 206, 207, 226, 228, 233, 236, 241, 249, 250, 256, 267, 268, 270, 275, 276, 277, 279, 283, 286, 289, 303, 305, 306, 310, 312, 314, 315, 316, 323, 326, 329, 349, 353, 355, 356, 357, 358, 361, 370, 372, 373, 376, 378, 382, ​​388, 391, 397, 399, 403, 405, 419, 421, 423, 424, 425, 427, 428, 430, 431, 432, 433, 449, 457, 473, 477, 480, 485, 487, 489, 49 one or more amino acid substitutions at 4, 496, 497, 498, 500, 502, 509, 511, 515, 518, 519, 520, 540, 545, 550, 555, 557, 570, 571, 580, 583, 585, 590, 594, 603, 607, 608, 611, 617, 620, 624, 636, 639, 641, 642, 644, 646, 655, 658, 660, 663, 665, 668, 672, 673, 678, 682, 685, 688, and 695;At least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 13, and at least one of positions 5, 10, 11, 26, 30, 35, 40, 42, 45, 46, 47, 58, 61, 65, 71, 72, 75, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 1 8, 80, 82, 83, 94, 98, 113, 115, 116, 117, 121, 128, 133, 138, 146, 148, 161, 171, 175, 177, 182, 184, 191, 193, 201, 203, 211, 212, 219, 225, 226, 232, 233, 235, 236, 237, 238, 240, 250, 274, 282, 286, 292, 295, 304, 307, 309, 312, 313, 315, 31 or one or more amino acid substitutions at SEQ ID NO: 14, including at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least at least 99%) identity to SEQ ID NO: 14 and one or more amino acid substitutions at positions 2, 9, 13, 14, 15, 34, 38, 42, 46, 50, 59, 60, 73, 75, 77, 82, 83, 85, 86, 97, 110, 115, 120, 124, 130, 132, 134, 140, 143, 145, 156, 159, 162, 164, 177, 199, 232, and 270 relative to SEQ ID NO: 14;

[0010] In some embodiments, the polypeptide has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:1, as well as any of the following amino acids relative to SEQ ID NO:1: A2T, T3I, L5S, T28A, A57T, F77L, Y80D, K107M, K107R, Y110C, Y110D, D116G, E122A, D142E, M155I, K161R, N166D, K1 one or more amino acid substitutions selected from the group consisting of 73E, Y177N, Y177D, C185R, D211Y, K216E, A227P, G230D, and G230S; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:2, and one or more amino acid substitutions selected from the group consisting of A2T, A2S, G5R, S22P, E24D, L25I, A29S, P75T, I141T, , V199I, S215R, D319V, Y347F, S364N, E370K, N383D, V439A, E454D, E454G, S458N, V485F, R509G, D533A, A538V, H565Y, A581T, H586L, N595K, D596N, D597N, D597Y, and I600V; one or more amino acid substitutions of at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 98%) relative to SEQ ID NO: 3; 7%, at least 98%, or at least 99%) identity with respect to SEQ ID NO: 4, and one or more amino acid substitutions of I9V, A15V, F16Y, S18F, S21N, N64D, H81Y, D86Y, N87K, V99I, E109D, E142K, V147I, N153D, I168M, A180E, A216S, L230F, K285E, and R304R relative to SEQ ID NO: 3; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%,at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 4, as well as R4K, N5K, P9S, A10P, N12D, T21I, V23M, S25N, S25R, V26M, V26G, S31N, S32I, E34A, F35L, A37D, H41L, D45N, I47V, E48G, G51V, S52I, E55K, E55D, E60K, F61L, S65T, S65A, P67T, P67L, P67S, P67H, T69A, A72V, A72D, S75I, S75R, S75T, K79 E, T80P, K82E, K87R, P88L, P88T, P88A, S90F, K91N, K91E, A93T, A93S, S94N, L96P, R98Q, A99D, A99V, E100K, A103T, A106T, S108A, I113F, V116F, V116I, V 125M, V125A, N126T, I128V, I128L, L129P, L135M, S139N, S139G, G143V, G14 3C, G146D, G146S, I147V, K149E, K149T, K149R, S153I, S153R, S153N, F154C, H156R, H156L, S158N, S158R, G159V, V160A, K162R, N164D, I166L, S167I, S1 68I, S168R, S168N, Q169R, V170M, V170G, V170L, T177I, T177A, S179R, F180C , F180L, F182C, F182L, G183S, M185I, K187R, G188D, V190I, K191N, A192S, D193N, G195V, G195D, G195S, C196W, T200A, T204I, A207V, A207T, and T208I at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 5, as well as one or more amino acid substitutions of M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L,L15Q、K16N、K16R、P17T、P17L、P17S、T19I、T19S、T19A、T19P、P20S、P20L、T21A、Q22R、Y23H、V24M、K25R、L26M、D27A、D27G、D28N、D28Y、A29T、A29V、N30K、I32F、I32S、Q33H、L36M、D37A、D37Y、F39L、S40P、D41E、T42I、T42K、T42A、F43L、F43S、F43V、K44N、N45D、N45S、Q49R、K52Q、S55A、T56A、D58E、K60Q、S62T、R63K、R63G、Q67R、Q67H、Q67K、D71Y、K74R、E76K、F78C、K79R、G80V、G80D、G81S、G81V、G81D、D82N、V83G、V83M、V83A、V84A、V84G、R85G、R85K、P86L、N87S、R89C、V91G、V91A、A92V、A92T、R95K、K97R、E100D、S101A、D104V、A106D、A106T、D110N、N112H、H113Y、M115R、N117Y、T119A、N120D、N120K、N120S、G124V、D125N、D125E、K127R、F129L、D130N、K131M、E134D、E134G、A139S、A139T、P142S、I144V、A145S、A145T、T146A、A147V、Q149R、Y150H、I155L、V156A、V156L、V156M、K157V、E158A、N159S、V163G、E164A、E164G、E164D、G165D、I167V、I169L、I169T、N173S、N173H、N173T、A174S、A174T、N176D、A181S、I182L、I182V、I182T、A186E、A186T、V187G、V187A、A190T、A190S、F195S、A197P、D198G、D198N、A205S、V208M、P209T、T211I、E215D、E218D、P223S、P223H、L226V、I227V、D231N、E232K、I235V、I235T、R239G、I246V、V248E、V248M、S250I、S259N、Y260C、K261R、S262N、P263L、S267N、A269V、T273I、T273N、H274Y、K277N、K277R、P278S、S280T、L281M、D282E、D282N、A283T、A283S、N285S、E287D、L288M、N290K、F295S、F298I、F298S、V302I、V303M、A307S、N313S、H316R、A317V、S320N、S320R、I323L、I325V、R331K、K332E、I339V、V345L、V345M、E348K、Y349H、Y349D、Y349N、Y349C、P352S、P352T、E353Q、E353D、L354M、G356S、N361D、I362V、I362T、L363P、L363T、L363M、E364G、K365R、E366G、E367G、K369N、K369E、K369M、P370S、E371K、V372M、D373G、I375V、M376I、T380P、T380A、E383K、E383D、F385L、H386Y、I389V、A390V、A390I、V392I、D396N、D396G、D396K、S397P、S399N、S399G、T402I、R403G、R403I、R403K、R403S、I404T、I404V、K407R、K407E、R408K、Q410K、Q410H、Q410R、Q411H、G412V、F413L、D414N、A415V、A415T、Y416C、M421I、N422K、E423K、E423D、E424A、E425K、E426D、T427A、T427S、R428K、F429L、S430A、M431L、R434H、R434C、R434S、I435V、D437G、D437N、T440S、T440I、R443C、G445S、F446L、F446I、Y448C、E450D、E450G、M452I、T456P、T456A、T456I、A459T、D460N、K463N、H464N、H464R、H464S、E470K、V472M、V472A、K473D、K473N、E494D、E494G、S495A、E498A、E498K、C501Y、T502I、T502S、P504S、P504L、T505A、G506Y、G506D、G506L、G506S、T508A、D509E、D509Y、C510Y、S512N、I513L、I513V、I513F、Y514H、K517M、K517N、K517Q、K520R, K521N, I522T, I522V, I522F, E525K, V526E, V526M, I527V, S530N, S530R, K531T, D532G, D532Y, S533Y, G535D, A537T, K538R, K5 38N, R540K, R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K, Q549R, E550A, Q551K, E552D, E552K , V553I, F554V, E556K, E556G, S557A, K558R, T559P, T559I, T559A, K560R, A561T, A561G, K562R, K562N, I563L, T564I, A565S, A565V, K 567R, K568N, K568R, Q569K, Q569L, Q569R, A570V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S, I583V, K584R, V585 M, S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S589R, S589N, A590S, A590T, A591V, P592L, V593M, V593A, Q594L, K595R, K595N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N, D600G, D600V, N601S, N601K, S602A, S602P, S602Y, D603A, D603V, D60 one or more amino acid substitutions among: D604G, D604Y, D604N, D606A, D606V, D606Y, D607Y, D607E, D608N, A611T, E613D, R618I, T620P, and A656V; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 6; and based on SEQ ID NO: 6, M1L, M1V, N2S, A3T, T5P, T5A, T5S, E6D, I7S, I7V, I9F, Q11R, L12M, N14D, N14S, M21I, H22P, H22Y, K26N, K26R, T27I, M31I, L35R, N38S, S43P, D44N, D44G, Q46L, C47S, T54I, S59T, H60Y, T61A, H64Y, Y65H, K67N, K67R, R68Q, A71G, T72A, N74D, S76C, S76Y, T79I, M80I, P81S, V84L, R89L, A95D , A95T, A102T, E105D, E105K, S109N, S109R, S110P, Q111R, I112T, K113N, K1 13E, K114N, K114M, K114E, G116D, K118N, K118R, T119I, D120V, K123N, L129M , I130V, K131R, A132S, K134M, K134N, F142V, L145M, I146T, E147K, F148S, S 150F, R154K, Q155H, E166D, K169E, P178S, A180V, A181T, A181S, I183V, A184 S, A184T, A184V, P187S, A190T, A190V, V194M, V194A, R197I, Y201N, L204M, D207N, K209N, Q213H, Q213V, A219S, K221N, D225N, V226E, P227T, K229E, S23 2N, K233N, K233R, N234H, T236A, A238V, A238S, A241S, E246D, K251N, H252Y , H252R, E256D, A257S, A261V, S263I, S263N, N265D, Y267C, E269K, E269D, K2 71E, K271R, H272Y, I274V, F280L, D281N, D281G, K285G, K286N, K288R, S291 F, S291P, K292N, K296R, K296N, I299S, D301G, E303D, I304T, I304V, E306G, V 307L, V307G, V307A, V307D, V307G, I308N, N310S, Y313H, N314K, N316K, N31 6D, A317D, L318Q, D319N, P320S, P320L, M323I, L324M, D326N, V328M, V328A,one or more amino acid substitutions among A330D, I331V, V332G, S340L, T341A, A343G, S344N, I355V, F412V, V418F, Y427C, R514K, S1198L, A1201V, G1206S, C1212G, F1260L, and V1282M; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) relative to SEQ ID NO: 7 and one or more amino acid substitutions of M99I, S189N, H265Q, A266V, L336F, and V343A relative to SEQ ID NO:7; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:8, and one or more amino acid substitutions of Y119H, N134R, N134Q, D155N, Q180R, D183N, R one or more amino acid substitutions among 274L, N319D, V447I, A454S, E458G, D461N, A512T, D538K, and P580Q; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 9, and one or more amino acid substitutions among R28K, A82T, K144E, C151R, N162S, K182E, D273G, one or more amino acid substitutions of A327D and M346I; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10, and one or more amino acid substitutions of A21S and V90A relative to SEQ ID NO: 10; at least 70% (e.g., at least 75%, at least 80%, at least 85%,and at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 11, as well as the following sequences: A2T, F3S, P7R, A9S, A9G, A11G, F12I, D14N, S16Y, Y20H, S26N, F29S, S32N, E34K, G35V, G35S, G35D, I40S, E43D, H45P, E46K, A54S, R61W, V64M, Y65C, N70S, A77T, D101N, K103E, N105K, N105D, S106G, V1 08M, A109G, Y111N, L119M, R120S, R123S, A126T, E127G, V130M, D131N, Q148 R, S149Y, H151Y, A157D, T159I, A164V, L166M, T185A, S194G, A196T, T203A, K 211R, E217K, R218K, R218S, N219S, A236T, E242D, N257K, N267S, M279I, M27 9V, D283G, N286S, T288I, K291Q, I293V, D296N, S303I, S303G, K306N, S310Y, one or more amino acid substitutions among S310P, I313T, Y314F, A316T, E326G, T331I, A336V, A347T, A347S, T352S, Y361H, M374T, M374I, R377G, T395I, S396T, S396F, G398V, and A408V; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) relative to SEQ ID NO: 12 and the identity of the following based on SEQ ID NO: 12: K4N, E5K, L6M, L6I, E8K, E8D, I9T, D11N, T12A, T13I, D16G, R17C, R17S, R20K, R20E, R21E, R21K, S24K, S24Q, S24R, Y26S, Y26H, A28S, A28D, M29I, G34D, A37S, V38M, V38G, I41V, R49L, D54G, K59R, K60N, K63N, A65T, A65V, K67E, K74E, K77E, W81C, K88R, K88E, I92T, R93E, R93K, V94M,K96N、E102D、E102G、T105A、L106M、S108P、V110A、G121S、S126P、K128R、L134M、Y138S、Q142H、W147L、K150N、V151M、V151L、A153T、S156R、S156G、D157N、K160R、K160E、A162T、S165N、S165G、V170E、K171E、F173V、K174N、K174R、T179A、K181T、S183N、P185T、E186K、E186D、E187K、A188S、A188V、D191Y、D191E、R198H、R198C、R198S、R201K、D206G、G207D、A226T、I228V、R233K、N236T、R241E、A249S、A250S、I256T、S267G、S267N、K268N、H270P、S275N、S275G、R276G、A277D、A277S、A277T、K279N、G283D、V286G、V289M、G303D、I305T、F306S、A310D、A310T、A312G、A312D、A312T、K314N、Q315R、R316G、N323S、E326A、E326K、N329S、G349D、E353D、L355M、L355R、E356G、E356D、S357P、A358V、R361S、P370T、N372K、E373D、S376F、T378I、F382L、M388V、G391S、R397K、A399S、K403N、M405I、L419P、D421N、K423R、H424N、H424R、V425L、I427V、E428K、D430A、D431G、E432D、H433N、A449T、G457D、R473K、E477D、G480D、F485L、S487R、S487G、S489N、N494D、S496N、A497S、V498G、K500N、K502N、Q509R、A511T、A511E、R515S、R518S、P519T、G520D、G520V、Y540C、Q545H、K550N、K555E、H557Q、P570S、E571D、C580R、S583R、E585K、E585G、E590D、R594K、M603I、H607N、H607L、K608R、D611N、L617P、N620S、K624N、T636P、M639V、N641S、V642G、S644N、one or more amino acid substitutions among S644G, E646D, A655V, V658M, K660N, T663A, T665I, R668S, I672V, G673V, S678R, M682L, A685V, A685D, K688N, and V695M; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) relative to SEQ ID NO: 13; ), and based on SEQ ID NO: 13, the identity of N5K, N5T, D10N, R11K, D26N, V30E, D35N, R40L, P42A, G45S, G45V, F46V, T47R, T47S, N58T, P61L, T65I, T71I, T71R, T71D, L72M, C75S, V77A, P78L, N80T, E82D, H83Y, H83N, A94S, V98M, E113D, C121F, A128S, A155S, E116D, T117I, R133K, G138V, N146D, G148V, C161R, A171V, A171S, K175T, A177V, K182E, L184M, I191V, S193A, S193F, F201S, S203N, E211K, A212V, Y219R, N225S, N225T, D226Y, E232K, E 232Q, A233N, A233S, A233K, K235R, Q236R, Q236S, F237L, V238Q, V238M , A240T, A240V, S250A, R274G, A282V, I286N, I286T, I286F, P292S, S29 5N, K304R, E307D, Y309C, A312V, L313M, N315K, N315T, N315S, C316G, I317V, T318A, T318P, K320R, N321D, E322K, K323N, I328T, M340I, K343 E. and A350T; or at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 14, and one or more amino acid substitutions of Q2K, H9L, K13E, Q14K, A15G, K34N, E38K, V42I, A46D, S50I, ... The amino acid sequence includes an amino acid sequence having one or more amino acid substitutions selected from the group consisting of V59G, Y60H, A73S, A73T, F75L, D77G, G82S, F83L, F83V, F83C, K85E, V86I, E97, I110S, I110L, S115R, K120N, K124R, G130D, D132E, N134T, A140T, E143K, D145G, S156I, E159K, I162V, H164Y, H164F, Y177C, S199I, S232L, and L270S.

[0011] In some embodiments, the polypeptide has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:1 and at least one or both of positions 2 and 230; 107 and 166; 107, 166, and 2 and 227; 211 and 110 or 142; 110, 155 and 230; 122 and 155; or 155 and 177. amino acid substitutions; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:2, and at positions 2 and 597; 24 and 25; 24, 25, 458, 509, 565, and 600; 75 and 597; 141, 454, 533, and 595; 581, 370, and 454; 370 and 581; 370 and 454; 458 and 509; 458, 509, and and 565; 458, 509, 565, and 600; 565, 586, and 596; or amino acid substitutions at 565, 509, 458, 600, and 24, 25, 29, 215, 319, 364, 383, and 586; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:3, and amino acid substitutions at positions 142 and 216. at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:4 and at positions 108 and 47 or 208; 170 and 207; 88 and 147; 47, 88 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 88, 128, 147, 170, and 182;or amino acid substitutions at 170, 207, and 108; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 5, and at positions 4, 23, and 590; 19, 169, and 549; 43 and 415; 80 and 593; 80, 144, 593, and 606; 1, 42, 80, 593, and 606; 156 and 6 04;283, 349, and 365;283, 349, 365, 396, and 594;283, 349, 365, 396, 594, 596, and 131;352 and 390;390, 396, and 594;396 and 594;456 and 502;464 and 502;464 and 17;17, 235, 464, and 596;235, 352, 396, 456, and 606;415, 456, and 502;456, 502, and 549;169, 456, 502, and 549;80, 456, 502, 593, and 606;1, 42, 80, 456, 502, 5 93, and 606; 80, 144, 456, 502, 593, and 606; 19, 169, 456, 502, and 549; 43, 415, 456, and 502; 352, 390, 396, and 594; 352, 390, and 396; 283, 349, 396, and 594; 11, 55, 120, 362, 584, 600, and 604; 43, 84, 144, 349, and 517; 164 and 165; 164 and 173; 362 and 446; 352, 390, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 43, 34 one or more positions selected from 9, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526; 43, 352, 390, 396, 464, 549, 594, 410, 526, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 21;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 21, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 174, 208, 427, 456, and 504; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 139; 594, 410, 526, 415, 502, 339, and 446; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 19, 460, 569, and 596; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 460, 586, 588, and 608; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, and 460; 352, 390, 396, 549, 586, and 594; 63, 158, 352, 390, 396, 549, 586, and 594; 16 amino acid substitutions at 4, 165, 352, 363, 390, 396, 410, 549, 586, and 594; 164, 173, 352, 390, 396, 549, 586, and 594; 83, 352, 390, 396, 549, 586, and 594; 8, 43, 174, 349, 352, 390, 396, 427, 464, 549, and 594; or 283, 349, 365, 396, 594, 596, and 131; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%), and identity at positions 2, 67, 95, and 226; 6 and 316; 38, 95, and 303; 67, 95, and 226; 44 and 76; 44, 76, and 118; 130, 234, 303; 118 and 1201; 118, 1201, and 44; 118, 1201, and 76; 130, 234, and 303; 154 and 269; 221 and 44; 44, 76, 130, 234, and 303; 44, 76, 118, and 1201;amino acid substitutions at 197 and 314; 76, 181, and 194; 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; 59, 76, 306, and 316; or 280 and 340; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity and amino acid substitutions at positions 105, 109, 131, 148, 279, and 310; or at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) relative to SEQ ID NO: 12; 9%), and identity at positions 134, 179, 185, 540, 555, 624, and 646; 138, 250, 275, and 421; 303, 405, 520, and 590; 134, 179, 185, 540, 555, 646, 4 and 49; 4 and 388; 4 and 571; 4, 162, and 480; 4 and 315; 5 and 316; 17 and 156; 38, 108, 497 and 583; 59, 157 and 644; 96, 305, 550 and 642; 106, 160 and 228; 312, 424, 449 and 457; or 376 and 611. amino acid substitutions; at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 13, and at positions 30, 46, 240, 304, and 316; 30, 46, 240, and 316; 42 and 318; 184, 240, 315, and 345; 211 and 274; 237 and 237; 286 and 350; 317 and 347; 171, 286, and 315;or amino acid substitutions at 328 and 350; or amino acid sequences having at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 14 and amino acid substitutions at positions 82, 110, 115, 164, and 199; 82, 110, 115, 124, 164, and 199; 110, 115, and 164; 110, 115, 164, and 199; 110, 115, 164, 199, and 124; or 110, 115, 164, 199, and 82 or 124;

[0012] In some embodiments, the polypeptide has at least 70% identity to SEQ ID NO:1 and amino acid substitutions at positions 55; 122 and 155; or 107, 166, and 227 relative to SEQ ID NO:1; at least 70% identity to SEQ ID NO:2 and amino acid substitutions at positions 24, 25, 458, 509, 565, and 600; 22, 347, and 454; or 485 relative to SEQ ID NO:2; at least 70% identity to SEQ ID NO:4 and amino acid substitutions at positions 75 and 182; 88, 147, and 188 relative to SEQ ID NO:4. and 177; 88 and 147; 88, 116 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 75, 88, and 147; 47, 88, and 147; 88, 128, 147, 170, and 182; or 88, 93, and 147; at least 70% identity to SEQ ID NO:5, and amino acid substitutions at positions 352, 390, 396, 594, and 596; 352, 390, 396, 549, and 594; 352, 390, 396, 46 one or more positions selected from: 4, 549, and 594; 289, 352, 390, 396, 549, 594, and 596; 235, 352, 390, 396, 567, and 594; 352, 363, 390, 396, 549, 586, and 594; 352, 390, 396, 549, 580, and 594; 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526; 43, 349, 352, 390, 396, 464, 549, 594, 415, and and 502; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502 and 67; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502 and 21; or amino acid substitutions at 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, 21 and 67; or at least 70% identity to SEQ ID NO:6 and at positions 197, 314, and optionally one of positions 7, 12, or 114 relative to SEQ ID NO:6; or 197 and 314; 76, 181, and 194;and amino acid sequences having amino acid substitutions at 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; or 59, 76, 306, and 316.

[0013] In some embodiments, the polypeptide has at least 70% identity to SEQ ID NO:1, and amino acid substitutions M155I; E122A and M155I; or K107M, N166D, and A227P relative to SEQ ID NO:1; at least 70% identity to SEQ ID NO:2, and amino acid substitutions at positions E24D, L25I, S458N, R509G, H565Y, and I600V; S22P, Y347F, and E454G; or V485F relative to SEQ ID NO:2; Identities and amino acid substitutions relative to SEQ ID NO:4 include S75I; F182L; P88T, I147V, and T177I; P88T and I147V; P88T, V116I and I147V; P88T, I147V, V170L, and F182L; P88T, I147V, V170L, F180L, and F182L; G51V, P88T, I147V, V170L, and F182L; P88T, I147V, and F154C; S75I, P88T, and I147V; or P88T, A93T, and I147V; relative to SEQ ID NO:5. and amino acid substitutions P352T, A390V, D396N, Q594L, and H596Y; P352S, A390V, D396N, Q549R, and Q594L; P352T, A390V, D396N, H464R, Q549R, and Q594L; Q289H, P352T, A390V, D396N, Q549R, Q594L, and H596Y; I235T, P352T, A390V, D396N, K567R, and Q594L; P352T, L363P, A390V, one or more substitutions selected from D396N, Q549R, S586A, and Q594L; P352T, A390V, D396N, Q549R, and Q594L; P352T, A390V, D396N, Q549R, T580I, and Q594L; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L and R63G, A145S, A174S, I182R, V208M, Q410K, T427S, T456I or T456P, P504S, and V526E;F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, and T502I; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and T21A; F43S, Y349N or Y349D, P35 2T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K; or F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, T21A, and Q67K; or at least 70% identity to SEQ ID NO: 6, and Based on SEQ ID NO: 6, positions R197I, N314K, and optionally one of I7S, L12M, or K114M; R197I and N314K; S76Y, A181S, and V194M; S76Y, K118R, H252R, and K292N; S76Y and I274V; S76Y, A102T, K118R, and V307G; L12M and S76Y; K67N , A95D, and V226E; K26N and S76Y; H22Y, S76Y, and D319N; R154K and E269D; S76Y and A238S; S76Y, A238S, K296N, and V328M; I7V and S76Y; S76Y and S263N; or amino acid sequences having amino acid substitutions at S59T, S76Y, E306G, and N316D.

[0014] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 13 and at least one amino acid substitution with a positively charged amino acid. In some embodiments, the positively charged amino acid is arginine or lysine. In selected embodiments, the positively charged amino acid is arginine. In some embodiments, the at least one amino acid substitution is at position 2, 5, 6, 7, 8, 9, 10, 12, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 64, 65, 66, 67, 68, 69, 70, 222, 224, 225, 227, 228, 229, 231, 232, 233, 234, 235, 255, 256, 257, 258, 277, 286, 287, 337, 338, 339, 340, 345, 346, 347, 348, 349, 350, or a combination thereof. In some embodiments, the at least one amino acid substitution is at positions 346 and 348; 346, 348 and 349; 346, 348, 349, and 350; 350 and 351; 350, 351, and 352; 350, 351, 352, and 353; 235 and 227; 235 and 345; 235 and 346; 235 and 347; 235 and 348; 235 and 349; 235 and 350; 235, 227, and 349; 5, 235 and 346; 5, 235 and 348; 5, 235 and 349; 227, 235, and 346; or 227, 235, and 348.

[0015] In some embodiments, the polypeptide is a fusion polypeptide comprising a first amino acid sequence and a second amino acid sequence. In some embodiments, the fusion polypeptide comprises a first amino acid sequence from one of the disclosed Cas or transposase proteins having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any of SEQ ID NOs: 1-14, with one or more amino acid substitutions, deletions, or additions relative to any of SEQ ID NOs: 1-14. In some embodiments, the fusion polypeptide further comprises a second amino acid sequence from one of the disclosed Cas or transposase proteins having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any of SEQ ID NOs: 1-14, with one or more amino acid substitutions, deletions, or additions relative to any of SEQ ID NOs: 1-14.

[0016] In some embodiments, the fusion polypeptide can comprise two or more of the disclosed transposase proteins (e.g., a first sequence having a sequence encoding a TnsA protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:1, and a second sequence having a sequence encoding a TnsB protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:2).

[0017] In some embodiments, the first amino acid sequence encodes a TnsA protein and has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 1, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 1, and the second amino acid sequence encodes a TnsB protein and has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 2, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 2.

[0018] In some embodiments, the first amino acid sequence comprises one or more amino acid substitutions at positions 2, 3, 5, 28, 57, 77, 80, 107, 110, 116, 122, 142, 155, 161, 166, 173, 177, 185, 211, 216, 227, and 230 relative to SEQ ID NO: 1. In some embodiments, the second amino acid sequence comprises one or more amino acid substitutions at positions 2, 5, 22, 24, 25, 29, 75, 141, 199, 215, 319, 347, 364, 370, 383, 439, 454, 458, 485, 509, 533, 538, 565, 581, 586, 595, 596, 597, and 600 relative to SEQ ID NO: 2.

[0019] In some embodiments, the first amino acid sequence comprises one or more amino acid substitutions of A2T, T3I, L5S, T28A, A57T, F77L, Y80D, K107M, K107R, Y110C, Y110D, D116G, E122A, D142E, M155I, K161R, N166D, K173E, Y177N, Y177D, C185R, D211Y, K216E, A227P, G230D, and G230S relative to SEQ ID NO:1. In some embodiments, the second amino acid sequence comprises one or more amino acid substitutions of A2T, A2S, G5R, S22P, E24D, L25I, A29S, P75T, I141T, V199I, S215R, D319V, Y347F, S364N, E370K, N383D, V439A, E454D, E454G, S458N, V485F, R509G, D533A, A538V, H565Y, A581T, H586L, N595K, D596N, D597N, D597Y, and I600V with respect to SEQ ID NO:2.

[0020] In some embodiments, the first amino acid sequence includes amino acid substitutions at positions 2 and 230; 107 and 166; one or both of 107, 166, and 2 and 227; 211 and 110 or 142; 110, 155 and 230; 122 and 155; or 155 and 177, relative to SEQ ID NO:1. In some embodiments, the second amino acid sequence comprises amino acid substitutions at positions 2 and 597; 24 and 25; 24, 25, 458, 509, 565, and 600; 75 and 597; 141, 454, 533, and 595; 581, 370, and 454; 370 and 581; 370 and 454; 458 and 509; 458, 509, and 565; 458, 509, 565, and 600; 565, 586, and 596; or 565, 509, 458, 600, and 24, 25, 29, 215, 319, 364, 383, and 586 relative to SEQ ID NO:2.

[0021] In some embodiments, the first amino acid sequence comprises amino acid substitutions at positions 107, 166, and 227 relative to SEQ ID NO:1, and the second amino acid sequence comprises amino acid substitutions at positions 24, 25, 458, 509, 565, and 600 relative to SEQ ID NO:2; or the first amino acid sequence comprises amino acid substitutions at position 155 relative to SEQ ID NO:1, and the second amino acid sequence comprises amino acid substitutions at positions 22, 347, and 454 relative to SEQ ID NO:2; or the first amino acid sequence comprises amino acid substitutions at positions 122 and 155 relative to SEQ ID NO:1, and the second amino acid sequence comprises amino acid substitution at position 485 relative to SEQ ID NO:2.

[0022] In some embodiments, the first amino acid sequence comprises the amino acid substitutions K107M, N166D, and A227P relative to SEQ ID NO:1, and the second amino acid sequence comprises the amino acid substitutions E24D, L25I, S458N, R509G, H565Y, and I600V relative to SEQ ID NO:2; or the first amino acid sequence comprises the amino acid substitution M155I relative to SEQ ID NO:1, and the second amino acid sequence comprises the amino acid substitutions S22P, Y347F, and E454G relative to SEQ ID NO:2; or the first amino acid sequence comprises the amino acid substitutions E122A and M155I relative to SEQ ID NO:1, and the second amino acid sequence comprises the amino acid substitution V485F relative to SEQ ID NO:2.

[0023] In some embodiments, the first amino acid sequence encodes a TnsA protein and has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 4, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 4. In some embodiments, the second amino acid sequence encodes a TnsA protein and has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 5, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 5.

[0024] In some embodiments, the first amino acid sequence is selected from the group consisting of positions 4, 5, 9, 10, 12, 21, 23, 25, 26, 31, 32, 34, 35, 37, 41, 45, 47, 48, 51, 52, 55, 60, 61, 65, 67, 69, 72, 75, 79, 80, 82, 87, 88, 90, 91, 93, 94, 96, 98, 99, 100, 103, 106, 108, 113, 116, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 2 126, 128, 129, 135, 139, 143, 146, 147, 149, 153, 154, 156, 158, 159, 160, 162, 164, 166, 167, 168, 169, 170, 177, 179, 180, 182, 183, 185, 187, 188, 190, 191, 192, 193, 195, 196, 200, 204, 207, and 208. In some embodiments, the second amino acid sequence is at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 49, 52, 55, 56, 58, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 1 0, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 110, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 145, 146 , 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169, 173, 174, 176, 181, 182, 186, 187, 190, 195, 197, 198, 205, 208, 209, 211, 215, 218, 223, 226, 227, 231, 232, 235, 239, 246, 248, 2 50, 259, 260, 261, 262, 263, 267, 269, 273, 274, 277, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298, 302, 303, 307, 313, 316, 317, 320, 323, 325, 331, 332, 339, 345, 348, 349, 352, 353, 354,356, 361, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373, 375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 404, 407, 408, 410, 411, 412, 413, 414, 415, 416, 421, 422, 423, 424 , 425, 426, 427, 428, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456, 459, 460, 463, 464, 470, 472, 473, 494, 495, 498, 501, 502, 504, 505, 506, 508, 509, 510, 512, 513, 514, 517, 520 , 521, 522, 525, 526, 527, 530, 531, 532, 533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 567, 568, 569, 570 , 571, 574, 575, 576, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656.

[0025] In some embodiments, the first amino acid sequence is selected from the group consisting of R4K, N5K, P9S, A10P, N12D, T21I, V23M, S25N, S25R, V26M, V26G, S31N, S32I, E34A, F35L, A37D, H41L, D45N, I47V, E48G, G51V, S52I, E55K, E55D, E60K, F61L, S65T, S65A, P67T, P67L, P67S, based on SEQ ID NO:4. , P67H, T69A, A72V, A72D, S75I, S75R, S75T, K79E, T80P, K82E, K87R, P88L, P88T, P88A, S90F, K91N, K91E, A93T, A9 3S, S94N, L96P, R98Q, A99D, A99V, E100K, A103T, A106T, S108A, I113F, V116F, V116I, V125M, V125A, N126T, I128V , I128L, L129P, L135M, S139N, S139G, G143V, G143C, G146D, G146S, I147V, K149E, K149T, K149R, S153I, S153R, S1 53N, F154C, H156R, H156L, S158N, S158R, G159V, V160A, K162R, N164D, I166L, S167I, S168I, S168R, S168N, Q169R , V170M, V170G, V170L, T177I, T177A, S179R, F180C, F180L, F182C, F182L, G183S, M185I, K187R, G188D, V190I, K191N, A192S, D193N, G195V, G195D, G195S, C196W, T200A, T204I, A207V, A207T, and T208I. In some embodiments, the second amino acid sequence is selected from the group consisting of M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L, L15Q, K16N, K16R, P17T, P17L, P17S, T19I, T19S, T19A, T19P, P20S, P20L, T21A, Q22R, Y23H, V24M, K25R, L26M, D27A, D27G, D28N, D28Y, A29T, A29V,N30K、I32F、I32S、Q33H、L36M、D37A、D37Y、F39L、S40P、D41E、T42I、T42K、T42A、F43L、F43S、F43V、K44N、N45D、N45S、Q49R、K52Q、S55A、T56A、D58E、K60Q、S62T、R63K、R63G、Q67R、Q67H、Q67K、D71Y、K74R、E76K、F78C、K79R、G80V、G80D、G81S、G81V、G81D、D82N、V83G、V83M、V83A、V84A、V84G、R85G、R85K、P86L、N87S、R89C、V91G、V91A、A92V、A92T、R95K、K97R、E100D、S101A、D104V、A106D、A106T、D110N、N112H、H113Y、M115R、N117Y、T119A、N120D、N120K、N120S、G124V、D125N、D125E、K127R、F129L、D130N、K131M、E134D、E134G、A139S、A139T、P142S、I144V、A145S、A145T、T146A、A147V、Q149R、Y150H、I155L、V156A、V156L、V156M、K157V、E158A、N159S、V163G、E164A、E164G、E164D、G165D、I167V、I169L、I169T、N173S、N173H、N173T、A174S、A174T、N176D、A181S、I182L、I182V、I182T、A186E、A186T、V187G、V187A、A190T、A190S、F195S、A197P、D198G、D198N、A205S、V208M、P209T、T211I、E215D、E218D、P223S、P223H、L226V、I227V、D231N、E232K、I235V、I235T、R239G、I246V、V248E、V248M、S250I、S259N、Y260C、K261R、S262N、P263L、S267N、A269V、T273I、T273N、H274Y、K277N、K277R、P278S、S280T、L281M、D282E、D282N、A283T、A283S、N285S、E287D、L288M、N290K、F295S、F298I、F298S、V302I、V303M、A307S、N313S、H316R、A317V、S320N、S320R、I323L、I325V、R331K、K332E、I339V、V345L、V345M、E348K、Y349H、Y349D、Y349N、Y349C、P352S、P352T、E353Q、E353D、L354M、G356S、N361D、I362V、I362T、L363P、L363T、L363M、E364G、K365R、E366G、E367G、K369N、K369E、K369M、P370S、E371K、V372M、D373G、I375V、M376I、T380P、T380A、E383K、E383D、F385L、H386Y、I389V、A390V、A390I、V392I、D396N、D396G、D396K、S397P、S399N、S399G、T402I、R403G、R403I、R403K、R403S、I404T、I404V、K407R、K407E、R408K、Q410K、Q410H、Q410R、Q411H、G412V、F413L、D414N、A415V、A415T、Y416C、M421I、N422K、E423K、E423D、E424A、E425K、E426D、T427A、T427S、R428K、F429L、S430A、M431L、R434H、R434C、R434S、I435V、D437G、D437N、T440S、T440I、R443C、G445S、F446L、F446I、Y448C、E450D、E450G、M452I、T456P、T456A、T456I、A459T、D460N、K463N、H464N、H464R、H464S、E470K、V472M、V472A、K473D、K473N、E494D、E494G、S495A、E498A、E498K、C501Y、T502I、T502S、P504S、P504L、T505A、G506Y、G506D、G506L、G506S、T508A、D509E、D509Y、C510Y、S512N、I513L、I513V、I513F、Y514H、K517M、K517N、K517Q、K520R、K521N、I522T、I522V、I522F、E525K、V526E、V526M、I527V、S530N、S530R、K531T、D532G、D532Y、S533Y、G535D、A537T、K538R、K538N、R540K、R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K, Q549R, E5 50A, Q551K, E552D, E552K, V553I, F554V, E556K, E556G, S557A, K558R, T559P, T559I, T559A, K560R , A561T, A561G, K562R, K562N, I563L, T564I, A565S, A565V, K567R, K568N, K568R, Q569K, Q569L, Q5 69R, A570V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S, I583V, K584R, V585M , S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S589R, S589N, A590S, A590T, A591V, P5 92L, V593M, V593A, Q594L, K595R, K595N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N , D600G, D600V, N601S, N601K, S602A, S602P, S602Y, D603A, D603V, D604G, D604Y, D604N, D606A, D606V, D606Y, D607Y, D607E, D608N, A611T, E613D, R618I, T620P, and A656V.

[0026] In some embodiments, the first amino acid sequence comprises amino acid substitutions at positions 108 and 47 or 208; 170 and 207; 88 and 147; 47, 88 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 88, 128, 147, 170, and 182; or 170, 207, and 108 relative to SEQ ID NO:4. In some embodiments, the second amino acid sequence is selected from the group consisting of positions 4, 23, and 590; 19, 169, and 549; 43 and 415; 80 and 593; 80, 144, 593, and 606; 1, 42, 80, 593, and 606; 42, 80, 593, and 606; 156 and 604; 283, 349, and 365; 283, 349, 365, 396, and 594; 283, 349, 365, 396, 594, 596, and 596, based on SEQ ID NO:5. , and 131; 352 and 390; 390, 396, and 594; 396 and 594; 456 and 502; 464 and 502; 464 and 17; 17, 235, 464, and 596; 235, 352, 396, 456, and 606; 415, 456, and 502; 456, 502, and 549; 169, 456, 502, and 549; 80, 456, 502, 593, and 606; 1, 42, 80, 456, 502, 593, and 60 6; 80, 144, 456, 502, 593, and 606; 19, 169, 456, 502, and 549; 43, 415, 456, and 502; 352, 390, 396, and 594; 352, 390, and 396; 283, 349, 396, and 594; 11, 55, 120, 362, 584, 600, and 604; 43, 84, 144, 349, and 517; 164 and 165; 164 and 173; 362 and 446; 352, 39 one or more positions selected from: 0, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526; 43, 352, 390, 396, 464, 549, and 594;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 21; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 67; 02, 21, and 67;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 174, 208, 427, 456, and 504;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 139;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 339, and 446;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 339, and 446; 96, 464, 549, 594, 410, 526, 19, 460, 569, and 596; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 460, 586, 588, and 608; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, and 460; 352, 390, 396, 549, 586, and 594; 63, 158, 352, 390, 396, 549, 586, and 5 94; 164, 165, 352, 363, 390, 396, 410, 549, 586, and 594; 164, 173, 352, 390, 396, 549, 586, and 594; 83, 352, 390, 396, 549, 586, and 594; 8, 43, 174, 349, 352, 390, 396, 427, 464, 549, and 594; or 283, 349, 365, 396, 594, 596, and 131.

[0027] In some embodiments, the first amino acid sequence comprises an amino acid substitution at position 182 relative to SEQ ID NO:4 and the second amino acid sequence comprises amino acid substitutions at positions 352, 390, 396, 594, and 596 relative to SEQ ID NO:5; or the first amino acid sequence comprises amino acid substitutions at positions 88, 147, and 177 relative to SEQ ID NO:4 and the second amino acid sequence comprises amino acid substitutions at positions 352, 390, 396, 549, and 594 .... , and amino acid substitutions at positions 88 and 147, and the second amino acid sequence comprises amino acid substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5; or the first amino acid sequence comprises amino acid substitutions at positions 88, 116, and 147 relative to SEQ ID NO:4, and the second amino acid sequence comprises amino acid substitutions at positions 289, 352, 390, 396, 549, 594, and 596 relative to SEQ ID NO:5; or the first amino acid sequence comprises an amino acid substitution at position 75 relative to SEQ ID NO:4; The second amino acid sequence comprises amino acid substitutions at positions 235, 352, 390, 396, 567, and 594 relative to SEQ ID NO:5, or the first amino acid sequence comprises amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO:4, and the second amino acid sequence comprises amino acid substitutions at positions 352, 363, 390, 396, 549, 586, and 594 relative to SEQ ID NO:5, or the first amino acid sequence comprises amino acid substitutions at positions 88, 147, 170, 182, and 51 or 180 relative to SEQ ID NO:4. wherein the second amino acid sequence comprises amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO:5, or wherein the first amino acid sequence comprises amino acid substitutions at positions 75, 88, and 147 relative to SEQ ID NO:4, and the second amino acid sequence comprises amino acid substitutions at positions 352, 390, 396, 549, and 594 relative to SEQ ID NO:5, or wherein the first amino acid sequence comprises amino acid substitutions at positions 88, 93, and 147 relative to SEQ ID NO:4;The second amino acid sequence includes amino acid substitutions at positions 352, 390, 396, 549, 580, and 594 relative to SEQ ID NO:5.

[0028] In some embodiments, the first amino acid sequence comprises the amino acid substitution F182L with respect to SEQ ID NO:4 and the second amino acid sequence comprises the amino acid substitutions P352T, A390V, D396N, Q594L, and H596Y with respect to SEQ ID NO:5; or the first amino acid sequence comprises the amino acid substitutions P88T, I147V, and T177I with respect to SEQ ID NO:4 and the second amino acid sequence comprises the amino acid substitutions P352S, A390V, D396N, Q549R, and Q594L ...4. and the second amino acid sequence, based on SEQ ID NO:5, comprises amino acid substitutions P352T, A390V, D396N, H464R, Q549R, and Q594L; or the first amino acid sequence, based on SEQ ID NO:4, comprises amino acid substitutions P88T, V116I, and I147V; the second amino acid sequence, based on SEQ ID NO:5, comprises amino acid substitutions Q289H, P352T, A390V, D396N, Q549R, Q594L, and H596Y; or the first amino acid sequence, based on SEQ ID NO:4, comprises amino acid substitutions P88T, V116I, and I147V; the first amino acid sequence comprises the amino acid substitution S75I, and the second amino acid sequence comprises the amino acid substitutions I235T, P352T, A390V, D396N, K567R, and Q594L based on SEQ ID NO:5; or the first amino acid sequence comprises the amino acid substitutions P88T, I147V, V170L, and F182L based on SEQ ID NO:5; the second amino acid sequence comprises the amino acid substitutions P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L based on SEQ ID NO:4; the first amino acid sequence comprises the amino acid substitutions S75I, P88T, and I147V with reference to SEQ ID NO:4, and the second amino acid sequence comprises the amino acid substitutions P352T, A390V, D396N, Q410K, H464R, V526E, Q549R, and Q594L with reference to SEQ ID NO:5; or the first amino acid sequence comprises the amino acid substitutions S75I, P88T, and I147V with reference to SEQ ID NO:4, and the second amino acid sequence comprises the amino acid substitutions P352T, A390V, D396N, Q549R, and Q594L with reference to SEQ ID NO:5; orThe first amino acid sequence includes amino acid substitutions P88T, A93T, and I147V based on SEQ ID NO:4, and the second amino acid sequence includes amino acid substitutions P352T, A390V, D396N, Q549R, T580I, and Q594L based on SEQ ID NO:5.

[0029] In some embodiments, the polypeptide further comprises one or more peptides fused to the polypeptide. In some embodiments, the one or more peptides comprise a linker peptide fusing the first amino acid sequence to the second amino acid sequence. In some embodiments, the one or more peptides comprise a nuclear localization sequence. In some embodiments, the nuclear localization sequence is a mono- or bi-articular sequence. In some embodiments, the one or more peptides comprise a tag or detectable label.

[0030] Also provided herein are nucleic acids comprising sequences encoding the disclosed polypeptides, and vectors comprising the disclosed nucleic acids.

[0031] Additionally, compositions comprising one or more of the disclosed transposon-associated protein or Cas protein polypeptides, or one or more nucleic acids encoding the polypeptides, are provided. In some embodiments, the compositions comprise two or more of the disclosed polypeptides, or one or more nucleic acids encoding the polypeptides described herein.

[0032] In some embodiments, the composition comprises two or all of a first polypeptide, a second polypeptide, and a third polypeptide (e.g., a first polypeptide having a sequence encoding a TnsA protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 1 or 4, a second polypeptide having a sequence encoding a TnsB protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 2 or 5, and / or a third polypeptide having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%) identity to SEQ ID NO: 3 or 6). %, or at least 99%) identity to SEQ ID NO: 8 or 12; or alternatively, a first polypeptide having a sequence encoding a Cas8 protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 9 or 13; a second polypeptide having a sequence encoding a Cas7 protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10 or 14, and / or a second polypeptide having a sequence encoding a Cas7 protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10 or 14,or a third polypeptide having a sequence encoding a Cas6 protein having at least 99% identity thereto).

[0033] In some embodiments, the first polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 1, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 1. In some embodiments, the second polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 2, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 2. In some embodiments, the third polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:3, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO:3.

[0034] In some embodiments, the first polypeptide comprises one or more amino acid substitutions at positions 2, 3, 5, 28, 57, 77, 80, 107, 110, 116, 122, 142, 155, 161, 166, 173, 177, 185, 211, 216, 227, and 230 relative to SEQ ID NO: 1. In some embodiments, the second polypeptide comprises one or more amino acid substitutions at positions 2, 5, 22, 24, 25, 29, 75, 141, 199, 215, 319, 347, 364, 370, 383, 439, 454, 458, 485, 509, 533, 538, 565, 581, 586, 595, 596, 597, and 600 relative to SEQ ID NO: 2. In some embodiments, the third polypeptide comprises one or more amino acid substitutions at positions 9, 15, 16, 18, 21, 64, 81, 86, 87, 99, 109, 142, 147, 153, 168, 180, 216, 230, 285, and 304 relative to SEQ ID NO:3.

[0035] In some embodiments, the first polypeptide comprises one or more amino acid substitutions of A2T, T3I, L5S, T28A, A57T, F77L, Y80D, K107M, K107R, Y110C, Y110D, D116G, E122A, D142E, M155I, K161R, N166D, K173E, Y177N, Y177D, C185R, D211Y, K216E, A227P, G230D, and G230S relative to SEQ ID NO:1. In some embodiments, the second polypeptide comprises one or more amino acid substitutions of A2T, A2S, G5R, S22P, E24D, L25I, A29S, P75T, I141T, V199I, S215R, D319V, Y347F, S364N, E370K, N383D, V439A, E454D, E454G, S458N, V485F, R509G, D533A, A538V, H565Y, A581T, H586L, N595K, D596N, D597N, D597Y, and I600V relative to SEQ ID NO:2. In some embodiments, the third polypeptide comprises one or more amino acid substitutions of I9V, A15V, F16Y, S18F, S21N, N64D, H81Y, D86Y, N87K, V99I, E109D, E142K, V147I, N153D, I168M, A180E, A216S, L230F, K285E, and R304R relative to SEQ ID NO:3.

[0036] In some embodiments, the first polypeptide comprises amino acid substitutions at positions 2 and 230; 107 and 166; one or both of 107, 166, and 2 and 227; 211 and 110 or 142; 110, 155 and 230; 122 and 155; or 155 and 177 relative to SEQ ID NO:1. In some embodiments, the second polypeptide comprises amino acid substitutions at positions 2 and 597; 24 and 25; 24, 25, 458, 509, 565, and 600; 75 and 597; 141, 454, 533 and 595; 581, 370, and 454; 370 and 581; 370 and 454; 458 and 509; 458, 509 and 565; 458, 509, 565, and 600; 565, 586, and 596; or 565, 509, 458, 600 and 24, 25, 29, 215, 319, 364, 383, and 586 relative to SEQ ID NO: 2. In some embodiments, the third polypeptide comprises amino acid substitutions at positions 142 and 216 relative to SEQ ID NO: 3.

[0037] In some embodiments, the first polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 4, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 4. In some embodiments, the second polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 5, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 5. In some embodiments, the third polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 6, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 6.

[0038] In some embodiments, the first polypeptide comprises a sequence selected from the group consisting of: 4, 5, 9, 10, 12, 21, 23, 25, 26, 31, 32, 34, 35, 37, 41, 45, 47, 48, 51, 52, 55, 60, 61, 65, 67, 69, 72, 75, 79, 80, 82, 87, 88, 90, 91, 93, 94, 96, 98, 99, 100, 103, 106, 108, 113, 116, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 2 126, 128, 129, 135, 139, 143, 146, 147, 149, 153, 154, 156, 158, 159, 160, 162, 164, 166, 167, 168, 169, 170, 177, 179, 180, 182, 183, 185, 187, 188, 190, 191, 192, 193, 195, 196, 200, 204, 207, and 208. In some embodiments, the second polypeptide is at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 49, 52, 55, 56, 58, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135 0, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 110, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 145, 146 , 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169, 173, 174, 176, 181, 182, 186, 187, 190, 195, 197, 198, 205, 208, 209, 211, 215, 218, 223, 226, 227, 231, 232, 235, 239, 246, 248, 2 50, 259, 260, 261, 262, 263, 267, 269, 273, 274, 277, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298, 302, 303, 307, 313, 316, 317, 320, 323, 325, 331, 332, 339, 345, 348, 349, 352, 353, 354,356, 361, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373, 375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 404, 407, 408, 410, 411, 412, 413, 414, 415, 416, 421, 422, 423, 424 , 425, 426, 427, 428, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456, 459, 460, 463, 464, 470, 472, 473, 494, 495, 498, 501, 502, 504, 505, 506, 508, 509, 510, 512, 513, 514, 517, 520 , 521, 522, 525, 526, 527, 530, 531, 532, 533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 567, 568, 569, 570 , 571, 574, 575, 576, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656. In some embodiments, the third polypeptide is located at positions 1, 2, 3, 5, 6, 7, 9, 11, 12, 14, 21, 22, 26, 27, 31, 35, 38, 43, 44, 46, 47, 54, 59, 60, 61, 64, 65, 67, 68, 71, 72, 74, 76, 79, 80, 81, 84, 89, 95, 102, 105, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 113, 114, 116, 118, 119, 120, 123, 129, 130, 131, 132, 134, 142, 145, 146, 147, 148, 150, 154, 155, 166, 169, 178, 180, 181, 183, 184, 187, 190, 194, 197, 201, 204, 207, 209, 213, 219, 221, 225, 226, 227, 229, 232,233, 234, 236, 238, 241, 246, 251, 252, 256, 257, 261, 263, 265, 267, 269, 271, 272, 274, 280, 281, 285, 286, 288, 291, 292, 296, 299, 301, 303, 304, 306, 307, 308, 310, 313, 314, 316, 317, 318, 319, 320, 323, 324, 326, 328, 330, 331, 332, 340, 341, 343, 344, 355, 412, 418, 427, 514, 1198, 1201, 1206, 1212, 1260, and 1282.

[0039] In some embodiments, the first polypeptide is selected from the group consisting of R4K, N5K, P9S, A10P, N12D, T21I, V23M, S25N, S25R, V26M, V26G, S31N, S32I, E34A, F35L, A37D, H41L, D45N, I47V, E48G, G51V, S52I, E55K, E55D, E60K, F61L, S65T, S65A, P67T, P67L, P67S, based on SEQ ID NO:4. , P67H, T69A, A72V, A72D, S75I, S75R, S75T, K79E, T80P, K82E, K87R, P88L, P88T, P88A, S90F, K91N, K91E, A93T, A9 3S, S94N, L96P, R98Q, A99D, A99V, E100K, A103T, A106T, S108A, I113F, V116F, V116I, V125M, V125A, N126T, I128V , I128L, L129P, L135M, S139N, S139G, G143V, G143C, G146D, G146S, I147V, K149E, K149T, K149R, S153I, S153R, S1 53N, F154C, H156R, H156L, S158N, S158R, G159V, V160A, K162R, N164D, I166L, S167I, S168I, S168R, S168N, Q169R , V170M, V170G, V170L, T177I, T177A, S179R, F180C, F180L, F182C, F182L, G183S, M185I, K187R, G188D, V190I, K191N, A192S, D193N, G195V, G195D, G195S, C196W, T200A, T204I, A207V, A207T, and T208I. In some embodiments, the second polypeptide is selected from the group consisting of M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L, L15Q, K16N, K16R, P17T, P17L, P17S, T19I, T19S, T19A, T19P, P20S, P20L, T21A, Q22R, Y23H, V24M, K25R, L26M, D27A, D27G, D28N, D28Y, A29T, A29V,N30K、I32F、I32S、Q33H、L36M、D37A、D37Y、F39L、S40P、D41E、T42I、T42K、T42A、F43L、F43S、F43V、K44N、N45D、N45S、Q49R、K52Q、S55A、T56A、D58E、K60Q、S62T、R63K、R63G、Q67R、Q67H、Q67K、D71Y、K74R、E76K、F78C、K79R、G80V、G80D、G81S、G81V、G81D、D82N、V83G、V83M、V83A、V84A、V84G、R85G、R85K、P86L、N87S、R89C、V91G、V91A、A92V、A92T、R95K、K97R、E100D、S101A、D104V、A106D、A106T、D110N、N112H、H113Y、M115R、N117Y、T119A、N120D、N120K、N120S、G124V、D125N、D125E、K127R、F129L、D130N、K131M、E134D、E134G、A139S、A139T、P142S、I144V、A145S、A145T、T146A、A147V、Q149R、Y150H、I155L、V156A、V156L、V156M、K157V、E158A、N159S、V163G、E164A、E164G、E164D、G165D、I167V、I169L、I169T、N173S、N173H、N173T、A174S、A174T、N176D、A181S、I182L、I182V、I182T、A186E、A186T、V187G、V187A、A190T、A190S、F195S、A197P、D198G、D198N、A205S、V208M、P209T、T211I、E215D、E218D、P223S、P223H、L226V、I227V、D231N、E232K、I235V、I235T、R239G、I246V、V248E、V248M、S250I、S259N、Y260C、K261R、S262N、P263L、S267N、A269V、T273I、T273N、H274Y、K277N、K277R、P278S、S280T、L281M、D282E、D282N、A283T、A283S、N285S、E287D、L288M、N290K、F295S、F298I、F298S、V302I、V303M、A307S、N313S、H316R、A317V、S320N、S320R、I323L、I325V、R331K、K332E、I339V、V345L、V345M、E348K、Y349H、Y349D、Y349N、Y349C、P352S、P352T、E353Q、E353D、L354M、G356S、N361D、I362V、I362T、L363P、L363T、L363M、E364G、K365R、E366G、E367G、K369N、K369E、K369M、P370S、E371K、V372M、D373G、I375V、M376I、T380P、T380A、E383K、E383D、F385L、H386Y、I389V、A390V、A390I、V392I、D396N、D396G、D396K、S397P、S399N、S399G、T402I、R403G、R403I、R403K、R403S、I404T、I404V、K407R、K407E、R408K、Q410K、Q410H、Q410R、Q411H、G412V、F413L、D414N、A415V、A415T、Y416C、M421I、N422K、E423K、E423D、E424A、E425K、E426D、T427A、T427S、R428K、F429L、S430A、M431L、R434H、R434C、R434S、I435V、D437G、D437N、T440S、T440I、R443C、G445S、F446L、F446I、Y448C、E450D、E450G、M452I、T456P、T456A、T456I、A459T、D460N、K463N、H464N、H464R、H464S、E470K、V472M、V472A、K473D、K473N、E494D、E494G、S495A、E498A、E498K、C501Y、T502I、T502S、P504S、P504L、T505A、G506Y、G506D、G506L、G506S、T508A、D509E、D509Y、C510Y、S512N、I513L、I513V、I513F、Y514H、K517M、K517N、K517Q、K520R、K521N、I522T、I522V、I522F、E525K、V526E、V526M、I527V、S530N、S530R、K531T、D532G、D532Y、S533Y、G535D、A537T、K538R、K538N、R540K、R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K, Q549R, E5 50A, Q551K, E552D, E552K, V553I, F554V, E556K, E556G, S557A, K558R, T559P, T559I, T559A, K560R , A561T, A561G, K562R, K562N, I563L, T564I, A565S, A565V, K567R, K568N, K568R, Q569K, Q569L, Q5 69R, A570V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S, I583V, K584R, V585M , S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S589R, S589N, A590S, A590T, A591V, P5 92L, V593M, V593A, Q594L, K595R, K595N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N , D600G, D600V, N601S, N601K, S602A, S602P, S602Y, D603A, D603V, D604G, D604Y, D604N, D606A, D606V, D606Y, D607Y, D607E, D608N, A611T, E613D, R618I, T620P, and A656V. In some embodiments, the third polypeptide is selected from the group consisting of M1L, M1V, N2S, A3T, T5P, T5A, T5S, E6D, I7S, I7V, I9F, Q11R, L12M, N14D, N14S, M21I, H22P, H22Y, K26N, K26R, T27I, M31I, L35R, N38S, S43P, D44N, D44G, Q46L, C47S, T54I, S5 9T, H60Y, T61A, H64Y, Y65H, K67N, K67R, R68Q, A71G, T72A, N74D, S76C, S76Y, T79I, M80I, P81S, V84L, R89L, A95D, A95T, A102T, E105D, E105K, S109N, S109R, S110P, Q111R, I112T, K113N, K113E, K114N, K114M, K114E,G116D, K118N, K118R, T119I, D120V, K123N, L129M, I130V, K131R, A132S, K134M, K134N, F142V, L145M, I146T, E147K, F148 S, S150F, R154K, Q155H, E166D, K169E, P178S, A180V, A181T, A181S, I183V, A184S, A184T, A184V, P187S, A190T, A190V, V1 94M, V194A, R197I, Y201N, L204M, D207N, K209N, Q213H, Q213V, A219S, K221N, D225N, V226E, P227T, K229E, S232N, K233N, K233R, N234H, T236A, A238V, A238S, A241S, E246D, K251N, H252Y, H252R, E256D, A257S, A261V, S263I, S263N, N265D, Y267C , E269K, E269D, K271E, K271R, H272Y, I274V, F280L, D281N, D281G, K285G, K286N, K288R, S291F, S291P, K292N, K296R, K29 6N, I299S, D301G, E303D, I304T, I304V, E306G, V307L, V307G, V307A, V307D, V307G, I308N, N310S, Y313H, N314K, N316K, N3 16D, A317D, L318Q, D319N, P320S, P320L, M323I, L324M, D326N, V328M, V328A, A330D, I331V, V332G, S340L, T341A, A343G, S344N, I355V, F412V, V418F, Y427C, R514K, S1198L, A1201V, G1206S, C1212G, F1260L, and V1282M.

[0040] In some embodiments, the first polypeptide comprises amino acid substitutions at positions 108 and 47 or 208; 170 and 207; 88 and 147; 47, 88 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 88, 128, 147, 170, and 182; or 170, 207, and 108, relative to SEQ ID NO:4. In some embodiments, the second polypeptide is located at positions 4, 23, and 590; 19, 169, and 549; 43 and 415; 80 and 593; 80, 144, 593, and 606; 1, 42, 80, 593, and 606; 42, 80, 593, and 606; 156 and 604; 283, 349, and 365; 283, 349, 365, 396, and 594; 283, 349, 365, 396, 594, 596 , and 131; 352 and 390; 390, 396, and 594; 396 and 594; 456 and 502; 464 and 502; 464 and 17; 17, 235, 464, and 596; 235, 352, 396, 456, and 606; 415, 456, and 502; 456, 502, and 549; 169, 456, 502, and 549; 80, 456, 502, 593, and 606; 1, 42, 80, 456, 502, 593, and 60 6; 80, 144, 456, 502, 593, and 606; 19, 169, 456, 502, and 549; 43, 415, 456, and 502; 352, 390, 396, and 594; 352, 390, and 396; 283, 349, 396, and 594; 11, 55, 120, 362, 584, 600, and 604; 43, 84, 144, 349, and 517; 164 and 165; 164 and 173; 362 and 446; 352, 39 one or more positions selected from: 0, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526; 43, 352, 390, 396, 464, 549, and 594;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 21; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 67; 02, 21, and 67;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 174, 208, 427, 456, and 504;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 139;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 339, and 446;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 339, and 446; 96, 464, 549, 594, 410, 526, 19, 460, 569, and 596; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 460, 586, 588, and 608; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, and 460; 352, 390, 396, 549, 586, and 594; 63, 158, 352, 390, 396, 549, 586, and 5 94; 164, 165, 352, 363, 390, 396, 410, 549, 586, and 594; 164, 173, 352, 390, 396, 549, 586, and 594; 83, 352, 390, 396, 549, 586, and 594; 8, 43, 174, 349, 352, 390, 396, 427, 464, 549, and 594; or 283, 349, 365, 396, 594, 596, and 131. In some embodiments, the third polypeptide is located at positions 2, 67, 95, and 226; 6 and 316; 38, 95, 303; 67, 95, and 226; 44 and 76; 44, 76, and 118; 130, 234, 303; 118 and 1201; 118, 1201, and 44; 118, 1201, and 76; 130, 234, and 303; 154 and 269; 221 and 44; 44, 76, 130, 234, and 303;Containing amino acid substitutions at 44, 76, 118, and 1201; 197 and 314; 76, 181, and 194; 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; 59, 76, 306, and 316; or 280 and 340.

[0041] In some embodiments, the first polypeptide comprises amino acid substitutions at positions 88 and 147 relative to SEQ ID NO:4, the second polypeptide comprises amino acid substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5, and the third polypeptide comprises amino acid substitutions at positions 197, 314, and optionally one of positions 7, 12, or 114 relative to SEQ ID NO:6; or the first polypeptide comprises amino acid substitutions at positions 88 and 147 relative to SEQ ID NO:4. The first polypeptide comprises amino acid substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5, and the third polypeptide comprises amino acid substitutions at positions 76, 181, and 194 relative to SEQ ID NO:6; or the first polypeptide comprises amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO:4, and the second polypeptide comprises amino acid substitutions at positions 352, 363, 390, 396, 549, 586, and 594 relative to SEQ ID NO:5. and 594, and the third polypeptide comprises amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114 relative to SEQ ID NO:6; or the first polypeptide comprises amino acid substitutions at positions 88, 147, 170, 180, and 182 relative to SEQ ID NO:4, and the second polypeptide comprises amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO:5; and the third polypeptide includes amino acid substitutions at positions 197 and 314 relative to SEQ ID NO:6, or alternatively, the first polypeptide includes amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO:4, the second polypeptide includes amino acid substitutions at positions 352, 363, 390, 396, 549, 586, and 594 relative to SEQ ID NO:5, and the third polypeptide includes amino acid substitutions at positions 76, 181, and 194 relative to SEQ ID NO:6.

[0042] In some embodiments, the first polypeptide comprises the amino acid substitutions P88T and I147V relative to SEQ ID NO:4; the second polypeptide comprises the amino acid substitutions P352T, A390V, D396N, H464R, Q549R, and Q594L relative to SEQ ID NO:5; and the third polypeptide comprises the amino acid substitutions R197I, N314K, and optionally one of I7S, L12M, or K114M relative to SEQ ID NO:6; or the first polypeptide comprises the amino acid substitutions P88T and I147V relative to SEQ ID NO:4. wherein the second polypeptide comprises amino acid substitutions P352T, A390V, D396N, H464R, Q549R, and Q594L relative to SEQ ID NO:5, and the third polypeptide comprises amino acid substitutions S76Y, A181S, and V194M relative to SEQ ID NO:6; or the first polypeptide comprises amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO:4, and the second polypeptide comprises amino acid substitutions P352T, L363P, A390V, D396N, Q549R, S76Y, A181S, and V194M relative to SEQ ID NO:5. and the third polypeptide comprises, with reference to SEQ ID NO:6, amino acid substitutions R197I, N314K, and optionally one of I7S, L12M, or K114M; or the first polypeptide comprises, with reference to SEQ ID NO:4, amino acid substitutions P88T, I147V, V170L, F180L, and F182L; and the second polypeptide comprises, with reference to SEQ ID NO:5, amino acid substitutions F43S, Y349N, P352T, A390V, D396N, Q410K, H464R, V526E, Q549R, and Q594L. 4L, and the third polypeptide comprises amino acid substitutions R197I and N314K with reference to SEQ ID NO:6; or the first polypeptide comprises amino acid substitutions at positions 88, 147, 170, and 182 with reference to SEQ ID NO:4, the second polypeptide comprises amino acid substitutions P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L with reference to SEQ ID NO:5, and the third polypeptide comprises amino acid substitutions S76Y, A181S, and V194M with reference to SEQ ID NO:6.

[0043] In some embodiments, the first polypeptide comprises the amino acid sequence of SEQ ID NO:4, the second polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:5, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO:5, and the third polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:6, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO:6. In some embodiments, the second polypeptide is at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 49, 52, 55, 56, 58, 60, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 240, 241, 242, 243, 24 4, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 110, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 145, 146, 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169, 173, 174, 176, 181, 182, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 2 87, 190, 195, 197, 198, 205, 208, 209, 211, 215, 218, 223, 226, 227, 231, 232, 235, 239, 246, 248, 250, 259, 260, 261, 262, 263, 267, 269, 273, 274, 277, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298, 302, 303, 307, 313, 316, 317, 320, 32 3, 325, 331, 332, 339, 345, 348, 349, 352, 353, 354, 356, 361, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373, 375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 404, 407, 408, 410, 411, 412, 413, 414, 415, 416, 421, 422,423, 424, 425, 426, 427, 428, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456, 459, 460, 463, 464, 470, 472, 473, 494, 495, 498, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 02, 504, 505, 506, 508, 509, 510, 512, 513, 514, 517, 520, 521, 522, 525, 526, 527, 530, 531, 532, 533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547 548, 549, 550, 551, 552, 553, 554, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 567, 568, 569, 570, 571, 574, 575, 576, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656, and / or the third polypeptide may comprise a sequence selected from the group consisting of: 1, 2, 3, 5, 6, 7, 9, 11, 12, 14, 21, 22, 26, 27, 31, 35, 38, 43, 44, 46, 47, 54, 59, 60, 61, 64, 65, 67, 68, 71, 72, 74, 76, 79, 80, 81, 84, 89, 95, 102, 105, 109, 110, 111, 112, 113, 114, 116, 118, 119, 120, 123, 129, 130, 131, 132, 134, 142, 145, 146, 147, 148, 150, 154, 155, 160, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 6, 169, 178, 180, 181, 183, 184, 187, 190, 194, 197, 201, 204, 207, 209, 213, 219, 221, 225, 226, 227, 229, 232, 233, 234, 236, 238, 241, 246, 251, 252, 256 , 257, 261, 263, 265, 267, 269, 271, 272, 274, 280, 281, 285, 286, 288, 291, 292, 296, 299, 301, 303, 304, 306, 307, 308, 310, 313, 314, 316, 317, 318, 319,Contains one or more amino acid substitutions at 320, 323, 324, 326, 328, 330, 331, 332, 340, 341, 343, 344, 355, 412, 418, 427, 514, 1198, 1201, 1206, 1212, 1260, and 1282.

[0044] In some embodiments, the second polypeptide is selected from the group consisting of M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L, L15Q, K16N, K16R, P17T, P17L, P17S, T19I, T19S, T19A, T19P, P20S, P20L, T21A, Q22R, Y23H, V24M, K25R, L26M, D27A, D27G, D28N, D28Y, A29T ... 29V, N30K, I32F, I32S, Q33H, L36M, D37A, D37Y, F39L, S40P, D41E, T42I, T42 K, T42A, F43L, F43S, F43V, K44N, N45D, N45S, Q49R, K52Q, S55A, T56A, D58E, K 60Q, S62T, R63K, R63G, Q67R, Q67H, Q67K, D71Y, K74R, E76K, F78C, K79R, G80 V, G80D, G81S, G81V, G81D, D82N, V83G, V83M, V83A, V84A, V84G, R85G, R85K, P 86L, N87S, R89C, V91G, V91A, A92V, A92T, R95K, K97R, E100D, S101A, D104V, A106D, A106T, D110N, N112H, H113Y, M115R, N117Y, T119A, N120D, N120K, N12 0S, G124V, D125N, D125E, K127R, F129L, D130N, K131M, E134D, E134G, A139S , A139T, P142S, I144V, A145S, A145T, T146A, A147V, Q149R, Y150H, I155L, V1 56A, V156L, V156M, K157V, E158A, N159S, V163G, E164A, E164G, E164D, G165 D, I167V, I169L, I169T, N173S, N173H, N173T, A174S, A174T, N176D, A181S, I 182L, I182V, I182T, A186E, A186T, V187G, V187A, A190T, A190S, F195S, A19 7P, D198G, D198N, A205S, V208M, P209T, T211I, E215D, E218D, P223S, P223H,L226V, I227V, D231N, E232K, I235V, I235T, R239G, I246V, V248E, V248M, S250I, S259N, Y260C, K261R, S262N, P263L, S267N, A269V, T273I, T273N, H274 Y, K277N, K277R, P278S, S280T, L281M, D282E, D282N, A283T, A283S, N285S, E287D, L288M, N290K, F295S, F298I, F298S, V302I, V303M, A307S, N313S, H31 6R, A317V, S320N, S320R, I323L, I325V, R331K, K332E, I339V, V345L, V345M, E348K, Y349H, Y349D, Y349N, Y349C, P352S, P352T, E353Q, E353D, L354M, G 356S, N361D, I362V, I362T, L363P, L363T, L363M, E364G, K365R, E366G, E367G, K369N, K369E, K369M, P370S, E371K, V372M, D373G, I375V, M376I, T380P T380A, E383K, E383D, F385L, H386Y, I389V, A390V, A390I, V392I, D396N, D396G, D396K, S397P, S399N, S399G, T402I, R403G, R403I, R403K, R403S, I404 T、I404V、K407R、K407E、R408K、Q410K、Q410H、Q410R、Q411H、G412V、F413L、 D414N、A415V、A415T、Y416C、M421I、N422K、E423K、E423D、E424A、E425K、E42 6D、T427A、T427S、R428K、F429L、S430A、M431L、R434H、R434C、R434S、I435V 、D437G、D437N、T440S、T440I、R443C、G445S、F446L、F446I、Y448C、E450D、E 450G、M452I、T456P、T456A、T456I、A459T、D460N、K463N、H464N、H464R、H46 4S、E470K、V472M、V472A、K473D、K473N、E494D、E494G、S495A、E498A、E498K、C501Y, T502I, T502S, P504S, P504L, T505A, G506Y, G506D, G506L, G506S, T5 08A, D509E, D509Y, C510Y, S512N, I513L, I513V, I513F, Y514H, K517M, K517N , K517Q, K520R, K521N, I522T, I522V, I522F, E525K, V526E, V526M, I527V, S 530N, S530R, K531T, D532G, D532Y, S533Y, G535D, A537T, K538R, K538N, R540 K, R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K, Q549R, E550A, Q551K, E552D, E552K, V553I, F554V, E556K, E55 6G, S557A, K558R, T559P, T559I, T559A, K560R, A561T, A561G, K562R, K562N , I563L, T564I, A565S, A565V, K567R, K568N, K568R, Q569K, Q569L, Q569R, A5 70V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S, I583 V, K584R, V585M, S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S 589R, S589N, A590S, A590T, A591V, P592L, V593M, V593A, Q594L, K595R, K59 5N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N, D600G, D600V, and / or the third polypeptide comprises one or more amino acid substitutions selected from the group consisting of M1L, M1V, N2S, A3T, T5P, T5A, T5S, E6D, I7S, I7V, I9F, Q11R, L12M, N14D, N14S, M21I,H22P、H22Y、K26N、K26R、T27I、M31I、L35R、N38S、S43P、D44N、D44G、Q46L、C47S、T54I、S59T、H60Y、T61A、H64Y、Y65H、K67N、K67R、R68Q、A71G、T72A、N74D、S76C、S76Y、T79I、M80I、P81S、V84L、R89L、A95D、A95T、A102T、E105D、E105K、S109N、S109R、S110P、Q111R、I112T、K113N、K113E、K114N、K114M、K114E、G116D、K118N、K118R、T119I、D120V、K123N、L129M、I130V、K131R、A132S、K134M、K134N、F142V、L145M、I146T、E147K、F148S、S150F、R154K、Q155H、E166D、K169E、P178S、A180V、A181T、A181S、I183V、A184S、A184T、A184V、P187S、A190T、A190V、V194M、V194A、R197I、Y201N、L204M、D207N、K209N、Q213H、Q213V、A219S、K221N、D225N、V226E、P227T、K229E、S232N、K233N、K233R、N234H、T236A、A238V、A238S、A241S、E246D、K251N、H252Y、H252R、E256D、A257S、A261V、S263I、S263N、N265D、Y267C、E269K、E269D、K271E、K271R、H272Y、I274V、F280L、D281N、D281G、K285G、K286N、K288R、S291F、S291P、K292N、K296R、K296N、I299S、D301G、E303D、I304T、I304V、E306G、V307L、V307G、V307A、V307D、V307G、I308N、N310S、Y313H、N314K、N316K、N316D、A317D、L318Q、D319N、P320S、P320L、M323I、L324M、D326N、V328M、V328A、A330D、I331V、V332G、S340L、T341A、A343G、S344N、I355V、F412V、V418F、Y427C、R514K、S1198L、A1201V、Contains one or more amino acid substitutions of G1206S, C1212G, F1260L, and V1282M.

[0045] In some embodiments, the second polypeptide is located at one or more positions selected from positions 43, 349, 352, 390, 396, 464, 549, 594, and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526, relative to SEQ ID NO:5; 43, 349, 352, 390, 396, 464, 549, 594, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 67; 43, and / or the third polypeptide comprises amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 76 and 7, 12, or 263; or 76, 238, 296, or 328, relative to SEQ ID NO:6.In some embodiments, the second polypeptide comprises, based on SEQ ID NO: 5, F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, and R63G, A145S, A174S, I182R, V208M, Q410K, T427S, T456I or T456P, P504S, and one or more substitutions selected from V526E; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V and T502I; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and T21A; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K; or F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T5 and / or the third polypeptide includes substitutions of R197I, N314K, and optionally one of I7S, L12M, or K114M; S76Y and I7V, L12M, or S263N; or S76Y, A238S, K296N, or V328M, relative to SEQ ID NO: 6.

[0046] In some embodiments, the first polypeptide and the second polypeptide are linked in a fusion protein.

[0047] In some embodiments, the composition comprises two or more of a first polypeptide, a second polypeptide, a third polypeptide, and a fourth polypeptide.

[0048] In some embodiments, the first polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 7, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 7. In some embodiments, the second polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 8, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 8. In some embodiments, the third polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 9, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 9. In some embodiments, the fourth polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 10.

[0049] In some embodiments, the first polypeptide comprises one or more amino acid substitutions at positions 99, 133, 189, 265, 266, 336, and 343 relative to SEQ ID NO: 7. In some embodiments, the second polypeptide comprises one or more amino acid substitutions at positions 119, 134, 155, 180, 183, 274, 319, 447, 454, 458, 461, 512, 538, and 580 relative to SEQ ID NO: 8. In some embodiments, the third polypeptide comprises one or more amino acid substitutions at positions 28, 82, 144, 151, 162, 182, 273, 327, and 346 relative to SEQ ID NO: 9. In some embodiments, the fourth polypeptide comprises one or more amino acid substitutions at positions 21 and 90 relative to SEQ ID NO: 10.

[0050] In some embodiments, the first polypeptide comprises one or more amino acid substitutions of M99I, S189N, H265Q, A266V, L336F, and V343A relative to SEQ ID NO: 7. In some embodiments, the second polypeptide comprises one or more amino acid substitutions of Y119H, N134R, N134Q, D155N, Q180R, D183N, R274L, N319D, V447I, A454S, E458G, D461N, A512T, D538K, and P580Q relative to SEQ ID NO: 8. In some embodiments, the third polypeptide comprises one or more amino acid substitutions of R28K, A82T, K144, C151R, N162S, K182E, D273G, A327D, and M346I relative to SEQ ID NO: 9. In some embodiments, the fourth polypeptide comprises one or more of the following amino acid substitutions relative to SEQ ID NO:10: A21S and V90A.

[0051] In some embodiments, the first polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 11, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 11. In some embodiments, the second polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 12, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 12. In some embodiments, the third polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 13, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 13. In some embodiments, the fourth polypeptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 14, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 14.

[0052] In some embodiments, the first polypeptide comprises a sequence selected from the group consisting of: 1) a sequence selected from the group consisting of: 2, 3, 7, 9, 11, 12, 14, 16, 20, 26, 29, 32, 34, 35, 40, 43, 45, 46, 54, 61, 64, 65, 70, 77, 101, 103, 105, 106, 108, 109, 111, 119, 120, 123, 126, 127, 130, 131, 148, 149, 151, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, and 408.

[0053] In some embodiments, the second polypeptide is at positions 4, 5, 6, 8, 9, 11, 12, 13, 16, 17, 20, 21, 24, 26, 28, 29, 34, 37, 38, 41, 49, 54, 59, 60, 63, 65, 67, 74, 77, 81, 88, 92, 93, 94, 96, 102, 105, 106, 108, 110, 121, 126, 128, 134, 138, 142, 147, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, , 153, 156, 157, 160, 162, 165, 170, 171, 173, 174, 179, 181, 183, 185, 186, 187, 188, 191, 198, 201, 206, 207, 226, 228, 233, 236, 241, 249, 250, 256, 267, 268, 270, 275, 276, 277, 279, 283, 286, 289, 303, 305, 306, 310, 312, 314, 315, 316, 323, 326, 329, 349, 353, 355, 356, 357, 358, 361, 370, 372, 373, 376, 378, 382, ​​388, 391, 397, 399, 403, 405, 419, 421, 423, 424, 425, 427, 428, 430, 431, 432, 433, 449, 457, 473, 477, 480, 485, 487, 489, 494, 496, 497, 498, 500, 502, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 530, 531, 532, 533, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 11, 515, 518, 519, 520, 540, 545, 550, 555, 557, 570, 571, 580, 583, 585, 590, 594, 603, 607, 608, 611, 617, 620, 624, 636, 639, 641, 642, 644, 646, 655, 658, 660, 663, 665, 668, 672, 673, 678, 682, 685, 688, and 695.

[0054] In some embodiments, the third polypeptide is at positions 5, 10, 11, 26, 30, 35, 40, 42, 45, 46, 47, 58, 61, 65, 71, 72, 75, 77, 78, 80, 82, 83, 94, 98, 113, 115, 116, 117, 121, 128, 133, 138, 146, 148, 161, 171, 175, 177, 182, 184, 191, 193, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 240, 241, 242, 243, 244, 245, 246, 247, 248, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 03, 211, 212, 219, 225, 226, 232, 233, 235, 236, 237, 238, 240, 250, 274, 282, 286, 292, 295, 304, 307, 309, 312, 313, 315, 316, 317, 318, 320, 321, 322, 323, 328, 340, 343, 344, 345, 347, 348, 349, and 350.

[0055] In some embodiments, the third polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 13 and at least one amino acid substitution with a positively charged amino acid. In some embodiments, the positively charged amino acid is arginine. In some embodiments, the at least one amino acid substitution is at position 2, 5, 6, 7, 8, 9, 10, 12, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 64, 65, 66, 67, 68, 69, 70, 222, 224, 225, 227, 228, 229, 231, 232, 233, 234, 235, 255, 256, 257, 258, 277, 286, 287, 337, 338, 339, 340, 345, 346, 347, 348, 349, 350, or a combination thereof. In some embodiments, the at least one amino acid substitution is at positions 346 and 348; 346, 348 and 349; 346, 348, 349, and 350; 350 and 351; 350, 351, and 352; 350, 351, 352, and 353; 235 and 227; 235 and 345; 235 and 346; 235 and 347; 235 and 348; 235 and 349; 235 and 350; 235, 227, and 349; 5, 235 and 346; 5, 235 and 348; 5, 235 and 349; 227, 235, and 346; or 227, 235, and 348.

[0056] In some embodiments, the fourth polypeptide comprises one or more amino acid substitutions at positions 2, 9, 13, 14, 15, 34, 38, 42, 46, 50, 59, 60, 73, 75, 77, 82, 83, 85, 86, 97, 110, 115, 120, 124, 130, 132, 134, 140, 143, 145, 156, 159, 162, 164, 177, 199, 232, and 270 relative to SEQ ID NO:14.

[0057] In some embodiments, the first polypeptide is selected from the group consisting of A2T, F3S, P7R, A9S, A9G, A11G, F12I, D14N, S16Y, Y20H, S26N, F29S, S32N, E34K, G35V, G35S, G35D, I40S, E43D, H45P, E46K, A54S, R61W, V64M, Y65C, N70S, A77T, D101N, K103E, N105K, N105D, S106G, V108M, A109G, Y111N, L119M , R120S, R123S, A126T, E127G, V130M, D131N, Q148R, S149Y, H151Y, A157D, T159I, A164V, L166M, T185A, S194G, A196T, T203A, K211R, E217K, R218K, R218S, N219S, A236T, E242D, N 257K, N267S, M279I, M279V, D283G, N286S, T288I, K291Q, I293V, D296N, S303I, S303G, K3 The amino acid sequence of the present invention includes one or more amino acid substitutions selected from the group consisting of 06N, S310Y, S310P, I313T, Y314F, A316T, E326G, T331I, A336V, A347T, A347S, T352S, Y361H, M374T, M374I, R377G, T395I, S396T, S396F, G398V, and A408V. In some embodiments, the second polypeptide is selected from the group consisting of K4N, E5K, L6M, L6I, E8K, E8D, I9T, D11N, T12A, T13I, D16G, R17C, R17S, R20K, R20E, R21E, R21K, S24K, S24Q, S24R, Y26S, Y26H, A28S, A28D, M29I, G34D, A37S, V38M, V38G, I41V, R49L, D54G, K59R, K60N, K63N, A6 5T, A65V, K67E, K74E, K77E, W81C, K88R, K88E, I92T, R93E, R93K, V94M, K96N, E102D, E102G, T105A, L106M, S108P, V110A, G121 S, S126P, K128R, L134M, Y138S, Q142H, W147L, K150N, V151M, V151L, A153T, S156R, S156G, D157N, K160R, K160E, A162T, S165N,S165G, V170E, K171E, F173V, K174N, K174R, T179A, K181T, S183N, P185T, E186K, E186D, E187K, A188S, A188V, D191Y, D191E, R198H, R198C, R198S, R 201K, D206G, G207D, A226T, I228V, R233K, N236T, R241E, A249S, A250S, I 256T, S267G, S267N, K268N, H270P, S275N, S275G, R276G, A277D, A277S, A2 77T, K279N, G283D, V286G, V289M, G303D, I305T, F306S, A310D, A310T, A3 12G, A312D, A312T, K314N, Q315R, R316G, N323S, E326A, E326K, N329S, G34 9D, E353D, L355M, L355R, E356G, E356D, S357P, A358V, R361S, P370T, N372 K, E373D, S376F, T378I, F382L, M388V, G391S, R397K, A399S, K403N, M405I , L419P, D421N, K423R, H424N, H424R, V425L, I427V, E428K, D430A, D431G , E432D, H433N, A449T, G457D, R473K, E477D, G480D, F485L, S487R, S487G, S489N, N494D, S496N, A497S, V498G, K500N, K502N, Q509R, A511T, A511E, R 515S, R518S, P519T, G520D, G520V, Y540C, Q545H, K550N, K555E, H557Q, P5 In some embodiments, the third polypeptide comprises one or more amino acid substitutions: 70S, E571D, C580R, S583R, E585K, E585G, E590D, R594K, M603I, H607N, H607L, K608R, D611N, L617P, N620S, K624N, T636P, M639V, N641S, V642G, S644N, S644G, E646D, A655V, V658M, K660N, T663A, T665I, R668S, I672V, G673V, S678R, M682L, A685V, A685D, K688N, and V695M.Based on SEQ ID NO: 13, N5K, N5T, D10N, R11K, D26N, V30E, D35N, R40L, P42A, G45S, G45V, F46V, T47R, T47S, N58T, P61L, T65I, T71I, T71R, T71D, L72M, C75S, V77A, P78L, N80T, E82D, H83Y, H83N, A94S, V98M, E113D, C121F, A128S , A155S, E116D, T117I, R133K, G138V, N146D, G148V, C161R, A171V, A171S, K175T, A177V, K182E, L184M, I191 V, S193A, S193F, F201S, S203N, E211K, A212V, Y219R, N225S, N225T, D226Y, E232K, E232Q, A233N, A233S, A23 3K, K235R, Q236R, Q236S, F237L, V238Q, V238M, A240T, A240V, S250A, R274G, A282V, I286N, I286T, I286F, P2 92S, S295N, K304R, E307D, Y309C, A312V, L313M, N315K, N315T, N315S, C316G, I317V, T318A, T318P, K320R, N and one or more amino acid substitutions selected from the group consisting of 321D, E322K, K323N, I328T, M340I, K343E, K343R, K344E, K344R, A345T, A345D, A345S, A345Y, A345R, A345K, A345E, A345G, A347K, A347S, A347D, K348N, K349R, A350K, A350D, A350V, and A350T. In some embodiments, the fourth polypeptide is selected from the group consisting of Q2K, H9L, K13E, Q14K, A15G, K34N, E38K, V42I, A46D, S50I, V59G, Y60H, A73S, A73T, F75L, D77G, G82S, F83L, F83V, F83C, K85E, V86I, E97, I110S, I110L, S115R, K120N, K124R, G130D, D132E, N134T, A140T, E143K, D145G, S156I, E159K, I162V, H164Y, H164F, Y177C, S199I, S232L,and L270S.

[0058] In some embodiments, the first polypeptide comprises one or more amino acid substitutions at positions 105, 109, 131, 148, 279, and 310; or 9, 105, 109, 131, 148, 279, and 310 relative to SEQ ID NO:11. In some embodiments, the second polypeptide comprises one or more amino acid substitutions at positions 134, 179, 185, 540, 555, 624, and 646; 138, 250, 275, and 421; 303, 405, 520, and 590; 134, 179, 185, 540, 555, and 646, 4 and 49; 4 and 388; 4 and 571; 4, 162, and 480; 4 and 315; 5 and 316; 17 and 156; 38, 108, 497 and 583; 59, 157 and 644; 96, 305, 550 and 642; 106, 160 and 228; 312, 424, 449 and 457; or 376 and 611 relative to SEQ ID NO:12. In some embodiments, the third polypeptide comprises one or more amino acid substitutions at positions 30, 46, 240, 304, and 316; 30, 46, 240, and 316; 42 and 318; 184, 240, 315, and 345; 211 and 274; 237 and 237; 286 and 350; 317 and 347; 171, 286, and 315; or 328 and 350 relative to SEQ ID NO:13. In some embodiments, the fourth polypeptide comprises one or more amino acid substitutions at positions 82, 110, 115, 164, and 199; 82, 110, 115, 124, 164, and 199; 110, 115, and 164; 110, 115, 164, and 199; 110, 115, 164, 199, and 124; or 110, 115, 164, 199, and 82 or 124, relative to SEQ ID NO: 14.

[0059] In some embodiments, the composition further comprises one or more Cas proteins, hi some embodiments, the one or more Cas proteins are selected from the group consisting of Cas5, Cas6, Cas7, Cas8, Cas9, Cas11, Cas12, and variants thereof.

[0060] In some embodiments, the composition further comprises at least one unfoldase protein. In some embodiments, the at least one unfoldase protein comprises ClpX.

[0061] Further provided herein are modified Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CRISPR-Tn or CAST) systems, or systems comprising one or more nucleic acids encoding the modified CRISPR-Tn systems. In some embodiments, the CRISPR-Tn system comprises at least one or both of: a) one or more Cas proteins selected from Cas5, Cas6, Cas7, Cas8, Cas9, Cas11, and combinations thereof; and b) one or more transposon-associated proteins selected from TnsA, TnsB, TnsC, TnsD, TniQ, and combinations thereof. In some embodiments, at least one of the one or more Cas proteins comprises Cas6, Cas7, or Cas8 described herein, or at least one of the one or more transposon-associated proteins comprises TnsA, TnsB, TnsC, or TniQ described herein.

[0062] In some embodiments, at least one of the one or more Cas proteins comprises a Cas6 protein comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10 or 14, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 10 or 14; a Cas7 protein comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 9 or 13, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 9 or 13; or a Cas8-Cas5 fusion protein comprising an amino acid sequence that is at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to SEQ ID NO: 8 or 12, but that has one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 8 or 12.In some embodiments, at least one of the one or more transposon-associated proteins is a TnsA protein comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 1 or 4, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 1 or 4; or a TnsA protein having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 2 or 5, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 2 or 5. These include TnsB proteins comprising an amino acid sequence; TnsC proteins comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 3 or 6, and having one or more amino acid substitutions, deletions, or additions based on SEQ ID NO: 3 or 6; or TniQ proteins comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 7 or 11, and having one or more amino acid substitutions, deletions, or additions based on SEQ ID NO: 7 or 11.

[0063] In some embodiments, the TniQ protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:7, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO:7. In some embodiments, the Cas6 protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:10, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO:10, and the Cas7 protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least and / or the Cas8-Cas5 fusion protein comprises an amino acid sequence that is at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to SEQ ID NO: 8, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 8.

[0064] In some embodiments, the TniQ protein comprises an amino acid sequence having one or more amino acid substitutions at positions 99, 133, 189, 265, 266, 336, and 343 relative to SEQ ID NO: 7. In some embodiments, the Cas6 protein comprises amino acids having one or more amino acid substitutions at positions 21 and 90 relative to SEQ ID NO: 10. In some embodiments, the Cas7 protein comprises amino acids having one or more amino acid substitutions at positions 28, 82, 144, 151, 162, 182, 273, 327, 346 relative to SEQ ID NO: 9. In some embodiments, the Cas8-Cas5 fusion protein comprises amino acids having one or more amino acid substitutions at positions 119, 134, 155, 180, 183, 274, 319, 447, 454, 458, 461, 512, 538, and 580 relative to SEQ ID NO: 8.

[0065] In some embodiments, the TniQ protein comprises an amino acid sequence having one or more amino acid substitutions of M99I, S189N, H265Q, A266V, L336F, and V343A relative to SEQ ID NO: 7. In some embodiments, the Cas6 protein comprises an amino acid sequence having one or more amino acid substitutions of A21S and V90A relative to SEQ ID NO: 10. In some embodiments, the Cas7 protein comprises an amino acid sequence having one or more amino acid substitutions of R28K, A82T, K144, C151R, N162S, K182E, D273G, A327D, and M346I relative to SEQ ID NO: 9. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO:8: Y119H, N134R, N134Q, D155N, Q180R, D183N, R274L, N319D, V447I, A454S, E458G, D461N, A512T, D538K, and P580Q.

[0066] In some embodiments, the TniQ protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 11, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 11. In some embodiments, the Cas6 protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 14, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 14. In some embodiments, the Cas7 protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 13, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 13. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 12, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 12.

[0067] In some embodiments, the TniQ protein is located at positions 2, 3, 7, 9, 11, 12, 14, 16, 20, 26, 29, 32, 34, 35, 40, 43, 45, 46, 54, 61, 64, 65, 70, 77, 101, 103, 105, 106, 108, 109, 111, 119, 120, 123, 126, 127, 130, 131, 148, 149, 151, 157, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, and 408. In some embodiments, the Cas6 protein comprises an amino acid sequence having one or more amino acid substitutions at positions 2, 9, 13, 14, 15, 34, 38, 42, 46, 50, 59, 60, 73, 75, 77, 82, 83, 85, 86, 97, 110, 115, 120, 124, 130, 132, 134, 140, 143, 145, 156, 159, 162, 164, 177, 199, 232, and 270 relative to SEQ ID NO: 14. In some embodiments, the Cas7 protein is located at positions 5, 10, 11, 26, 30, 35, 40, 42, 45, 46, 47, 58, 61, 65, 71, 72, 75, 77, 78, 80, 82, 83, 94, 98, 113, 115, 116, 117, 121, 128, 133, 138, 146, 148, 161, 171, 175, 177, 182, 184, 191, 193, 201, 203, 21 and amino acid sequences having one or more amino acid substitutions at positions 1, 212, 219, 225, 226, 232, 233, 235, 236, 237, 238, 240, 250, 274, 282, 286, 292, 295, 304, 307, 309, 312, 313, 315, 316, 317, 318, 320, 321, 322, 323, 328, 340, 343, 344, 345, 347, 348, 349, and 350.In some embodiments, the Cas8-Cas5 fusion protein is located at positions 4, 5, 6, 8, 9, 11, 12, 13, 16, 17, 20, 21, 24, 26, 28, 29, 34, 37, 38, 41, 49, 54, 59, 60, 63, 65, 67, 74, 77, 81, 88, 92, 93, 94, 96, 102, 105, 106, 108, 110, 121, 126, 128, 134, 138, 142, 147, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 205, 206, 207, 208, 209, 210, 211, 212, 21 151, 153, 156, 157, 160, 162, 165, 170, 171, 173, 174, 179, 181, 183, 185, 186, 187, 188, 191, 198, 201, 206, 207, 226, 228, 233, 236, 241, 249, 250, 256, 267, 268, 270, 275, 276, 277, 279, 283, 286, 289, 303, 305, 306, 310, 312, 314, 315, 316, 32 3, 326, 329, 349, 353, 355, 356, 357, 358, 361, 370, 372, 373, 376, 378, 382, ​​388, 391, 397, 399, 403, 405, 419, 421, 423, 424, 425, 427, 428, 430, 431, 432, 433, 449, 457, 473, 477, 480, 485, 487, 489, 494, 496, 497, 498, 500, 502, 509, 511, 515 , 518, 519, 520, 540, 545, 550, 555, 557, 570, 571, 580, 583, 585, 590, 594, 603, 607, 608, 611, 617, 620, 624, 636, 639, 641, 642, 644, 646, 655, 658, 660, 663, 665, 668, 672, 673, 678, 682, 685, 688, and 695.

[0068] In some embodiments, the TniQ protein is selected from the group consisting of A2T, F3S, P7R, A9S, A9G, A11G, F12I, D14N, S16Y, Y20H, S26N, F29S, S32N, E34K, G35V, G35S, G35D, I40S, E43D, H45P, E46K, A54S, R61W, V6 4M, Y65C, N70S, A77T, D101N, K103E, N105K, N105D, S106G, V108M, A109G, Y111N, L119M, R1 20S, R123S, A126T, E127G, V130M, D131N, Q148R, S149Y, H151Y, A157D, T159I, A164V, L166M , T185A, S194G, A196T, T203A, K211R, E217K, R218K, R218S, N219S, A236T, E242D, N257K, N 267S, M279I, M279V, D283G, N286S, T288I, K291Q, I293V, D296N, S303I, S303G, K306N, S310 and an amino acid sequence having one or more amino acid substitutions selected from the group consisting of Y, S310P, I313T, Y314F, A316T, E326G, T331I, A336V, A347T, A347S, T352S, Y361H, M374T, M374I, R377G, T395I, S396T, S396F, G398V, and A408V. In some embodiments, the Cas6 protein is selected from the group consisting of Q2K, H9L, K13E, Q14K, A15G, K34N, E38K, V42I, A46D, S50I, V59G, Y60H, A73S, A73T, F75L, D77G, G82S, F83L, F83V, F83C, K85E, V86I, E97, I1 In some embodiments, the Cas7 protein comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 13: N5K, N5T, D10N, R11K, D26N, V30E, D35N, R40L, P42A, G45S, G45V, F46V, T47R, T47S, ...N58T, P61L, T65I, T71I, T71R, T71D, L72M, C75S, V77A, P78L, N80T, E82D, H83Y, H83N, A94S, V98M, E113D, C121F, A128S, A155S, E116D, T117I, R133K, G138V, N146D, G148V, C161R, A171V, A171S, K1 75T, A177V, K182E, L184M, I191V, S193A, S193F, F201S, S203N, E211K, A212V, Y219R, N225S, N225 T, D226Y, E232K, E232Q, A233N, A233S, A233K, K235R, Q236R, Q236S, F237L, V238Q, V238M, A240T, A240V, S250A, R274G, A282V, I286N, I286T, I286F, P292S, S295N, K304R, E307D, Y309C, A312V, L3 13M, N315K, N315T, N315S, C316G, I317V, T318A, T318P, K320R, N321D, E322K, K323N, I328T, M340 and an amino acid sequence having one or more amino acid substitutions selected from the group consisting of I, K343E, K343R, K344E, K344R, A345T, A345D, A345S, A345Y, A345R, A345K, A345E, A345G, A347K, A347S, A347D, K348N, K349R, A350K, A350D, A350V, and A350T. In some embodiments, the Cas8-Cas5 fusion protein comprises any of the following sequences, based on SEQ ID NO: 12: K4N, E5K, L6M, L6I, E8K, E8D, I9T, D11N, T12A, T13I, D16G, R17C, R17S, R20K, R20E, R21E, R21K, S24K, S24Q, S24R, Y26S, Y26H, A28S, A28D, M29I, G34D, A37S, V38M, V38G , I41V, R49L, D54G, K59R, K60N, K63N, A65T, A65V, K67E, K74E, K77E, W81C, K88R, K88E, I92T, R93E, R93K, V94M , K96N, E102D, E102G, T105A, L106M, S108P, V110A, G121S, S126P, K128R, L134M, Y138S, Q142H, W147L, K150N,V151M、V151L、A153T、S156R、S156G、D157N、K160R、K160E、A162T、S165N、S1 65G、V170E、K171E、F173V、K174N、K174R、T179A、K181T、S183N、P185T、E186 K、E186D、E187K、A188S、A188V、D191Y、D191E、R198H、R198C、R198S、R201K、 D206G、G207D、A226T、I228V、R233K、N236T、R241E、A249S、A250S、I256T、S26 7G, S267N, K268N, H270P, S275N, S275G, R276G, A277D, A277S, A277T, K279N, G283D, V286G, V289M, G303D, I305T, F306S, A310D, A310T, A312G, A312D, A 312T, K314N, Q315R, R316G, N323S, E326A, E326K, N329S, G349D, E353D, L355M, L355R, E356G, E356D, S357P, A358V, R361S, P370T, N372K, E373D, S376F T378I, F382L, M388V, G391S, R397K, A399S, K403N, M405I, L419P, D421N, K423R, H424N, H424R, V425L, I427V, E428K, D430A, D431G, E432D, H433N, A449 T、G457D、R473K、E477D、G480D、F485L、S487R、S487G、S489N、N494D、S496N、 A497S、V498G、K500N、K502N、Q509R、A511T、A511E、R515S、R518S、P519T、G52 0D、G520V、Y540C、Q545H、K550N、K555E、H557Q、P570S、E571D、C580R、S583R 、E585K、E585G、E590D、R594K、M603I、H607N、H607L、K608R、D611N、L617P、N 620S、K624N、T636P、M639V、N641S、V642G、S644N、S644G、E646D、A655V、V65 8M、K660N、T663A、T665I、R668S、I672V、G673V、S678R、M682L、A685V、A685D、It comprises an amino acid sequence having one or more amino acid substitutions of K688N, K688N, and V695M.

[0069] In some embodiments, the TniQ protein comprises an amino acid sequence with amino acid substitutions at positions 105, 109, 131, 148, 279, and 310, or 9, 105, 109, 131, 148, 279, and 310, relative to SEQ ID NO: 11. In some embodiments, the Cas6 protein comprises an amino acid sequence with amino acid substitutions at positions 110, 115, 164, 199, and 82 or 124, relative to SEQ ID NO: 14. In some embodiments, the Cas7 protein comprises an amino acid sequence having amino acid substitutions at positions 30, 46, 240, 304, and 316; 30, 46, 240, and 316; 42 and 318; 184, 240, 315, and 345; 211 and 274; 237 and 237; 286 and 350; 317 and 347; 171, 286, and 315; or 328 and 350 relative to SEQ ID NO: 13. In some embodiments, the Cas8-Cas5 fusion protein comprises an amino acid sequence having amino acid substitutions at positions 134, 179, 185, 540, 555, 624, and 646; 138, 250, 275, and 421; 303, 405, 520, and 590; 134, 179, 185, 540, 555, and 646, 4 and 49; 4 and 388; 4 and 571; 4, 162, and 480; 4 and 315; 5 and 316; 17 and 156; 38, 108, 497 and 583; 59, 157 and 644; 96, 305, 550 and 642; 106, 160 and 228; 312, 424, 449 and 457; or 376 and 611 relative to SEQ ID NO: 12.

[0070] In some embodiments, the Cas7 protein comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 13 and at least one amino acid substitution with a positively charged amino acid. In some embodiments, the positively charged amino acid is arginine. In some embodiments, the at least one amino acid substitution is at position 2, 5, 6, 7, 8, 9, 10, 12, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 64, 65, 66, 67, 68, 69, 70, 222, 224, 225, 227, 228, 229, 231, 232, 233, 234, 235, 255, 256, 257, 258, 277, 286, 287, 337, 338, 339, 340, 345, 346, 347, 348, 349, 350, or a combination thereof. In some embodiments, the at least one amino acid substitution is at positions 346 and 348; 346, 348 and 349; 346, 348, 349, and 350; 350 and 351; 350, 351, and 352; 350, 351, 352, and 353; 235 and 227; 235 and 345; 235 and 346; 235 and 347; 235 and 348; 235 and 349; 235 and 350; 235, 227, and 349; 5, 235 and 346; 5, 235 and 348; 5, 235 and 349; 227, 235, and 346; or 227, 235, and 348.

[0071] In some embodiments, the system comprises a TnsA protein and a TnsB protein. In some embodiments, the TnsA protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 1, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 1. In some embodiments, the TnsB protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 2, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 2.

[0072] In some embodiments, the TnsA protein comprises an amino acid sequence having one or more amino acid substitutions at positions 2, 3, 5, 28, 57, 77, 80, 107, 110, 116, 122, 142, 155, 161, 166, 173, 177, 185, 211, 216, 227, and 230 relative to SEQ ID NO: 1. In some embodiments, the TnsB protein comprises an amino acid sequence having one or more amino acid substitutions at positions 2, 5, 22, 24, 25, 29, 75, 141, 199, 215, 319, 347, 364, 370, 383, 439, 454, 458, 485, 509, 533, 538, 565, 581, 586, 595, 596, 597, and 600 relative to SEQ ID NO: 2.

[0073] In some embodiments, the TnsA protein comprises an amino acid sequence having one or more amino acid substitutions of A2T, T3I, L5S, T28A, A57T, F77L, Y80D, K107M, K107R, Y110C, Y110D, D116G, E122A, D142E, M155I, K161R, N166D, K173E, Y177N, Y177D, C185R, D211Y, K216E, A227P, G230D, and G230S relative to SEQ ID NO:1. In some embodiments, the TnsB protein comprises an amino acid sequence having one or more amino acid substitutions of A2T, A2S, G5R, S22P, E24D, L25I, A29S, P75T, I141T, V199I, S215R, D319V, Y347F, S364N, E370K, N383D, V439A, E454D, E454G, S458N, V485F, R509G, D533A, A538V, H565Y, A581T, H586L, N595K, D596N, D597N, D597Y, and I600V relative to SEQ ID NO:2.

[0074] In some embodiments, the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 2 and 230; 107 and 166; one or both of 107, 166, and 2 and 227; 211 and 110 or 142; 110, 155 and 230; 122 and 155; or 155 and 177, relative to SEQ ID NO:1. In some embodiments, the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 2 and 597; 24 and 25; 24, 25, 458, 509, 565, and 600; 75 and 597; 141, 454, 533 and 595; 581, 370, and 454; 370 and 581; 370 and 454; 458 and 509; 458, 509 and 565; 458, 509, 565, and 600; 565, 586, and 596; or 565, 509, 458, 600 and 24, 25, 29, 215, 319, 364, 383, and 586 relative to SEQ ID NO:2.

[0075] In some embodiments, the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 107, 166, and 227 relative to SEQ ID NO: 1, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 24, 25, 458, 509, 565, and 600 relative to SEQ ID NO: 2; or the TnsA protein comprises an amino acid sequence having an amino acid substitution at position 155 relative to SEQ ID NO: 1, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 22, 347, and 454 relative to SEQ ID NO: 2; or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 122 and 155 relative to SEQ ID NO: 1, and the TnsB protein comprises an amino acid sequence having an amino acid substitution at position 485 relative to SEQ ID NO: 2.

[0076] In some embodiments, the TnsA protein comprises an amino acid sequence having the amino acid substitutions K107M, N166D, and A227P relative to SEQ ID NO: 1, and the TnsB protein comprises an amino acid sequence having the amino acid substitutions E24D, L25I, S458N, R509G, H565Y, and I600V relative to SEQ ID NO: 2; or the TnsA protein comprises an amino acid sequence having the amino acid substitution M155I relative to SEQ ID NO: 1, and the TnsB protein comprises an amino acid sequence having the amino acid substitutions S22P, Y347F, and E454G relative to SEQ ID NO: 2; or the TnsA protein comprises an amino acid sequence having the amino acid substitutions E122A and M155I relative to SEQ ID NO: 1, and the TnsB protein comprises an amino acid sequence having the amino acid substitution V485F relative to SEQ ID NO: 2.

[0077] In some embodiments, the system further includes a TnsC protein comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 3, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 3. In some embodiments, the TnsC protein includes an amino acid sequence with one or more amino acid substitutions at positions 9, 15, 16, 18, 21, 64, 81, 86, 87, 99, 109, 142, 147, 153, 168, 180, 216, 230, 285, and 304 relative to SEQ ID NO: 3. In some embodiments, the TnsC protein comprises an amino acid sequence with one or more amino acid substitutions of I9V, A15V, F16Y, S18F, S21N, N64D, H81Y, D86Y, N87K, V99I, E109D, E142K, V147I, N153D, I168M, A180E, A216S, L230F, K285E, and R304R relative to SEQ ID NO: 3. In some embodiments, the TnsC protein comprises an amino acid sequence with amino acid substitutions at positions 142 and 216 relative to SEQ ID NO: 3.

[0078] In some embodiments, the TnsA protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 4, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 4. In some embodiments, the TnsB protein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 5, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 5.

[0079] In some embodiments, the TnsA protein is located at positions 4, 5, 9, 10, 12, 21, 23, 25, 26, 31, 32, 34, 35, 37, 41, 45, 47, 48, 51, 52, 55, 60, 61, 65, 67, 69, 72, 75, 79, 80, 82, 87, 88, 90, 91, 93, 94, 96, 98, 99, 100, 103, 106, 108, 113, 116, 125, 126, 127, 128, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, and amino acid sequences having one or more amino acid substitutions at positions 28, 129, 135, 139, 143, 146, 147, 149, 153, 154, 156, 158, 159, 160, 162, 164, 166, 167, 168, 169, 170, 177, 179, 180, 182, 183, 185, 187, 188, 190, 191, 192, 193, 195, 196, 200, 204, 207, and 208. In some embodiments, the TnsB protein is located at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 49, 52, 55, 56, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 106, 107, 108, 109, 110, 111, 120, 130, 140, 141, 142, 143, 144, 145, 149, 152, 155, 156, 158, 160, 161, 162, 163, 164, 165, 166, 167, 168, 170, 171, 172, 173, 1 , 60, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 110, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 145 , 146, 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169, 173, 174, 176, 181, 182, 186, 187, 190, 195, 197, 198, 205, 208, 209, 211, 215, 218, 223, 226, 227, 231, 232, 235, 239, 246 , 248, 250, 259, 260, 261, 262, 263, 267, 269, 273, 274, 277, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298, 302, 303, 307, 313, 316, 317, 320, 323, 325, 331, 332, 339, 345, 348, 349, 352,353, 354, 356, 361, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373, 375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 404, 407, 408, 410, 411, 412, 413, 414, 415, 416, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456, 459, 460, 463, 464, 470, 472, 473, 494, 495, 498, 501, 502, 504, 505, 506, 508, 509, 510, 512, 513, 514, 517, 520, 521, 522, 525, 526, 527, 530, 531, 532, 533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 567, 568, 569, 570, 571, and amino acid sequences having one or more amino acid substitutions at 574, 575, 576, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656.

[0080] In some embodiments, the TnsA protein is selected from the group consisting of R4K, N5K, P9S, A10P, N12D, T21I, V23M, S25N, S25R, V26M, V26G, S31N, S32I, E34A, F35L, A37D, H41L, D45N, I47V, E48G, G51V, S52I, E55K, E55D, E60K, F61L, S65T, S65A, P67T, P67L, P67S, P 67H, T69A, A72V, A72D, S75I, S75R, S75T, K79E, T80P, K82E, K87R, P88L, P88T, P88A, S90F, K91N, K91E, A93T, A93S, S94N, L96P, R98Q, A99D, A99V, E100K, A103T, A106T, S108A, I113F, V116F, V116I, V125M, V125A, N126T, I128V, I128 L, L129P, L135M, S139N, S139G, G143V, G143C, G146D, G146S, I147V, K149E, K149T, K149R, S153I, S153R, S153N, F1 54C, H156R, H156L, S158N, S158R, G159V, V160A, K162R, N164D, I166L, S167I, S168I, S168R, S168N, Q169R, V170M, The amino acid sequence includes an amino acid sequence having one or more amino acid substitutions selected from the group consisting of V170G, V170L, T177I, T177A, S179R, F180C, F180L, F182C, F182L, G183S, M185I, K187R, G188D, V190I, K191N, A192S, D193N, G195V, G195D, G195S, C196W, T200A, T204I, A207V, A207T, and T208I. In some embodiments, the TnsB protein is selected from the group consisting of M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L, L15Q, K16N, K16R, P17T, P17L, P17S, T19I, T19S, T19A, T19P, P20S, P20L, T21A, Q22R, Y23H, V24M, K25R, L26M, D27A, D27G, D28N, D28Y,A29T、A29V、N30K、I32F、I32S、Q33H、L36M、D37A、D37Y、F39L、S40P、D41E、T42I、T42K、T42A、F43L、F43S、F43V、K44N、N45D、N45S、Q49R、K52Q、S55A、T56A、D58E、K60Q、S62T、R63K、R63G、Q67R、Q67H、Q67K、D71Y、K74R、E76K、F78C、K79R、G80V、G80D、G81S、G81V、G81D、D82N、V83G、V83M、V83A、V84A、V84G、R85G、R85K、P86L、N87S、R89C、V91G、V91A、A92V、A92T、R95K、K97R、E100D、S101A、D104V、A106D、A106T、D110N、N112H、H113Y、M115R、N117Y、T119A、N120D、N120K、N120S、G124V、D125N、D125E、K127R、F129L、D130N、K131M、E134D、E134G、A139S、A139T、P142S、I144V、A145S、A145T、T146A、A147V、Q149R、Y150H、I155L、V156A、V156L、V156M、K157V、E158A、N159S、V163G、E164A、E164G、E164D、G165D、I167V、I169L、I169T、N173S、N173H、N173T、A174S、A174T、N176D、A181S、I182L、I182V、I182T、A186E、A186T、V187G、V187A、A190T、A190S、F195S、A197P、D198G、D198N、A205S、V208M、P209T、T211I、E215D、E218D、P223S、P223H、L226V、I227V、D231N、E232K、I235V、I235T、R239G、I246V、V248E、V248M、S250I、S259N、Y260C、K261R、S262N、P263L、S267N、A269V、T273I、T273N、H274Y、K277N、K277R、P278S、S280T、L281M、D282E、D282N、A283T、A283S、N285S、E287D、L288M、N290K、F295S、F298I、F298S、V302I、V303M、A307S、N313S, H316R, A317V, S320N, S320R, I323L, I325V, R331K, K332E, I339V, V345L, V345M, E348K, Y349H, Y349D, Y349N, Y349C, P352S, P352T, E353Q, E353 D, L354M, G356S, N361D, I362V, I362T, L363P, L363T, L363M, E364G, K365R, E366G, E367G, K369N, K369E, K369M, P370S, E371K, V372M, D373G, I375V, M37 6I、T380P、T380A、E383K、E383D、F385L、H386Y、I389V、A390V、A390I、V392I 、D396N、D396G、D396K、S397P、S399N、S399G、T402I、R403G、R403I、R403K、R 403S、I404T、I404V、K407R、K407E、R408K、Q410K、Q410H、Q410R、Q411H、G41 2V、F413L、D414N、A415V、A415T、Y416C、M421I、N422K、E423K、E423D、E424A、 E425K、E426D、T427A、T427S、R428K、F429L、S430A、M431L、R434H、R434C、R4 34S、I435V、D437G、D437N、T440S、T440I、R443C、G445S、F446L、F446I、Y448 C、E450D、E450G、M452I、T456P、T456A、T456I、A459T、D460N、K463N、H464N、 H464R、H464S、E470K、V472M、V472A、K473D、K473N、E494D、E494G、S495A、E49 8A、E498K、C501Y、T502I、T502S、P504S、P504L、T505A、G506Y、G506D、G506L 、G506S、T508A、D509E、D509Y、C510Y、S512N、I513L、I513V、I513F、Y514H、K 517M, K517N, K517Q, K520R, K521N, I522T, I522V, I522F, E525K, V526E, V526M, I527V, S530N, S530R, K531T, D532G, D532Y, S533Y, G535D, A537T, K538RK538N, R540K, R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K , Q549R, E550A, Q551K, E552D, E552K, V553I, F554V, E556K, E556G, S557A, K558R, T559P, T559I, T559A , K560R, A561T, A561G, K562R, K562N, I563L, T564I, A565S, A565V, K567R, K568N, K568R, Q569K, Q569 L, Q569R, A570V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S, I583V, K584R, V585 M, S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S589R, S589N, A590S, A590T, A591V, P59 2L, V593M, V593A, Q594L, K595R, K595N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N, D60 and A656V.

[0081] In some embodiments, the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 108 and 47 or 208; 170 and 207; 88 and 147; 47, 88 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 88, 128, 147, 170, and 182; or 170, 207, and 108, relative to SEQ ID NO:4. In some embodiments, the TnsB protein is located at positions 4, 23, and 590; 19, 169, and 549; 43 and 415; 80 and 593; 80, 144, 593, and 606; 1, 42, 80, 593, and 606; 42, 80, 593, and 606; 156 and 604; 283, 349, and 365; 283, 349, 365, 396, and 594; 283, 349, 365, 396, 594, 596 , and 131; 352 and 390; 390, 396, and 594; 396 and 594; 456 and 502; 464 and 502; 464 and 17; 17, 235, 464, and 596; 235, 352, 396, 456, and 606; 415, 456, and 502; 456, 502, and 549; 169, 456, 502, and 549; 80, 456, 502, 593, and 606; 1, 42, 80, 456, 502, 593, and 60 6; 80, 144, 456, 502, 593, and 606; 19, 169, 456, 502, and 549; 43, 415, 456, and 502; 352, 390, 396, and 594; 352, 390, and 396; 283, 349, 396, and 594; 11, 55, 120, 362, 584, 600, and 604; 43, 84, 144, 349, and 517; 164 and 165; 164 and 173; 362 and 446; 352, 39 one or more positions selected from: 0, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526; 43, 352, 390, 396, 464, 549, and 594;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 21; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 67; 21, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 174, 208, 427, 456, and 504; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 139; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 339, and 446; 43, 349, 352, 390, 396, 46 4, 549, 594, 410, 526, 19, 460, 569, and 596; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 460, 586, 588, and 608; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, and 460; 352, 390, 396, 549, 586, and 594; 63, 158, 352, 390, 396, 549, 586, and 594; 164, or amino acid sequences having amino acid substitutions at 283, 349, 365, 396, 594, 596, and 131;

[0082] In some embodiments, the TnsA protein comprises the amino acid sequence of SEQ ID NO: 4, and the TnsB protein comprises one or more positions selected from positions 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526 relative to SEQ ID NO: 5; or amino acid sequences having amino acid substitutions at 4, 549, 594, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 21; or 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, 21, and 67. In some embodiments, the TnsB protein is selected from the group consisting of: 43, 349, 352, 390, 396, 464, 549, 594, and 456; 43, 349, 352, 390, 396, 464, 549, 594, 456, and 526; 43, 349, 352, 390, 396, 464, 549, 594, and 5 04;43, 349, 352, 390, 396, 464, 549, 594, and 526;43, 349, 352, 390, 396, 464, 549, 594, 410, and 526;43, 349, 352, 390, 396, 464, 549, 594, 174, and 427;43, 349, 352, 39 0, 396, 464, 549, 594, and 208;43, 349, 352, 390, 396, 464, 549, 594, 63, 145, 182, and 526;43, 349, 352, 390, 396, 464, 549, 594, 415, and 502;43, ​​349, 352, 390, 396, 464, 549, 594, 415, 502 and 67; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502 and 21; or amino acid sequences having amino acid substitutions at 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, 21 and 67.

[0083] In some embodiments, the TnsA protein comprises the amino acid sequence of SEQ ID NO: 4, and the TnsB protein comprises one or more substitutions selected from F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, and R63G, A145S, A174S, I182R, V208M, Q410K, T427S, T456I or T456P, P504S, and V526E; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V and T502I; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and T21A; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K; or an amino acid sequence having the amino acid substitutions F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, T21A, and Q67K.In some embodiments, the TnsB protein has the following sequence: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, and T456I; F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, T456P, and V526E; F43S, Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, and P504S; F43S, Y349N , P352T, A390V, D396N, H464R, Q549R, Q594L, and V526E; F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, Q410K, and V526E; F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, A174S, and T427S; F43S, Y349N, P352T, A390V, D396N, H464R, Q5 49R, Q594L, and V208M; F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, R63G, A145S, I182T, and V526E; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, and T502I; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q54 and T21A; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K; or amino acid sequences having the amino acid substitutions F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, T21A, and Q67K.

[0084] In some embodiments, the TnsA protein comprises an amino acid sequence with an amino acid substitution at position 182 relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence with amino acid substitutions at positions 352, 390, 396, 594, and 596 relative to SEQ ID NO:5; or the TnsA protein comprises an amino acid sequence with amino acid substitutions at positions 88, 147, and 177 relative to SEQ ID NO:4, and the TnsB protein comprises amino acid substitutions at positions 352, 390, 396, 549, and 594 relative to SEQ ID NO:5. The TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88 and 147 relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5, or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 116, and 147 relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 289, 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5. The TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 549, 594, and 596, or the TnsA protein comprises an amino acid sequence having an amino acid substitution at position 75 relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 235, 352, 390, 396, 567, and 594 relative to SEQ ID NO:5, or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 235, 352, 390, 396, 567, and 594 relative to SEQ ID NO:5. the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 352, 363, 390, 396, 549, 586, and 594 relative to SEQ ID NO:4; the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, 182, and 51 or 180 relative to SEQ ID NO:4; the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO:5; the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 75, 76, 78, 79, 80, 81, 82, and 83 or 84 relative to SEQ ID NO:4;The TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 93, and 147 relative to SEQ ID NO: 5, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 352, 390, 396, 549, and 594 relative to SEQ ID NO: 5, or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 93, and 147 relative to SEQ ID NO: 4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 352, 390, 396, 549, 580, and 594 relative to SEQ ID NO: 5.

[0085] In some embodiments, the TnsA protein comprises an amino acid sequence with the amino acid substitution F182L relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence with the amino acid substitutions P352T, A390V, D396N, Q594L, and H596Y relative to SEQ ID NO:5; or the TnsA protein comprises an amino acid sequence with the amino acid substitutions P88T, I147V, and T177I relative to SEQ ID NO:4, and the TnsB protein comprises the amino acid substitutions P352S, A390V, D396N, Q549R, and Q594L, or the TnsA protein comprises an amino acid sequence with the amino acid substitutions P88T and I147V relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence with the amino acid substitutions P352T, A390V, D396N, H464R, Q549R, and Q594L relative to SEQ ID NO:5, or the TnsA protein comprises an amino acid sequence with the amino acid substitutions P88T, V116I, and I147V relative to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence with the amino acid substitutions Q28 the TnsA protein comprises an amino acid sequence having the amino acid substitutions S75I, S75I, P352T, A390V, D396N, Q549R, Q594L, and H596Y, based on SEQ ID NO: 4; the TnsB protein comprises an amino acid sequence having the amino acid substitutions I235T, P352T, A390V, D396N, K567R, and Q594L, based on SEQ ID NO: 5; or the TnsA protein comprises an amino acid sequence having the amino acid substitutions P88T, I147V, V170L, and F182L, based on SEQ ID NO: 4. The TnsB protein comprises an amino acid sequence having the amino acid substitutions P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L based on SEQ ID NO: 5; the TnsA protein comprises an amino acid sequence having the amino acid substitutions P88T, I147V, V170L, F182L, and G51V or F180L based on SEQ ID NO: 4; the TnsB protein comprises an amino acid sequence having the amino acid substitutions F43S, Y349N, P352T, A390V, D396N, Q410K, H464R, V526E, Q549R,and Q594L, or the TnsA protein comprises an amino acid sequence with the amino acid substitutions S75I, P88T, and I147V based on SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence with the amino acid substitutions P352T, A390V, D396N, Q549R, and Q594L based on SEQ ID NO:5, or the TnsA protein comprises an amino acid sequence with the amino acid substitutions P88T, A93T, and I147V based on SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence with the amino acid substitutions P352T, A390V, D396N, Q549R, T580I, and Q594L based on SEQ ID NO:5.

[0086] In some embodiments, the system further includes a TnsC protein comprising an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 6, with one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 6. In some embodiments, the TnsC protein is located at positions 1, 2, 3, 5, 6, 7, 9, 11, 12, 14, 21, 22, 26, 27, 31, 35, 38, 43, 44, 46, 47, 54, 59, 60, 61, 64, 65, 67, 68, 71, 72, 74, 76, 79, 80, 81, 84, 89, 95, 102, 105, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 1 , 110, 111, 112, 113, 114, 116, 118, 119, 120, 123, 129, 130, 131, 132, 134, 142, 145, 146, 147, 148, 150, 154, 155, 166, 169, 178, 180, 181, 183, 184, 187, 190, 194, 197, 201, 204, 207, 209, 213, 21 9, 221, 225, 226, 227, 229, 232, 233, 234, 236, 238, 241, 246, 251, 252, 256, 257, 261, 263, 265, 267, 269, 271, 272, 274, 280, 281, 285, 286, 288, 291, 292, 296, 299, 301, 303, 304, 306, 307, 308, and amino acid sequences having one or more amino acid substitutions at 310, 313, 314, 316, 317, 318, 319, 320, 323, 324, 326, 328, 330, 331, 332, 340, 341, 343, 344, 355, 412, 418, 427, 514, 1198, 1201, 1206, 1212, 1260, and 1282. In some embodiments, the TnsC protein is selected from the group consisting of M1L, M1V, N2S, A3T, T5P, T5A, T5S, E6D, I7S, I7V, I9F, Q11R, L12M, N14D, N14S, M21I, H22P, H22Y, K26N, K26R, T27I, M31I, L35R, N38S, S43P, D44N, D44G,Q46L、C47S、T54I、S59T、H60Y、T61A、H64Y、Y65H、K67N、K67R、R68Q、A71G、T72A、N74D、S76C、S76Y、T79I、M80I、P81S、V84L、R89L、A95D、A95T、A102T、E105D、E105K、S109N、S109R、S110P、Q111R、I112T、K113N、K113E、K114N、K114M、K114E、G116D、K118N、K118R、T119I、D120V、K123N、L129M、I130V、K131R、A132S、K134M、K134N、F142V、L145M、I146T、E147K、F148S、S150F、R154K、Q155H、E166D、K169E、P178S、A180V、A181T、A181S、I183V、A184S、A184T、A184V、P187S、A190T、A190V、V194M、V194A、R197I、Y201N、L204M、D207N、K209N、Q213H、Q213V、A219S、K221N、D225N、V226E、P227T、K229E、S232N、K233N、K233R、N234H、T236A、A238V、A238S、A241S、E246D、K251N、H252Y、H252R、E256D、A257S、A261V、S263I、S263N、N265D、Y267C、E269K、E269D、K271E、K271R、H272Y、I274V、F280L、D281N、D281G、K285G、K286N、K288R、S291F、S291P、K292N、K296R、K296N、I299S、D301G、E303D、I304T、I304V、E306G、V307L、V307G、V307A、V307D、V307G、I308N、N310S、Y313H、N314K、N316K、N316D、A317D、L318Q、D319N、P320S、P320L、M323I、L324M、D326N、V328M、V328A、A330D、I331V、V332G、S340L、T341A、A343G、S344N、I355V、F412V、V418F、Y427C、R514K、S1198L、A1201V、G1206S、C1212G、F1260L、and V1282M. In some embodiments, the TnsC protein comprises an amino acid sequence having one or more amino acid substitutions at positions 2, 67, 95, and 226; 6 and 316; 38, 95, 303; 67, 95, and 226; 44 and 76; 44, 76, and 118; 130, 234, 303; 118 and 1201; 118, 1201, and 44; 118, 1201, and 76; 130, 234, and 303; 154 and 269; 221 and 44; 44, 76, 130, 234, and 303; 44, 76, 118; 8, and 1201; 197 and 314; 76, 181, and 194; 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; 59, 76, 306, and 316; or 280 and 340.

[0087] In some embodiments, the TnsA protein comprises an amino acid sequence with amino acid substitutions at positions 88 and 147 relative to SEQ ID NO:4; the TnsB protein comprises an amino acid sequence with amino acid substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5; and the TnsC protein comprises an amino acid sequence with amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 76 and 7, 12, or 263; or 76, 238, 296, or 328 relative to SEQ ID NO:6; or the TnsA protein comprises an amino acid sequence with amino acid substitutions at positions 88 and 147 relative to SEQ ID NO:4; and the TnsB protein comprises an amino acid sequence with amino acid substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5; and the TnsC protein comprises an amino acid sequence with amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 76 and 7, 12, or 263; or 76, 238, 296, or 328 relative to SEQ ID NO:6. and the TnsC protein comprises an amino acid sequence having amino acid substitutions at positions 76, 181, and 194 relative to SEQ ID NO: 6; or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO: 4; the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 352, 363, 390, 396, 549, 586, and 594 relative to SEQ ID NO: 5; and the TnsC protein comprises an amino acid sequence having amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 76 and 7, 12, or 263 relative to SEQ ID NO: 6.or an amino acid sequence having amino acid substitutions at positions 76, 238, 296, or 328, or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, and 182 relative to SEQ ID NO:4; the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 352, 363, 390, 396, 549, 586, and 594 relative to SEQ ID NO:5; and the TnsC protein comprises an amino acid sequence having amino acid substitutions at positions 76, 181, and 194 relative to SEQ ID NO:6. The TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, 180, and 182 relative to SEQ ID NO:4; the TnsB protein comprises an amino acid sequence having amino acid substitutions at positions 43, 349, 352, 390, 396, 410, 464, 526, 549, and 594 relative to SEQ ID NO:5; and the TnsC protein comprises an amino acid sequence having amino acid substitutions at positions 197 and 314 relative to SEQ ID NO:6. Alternatively, the TnsA protein comprises the amino acid sequence of SEQ ID NO: 4, and the TnsB protein comprises one or more positions selected from positions 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526, relative to SEQ ID NO: 5; 43, 349, 352, 390, 396, 464, 549, 594, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 6 7; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 21; or 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, 21, and 67, and the TnsC protein comprises an amino acid sequence having amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 76 and 7, 12, or 263; or 76, 238, 296, or 328, relative to SEQ ID NO: 6;

[0088] In some embodiments, the TnsA protein comprises an amino acid sequence with amino acid substitutions P88T and I147V relative to SEQ ID NO:4, the TnsB protein comprises an amino acid sequence with amino acid substitutions P352T, A390V, D396N, H464R, Q549R, and Q594L relative to SEQ ID NO:5, and the TnsC protein comprises an amino acid sequence with amino acid substitutions R197I, N314K, and optionally one of I7S, L12M, or K114M relative to SEQ ID NO:6, or the TnsA protein comprises an amino acid sequence with amino acid substitutions P88T and I147V relative to SEQ ID NO:4, the TnsB protein comprises an amino acid sequence with amino acid substitutions P352T, A390V, D396N, H464R, Q549R, and Q594L relative to SEQ ID NO:5, and the TnsC protein comprises an amino acid sequence with amino acid substitutions R197I, N314K, and optionally one of I7S, L12M, or K114M relative to SEQ ID NO:6. The TnsB protein comprises an amino acid sequence having amino acid substitutions of P352T, A390V, D396N, H464R, Q549R, and Q594L based on SEQ ID NO:5, and the TnsC protein comprises an amino acid sequence having amino acid substitutions of S76Y, A181S, and V194M based on SEQ ID NO:6; or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, and 182 based on SEQ ID NO:4. The TnsB protein comprises an amino acid sequence having amino acid substitutions P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L with reference to SEQ ID NO:5, and the TnsC protein comprises an amino acid sequence having amino acid substitutions R197I, N314K, and optionally one of I7S, L12M, or K114M with reference to SEQ ID NO:6; or the TnsA protein comprises an amino acid sequence having amino acid substitutions at positions 88, 147, 170, and 182 with reference to SEQ ID NO:4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L with reference to SEQ ID NO:6. The TnsC protein comprises an amino acid sequence having amino acid substitutions of P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L based on SEQ ID NO: 5, and the TnsC protein comprises an amino acid sequence having amino acid substitutions of S76Y, A181S, and V194M based on SEQ ID NO: 6; the TnsA protein comprises an amino acid sequence having amino acid substitutions of P88T, I147V, V170L, F180L, and F182L based on SEQ ID NO: 4; and the TnsB protein comprises an amino acid sequence having amino acid substitutions of F43S, Y349N,The TnsC protein comprises an amino acid sequence having amino acid substitutions P352T, A390V, D396N, Q410K, H464R, V526E, Q549R, and Q594L based on SEQ ID NO: 6, or the TnsA protein comprises an amino acid sequence of SEQ ID NO: 4, and the TnsB protein comprises an amino acid sequence having amino acid substitutions F43S, Y349N, or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L and one or more substitutions selected from R63G, A145S, A174S, I182R, V208M, Q410K, T427S, T456I or T456P, P504S, and V526E; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V and T50 2I; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and T21A; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K; or F43S, Y349N or Y349D, P352T, A390V, D396N, H464 The TnsC protein comprises an amino acid sequence having the amino acid substitutions R, Q549R, Q594L, A415V, T502I, T21A, and Q67K, and the TnsC protein comprises an amino acid sequence having the amino acid substitutions R197I, N314K, and optionally one of I7S, L12M, or K114M; S76Y and I7V, L12M, or S263N; or S76Y, A238S, K296N, or V328M, based on SEQ ID NO: 6.

[0089] In some embodiments, the TnsA protein comprises an amino acid sequence with substitutions at positions 88 and 147 relative to SEQ ID NO:4; the TnsB protein comprises an amino acid sequence with substitutions at positions 352, 390, 396, 464, 549, and 594 relative to SEQ ID NO:5; the TnsC protein comprises an amino acid sequence with substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 76 and 7, 12, or 263; or 76, 238, 296, or 328 relative to SEQ ID NO:6; the Cas7 protein comprises an amino acid sequence with an amino acid substitution at position 345 relative to SEQ ID NO:13; and / or the Cas8-Cas5 fusion protein comprises an amino acid sequence with an amino acid substitution at position 198 relative to SEQ ID NO:12.

[0090] In some embodiments, the TnsA protein comprises an amino acid sequence with substitutions of P88T and I147V relative to SEQ ID NO:4; the TnsB protein comprises an amino acid sequence with substitutions of P352T, A390V, D396N, H464R, Q549R, and Q594L relative to SEQ ID NO:5; the TnsC protein comprises an amino acid sequence with substitutions of R197I, N314K, and optionally one of I7S, L12M, or K114M relative to SEQ ID NO:6; the Cas7 protein comprises an amino acid sequence with substitutions of A345R relative to SEQ ID NO:13; and the Cas8-Cas5 fusion protein comprises an amino acid sequence with substitutions of R198H relative to SEQ ID NO:12.

[0091] In some embodiments, the one or more Cas proteins are encoded by a single nucleic acid. In some embodiments, the one or more transposon-associated proteins are encoded by a single nucleic acid. In some embodiments, the one or more Cas proteins and the one or more transposon-associated proteins are encoded on a single nucleic acid. In some embodiments, the one or more Cas proteins and the one or more transposon-associated proteins are encoded by different nucleic acids. In some embodiments, the one or more nucleic acids comprise one or more messenger RNAs, one or more vectors, or a combination thereof.

[0092] In some embodiments, at least one of the one or more Cas proteins and the one or more transposon-associated proteins comprises a nuclear localization signal (NLS).

[0093] In some embodiments, TnsA and TnsB are linked in a TnsA-TnsB fusion protein. In some embodiments, the TnsA-TnsB fusion protein further comprises an amino acid linker between the TnsA and TnsB. In some embodiments, the linker is a flexible linker. In some embodiments, the linker comprises an NLS.

[0094] In some embodiments, the one or more Cas proteins comprise a Cas8-Cas5 fusion protein.

[0095] In some embodiments, one or more of the at least one Cas protein and the at least one transposon-associated protein are part of a single fusion protein, hi some embodiments, each of the at least one Cas protein and the at least one transposon-associated protein are part of a single fusion protein.

[0096] In some embodiments, the system further includes at least one guide RNA (gRNA) complementary to at least a portion of the target nucleic acid, or at least one nucleic acid encoding the same. In some embodiments, the one or more Cas proteins, the one or more transposon-associated proteins, and the at least one gRNA are encoded by different nucleic acids. In some embodiments, at least one of the one or more Cas proteins and the one or more transposon-associated proteins, and the at least one gRNA are encoded by a single nucleic acid.

[0097] In some embodiments, at least one gRNA is a non-naturally occurring gRNA. In some embodiments, at least one gRNA is encoded by a CRISPR RNA (crRNA) array. In some embodiments, at least one of the one or more Cas proteins is part of a ribonucleoprotein complex with at least one gRNA.

[0098] In some embodiments, the system further comprises at least one unfoldase protein, or a nucleic acid encoding same, hi some embodiments, the at least one unfoldase protein comprises ClpX.

[0099] In some embodiments, the system further comprises a donor nucleic acid, wherein the donor nucleic acid comprises a cargo nucleic acid sequence flanked by at least one transposon end sequence. In some embodiments, the system further comprises a target nucleic acid.

[0100] In some embodiments, the system is a cell-free system.

[0101] Also provided are compositions and cells comprising the disclosed systems. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are eukaryotic cells (e.g., mammalian cells, human cells).

[0102] Additionally, methods for nucleic acid modification and incorporation are provided. In some embodiments, the methods include contacting a target nucleic acid with a system, composition, or polypeptide disclosed herein.

[0103] In some embodiments, the target nucleic acid sequence is in a cell. In some embodiments, contacting the target nucleic acid sequence comprises introducing the system into the cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell, a human cell).

[0104] In some embodiments, introducing the system into the cell comprises administering the system to the subject. In some embodiments, administering comprises in vivo administration. In some embodiments, administering comprises transplantation of ex vivo treated cells comprising the system. In some embodiments, the system, composition, or polypeptide(s) are provided in one or more delivery vehicles. In some embodiments, the one or more delivery vehicles are selected from the group consisting of viral particles, virus-like particles, liposomes, nanoparticles, and combinations thereof.

[0105] Another aspect provided by the present disclosure is a method for generating and analyzing variant Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CRISPR-Tn) polypeptides.

[0106] In some embodiments, the method comprises: a) exposing nucleic acid sequences encoding two or more different CRISPR-Tn polypeptides to mutagenic conditions; b) encoding one or more of TnsA, TnsB, and TnsC polypeptides on a selection phage; c) encoding crRNA, TniQ, a Cas8-Cas5 fusion, Cas7, Cas6, and any of the TnsA, TnsB, and TnsC polypeptides not contained in the selection phage on one or more complementation plasmids; d) encoding a phage coat protein on an accessory plasmid; e) introducing the selection phage, complementation plasmid, and accessory plasmid into a host cell; and f) screening for one or more variant CRISPR-Tn polypeptides expressed by the host.

[0107] In some embodiments, crRNA, TniQ, the Cas8-Cas5 fusion, Cas7, and Cas6 are encoded on a single complementing plasmid. In some embodiments, crRNA is encoded on a first complementing plasmid and TniQ, the Cas8-Cas5 fusion, Cas7, and Cas6 are encoded on a second complementing plasmid.

[0108] In some embodiments, the first complementation plasmid further encodes a ribosome binding site (RBS), a crRNA target, and a T7 RNA polymerase (RNAP) downstream of the crRNA target and RBS. In some embodiments, the first complementation plasmid further encodes an Npu intein (gIII) downstream of a T7 promoter. N In some embodiments, the accessory plasmid further encodes an N-terminal gill fragment linked to an Npu intein (-Npu). In some embodiments, the phage coat protein is gene III (gIII), and the accessory plasmid comprises a C-terminal gill fragment linked to a crRNA target and an Npu intein encoded downstream of the RBS. In some embodiments, the second complementation plasmid further comprises a donor cassette.

[0109] In some embodiments, the first complementation plasmid further encodes a ribosome binding site (RBS) and a crRNA target. In some embodiments, the first complementation plasmid further encodes an Npu intein (gIII N In some embodiments, the accessory plasmid further encodes an N-terminal gill fragment linked to an Npu intein (-Npu). In some embodiments, the phage coat protein is gene III (gIII), and the accessory plasmid comprises a C-terminal gill fragment linked to a crRNA target and an Npu intein encoded downstream of the RBS. In some embodiments, the second complementation plasmid further comprises a donor cassette.

[0110] In some embodiments, the second complementing plasmid comprises a donor cassette.

[0111] In some embodiments, the method comprises: a) exposing nucleic acid sequences encoding two or more different CRISPR-Tn polypeptides to mutagenic conditions; b) encoding one or more of Cas6, Cas7, Cas8-Cas5 fusions, and TniQ polypeptides on a selection phage; c) encoding crRNA, TnsA, TnsB, TnsC, and any of Cas6, Cas7, Cas8-Cas5, and TniQ polypeptides not contained in the selection phage on one or more complementation plasmids; d) encoding a phage coat protein on an accessory plasmid; e) introducing the selection phage, complementation plasmid, and accessory plasmid into a host cell; and f) screening for one or more variant CRISPR-Tn polypeptides expressed by the host.

[0112] In some embodiments, crRNA, TnsA, TnsB, and TnsC are encoded on a single complementation plasmid. In some embodiments, the accessory plasmid encodes a C-terminal phage coat protein fragment linked to an intein. In some embodiments, the complementation plasmid further encodes an N-terminal phage coat protein fragment linked to an intein downstream of T7 RNA polymerase (RNAP). In some embodiments, crRNA is encoded on a Plasmid Donor (PD).

[0113] In some embodiments, the Plasmid Donor comprises a Donor Cassette.

[0114] In some embodiments, the ribosome binding site (RBS) is encoded on the accessory plasmid or the accessory and complementing plasmids.

[0115] Also provided are methods for treating a disease or disorder in a subject, the methods comprising administering to a subject in need of such treatment a polypeptide, system, or composition, or cells comprising the same. In some embodiments, the subject is a human. In some embodiments, the system or composition comprises a donor nucleic acid encoding a therapeutic gene product or a wild-type or modified version of a disease-associated gene.

[0116] Also provided are methods for inactivating a microbial gene, the methods comprising introducing a system or composition described herein into one or more cells. In some embodiments, the gRNA is specific for a target site proximal to the microbial gene, and the system or composition modifies the microbial gene. In some embodiments, the system or composition inserts a donor nucleic acid into the microbial gene. In some embodiments, the microbial gene is a bacterial antibiotic resistance gene, virulence gene, or metabolic gene. In some embodiments, the one or more cells are bacterial cells.

[0117] Also provided are methods for modifying a target nucleic acid in a plant cell, the methods comprising providing a system or composition described herein to a plant, or a plant cell, seed, fruit, plant part, or propagation material of the plant. In some embodiments, the system or composition inserts a donor nucleic acid into the target nucleic acid. In some embodiments, the donor nucleic acid comprises a gene product.

[0118] In some embodiments, the plant is a monocotyledonous or dicotyledonous plant. In some embodiments, the plant is a grain crop, fruit crop, forage crop, root crop, leafy vegetable crop, flower crop, conifer, oil crop, plant used in phytoremediation, industrial crop, medicinal crop, or experimental model plant. In some embodiments, the system or composition is provided via Agrobacterium-mediated transformation. In some embodiments, the method confers one or more of the following traits to a plant or plant cell, seed, fruit, plant part, or propagation material of a plant: herbicide tolerance, drought tolerance, male sterility, pest resistance, abiotic stress tolerance, modified fatty acid metabolism, modified carbohydrate metabolism, modified seed yield, modified oil content, modified protein content, disease resistance, cold and frost tolerance, improved taste, increased germination, increased micronutrient absorption, improved flower longevity, modified fragrance, modified nutritional value, modified fruit or flower size or number, modified growth, and modified plant size.

[0119] Other aspects and embodiments of the present disclosure will become apparent in light of the following detailed description. [Brief explanation of the drawings]

[0120] [Figure 1]Figures A-D show exemplary vector circuit designs for phage-assisted evolution of TnsABC. In Figure A, TnsA, TnsB, and TnsC are genes to be evolved encoded on the selection phage (SP). TnsA and TnsB are encoded in a single coding region linked by a mammalian nuclear localization signal (NLS), also abbreviated as TnsAB or TnsA-bpNLS-TnsB. The crRNA, TniQ, Cas8, Cas7, Cas6, and promoter-containing donor cassette are encoded on the complementation plasmid (CP). The crRNA target, RBS, and gene III (gIII) are encoded on the accessory plasmid (AP). The INTEGRATE system (TnsA, TnsB, TnsC, TniQ, Cas8, Cas7, Cas6, and crRNA) catalyzes the integration of the donor cassette downstream of the crRNA target on the AP, resulting in gIII expression and SP propagation. In B, the circuit is a modified version of the circuit shown in A, where crRNA, TniQ, Cas8, Cas7, Cas6, and crRNA are encoded on Complementation Plasmid 1 (CP1), and the donor cassette is encoded on Complementation Plasmid 2 (CP2), also known as Plasmid Donor (PD). In C, the circuit is a modified version of the circuit shown in B, where a C-terminal gIII linked to an Npu intein (gIIIC-Npu) is encoded downstream of the crRNA target and RBS on the AP, an N-terminal gIII linked to an Npu intein (gIIIN-Npu) is encoded downstream of the crRNA target and RBS on the CP, and the donor cassette and crRNA are encoded on Plasmid Donor (PD). The INTEGRATE system catalyzes the integration of the donor cassette downstream of the crRNA target on the AP and downstream of the crRNA target on the CP, resulting in expression of both gIII halves and reconstitution of the full-length pIII protein. This circuit splits gill between two plasmids, minimizing the possibility that the SP acquires full-length gill.In D, the circuit is a modified version of the circuit shown in C, in which T7 RNA polymerase (RNAP) is encoded downstream of the crRNA target and RBS on the CP, and an N-terminal gIII linked to an Npu intein (gIIIN-Npu) is encoded downstream of the T7 promoter on the CP. Integration at the crRNA target on the CP promotes T7 RNAP expression, which in turn drives gIIIN-Npu expression. This circuit increases the amount of gIIIN-Npu expressed per CP integration event, thereby reducing selection stringency. [Figure 2] A and B show that variants of TnsA, TnsB, and TnsC from Tn6677 derived from initial phage-assisted non-continuous evolution (PANCE) propagation rounds (clones 1–4) propagate more efficiently on selection circuits when programmed with targeted crRNA, and this propagation correlates with donor integration on APs measured by qPCR. [Figure 3] A schematic diagram of the inter-plasmid mammalian cell editing system used to evaluate the efficiency of evolved variants is shown. Evolved variants were cloned into expression vectors and co-transfected with a donor transposon (pDonor Mini-Tn) and a target plasmid (pTarget), along with other components of the CRISPR system, as needed. After 72 hours of incubation, cells were lysed, and integrated target plasmid was measured by qPCR using an integration probe 49 bp downstream of the target site. [Figure 4] (A) Variants of TnsA, TnsB, and TnsC from Tn6677 derived from an initial phage-assisted discontinuous evolution (PANCE) propagation round show increased interplasmid editing in mammalian cells. (B) Comparison of TnsA, TnsB, and TnsC variants from Vibrio cholerae Tn6677 with a system derived from Tn7016, a transposon encoded by Pseudoalteromonas sp. S983. [Figure 5]Figures 5A and 5B show that variants of Tn7016-derived TnsA, TnsB, and TnsC from an early phage-assisted continuous evolution (PACE) propagation round improved transposition in E. coli compared to the wild type (Figure 5A), but did not have improved transposition efficiency in mammalian cells (Figure 5B). Figures 5C–5E show that variants of Tn7016-derived TnsA, TnsB, and TnsC from an early phage-assisted discontinuous evolution (PANCE) propagation round had improved integration in E. coli (Figure 5C) and in plasmid and genomic targets in mammalian cells (Figures 5D and 5E). [Figure 6] A-D show that variants from the first round of PANCE were used in further propagation of PACE and PANCE to generate a series of variants that improve editing in mammalian cells. A shows those genotypes that allow for the highest editing efficiency. B-D show the plasmid and genomic targets as indicated. E shows that the series of variants also improves editing efficiency in bacteria. [Figure 7] 1 shows editing efficiencies from reversion of exemplary mutation variants at multiple genomic sites. [Figure 8] 1 is a graph of the editing efficiency of variants recovered at different time points during a single round of PACE / PANCE amplification. [Figure 9]A and B are exemplary vector circuit designs for phage-assisted evolution of the QCascade components Cas6, Cas7, Cas8, and TniQ. In A, TniQ, Cas8, Cas7, and Cas6 are genes to be evolved encoded on the selection phage (SP). The crRNA, TnsAB, and TnsC are encoded on the complementation plasmid (CP). TnsA and TnsB are encoded in a single coding region linked by a mammalian nuclear localization signal (NLS), also abbreviated as TnsAB or TnsA-bpNLS-TnsB. The donor cassette is encoded on the plasmid donor (PD). The crRNA target, RBS, and gene III (gIII) are encoded on the accessory plasmid (AP). The system catalyzes integration of the donor cassette downstream of the crRNA target on the AP, resulting in gIII expression. In B, the circuit from A was modified by encoding TnsAB, TnsC crRNA target sites, T7 RNAP, and an N-terminal gIII linked to an Npu intein (gIIIN-Npu) on a complementation plasmid, donor cassette and crRNA on a donor plasmid (PD), and C-terminal gIII linked to an Npu intein (gIIIC-Npu) on an accessory plasmid (AP). The system catalyzes the integration of the donor cassette downstream of the crRNA target on the AP and downstream of the crRNA target on the CP, resulting in expression of both gIII halves and reconstitution of the full-length pIII protein. [Figure 10] A and B show that TnsC can acquire mutations that inhibit mammalian activity during evolution. Editing efficiencies of PANCE N23 and PACE P9 variants against plasmid (FIG. 10A) and genomic (FIG. 10B) targets were tested with evolved TnsAB in combination with wild-type TnsC, and evolved TnsC in combination with wild-type TnsAB, as shown. In many cases, PACE P9 variants performed best when evolved TnsAB was combined with wild-type TnsC. Plasmid: 15 cycles of PCR1. Genome: 25 cycles of PCR1. [Figure 11]A is a schematic diagram of the TnsAB single integration circuit for Tns PACE circuit 4 (TnsAB evolution). Compared to Tns circuit 3, the circuit has the following modifications: TnsC is removed from the SP and encoded on the CP; the CP target site is removed (returning to a single integration circuit); the AP backbone size is increased (preventing gIII acquisition by the SP); and pDonor contains either the wild-type sequence or a transposon left end containing a mutated binding site for a putative bacterial integration host factor (referred to as "s-IBS") to prevent the SP from evolving bacterial-specific fitness. The single integration circuit reduces the selection stringency of TnsAB evolution and simplifies PACE circuit design. Removing TnsC from the SP reduces the accumulation of deleterious mutations for mammalian activity. B and C show TnsAB PANCE N25 in Tns circuit 4. SPs encoded the P8-L5-8 or N23-P16-L1-2 TnsABs, the best-functioning TnsABs from previous TnsABC evolution. Variants isolated at P13 and P25. * indicates selection-free drift passage. [Figure 12] A-C show that the TnsAB PANCE N25-P13 variant is significantly less efficient than the starting genotype. Graphs show editing efficiency at the plasmid and genomic target, as indicated. Arrows indicate the starting TnsAB variants (P8-L5-8, N23-P16-L1-2) that yielded the variant on the right. All TnsABs were tested with the P8-L5-8 TnsC, which was the best TnsC at the time of characterization. [Figure 13]Figures A-C show that TnsAB PANCE N25-P25 variants demonstrate improved mammalian activity compared to the input variants. Graphs show editing efficiency at the plasmid and genomic target, as indicated. Arrows indicate the starting TnsAB variants (P8-L5-8, N23-P16-L1-2) that resulted in the variants on the right. All variants were tested with the N23-P16-L1-5 TnsC, the best TnsC at the time of characterization. N25 TnsAB variants represent some of the most active Tn7016 TnsABs. AAVS1 sites were quantified by HTS and ddPCR. [Figure 14] Measurement of N25 TnsAB editing by ddPCR and HTS is shown. The HTS strategy for measuring integration requires comparing integrated and unintegrated PCR amplicons, and therefore, % integration may be biased by PCR bias. ddPCR is an established method for measuring integration without PCR bias, and values ​​can be interpreted as the "ground truth" for % integration. Comparison of HTS and ddPCR shows that HTS values ​​are, on average, approximately 3.5-fold higher than ddPCR (top). Values ​​normalized to starting activity are consistent across ddPCR / HTS (bottom). Much of the data shown in these slides was obtained by HTS (shown in terms of PCR cycle number in the graphs), which allows for high-throughput characterization of the relative editing efficiencies of variants. Absolute editing variants are determined hereafter by ddPCR unless otherwise noted. [Figure 15] Analysis of N25-P25 TnsAB with wild-type or s-IBS mutant transposons in mammalian cells is shown. Editing in AAVS1 was tested using WT or IHF-binding mutant (s-IBS) transposon donors. Evolution with WT or s-IBS transposons did not result in transposon-specific activity. Arrows indicate the starting TnsAB variants (P8-L5-8, N23-P16-L1-2) that resulted in the variants on the right. All variants were tested with N23-P16-L1-5 TnsC. [Figure 16] A and B show PACE P11 of highly active N25-P25 TnsAB. Input SPs were the top two N25 TnsAB variants (Figure 16A) and the pooled N25 PANCE LAG. Evolution was performed with both WT transposons (L1-L3) and s-IBS transposons (L4-L6) (Figure 16B). L1 / L2 bottlenecked at approximately 144 hours, so genotypes were sampled from 168 and 120 hours. [Figure 17] A-D show that PACE(P11) of mammalian-active TnsAB failed to substantially improve editing. Boxed is the input N25 TnsAB variant to PACE. No PACE variant had significantly improved editing across sites. Higher selection stringency may further improve TnsAB mammalian activity. [Figure 18] Panels A and B show the PANCE of TnsAB PANCE N29, the top eight clonally isolated N25 TnsAB variants, and the N25 PANCE lagoon. All evolutions were performed on the s-IBS transposon, targeting the AAVS1 sequence on the AP (previous evolutions were performed on target sequences not found in mammalian cells). Multiple lagoons acquired gIII (CAST-independent recombination) and are highlighted in red. [Figure 19] A and B show TnsAB PACE P12 in Tns circuit 5. Tns circuit 5 (Figure 19A) has the following modifications compared to Tns circuit 4: introduction of a ribosome binding site between TnsA and TnsB and splitting the synthetic TnsA-TnsB fusion into its native TnsA+TnsB form. TnsAB PACE frequently evolved a stop codon within the bpNLS (splitting TnsA-TnsB into TnsA+TnsB) to improve circuit fitness. P12 PACE (Figure 19B) evolved the two best N25 TnsABs in Tns circuit 5 and evolved them on a 5 kb transposon to increase selection stringency (the previous Tn7016 evolution was on a 1 kb transposon). [Figure 20] 1 shows an overview of TnsAB and TnsC evolution to identify TnsAB / TnsC combinations. [Figure 21] Figures 21A-C show a TnsC screen using N25-P25-L5-5 TnsAB. TnsC variants cloned into mammalian vectors were tested (69 in total). Plasmid (Figure 21A) and AAVS1 (Figure 21B) editing efficiencies correlate. N14-5 TnsC (a variant from the first TnsABC PANCE) is preferred (Figure 21C). Arrows in each of A and B indicate WT TnsC. [Figure 22] A-B show ddPCR of the top TnsC variants from the screen. The top six TnsC variants, WT, and ΔTnsC, were quantified by ddPCR in addition to HTS. Comparing editing values: ddPCR shows approximately 2.25% editing of the T-RL insertion 48 bp downstream of the target (Figure 22A). Comparing WT normalized values: ddPCR and HTS are consistent in identifying the best TnsC for subsequent combination of beneficial mutations (Figure 22B). The 2.25% editing efficiency with N25-P25-L5-5 TnsAB + N14-5 TnsC is a higher editing efficiency than previously observed in AAVS1. [Figure 23] TnsC genotypes sorted by efficiency are shown. Mutations were classified by editing relative to WT (P2P and genomic average): green: >1.35-fold relative to WT; red: <1-fold relative to WT. All single mutants associated with >1.35-fold editing and mutants that appeared in more than one beneficial variant in N14-5 TnsC. 29 mutations (green) were cloned. [Figure 24]Figures A-D show repeats of a TnsC screen similar to those in 21A-21C in the presence and absence of ClpX to determine whether the addition of ClpX alters the TnsC fitness landscape. Transfection conditions modified from the previous screen include drug selection of transfected cells and harvesting at 4 days post-transfection (instead of 3 days post-transfection). Figures A and B show that editing efficiency correlates in the absence (Figure 24A) and presence (Figure 24B) of ClpX. Figure C shows that the absence of ClpX is consistent with results from the previous screen. Editing relative to WT is higher in this screen, likely due to the modified transfection conditions. Figure D shows that ClpX improves editing for nearly all TnsC variants. While ClpX improves moderately active variants, the best TnsC variants without ClpX (such as N14-5) lack significant improvement with ClpX. [Figure 25] A–F show the single-mutation TnsC screen. Twenty-nine point mutations were individually cloned into the N14-5 TnsC backbone and tested in AAVS1 (Figures 25A and 25C) and HEK3 (Figures 25B and 25D) in the presence and absence of ClpX, as indicated. Lines C and D show N14-5 activity. In AAVS1, activity with and without ClpX generally correlates. In HEK3, some improvement was observed without ClpX, but no improvement was observed in the presence of ClpX, indicating that the benefit from TnsC mutations may be redundant with the addition of ClpX. E and F show a summary of the single-mutation TnsC screen. Single mutations in N14-5 TnsC show significant improvement only in HEK3 without ClpX (which had the lowest starting editing). Single mutations do not significantly improve editing by ClpX. Stacking of multiple mutations can be used to further improve activity. The best single mutation of N14-5 TnsC is shown in the upper right quadrant of E. [Figure 26]ClpX titration with and without puromycin selection is shown. ClpX was titrated with WT TnsABC (pink), P8-L5-8 (purple), and N25-P25-L5-5 TnsAB + N23-P16-L1-5 TnsC (blue). Toxicity was observed at high doses of ClpX. Puromycin selection was tested to determine whether selection of transfected cells alleviates the low editing observed at high doses of ClpX. Puromycin selection of transfected cells did not substantially change the propensity for plasmid editing but may allow for higher ClpX concentrations for genome editing. High doses of ClpX may cause TnsB degradation prior to transposition or stress cells, reducing transgene expression, both of which reduce editing. [Figure 27] Analysis of a representative set of evolved TnsABCs, including previously successful (N14-1, P8-L5-8, and N25 variants) and previously unsuccessful (P9-144h variant) variants, in the presence and absence of ClpX is shown. Addition of ClpX generally did not affect the relative efficiency of previously evolved TnsABCs and did not rescue the P9-144h variant. The fold improvement with ClpX inclusion was much greater for WT and weakly evolved variants compared to highly evolved variants, suggesting that evolved mutations from TnsPACE may be addressing the same bottleneck that ClpX addition improves. [Figure 28] Analysis of the best evolved TnsAB (x-axis) and best evolved TnsC (y-axis) in a different AAVS1 strain than previously shown, in the presence and absence of ClpX, shows the same trend seen previously, with ClpX improving the efficiency of WT and less evolved TnsABC more than highly evolved TnsABC. pBK17 TnsC is a combination of PACE / PANCE TnsC mutations, genotypes described in the TnsC screen. [Figure 29]Figures A-C show the effect of transfection stoichiometry on one of the best evolved TnsABC variants in mammalian cells. The stoichiometry of the plasmid components was optimized with N23-P16-L1-5 TnsABC. All non-titrated components were held constant according to the previous stoichiometry. Plasmid editing was completed in parallel with reoptimization of WT TnsABC—the opposite trend for TnsC. [Figure 30] We demonstrate that modifying the transfection stoichiometry of PACE9 TnsABC variants did not restore mammalian activity. Representative PACE9 144h TnsABC variants were titrated to modify stoichiometry and assess whether activity could be restored. Each titrated variant was tested with the coevolved subunit (L3-1 TnsAB titration was tested with L3-1 TnsC). No stoichiometry allowed editing beyond that of N23-P16-L1-5 TnsABC. [Figure 31] Figures A-C show the N23-P16-L1-5 TnsABC gene tested with larger transposons in mammalian cells. Integration of two cargoes per transposon size (5 kb, 10 kb) was tested in plasmid and genomic targets, as indicated. Efficiency decreased as a function of transposon size, although activity declined less with inter-plasmid editing. [Figure 32] Analysis of split TnsA / TnsB usage in mammalian cells. The Tn7016 TnsAB fusion is an artificial construct inspired by the natural TnsAB fusion in the ortholog CAST (see Vo, et al. Mobile DNA 2021). TnsA-bpNLS and bpNLS-TnsB were tested for N23-P16-L1-5 TnsABC. Adjusting the stoichiometry of the split TnsA-NLS and NLS-TnsB allowed editing to approximate the efficiency of the TnsA-NLS-TnsB fusion (shown in the bottom right), but did not substantially improve mammalian activity. [Figure 33]Panels A and B show a comparison of the TnsAB and TnsC backbones in the presence and absence of ClpX. The Sternberg and Liu constructs used different mammalian expression backbones for TnsAB and TnsC: the Sternberg backbone contained an SV40 onset, while the Sternberg TnsC backbone contained a consensus Kozak sequence for TnsC. All four combinations of the Liu / Sternberg TnsAB / TnsC backbones were tested with WT and the current best TnsABC, with and without ClpX. The Sternberg backbone enabled optimal editing, both with and without ClpX. The Sternberg TnsC backbone significantly improved the editing efficiency of WT TnsC. WT TnsC performed better than the evolved TnsC in the Sternberg backbone. This difference was likely due to the different stoichiometry caused by the SV40 onset, as transfected cells were able to replicate the TnsAB and TnsC vectors. [Figure 34] (A-F) Evolution of the Tn6677 QCascade complex on circuit 1.0 results in improved interplasmid integration efficiency in bacterial cells. A: Schematic of PACE circuit 1.0 adapted from the TnsABC circuit. B: Overnight growth and Tn integration with WT and evolved TnsABC. C: Phage titer and lagoon flux over time for Tn6677 PACE1. D: Schematic of bacterial interplasmid integration assay. E: Table of selected mutations from PACE1. F: E. coli interplasmid integration results for selected clones. [Figure 35]Figures 35A-E show that evolution of the Tn7016 QCascade complex on Circuit 1.0 results in improved inter-plasmid integration efficiency in bacterial cells. A is a schematic of PACE Circuit 1.0 adapted from the TnsABC circuit. B shows overnight growth and Tn integration under the indicated conditions. C shows phage titer and lagoon flux over time for Tn7016 PANCE. D and E show overnight growth (left), PACE (center), and E. coli inter-plasmid integration results (right) for selected clones harboring P2-L3-2 TnsABC (Figure 35D) or N14-1 TnsABC (Figure 35E). [Figure 36] A-C show that Tn7016 QCascade variants have improved E. coli genome integration efficiency (Figure 36A) and improved plasmid editing (P2P) in mammalian cells (Figure 36B), but reduced mammalian genome integration efficiency measured in HEK3-2 (Figure 36C). [Figure 37] A-E show the construction of Circuit 2.0 for the evolution of the Tn7016 QCascade complex. A is a schematic showing the change from PACE Circuit 1.0 to PACE Circuit 2.0 single integration. B shows a schematic of the evolution of different PAM preferences. C shows that CRISPR repeats affect integration efficiency. D shows integration with an improved TnsABC variant (N20 / P8). E shows the toxicity of TnsABC variants in bacterial cells. [Figure 38] Figures 38A and 38B show that evolution in Circuit 2.0 is possible with regular monitoring of PACE and Cheetah phage. The Cheetah lagoon was discontinued, and a new lagoon was inoculated with phage from either one of the non-Cheetah lagoons or a pool of phage from non-Cheetah lagoons (Figure 38A). Phage propagation increased, but the number of distinct genotypes decreased. Five PACE attempts failed in Circuit 2.0 (Figure 38B). [Figure 39]Figures 39A and 39B show that the evolution campaign with Circuit 2.0 resulted in novel, highly mutated QCascade variants with approximately 0% integration efficiency in HEK293T cells, both at the genomic site (Figure 39A) and via plasmid-to-plasmid transfer (Figure 39B). HTS performed with high PCR cycle numbers: values ​​may be skewed by PCR bias. [Figure 40] Integration at genomic sites in wild-type counterparts and independently evolved QCascade components is shown. HTS performed with high PCR cycle numbers: values ​​may be skewed by PCR bias. [Figure 41] Schematic showing the evolution of Circuit 4.0, which enables cheater-free evolution of the Tn7016 QCascade complex. [Figure 42] A and B show that phage grow (FIG. 42A) and integrate (FIG. 42B) more efficiently in Circuit 4.0 compared to previous circuits. [Figure 43] Panels A and B show results for variants evolved with Circuit 4.0 (v4). None of the variants evolved with v4 consistently exhibit higher integration efficiencies across multiple sites. Panel A shows integration efficiencies measured by HTS for AAVS1, HEK3-2 (25-cycle PCR1), and P2P (15-cycle PCR1). QCascade variants evolved from Circuit v4 are indicated by variant name (4V1-4V8). WT combinations include variant name-evolved components. Editing efficiencies are shown as fold improvements over WT QCascade. Variants from phage (v4, v8) that performed particularly well during PANCE are among the variants with the lowest editing efficiencies in mammalian cells. Panel B shows that editing efficiencies measured by ddPCR are approximately four-fold lower than low-cycle HTS values, but the relative values ​​are the same; therefore, ddPCR data correlate well with HTS data. 4V2 (mutations present only in Cas6) and 4V6-6 potentially improved integration at the AAVS1 site, and 4V6-6 can be further evolved in a single subunit evolution circuit. [Figure 44]Results are shown using WT combinations of evolved Tn7016 QCascade components. Conditions with more than one evolved QCascade component were among the lowest editing efficiencies, encouraging single-subunit evolution. Improvements were observed in combinations of evolved Cas6 with WT Cas7, 8, and TniQ. [Figure 45] This shows that the combination of potentially beneficial mutations and reversion of potentially deleterious mutations did not result in increased integration efficiency. Repeated experiments with evolved Cas6 variants showed no significant improvement at the AAVS1 site (blue arrow). Conserved mutations in Cas6 inhibit activity in mammalian contexts (red arrow). Cas7 mutations in the context of the N23 P16 L1-5 transposase only slightly improved integration efficiency (black arrow). [Figure 46] Panels A–C show that evolved QCascade variants exhibit different behavior in mammalian cells than in bacterial cells. For each of four representative genotypes from PACE circuits v2 and v4, two biological replicates, each with two technical replicates, were monitored for integration efficiency (Figure 46A). The integration efficiencies of WT, v4V5, and v4V6 were lower than expected, while v4V5 and v4V6 transformed poorly. Panel B shows the reduced integration efficiency of P8 L5-8 Tn. Potential reasons for the reduced integration efficiency include the consumption of available transposons for integration by integration with the crRNA cassette and toxicity. Transformation into freshly prepared competent cells rescues the activity of v4V5 and v4V6, but also improves WT activity (Figure 46C). [Figure 47] 1 shows an analysis of evolved QCascade components with evolved TnsABC in the presence and absence of ClpX ("SLF"). [Figure 48]Panels A–F show optimization of transfection with ClpX ("SLF") and re-evaluation of evolved QCascade variants. SLF significantly improves integration efficiency in 48-well plates, both with and without puromycin selection (Figure 48A, approximately 42k cells per well). Low-cell-density transfection (24-well plates (approximately 20k cells per well)) further increases integration efficiency to approximately 0.3%, approaching the value of the Sternberg lab (approximately 1.0%), but most cells (approximately 80%) die (Figure 48B). Panels C and D show results from v2 (circuit version 2), V5 (variant 5)-evolved constructs. In the context of evolved TnsAB and C variants, SLF only slightly improves integration efficiency (approximately 1-3 fold, depending on transfection conditions). The QCascade mutation A345T from variant v2V5-7 performed slightly better in the absence of SLF but not in its presence. E and F show results from v4 (circuit version 4) and V5 (variant 5)-evolved constructs. In the context of the evolved TnsAB and C variants, SLF only slightly improved the expression (approximately 1-3 fold, depending on transfection conditions). The evolved QCascade variant v4V6-7 from circuit v4 performed slightly better than WT in both + / - SLF conditions. C and E—48-well plates (approximately 42k cells per well). D and F—24-well plates (approximately 20k cells per well). [Figure 49] A shows that Cas7 A345T potentially increases DNA binding affinity. Red: Mutations after PANCE30 passages in circuit 2.0 (111 mutations in total). The alpha-fold Tn7016 structure was mapped to the Tn6677 structure (PDB 6PIJ). B shows the mutation table for QCascade circuit v2. [Figure 50]Panels A to C show structure-based rational modifications to improve DNA binding affinity. A shows Cas8 DNA-binding residues in Tn6677 QCascade and Tn7016 QCascade. Minor changes: R20K, R21K, S24Q, K88R, R93K, N134Q, R233K. Electrostatic mutations: S24K, S24R, H124R, N134R, R20E, R21E, K88E, R93E, R241E. B shows Cas9 DNA-binding residues in Tn6677 QCascade and Tn7016 QCascade. Minor changes: Q236S, K343R, K344R. Electrostatic mutations: N5K, N5R, T47R, T71R, Q236E, N5D, T47D, T71D, K343E, K344E. C shows rational modifications based on the Cas7 structure to improve DNA binding affinity. All mutants were tested with 20 ng of ClpX. Minor changes: Q236S, K343R, K344R. Electrostatic mutations: N5K, T47R, T71R, Q236E, T71D, K343E, K344E. [Figure 51] PACE-inspired rational mutagenesis of Cas7 mutants. All mutants were tested with 20 ng of ClpX. Minor changes: A345S, A345Y. Electrostatic mutations: A345R, A345K, A345D, A345E. [Figure 52]A-F show an arginine screen of DNA-binding residues to improve DNA / crRNA binding affinity. In A, the DNA / crRNA-binding residues of Cas7 (red, left) and Cas8 (red, right) were mutated to Arg. The Tn7017 QCascade structure was predicted in alpha fold and mapped to Tn6677 QCascade (PDB 6PIJ). B shows Cas7 arginine mutations that increased integration efficiency. All mutants were tested with 20 ng of ClpX. Values ​​depend on the ddPCR machine (BioRad vs. Qiagen). C shows that Cas7 double and triple mutants further improved integration efficiency. dPCR (% positive partitions) vs. ddPCR (% positive droplets). Optimized quantification workflow: 100-400 ng of crude lysate was loaded directly into the (d) dPCR machine. D–F show that the improvement is significant in the context of other TnsABC variants (P12 L2-6 TnsAB and N25 P15 L5-5 TnsAB), but does not apply to all genomic sites (Figure 52D—AAVS1, Figure 52E—HEK3-2, Figure 52F—FANCF). [Figure 53] QCascade rational mutagenesis to reduce crRNA binding affinity is shown. Top: Cas7 mutations predicted to interact with crRNA based on the alpha-fold Tn7016 structure. Bottom: None of the rationally engineered Cas7 mutations result in higher integration efficiency. Cas8 R198H mutation obtained through PACE in circuit v4. [Figure 54]A–E show that beneficial arginine residues are located within the flexible region of the alpha-fold Tn7016 QCascade structure. A shows clusters 1 and 2 from the flexible internal and C-terminal regions, respectively, and an additional beneficial mutation (N5R) with that structure. B shows stacking of arginine mutations between and within clusters. Mutations between clusters are stackable. Stacking mutations within cluster 2 reduces integration efficiency. Having multiple adjacent arginine residues is likely to be deleterious. C shows the site-dependence of rationally engineered Cas7 arginine residues due to the possibility of more favorable interactions with guanine. D shows improvement at the AAVS1-1 site through ortholog-inspired rational engineering. E shows a summary of rationally engineered Cas7 variants with evolved TnsB / C variants. 1 kb transposon integration in HEK293T cells. The x-axis label indicates Cas7 genotype. n=2 for FANCF, n=4 for HEK3 and AAVS1. [Figure 55] Figure 1 shows an overview of TnsABC evolution. Extensive evolution of TnsABC following N14-1 failed to further improve mammalian integration activity (1 kb transposon integration in HEK293T cells). [Figure 56] A-C show the efficiency of evolved subunits in mammalian cells and TnsC mutations that inhibit mammalian integration activity. A is a summary of mammalian integration activity (1 kb transposon integration in HEK293T cells). B shows a chart of TnsC mutations that identify mutations that inhibit mammalian activity. C shows reversion analysis of selected TnsCs (shown in Figure 56B) in HEK293T cells with 1 kb transposon integration. The dashed line indicates WT TnsC activity. The arrows indicate significant mammalian deleterious mutations. [Figure 57]Figures 57A-57F show the PACE of Tn7016 TnsAB. Figure 57A shows a schematic diagram of the TnsAB PACE (Tns circuit 4 / 5). To prevent the accumulation of mammalian deleterious mutations during evolution, TnsC was moved from the SP to the CP in the host E. coli. Figure 57B shows a summary of the characterization of PACE P12 with a 1 kb transposon integration in HEK293T cells. Figures 57C and 57D show the complete characterization of mammalian genome integration (1 kb transposon integration in HEK293T cells) at two different sites, AAVS1 (Figure 57C) and HEK3 (Figure 57D), in the presence and absence of ClpX. Figure 57E shows a mutation table showing the P12-L2-6 variants of TnsA and TnsB. Figure 57F shows that mutations in TnsB are the main driver of the improved mammalian efficiency (1 kb transposon integration in HEK293T cells). [Figure 58] A-D show an examination of the effect of ClpX on mammalian activity. ClpX enhances genome integration in WT Tn7016 (Figure 58A), whereas PACE reduced ClpX dependence for mammalian activity (Figure 58B). 1 kb transposon integration in HEK293T cells. C is a schematic and Western blot showing the establishment of a ΔclpX host strain for CAST PACE. Deletion of endogenous clpX from the PACE host strain (S2060) was achieved using lambda Red recombination. D shows that ΔclpX introduces new selective pressure for CAST PACE. [Figure 59]A-J show the Tn7016 TnsAB and TnsB PACEs. A is a schematic of Tns circuit 6 for TnsB PACE. Tns circuit 5 with the following modifications: removal of tnsA from the SP and addition of tnsA to the CP. This was modified to focus on TnsB (the primary evolutionary source of improved mammalian integration). B and C show the Tn7016 TnsAB and TnsB PACEs in ΔclpX E. coli. B shows the TnsAB PACE (Tns circuit 5 in a ΔclpX host). C shows the TnsB PACE (Tns circuit 6 in a ΔclpX host). The dashed lines in both B and C indicate P12-L2-6 activity (the input variant for ΔclpX evolution). D-G show the characterization of mammalian genome integration of PANCE N30, PACE P13, PANCE N31, and PACE P14, respectively, as outlined in the schematics shown in B and C. 1 kb transposon integration in HEK293T cells. X-axis labels indicate TnsAB genotype (Figures 59D and 59E) or TnsB genotype (Figures 59F and 59G). H is a schematic of the evolution leading to TnsB variants - PANCE 76 passages, PACE 300 hours, approximately 1000 evolutionary generations. I is a mutation table of TnsB for the major variants. J is a summary of the integration activity of the major variants shown in I compared to WT. PACE improved integration activity by more than 150-fold without ClpX and more than 20-fold with ClpX. [Figure 60]Figures A–C show PACE P15 for TnsB. Figure A shows a schematic of the design of PACE P15. TnsA-specific PCR of P15 lagoons (Figure 60B) showed that all P15 lagoons (where TnsB SPs are thought to have evolved) were contaminated with TnsAB SPs (presumably from the PACE equipment). Lagoons P15-L1, L2, and L3 had trace contaminants (all sequenced SPs were TnsB), while lagoons P15-L4, L5, and L6 had approximately 100% contaminants (all sequenced SPs were TnsAB). Because TnsAB contaminants outcompeted TnsB in P15 lagoons L4, L5, and L6, genotypes from these lagoons were tested in HEK293T cells (see Figure 60C). PACE P15 TnsB genotypes from L1, L2, and L3 were not tested due to the lack of new coding mutations acquired during PACE. (C) Summary of PACE P15 mammalian genome integration (1 kb transposon integration in HEK293T cells). Only evolved TnsB was tested (the contaminant TnsAB lacked the new consensus coding mutation in TnsA; see legend to Figure 60B). The contaminant P15 TnsB genotype did not have significantly more activity than P14-L4-5 TnsB. The x-axis label indicates the TnsB genotype. [Figure 61] A and B show rational combinations of PACE P14 TnsB mutations. Twelve mutations from the top eight TnsB variants were individually introduced into P14-L4-5 (Figure 61A). Mutations in yellow were not tested in the initial mammalian characterization. In all conditions, no point mutations significantly improved activity compared to P14-L4-5 (Figure 61B). [Figure 62] Characterization of evolved TnsABC in HeLa cells compared to HEK293T cells. HeLa cells were transfected with lipofectamine 2000 using the same protocol as HEK293T cells, with P12-L2-6 TnsB and N14-5 TnsC as WT, with all other CAST components as WT. [Figure 63]A–K show high-stringency evolution of TnsB (Tns circuit 6 in a ΔclpX host). A is a schematic of the PACE evolution of TnsB. Three TnsB variants from PACE P14 were evolved under higher selection stringency by reducing the strength of the transposon-encoded promoter and the ribosome binding site (RBS) upstream of gIII (Figure 63B). PACE P19, P21, and P22 all had severe SP titer bottlenecks early in evolution (within 72 h), suggesting that the previously evolved TnsB variants were unable to support robust SP growth under higher selection stringency. C shows P14-L4-5 TnsB in hosts of varying stringency. Brackets indicate the promoter strength-RBS strength ratio for each host. D shows characterization of PACE P19 mammalian genome integration (1 kb transposon integration in HEK293T cells; x-axis labels indicate TnsB genotype). Figure E shows an overview of the PACE P19 TnsB variant. TnsPACE allows for >15% integration (ddPCR) in AAVS1 and HEK3 in HEK293T cells. Figure F shows the phage titer and lagoon flow rate over time for PACE P17, P19, P21, and P22. Cloned SPs from PACE P19 (P19-L3-5) and P22 (P22-L1-4) show slightly improved activity-dependent overnight growth in the selected strain E. coli compared to the input SP (P14-L4-5) (Figure 63G). Evolution resulted in minimal improvement in SP fitness—often >1E3-fold activity-dependent growth improvements are observed after successful PACE campaigns, but here we observed an approximately 1E1-fold improvement. Figures H–K show mutation tables for PACE P17, P19, P21, and P22, respectively. [Figure 64]A-I show a summary of the characterization of evolved TnsBs with unique genotypes from PACE P19, P21, and P22 in HEK293T cells harboring WT TnsA, N14-5 TnsC, and WT QCascade. Only a few TnsB variants show significantly improved activity compared to P14-L4-5 at both target sites (Figure 64A). The dashed line represents the average of two P14-L4-5 edits. The dots represent the average of two TnsB variant edits, all without ClpX. Variants with slight improvements (top right quadrant of the graph) were selected for further characterization. B-G show the complete characterization of PACE P19, P21, and P22 at two genomic locations in HEK293T cells. 1 kb transposon integration; WT TnsA, N14-5 TnsC, WT QCascade; x-axis labels indicate TnsB genotypes. H shows replication of PACE P19 TnsB in HEK293T cells. The best PACE P19 variant is not significantly better than P14-L4-5 in additional replication. I shows replication of PACE P22 TnsB in HEK293T cells at four genomic locations. None of the variants is significantly better than P14-L4-5 (shown by the dashed line) at any target site. P14-L4-5 is the PACE-generated TnsB with the highest activity in HEK293T cells. [Figure 65] Figures 65A-65C show the characterization of rational combinations of PACE P14 TnsB mutations. Single mutations introduced into P14-L4-5 do not confer significantly improved integration activity in all conditions tested. Figure 65A shows a mutation table of TnsB and the introduced combination mutations ("5mut" and "6mut" in P14-L4-5). Figures 65B and 65C show the integration efficiency at two different genomic loci with and without ClpX. Combining mutations into P14-L4-5 did not significantly improve integration activity. [Figure 66]Panels A–K show the analysis of TnsABC combinations. The previous best-performing combination of TnsA, TnsB, and TnsC components is shown in A. To analyze the activity of P14-L4-5 TnsB and previously evolved TnsA and TnsC, we designed a screen in which TnsA with P14-L4-5 TnsB and N14-5 TnsC, and TnsC with WT TnsA and P14-L4-5 TnsB were tested separately at two genomic locations, AAVS1 and HEK3, all in the absence of ClpX. Panels B and C show the complete characterization of evolved TnsA with a 1 kb transposon integration, WT QCascade, and P14-L4-5 TnsB and N14-5 TnsC without ClpX. In B, the dark bars are the results for WT TnsA. In C, the dashed line represents the average of n=2 for WT TnsA, the dots represent the average of n=2 for TnsA variant editing, and the green dots labeled with the TnsA genotype indicate the TnsA selected for subsequent characterization. D and E show the complete characterization of evolved TnsCs with 1 kb transposon integration, WT QCascade, P14-L4-5 TnsB without ClpX, and WT TnsA. In D, the dark bars are results for WT TnsC, and the blue bars represent N14-5 TnsC. In E, the dashed line represents the average of n=2 for WT TnsC, the dots represent the average of n=2 for TnsC variant editing, and the green dots labeled with the TnsC genotype indicate the TnsC selected for subsequent characterization. F shows characterization of wild-type and the three best evolved TnsAs (indicated in the legend) at four genomic locations, as well as wild-type and the five best evolved TnsCs (x-axis) without a 1 kb transposon integration, P14-L4-5 TnsB, WT QCascade, or ClpX. G–I show an overview of TnsABC combinations in HEK293T cells. The combination of P12-L6-5 TnsA, P14-L4-5 TnsB, and N14-5 TnsC is the best-performing evoTnsABC combination tested. J–K are tables of evolved TnsAs and TnsCs, respectively. Those highlighted in green were highly performant in the initial screen. [Figure 67]Panels A and B show characterization of the evolved CAST system using P14-L4-5 TnsB and N14-5 TnsC at various target sites. Preliminary data measured by HTS; ND = no data (less than 5,000 total aligned reads by HTS). Averaged across all sites, the results show that evoCAST improves integration activity by 44-fold without ClpX and 15-fold with ClpX (based on HTS measurements), and averaged across the best sites at each locus, evoCAST improves integration activity by 67-fold without ClpX and 10-fold with ClpX (based on HTS measurements) (Figure 67B). [Figure 68] Figures 6A-C show the results of a guide RNA screening across six sites. The initial screening was quantified by HTS (Figures 68A and 68B), and the most highly edited sites were requantified via ddPCR using genome:transposon junction probes (methods outlined in Lampe, King, et al. Nature Biotechnology 2023) (Figure 68C). All experiments were performed using 1 kb transposon integrations, WT QCascade, WT TnsA, P14-L4-5 TnsB, and N14-5 TnsC. AAVS1-1 in this screening was previously referred to as "AAVS1." HTS and junction ddPCR were generally consistent for most sites, although most sites showed higher values ​​for HTS than for ddPCR, likely due to PCR bias toward the integration amplicon. [Figure 69]Figures A-D show the effect of crRNA architecture on integration efficiency. Atypical and typical crRNAs are responsible for similar integration efficiencies in E. coli for Tn7016. Previous mammalian characterization primarily used atypical crRNA architectures in mammalian cells and found that atypical and typical crRNAs had similar efficiencies for WT Tn7016 CAST in HEK293T cells. Except for the screening of 44 common transgene insertion sites using atypical crRNAs shown in Figure 68, all characterization of evolved variants was performed using typical crRNAs. Comparison of typical versus atypical crRNA architectures for the best edited site(s) from a target site screen performed in HEK293T cells (Figures 69A and 69B). Typical crRNAs outperform atypical crRNAs across all loci tested for evoCAST. Sequences of unprocessed crRNAs ("pre-crRNAs") are: typical Tn7016 Cascade crRNA: GTGACCTGCCGTATAGGCAGCTGAAAAT (SEQ ID NO: 22) [spacer] GTGACCTGCCGTATAGGCAGCTGAAAAT (SEQ ID NO: 22), atypical Tn7016 Cascade crRNA: GTGACCTGCCGTATAGGCAGCTGAAGAT (SEQ ID NO: 23) [spacer] AATTCTCGCCGAAAAGGCAGTGAGTAGT (SEQ ID NO: 24). Previous mammalian characterization primarily used a 33-nt spacer for the crRNA in mammalian cells and found that the 33-nt spacer length had slightly improved activity compared to the 32-nt spacer length for WT Tn7016 CAST in HEK293T cells (Lampe, King et al. Nature Biotechnology 2023). However, characterization of the evolved variants described above was performed using a 32-nt spacer for the crRNA. C and D show a comparison of 32 vs. 33 nt spacer lengths for the most edited sites at each locus from a target site screen performed in HEK293T cells.The 32-nt spacer is equivalent to or superior to the 33-nt spacer across all loci tested for evoCAST. 1 kb transposon integration, WT QCascade, WT TnsA. [Figure 70] Panels A-D show the effect of transfection conditions on integration efficiency. Panels A and B show the effect of transfection conditions on HEK293T cells. Transfection with Lipofectamine 3000 (previously Lipofectamine 2000) and increasing concentrations of puromycin (previously 1 μg / mL) may further increase the integration efficiency observed in HEK293T cells. Panels C and D show the effect of transfection conditions on HeLa cells. Transfection with Lipofectamine 3000 may also improve integration efficiency in HeLa cells (although the efficiency with Lipofectamine 2000 is unusually low). All efficiencies were measured by HTS. [Figure 71] A and B show specificity characterization of evoCAST. A: Schematic of UDiTaS-based detection of off-targets. B: UDiTaS in host E. coli (encoding WT QCascade / TnsA and N14-5 TnsC) after overnight incubation with SP encoding evoTnsB. [Figure 72] A. Schematic of the DNA binding circuit. B. DNA binding circuit using the TnsC-rpoZ fusion. [Figure 73] A-D show DNA-binding-independent phage propagation using Cas6-rpoZ fusions. A is a schematic diagram of Lux assay 1.0. B is a schematic diagram of PANCE 1.0. C and D show fold propagation of two hosts: evoCas78(p6): phage pool from PANCE passage 6; neg.: TnsABC phage; dCas8(R241A, P242A). Phage propagation is most likely independent of target DNA binding. [Figure 74]Figures A–L show the characterization of TniQ-rpoZ and TnsC-rpoZ fusion constructs. Figure A shows a schematic diagram of Lux assay 2.0, which, compared to lux assay 1.0 in Figure 73A, has the following differences: the P3 copy number was changed from p15A to SC101, the P2 promoter / RBS was changed from Jsd8 to pro1 SD8 to potentially avoid a possible hook effect, and the promoter on P1 was changed from Pbad to pro1, allowing for rpoZ-TniQ and TnsC-rpoZ fusions. The lac promoter was optimized for increased signal-to-noise ratio (*), and rpoZ was mutated (****). Figure B shows a schematic diagram of the constructs used in the screening. In this second round of screening, all constructs used the SC101, pro1, SD8 backbone. The rpoZ domain was fused to either Cas6, TniQ, Cas7, or TnsC. The distance between the protospacer and the lac promoter was increased in 2-bp increments to allow for maximum circuit activation upon RNAP recruitment. For each architecture, two different protospacers, AAVS1-1 and O155, were tested. C shows the excellent signal-to-noise ratio with the TnsC-rpoZ fusion on the O155 protospacer but not on the AAVS1-1 protospacer. D shows the signal-to-noise ratio with the rpoZ-TniQ fusion and the O155 spacer. Distance d: distance between the protospacer and Plac*. T: target host with the matching O155 protospacer / spacer sequence. NT: non-target host with the AAVS1-1 protospacer and the O155 spacer (TnsABC circuit spacer). E shows Lux expression with different spacer sequences using the rpoZ-TniQ fusion. F and G show phage encoding the Tn7016 Cascade complex grown in a host with a TnsC-rpoZ fusion. SP Cas678 (Figure 74F) or QCas (Figure 74G). H and I show that phage encoding Tn7016 QCas78 is propagated in a host carrying a TniQ-rpoZ fusion.T: target host with matching 0155 protospacer / spacer sequence; NT: non-target host with AAVS1-1 protospacer and 0155 spacer. J: Overnight propagation of Cas7 and Cas8 in the TniQ-rpoZ DNA binding circuit, demonstrating DNA-binding-dependent phage growth. dCas78: Cas8(R241A,P242A), with impaired DNA unwinding ability (negative control). K: Evolutionary trajectory of Cas7 and Cas8 in the TniQ-rpoZ DNA binding circuit. L: Overnight propagation of Cas7 and Cas8 in the TniQ-rpoZ DNA binding circuit improved phage growth with evolved Cas7 and Cas8. evoCas78: Phage pool after PANCE19 passages in the 0155 spacer. [Figure 75] A shows a schematic diagram of the Cas7 / 8 DNA-binding circuit. This DNA-binding circuit is referred to as the version 5 circuit (v5). Upon successful assembly and target binding of the QCascade complex, RNAP is recruited through the rpoZ (ω) domain, driving gill expression and phage propagation. Evolution for improved complex assembly, target search, and binding. B and C show improved lux signaling by evolved Cas7 / 8 variants. Modestly improved transcriptional activation by evoCas7 / 8 from v5 PACE1. Increased activity with the 0155 spacer correlates with the AAVS1-1 spacer. Transcriptional activation of the L2-2, L3-3, L3-5, and L4-3 variants is significantly above background levels. D and E show improved lux signaling by evolved Cas7 / 8 variants, including genotypes (L2-1, L2-6) containing the rationally identified mutation K235R. The increased lux signal in L4-3 is primarily driven by L4-3 Cas8. (F) Improved phage propagation with evolved Cas7 / 8 phages, i.e., L3-3 Cas78: clonal phage, L1-L3 Cas7 / 8: clonal phage pool, and dCas78:Cas8 (R241A, P242A). (G) and (H) are tables of evolved Cas8 and Cas7 mutations, respectively. In characterization assays, substitutions at K4 and E8 of Cas8 restored wild-type DNA. [Figure 76]A and B show that phage growth / transcriptional activation does not necessarily correlate with mammalian integration efficiency by evolved Cas7 / 8 variants. L4-3 (strongest transcriptional activation in bacterial cells) has among the lowest integration values ​​in mammalian cells. L3-3 (significantly improved activation in bacterial cells) and significantly improved integration. [Figure 77] This shows that evoCas7 and / or evoCas8 are involved in decreasing / increasing integration efficiency. Improvement by L3-3 at the AAVS1-1 site driven by evoCas7. Decreased integration efficiency of L4-3 caused by evoCas8. Decreased integration efficiency of L4-5 caused by evoCas7. [Figure 78] A and B show that the isolated conservative mutations show a significant increase in E. coli transcriptional activation (Figure 78A) but no change in mammalian (HEK293T cell) integration (Figure 78B). [Figure 79] A-D show evolved Cas7 / 8 variants with evoTnsABC across different target sites. L4-3 evoCas7 / 8: had the highest signal in the lux assay but significantly reduced activity in mammalian cells across all target sites tested (Figure 79A). Activity was partially rescued by WT Cas8 (Figure 79B). L4-3 Cas8 significantly reduced integration efficiency across target sites (Figure 79C). Cas7 L3-3 showed a slight improvement in activity across highly edited target sites (Figure 79D). [Figure 80] A and B show the identification of new Cas7 / 8 variants by high-stringency evolution at the sd2 RBS. A shows the genotype from PANCE at the sd2 RBS. B shows the genotype from PACE at the sd2 RBS. Improvement by few variants across the three target sites tested. C and D show mutation tables for evolved Cas8 and Cas7, respectively. In characterization assays, substitutions at K4, E5, L6, I9, D11, and T12 in Cas8 restored wild-type. [Figure 81]A–D show reversion analysis of P14-L4-5 TnsB in HEK293T cells. A shows the evolution of P14-L4-5. Each of the 10 mutations in P14-L4-5 was restored to its wild-type identity (Figure 81B). All mutations appear to contribute modestly to the efficiency of P14-L4-5 (1 kb transposon integration; WT QCascade, WT TnsA, WT TnsC), as each revertant has approximately 50% of the activity of P14-L4-5. Q549R and Q594L appear to contribute less to the increased activity, but reversion of these mutations does not result in variants with significantly higher activity than P14-L4-5. Reversion analysis was also performed with ClpX. Absolute editing efficiencies are shown in C, and the relative integration ClpX:no ClpX is shown in D. WT TnsB benefits substantially from ClpX (approximately 5.5-fold in AAVS1 and approximately 30-fold in HEK3), whereas P14-L4-5 and all single revertants benefit only slightly (approximately 1.5-fold in AAVS1 and HEK3). [Figure 82] 1 shows the characterization of evolved Tn7016 CAST in K562 cell conditions. [Figure 83] A-C show the Cas8 variants in QCascade tested with evoTnsABC. A shows a Cas8 variant containing mutations in two DNA contact interfaces of Cas8: the PAM interaction domain and the helix bundle. B shows the integration efficiency at six different genomic locations. The x-axis labels indicate the Cas8 genotype. C shows a summary of the fold change in T-RL integration relative to WT QCascade. All screens were completed using evoTnsABC (P12-L6-5 TnsA, P14-L4-5 TnsB, and N14-5 TnsC) without ClpX and with 1 kb transposon integration, as shown in Figure 66H. [Figure 84]A-C show the Cas7 variants in QCascade tested with evoTnsABC. A shows the Cas7 variants. B shows the integration efficiency at six different genomic locations. The x-axis labels indicate the Cas7 genotype. C shows a summary of the fold change in T-RL integration relative to WT QCascade. All screens were completed using evoTnsABC (P12-L6-5 TnsA, P14-L4-5 TnsB, and N14-5 TnsC) without ClpX and with 1 kb transposon integration, as shown in Figure 66H. [Figure 85] A and B show the QCascade NLS architecture variants tested with evoTnsABC. Four different architecture variants were tested: original architecture - 1X NLS TniQ + 1X NLS Cas6 + 1X NLS Cas7 + 1X NLS Cas8; NLS architecture 1-2X NLS TniQ + 2X NLS Cas6 + 1X NLS Cas7 + 2X NLS Cas8; NLS architecture 2-3X NLS TniQ + 2X NLS Cas6 + 1X NLS Cas7 + 3X NLS Cas8; and NLS architecture 3-3X NLS TniQ + 2X NLS Cas6 + 1X NLS Cas7 + 4X NLS Cas8. A shows the integration efficiency at six different genomic locations. The x-axis labels indicate the NLS architecture. B shows a summary of the fold change in T-RL integration relative to the original architecture. All screens were completed using evoTnsABC (P12-L6-5 TnsA, P14-L4-5 TnsB and N14-5 TnsC) with WT QCascade, no ClpX, and 1 kb transposon integration, as shown in Figure 66H. [Figure 86]Screening of guide RNAs targeting therapeutically relevant human genomic loci. Forty targets across eight therapeutically relevant loci (five sites per locus) were screened by HTS using evoTnsABC (P12-L6-5 TnsA, P14-L4-5 TnsB, and N14-5 TnsC), WT QCascade, no ClpX, and 1 kb transposon integration, as shown in Figure 66H. DETAILED DESCRIPTION OF THE INVENTION

[0121] In bacteria and archaea, the CRISPR / Cas system provides immunity by integrating fragments of invading phage, viral, and plasmid DNA into CRISPR loci and directing degradation of homologous sequences using the corresponding CRISPR RNA ("crRNA"). Transcription of the CRISPR locus produces a "pre-crRNA," which is processed to yield a crRNA containing a spacer repeat fragment that guides an effector nuclease complex to cleave dsDNA sequences complementary to the spacer. Several different types of CRISPR systems are known (e.g., Type I, Type II, or Type III), classified primarily based on the type of Cas protein and the use of a protospacer adjacent motif (PAM) for selection of the protospacer in the invading DNA.

[0122] Although RNA-guided targeting typically results in endonucleolytic cleavage of the bound substrate, recent studies have revealed a series of atypical pathways in which CRISPR protein-RNA effector complexes are naturally repurposed for alternative functions. For example, some Type I (Cascade) and Type II (Cas9) systems utilize truncated guide RNAs to achieve potent transcriptional repression without cleavage, and other Type I (Cascade) and Type V (Cas12) systems reside within unusual bacterial Tn7-like transposons and lack a nuclease component entirely.

[0123] The present disclosure provides transposon-associated and associated Cas proteins for use in CRISPR-Tn systems, such as type I (Cascade) and type V (Cas12) systems. The present disclosure also provides methods for making the transposon-associated and associated Cas proteins, as well as methods for using the transposon-associated and associated Cas proteins or nucleic acid molecules encoding the transposon-associated and associated Cas proteins in applications involving nucleic acid molecules, such as genome editing. Methods for modifying the transposon-associated and associated Cas proteins described herein can include phage-assisted continuous evolution (PACE) or phage-assisted discontinuous evolution (e.g., PANCE). The present disclosure also provides methods for nucleic acid modification (e.g., RNA-guided DNA integration) utilizing modified CRISPR-transposon systems containing one or more of the disclosed transposon-associated and associated Cas proteins.

[0124] The section headings used in this section and throughout this disclosure are for organizational purposes only and are not intended to be limiting.

[0125] A.Definition The terms "comprise(s)," "include(s)," "having," "has," "can," "contain(s)," and variations thereof, as used herein, are intended to be open-ended transitional phrases, terms, or phrases that do not exclude the possibility of additional acts or structures. As used herein, including a particular sequence or a particular SEQ ID NO: generally means that at least one copy of the sequence is present in the recited peptide or polynucleotide. However, more than one copy is also contemplated. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that "comprising," "consisting of," and "consisting essentially of" embodiments or elements present herein, whether explicitly stated or not.

[0126] When numerical ranges are recited herein, each intervening number to the same degree of precision is expressly contemplated. For example, in the range of 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and in the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.

[0127] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by those of ordinary skill in the art. For example, the nomenclature and techniques used in connection with cell and tissue culture, molecular biology, genetics, and protein and nucleic acid chemistry, and hybridization described herein are those known and commonly used in the art. The meaning and scope of terms should be clear, but in the event of any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definitions. Furthermore, unless otherwise required by context, singular terms shall include the plural and plural terms shall include the singular.

[0128] As used herein, the term "accessory plasmid" refers to a plasmid containing genes necessary for the generation of infectious viral particles under the control of a conditional promoter. In the context of continuous evolution described herein, transcription from the conditional promoter of the accessory plasmid is typically activated by the function of the protein(s) to be evolved. Thus, the accessory plasmid serves to confer a competitive advantage to those viral vectors in a given population of viral vectors that carry a gene of interest capable of activating the conditional promoter. Only viral vectors carrying an "activated" version of the protein(s) of interest are able to induce expression of genes necessary for the generation of infectious viral particles in host cells, thus enabling packaging and propagation of the viral genome in a stream of host cells. On the other hand, vectors carrying an inactivated version of the protein of interest do not induce expression of genes necessary for the generation of infectious viral vectors and therefore are not packaged into viral particles capable of infecting fresh host cells.

[0129] The term "contacting" as used herein refers to bringing into contact or being in contact, being in contact, or coming into contact. As used herein, the term "contact" refers to the state or condition of being in contact, or the state or condition of being in direct or local proximity. Contacting a composition with a target destination, such as, but not limited to, an organ, tissue, cell, or tumor, can occur by any means of administration known to those skilled in the art.

[0130] As used herein, the term "continuous evolution" refers to an evolutionary process in which a population of nucleic acids encoding a protein of interest is subjected to multiple rounds of (a) replication, (b) mutation, and (c) selection to produce a desired evolved protein that differs from the original protein of interest. Multiple rounds can be performed without researcher intervention, and steps (a) through (c) can be performed simultaneously. Typically, the evolution procedure is performed in vitro, e.g., using cells in culture as host cells. Generally, the continuous evolution process provided herein relies on a system in which a gene encoding the protein of interest is provided in a viral vector that undergoes a life cycle that includes replication in a host cell and transfer to another host cell, whereby a critical component of the life cycle, e.g., a gene essential for the production of infectious viral particles, is inactivated, and reactivation of this component relies on the activity of the protein of interest resulting from mutations in the viral vector.

[0131] The term "gene" refers to a DNA sequence that includes regulatory and coding sequences necessary for the production of an RNA having a non-coding function (e.g., ribosomal RNA or transfer RNA), a polypeptide, or a precursor of either of these. The RNA or polypeptide can be encoded by a full-length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Thus, a "gene" refers to a DNA or RNA, or portion thereof, that encodes a polypeptide or an RNA chain that has a functional role to play in an organism. For purposes of this disclosure, a gene may be considered to include regions that regulate the production of a gene product, regardless of whether such regulatory sequences are adjacent to the coding and / or transcribed sequence. Thus, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions.

[0132] A cell has been "genetically modified," "transformed," or "transfected" by exogenous DNA, e.g., a recombinant expression vector, when such DNA has been introduced inside the cell. The presence of the exogenous DNA results in a permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently linked) into the genome of the cell. For example, the transforming DNA may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones comprising a population of daughter cells containing the transforming DNA. A "clone" is a population of cells derived from a single cell or common ancestor by mitosis. A "cell line" is a clone of a primary cell capable of stable growth in vitro for many generations.

[0133] The terms "high copy number plasmid" and "low copy number plasmid" are art-recognized, and one of skill in the art can ascertain whether a given plasmid is a high copy number plasmid or a low copy number plasmid. In some embodiments, a low copy number accessory plasmid is a plasmid having an average copy number of plasmid per host cell in a host cell population of about 5 to about 100. In some embodiments, a very low copy number accessory plasmid is a plasmid having an average copy number of plasmid per host cell in a host cell population of about 1 to about 10. In some embodiments, a very low copy number accessory plasmid is a single copy plasmid per cell. In some embodiments, a high copy number accessory plasmid is a plasmid having an average copy number of plasmid per host cell in a host cell population of about 100 to about 5000.

[0134] The terms "homology" and "homologous" refer to the degree of identity. There may be partial or complete homology. A partially homologous sequence is one that is less than 100% identical to another sequence.

[0135] As used herein, the term "host cell" refers to a cell capable of hosting, replicating, and transferring a phage vector useful in the continuous evolution process provided herein. In embodiments in which the vector is a viral vector, a suitable host cell is one that can be infected with the viral vector, replicate it, and package it into viral particles that can infect fresh host cells. A cell can host a viral vector if it supports expression of the viral vector's genes, replication of the viral genome, and / or production of viral particles. One criterion for determining whether a cell is a suitable host cell for a given viral vector is whether the cell can support the viral life cycle of the wild-type viral genome from which the viral vector is derived. For example, as provided in some embodiments herein, if the viral vector is a modified M13 phage genome, a suitable host cell would be any cell capable of supporting the wild-type M13 phage life cycle. Suitable host cells for viral vectors useful in continuous evolution processes are known to those of skill in the art, and the disclosure is not limited in this respect. In some embodiments, the viral vector is a phage and the host cell is a bacterial cell. In some embodiments, the host cell is an E. coli cell. Suitable E. coli host strains will be apparent to those skilled in the art and include, but are not limited to, New England Biolabs (NEB) Turbo, ToplOF', DH12S, ER2738, ER2267, and XL1-Blue MRF'. These strain names are recognized in the art, and the genotypes of these strains have been well characterized. It should be understood that the above strains are merely exemplary, and the invention is not limited in this respect. The term "fresh," used interchangeably herein with the terms "uninfected" or "uninfected" in the context of host cells, refers to host cells that have not been infected with a viral vector containing the gene of interest used in the continuous evolution process provided herein.However, fresh host cells may also be infected with a viral vector unrelated to the vector to be evolved, or with a vector of the same or similar type but not carrying the gene of interest. In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the host cell is a eukaryotic cell, such as a yeast cell, an insect cell, or a mammalian cell. The type of host cell will, of course, depend on the viral vector used, and suitable host cell / viral vector combinations will be readily apparent to one of skill in the art.

[0136] As used herein, the term "hybridization" is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (e.g., the strength of the association between nucleic acids) depend on the degree of complementarity between the nucleic acids, the stringency of the conditions involved, and the T of the hybrid formed. m Hybridization is influenced by factors such as: temperature, ionic strength, and ionic strength. Hybridization methods involve annealing one nucleic acid to another complementary nucleic acid, e.g., a nucleic acid having a complementary nucleotide sequence. The ability of two nucleic acid polymers containing complementary sequences to find each other and "anneal" or "hybridize" through base-pairing interactions is a well-recognized phenomenon. Following the initial observation of the "hybridization" process by Marmur and Lane, Proc. Natl. Acad. Sci. USA, 46:453 (1960) and Doty et al., Proc. Natl. Acad. Sci. USA, 46:461 (1960), this process has been refined into an essential tool in modern biology. For example, hybridization and washing conditions are now well known and are exemplified in Sambrook et al., supra. Conditions of temperature and ionic strength determine the "stringency" of hybridization.

[0137] As used herein, the term "lagoon" refers to a culture vessel or bioreactor into which a stream of host cells is directed. As used in the continuous evolution process described herein, the lagoon typically holds a host cell population and a viral vector population that replicates within the host cell population, and the lagoon includes an outlet through which host cells are removed from the lagoon and an inlet through which fresh host cells are introduced into the lagoon, thereby replenishing the host cell population.

[0138] As used herein, "nucleic acid" or "nucleic acid sequence" refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982)). The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, as well as any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. The polymer or oligomer may be heterogeneous or homogeneous in composition and may be isolated from naturally occurring sources or artificially or synthetically produced. Furthermore, the nucleic acid may be DNA or RNA, or a mixture thereof, and may exist permanently or transiently in single- or double-stranded form, including homoduplexes, heteroduplexes, and hybrid states. In some embodiments, the nucleic acid or nucleic acid sequence comprises other types of nucleic acid structures, such as, for example, a DNA / RNA helix, a peptide nucleic acid (PNA), a morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002) and U.S. Patent No. 5,034,506), a locked nucleic acid (LNA, see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97:5633-5638 (2000)), a cyclohexenyl nucleic acid (see Wang, J. Am. Chem. Soc., 122:8595-8602 (2000)), and / or a ribozyme. Thus, the terms "nucleic acid" or "nucleic acid sequence" can also encompass strands that include non-naturally occurring nucleotides, modified nucleotides, and / or non-nucleotide components (e.g., "nucleotide analogs") that can perform the same function as natural nucleotides; furthermore, as used herein, the term "nucleic acid sequence" refers to oligonucleotides, nucleotides, or polynucleotides, and fragments or portions thereof, as well as DNA or RNA of genomic or synthetic origin, which may be single-stranded or double-stranded and represent the sense or antisense strand.The terms "nucleic acid," "polynucleotide," "nucleotide sequence," and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.

[0139] The "identity" of a nucleic acid or amino acid sequence described herein can be determined by comparing a subject nucleic acid or amino acid sequence to a reference nucleic acid or amino acid sequence. Numerous mathematical algorithms for obtaining optimal alignments and calculating identity between two or more sequences are known and are incorporated into many available software programs. Examples of such programs include CLUSTAL-W, T-Coffee, and ALIGN (for aligning nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST2.1, BL2SEQ, and later versions), and FASTA programs (e.g., FASTA3x, FAS™, and SSEARCH) (for sequence alignment and sequence similarity searches). Sequence alignment algorithms are described, for example, in Altschul et al., J. Molecular Biol., 215(3):403-410 (1990), Beigert et al., Proc. Natl. Acad. Sci. USA, 106(10):3770-3775 (2009), Durbin et al., eds., Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, Cambridge, UK (2009), Soding, Bioinformatics, 21(7):951-960 (2005), Altschul et al., Nucleic Acids Res., 25(17):3389-3402 (1997), and Gusfield, Algorithms on Strings, Trees and Sequences, Cambridge University Press, Cambridge It is also disclosed in UK (1997).

[0140] The terms "non-naturally occurring," "modified," and "synthetic" are used interchangeably and indicate the involvement of the hand of man. These terms, when referring to a nucleic acid molecule or polypeptide, mean that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated in nature and which is found in nature.

[0141] The term "phage," used interchangeably herein with the term "bacteriophage," refers to a virus that infects bacterial cells. Typically, phages consist of an outer protein capsid that encloses genetic material. The genetic material can be ssRNA, dsRNA, ssDNA, or dsDNA, in either linear or circular form. Phages and phage vectors are known to those of skill in the art; non-limiting examples of phages useful for practicing the methods provided herein include λ(Lysogen), T2, T4, T7, T12, R17, M13, MS2, G4, PI, P2, P4, Phi X174, N4, Φ6, and Φ29. In a specific embodiment, the phage used in the present invention is M13. Additional suitable phages and host cells will be apparent to those of skill in the art, and the present invention is not limited in this respect. For an exemplary description of additional suitable phages and host cells, see Elizabeth Kutter and Alexander Sulakvelidze: Bacteriophages: Biology and Applications. CRC Press; st edition(December 2004), ISBN:0849313368, Martha RJClokie and Andrew M.Kropinski:Bacteriophages:Methods and Protocols,Volume 1:Isolation,Characterization,and Interactions(Methods in Molecular Biology)Humana Press, 1 stedition(December,2008),ISBN:1588296822, Martha RJClokie and Andrew M.Kropinski:Bacteriophages:Methods and Protocols,Volume 2:Molecular and Applied Aspects(Methods in Molecular Biology)Humana Press, 1 st edition (December 2008), ISBN: 1603275649, all of which are incorporated herein by reference in their entireties for their disclosure of suitable phages and host cells, as well as methods and protocols for the isolation, culture, and manipulation of such phages.

[0142] The term "phage-assisted continuous evolution" or "PACE" as used herein refers to continuous evolution using phage as a viral vector. PACE technology has previously been described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010 as WO2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, and published June 28, 2012 as WO2012 / 088381; International PCT Application No. PCT / US2015 / 057012, filed October 22, 2015, and published May 19, 2016 as WO2016 / 077052; No. 9,267,127, filed on June 20, 2013, issued based on U.S. application Ser. No. 13 / 922,812, all of which are incorporated herein by reference.

[0143] As used herein, the term "phage-assisted discontinuous evolution" or "PANCE" refers to discontinuous evolution using phages as viral vectors. The general concept of PANCE technology is described, for example, in Suzuki T. et al., "Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase," Nat Chem Biol. 13(12):1261-1266 (2017), incorporated herein by reference in its entirety. Briefly, PANCE is a simplified technique for rapid in vivo directed evolution using serial flask transfers of evolving "selection phage" (SP) containing the gene of interest to be evolved between fresh E. coli host cells, thereby allowing the gene contained in the SP to be continuously evolved while maintaining the gene constant within the host E. coli. Following phage growth, an aliquot of infected cells is used to transfect subsequent flasks containing host E. coli. This process is continued for the required number of transfers until the desired phenotype has evolved. Serial flask transfer has long been a widely available approach for the laboratory evolution of microorganisms, and more recently, similar approaches have been developed for bacteriophage evolution. In general, the PANCE system is characterized by lower stringency than the PACE system.

[0144] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked by peptide bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, or a linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide may also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide may be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, modified, synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods of recombinant protein expression and purification are known, including those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0145] As used herein, the terms "providing," "administering," and "introducing" are used interchangeably herein and refer to placing a system of the present disclosure into a cell, organism, or subject by a method or route that results in at least partial localization of the system to a desired site. The system may be administered by any suitable route that results in delivery to a desired location within a cell, organism, or subject.

[0146] As used herein, the term "selection phage" is used interchangeably with the term "selection plasmid" and refers to a modified phage that contains a gene of interest to be evolved and lacks a full-length gene encoding a protein required for the generation of infectious phage particles. For example, some M13 selection phages provided herein contain a nucleic acid sequence encoding one or more transposases to be evolved, e.g., under the control of an M13 promoter, and lack all or a portion of the phage genes encoding proteins required for the generation of infectious phage particles, e.g., gI, gII, gIII, gIV, gV, gVI, gVII, gVIII, gIX, or gX, or any combination thereof. For example, some M13 selection phages provided herein contain a nucleic acid sequence encoding one or more transposases to be evolved, e.g., under the control of an M13 promoter, and lack all or a portion of the genes encoding proteins required for the generation of infectious phage particles, e.g., the gIII gene encoding the pIII protein.

[0147] A "subject" or "patient" may be human or non-human and may include, for example, animal strains or species used as "model systems" for research purposes, such as the mouse model described herein. Similarly, a patient may include either an adult or a juvenile (e.g., a child). Furthermore, a patient may refer to any organism, preferably a mammal (e.g., human or non-human), that can benefit from the administration of the compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the class Mammalia, i.e., humans, non-human primates such as chimpanzees, and other ape and monkey species; farm animals such as cows, horses, sheep, goats, pigs; domestic animals such as rabbits, dogs, and cats; and laboratory animals, including rodents such as rats, mice, and guinea pigs. Examples of non-mammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods and compositions provided herein, the mammal is a human.

[0148] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an "insert," can be attached or incorporated, so as to bring about the replication of the attached segment in a cell.

[0149] Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in the practice or testing of this disclosure. All publications, patent applications, patents, and other documents mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and are not intended to be limiting.

[0150] CRISPR transposon protein components Disclosed herein are modified transposon-associated proteins and Cas proteins. Additionally, disclosed are nucleic acids and vectors comprising sequences encoding the modified transposon-associated proteins and Cas proteins.

[0151] Modified transposon-associated proteins and / or Cas proteins may confer desirable properties (e.g., increased stability, increased activity) not found in the wild-type version of the protein. In some embodiments, the modified proteins exhibit increased activity or utility in modifying target nucleic acids compared to proteins without the disclosed modifications. In some embodiments, the modified proteins have increased target DNA binding activity compared to proteins without the disclosed modifications. In some embodiments, the modified proteins have increased nucleic acid integration activity at a target nucleic acid compared to proteins without the disclosed modifications. In some embodiments, the modified proteins have increased nucleic acid integration activity or efficiency at a target nucleic acid in vivo (e.g., in a prokaryotic or eukaryotic cell, in a subject) compared to proteins without the disclosed modifications. In some embodiments, the combination of modified transposon-associated proteins and / or Cas proteins confers desirable properties. In some embodiments, the combination of one or more modified transposon-associated proteins and / or Cas proteins with one or more wild-type transposon-associated proteins and / or Cas proteins confers desirable properties.

[0152] Provided herein are polypeptides comprising one or more amino acid sequences having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to any of SEQ ID NOs: 1-14. In some embodiments, the polypeptides have one or more amino acid substitutions, deletions, or additions relative to SEQ ID NOs: 1-14. In some embodiments, the polypeptides have one or more amino acid substitutions, deletions, or additions as shown in Tables 1-4 relative to SEQ ID NOs: 1-14.

[0153] Any of the proteins described or referenced herein may contain one or more amino acid substitutions compared to the listed sequences. An amino acid "substitution" or "substitution" refers to the replacement of one amino acid at a given position or residue within a polypeptide sequence with another amino acid at the same position or residue. Amino acids are broadly classified as "aromatic" or "aliphatic." Aromatic amino acids contain an aromatic ring. Examples of "aromatic" amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y or Tyr), and tryptophan (W or Trp). Non-aromatic amino acids are broadly classified as "aliphatic." Examples of "aliphatic" amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Val), leucine (L or Leu), isoleucine (I or Ile), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine ​​(C or Cys), proline (P or Pro), glutamic acid (E or Glu), aspartic acid (D or Asp), asparagine (N or Asn), glutamine (Q or Gln), lysine (K or Lys), and arginine (R or Arg).

[0154] Amino acid substitutions or substitutions can be conservative, semi-conservative, or non-conservative. The phrase "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid that shares common properties. A practical method for defining common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz and Schirmer, Principles of Protein Structure, Springer-Verlag, New York (1979)). Such analysis reveals that amino acids within a group preferentially exchange with each other, and thus groups of amino acids that are most similar to each other in their impact on the overall protein structure can be defined (Schulz and Schirmer, supra). Examples of conservative amino acid substitutions include amino acid substitutions within the subgroups described above, such as arginine for lysine and vice versa to maintain a positive charge, aspartic acid for glutamic acid and vice versa to maintain a negative charge, threonine for serine to maintain a free -OH, and asparagine for glutamine to maintain a free -NH. "Semi-conservative mutations" include amino acid substitutions within the same group listed above but not within the same subgroup. For example, asparagine for aspartic acid or lysine for asparagine includes amino acids within the same group but in a different subgroup. "Non-conservative mutations" include amino acid substitutions between different groups, such as tryptophan for lysine or serine for phenylalanine.

[0155] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:1 and one or more amino acid substitutions at positions 2, 3, 5, 28, 57, 77, 80, 107, 110, 116, 122, 142, 155, 161, 166, 173, 177, 185, 211, 216, 227, and 230 relative to SEQ ID NO:1. In some embodiments, the polypeptide comprises an amino acid sequence having one or more amino acid substitutions from SEQ ID NO:1, including A2T, T3I, L5S, T28A, A57T, F77L, Y80D, K107M, K107R, Y110C, Y110D, D116G, E122A, D142E, M155I, K161R, N166D, K173E, Y177N, Y177D, C185R, D211Y, K216E, A227P, G230D, and G230S.

[0156] In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions at positions 2 and 230; 107 and 166; one or both of 107, 166, and 2 and 227; 211 and 110 or 142; 110, 155 and 230; 122 and 155; 155 and 177 relative to SEQ ID NO: 1. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions of A2T and G230D or G230S relative to SEQ ID NO: 1. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions of K107M and N166D relative to SEQ ID NO: 1. In some embodiments, the polypeptide further comprises amino acid substitutions of A2T and / or Y177N or Y177D relative to SEQ ID NO: 1. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions of D211Y and Y110C, Y110D, or D142E relative to SEQ ID NO: 1. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 1: Y110C or Y110D, M155I, and G230D or G230S. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 1: E122A and M155I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 1: M155I and Y177N or Y177D.

[0157] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:2 and one or more amino acid substitutions at positions 2, 5, 22, 24, 25, 29, 75, 141, 199, 215, 319, 347, 364, 370, 383, 439, 454, 458, 485, 509, 533, 538, 565, 581, 586, 595, 596, 597, and 600 relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having one or more amino acid substitutions of A2T, A2S, G5R, S22P, E24D, L25I, A29S, P75T, I141T, V199I, S215R, D319V, Y347F, S364N, E370K, N383D, V439A, E454D, E454G, S458N, V485F, R509G, D533A, A538V, H565Y, A581T, H586L, N595K, D596N, D597N, D597Y, and I600V relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions at positions 2 and 597; 24 and 25; 24, 25, 458, 509, 565, and 600; 75 and 597; 141, 454, 533 and 595; 581, 370, and 454; 370 and 581; 370 and 454; 458 and 509; 458, 509 and 565; 458, 509, 565, and 600; 565, 586, and 596; or 565, 509, 458, 600 and 24, 25, 29, 215, 319, 364, 383, and 586 relative to SEQ ID NO:2.

[0158] In some embodiments, the polypeptide comprises an amino acid sequence having the amino acid substitutions A2T or A2S, and D597N or D597Y relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having the amino acid substitutions E24D and L25I relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having the amino acid substitutions E24D and L25I, S458N, R509G, H565Y, and I600V relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having the amino acid substitutions P75T and D597N or D597Y relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having the amino acid substitutions E454D or E454G, D533A, and N595K relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having the amino acid substitutions E370K, E454D or E454G, and A581T relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having E370K, and A581T, E454D, or E454G amino acid substitutions relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having S458N and R509G amino acid substitutions relative to SEQ ID NO:2. In some embodiments, the polypeptide further comprises H565Y and / or I600V amino acid substitutions. In some embodiments, the polypeptide comprises an amino acid sequence having H565Y, H586L, and D596N amino acid substitutions relative to SEQ ID NO:2. In some embodiments, the polypeptide comprises an amino acid sequence having H565Y, R509G, S458N, I600V, and at least one of E24D, L25I, A29S, S215R, D319V, S364N, N383D, and H586L amino acid substitutions relative to SEQ ID NO:2.

[0159] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:3 and one or more amino acid substitutions at positions 9, 15, 16, 18, 21, 64, 81, 86, 87, 99, 109, 142, 147, 153, 168, 180, 216, 230, 285, and 304 relative to SEQ ID NO:3. In some embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 3: I9V, A15V, F16Y, S18F, S21N, N64D, H81Y, D86Y, N87K, V99I, E109D, E142K, V147I, N153D, I168M, A180E, A216S, L230F, K285E, and R304R. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions at positions 142 and 216 relative to SEQ ID NO: 3. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 3: E142K and A216S.

[0160] As used herein, a sequence identity of at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO:4, as well as sequences at positions 4, 5, 9, 10, 12, 21, 23, 25, 26, 31, 32, 34, 35, 37, 41, 45, 47, 48, 51, 52, 55, 60, 61, 65, 67, 69, 72, 75, 79, 80, 82, 87, 88, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 1 , 91, 93, 94, 96, 98, 99, 100, 103, 106, 108, 113, 116, 125, 126, 128, 129, 135, 139, 143, 146, 147, 149, 153, 154, 156, 158, 159, 160, 162, 164, 166, 167, 168, 169, 170, 177, 179, 180, 182, 183, 185, 187, 188, 190, 191, 192, 193, 195, 196, 200, 204, 207, and 208.In some embodiments, the polypeptide is selected from the group consisting of R4K, N5K, P9S, A10P, N12D, T21I, V23M, S25N, S25R, V26M, V26G, S31N, S32I, E34A, F35L, A37D, H41L, D45N, I47V, E48G, G51V, S52I, E55K, E55D, E60K, F61L, S65T, S65A, P67T, P67L, P67S, P67 H, T69A, A72V, A72D, S75I, S75R, S75T, K79E, T80P, K82E, K87R, P88L, P88T, P88A, S90F, K91N, K91E, A93T, A93S, S9 4N, L96P, R98Q, A99D, A99V, E100K, A103T, A106T, S108A, I113F, V116F, V116I, V125M, V125A, N126T, I128V, I128L , L129P, L135M, S139N, S139G, G143V, G143C, G146D, G146S, I147V, K149E, K149T, K149R, S153I, S153R, S153N, F15 4C, H156R, H156L, S158N, S158R, G159V, V160A, K162R, N164D, I166L, S167I, S168I, S168R, S168N, Q169R, V170M, V The amino acid sequence includes an amino acid sequence having one or more amino acid substitutions selected from the group consisting of 170G, V170L, T177I, T177A, S179R, F180C, F180L, F182C, F182L, G183S, M185I, K187R, G188D, V190I, K191N, A192S, D193N, G195V, G195D, G195S, C196W, T200A, T204I, A207V, A207T, and T208I.

[0161] In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions at positions 108 and 47 or 208; 170 and 207; 88 and 147; 47, 88 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 88, 128, 147, 170, and 182; or 170, 207, and 108 relative to SEQ ID NO: 4. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions of S108A, and I47V, or T208I relative to SEQ ID NO: 4. In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions of V170M, and A207V, or A207T relative to SEQ ID NO: 4. In some embodiments, the polypeptide further comprises the following amino acid substitutions relative to SEQ ID NO:4: S108A, V170M, and A207V or A207T.

[0162] As used herein, a sequence identity of at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO:5, as well as sequences at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 11 , 49, 52, 55, 56, 58, 60, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 110, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 145, 146, 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169, 173, 174, 176, 181, 182, 1 86, 187, 190, 195, 197, 198, 205, 208, 209, 211, 215, 218, 223, 226, 227, 231, 232, 235, 239, 246, 248, 250, 259, 260, 261, 262, 263, 267, 269, 273, 274, 2 77, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298, 302, 303, 307, 313, 316, 317, 320, 323, 325, 331, 332, 339, 345, 348, 349, 352, 353, 354, 356, 36 1, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373, 375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 404, 407, 408, 410, 411, 412 , 413, 414, 415, 416, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456, 459, 460, 463, 464, 470,472, 473, 494, 495, 498, 501, 502, 504, 505, 506, 508, 509, 510, 512, 513, 514, 517, 520, 521, 522, 525, 526, 527, 530, 531, 532, 533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 556, 557, 558, 559, 560, 561, 562 Polypeptides having one or more amino acid substitutions at positions 2, 563, 564, 565, 567, 568, 569, 570, 571, 574, 575, 576, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656 are provided.

[0163] In some embodiments, the polypeptide is selected from the group consisting of M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L, L15Q, K16N, K16R, P17T, P17L, P17S, T19I, T19S, T19A, T19P, P20S, P20L, T21A, Q22R, Y23H, V24M, K25R, L26M, D27A, D27G, D28N, D28Y, A29T, A29L, A29R, A29S ... V, N30K, I32F, I32S, Q33H, L36M, D37A, D37Y, F39L, S40P, D41E, T42I, T42K, T42A, F43L, F43S, F43V, K44N, N45D, N45S, Q49R, K52Q, S55A, T56A, D58E, K6 0Q, S62T, R63K, R63G, Q67R, Q67H, Q67K, D71Y, K74R, E76K, F78C, K79R, G80V , G80D, G81S, G81V, G81D, D82N, V83G, V83M, V83A, V84A, V84G, R85G, R85K, P8 6L, N87S, R89C, V91G, V91A, A92V, A92T, R95K, K97R, E100D, S101A, D104V, A 106D, A106T, D110N, N112H, H113Y, M115R, N117Y, T119A, N120D, N120K, N12 0S, G124V, D125N, D125E, K127R, F129L, D130N, K131M, E134D, E134G, A139S , A139T, P142S, I144V, A145S, A145T, T146A, A147V, Q149R, Y150H, I155L, V1 56A, V156L, V156M, K157V, E158A, N159S, V163G, E164A, E164G, E164D, G165 D, I167V, I169L, I169T, N173S, N173H, N173T, A174S, A174T, N176D, A181S, I 182L, I182V, I182T, A186E, A186T, V187G, V187A, A190T, A190S, F195S, A19 7P, D198G, D198N, A205S, V208M, P209T, T211I, E215D, E218D, P223S, P223H,L226V, I227V, D231N, E232K, I235V, I235T, R239G, I246V, V248E, V248M, S250I, S259N, Y260C, K261R, S262N, P263L, S267N, A269V, T273I, T273N, H274 Y, K277N, K277R, P278S, S280T, L281M, D282E, D282N, A283T, A283S, N285S, E287D, L288M, N290K, F295S, F298I, F298S, V302I, V303M, A307S, N313S, H31 6R, A317V, S320N, S320R, I323L, I325V, R331K, K332E, I339V, V345L, V345M, E348K, Y349H, Y349D, Y349N, Y349C, P352S, P352T, E353Q, E353D, L354M, G 356S, N361D, I362V, I362T, L363P, L363T, L363M, E364G, K365R, E366G, E367G, K369N, K369E, K369M, P370S, E371K, V372M, D373G, I375V, M376I, T380P T380A, E383K, E383D, F385L, H386Y, I389V, A390V, A390I, V392I, D396N, D396G, D396K, S397P, S399N, S399G, T402I, R403G, R403I, R403K, R403S, I404 T、I404V、K407R、K407E、R408K、Q410K、Q410H、Q410R、Q411H、G412V、F413L、 D414N、A415V、A415T、Y416C、M421I、N422K、E423K、E423D、E424A、E425K、E42 6D、T427A、T427S、R428K、F429L、S430A、M431L、R434H、R434C、R434S、I435V 、D437G、D437N、T440S、T440I、R443C、G445S、F446L、F446I、Y448C、E450D、E 450G、M452I、T456P、T456A、T456I、A459T、D460N、K463N、H464N、H464R、H46 4S、E470K、V472M、V472A、K473D、K473N、E494D、E494G、S495A、E498A、E498K、C501Y, T502I, T502S, P504S, P504L, T505A, G506Y, G506D, G506L, G506S, T508A, D509E, D509Y, C510Y, S512N, I513L, I513V, I513F, Y514H , K517M, K517N, K517Q, K520R, K521N, I522T, I522V, I522F, E525K, V526E, V526M, I527V, S530N, S530R, K531T, D532G, D532Y, S533Y, G535D , A537T, K538R, K538N, R540K, R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K, Q549R, E550A, Q551 K, E552D, E552K, V553I, F554V, E556K, E556G, S557A, K558R, T559P, T559I, T559A, K560R, A561T, A561G, K562R, K562N, I563L, T564I, A565 S, A565V, K567R, K568N, K568R, Q569K, Q569L, Q569R, A570V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S, I583V, K58 4R, V585M, S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S589R, S589N, A590S, A590T, A591V, P592L, V593M, V593A, Q594L, K59 5R, K595N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N, D600G, D600V, N601S, N601K, S602A, S602P, S602Y, D603A, D603V, D604G, D604Y, D604N, D606A, D606V, D606Y, D607Y, D607E, D608N, A611T, E613D, R618I, T620P, and A656V.

[0164] In some embodiments, the polypeptide is located at positions 4, 23, and 590; 19, 169, and 549; 43 and 415; 80 and 593; 80, 144, 593, and 606; 1, 42, 80, 593, and 606; 42, 80, 593, and 606; 156 and 604; 283, 349, and 365; 283, 349, 365, 396, and 594; 283, 349, 365, 396, 594, 596, and 131; 352 and 390; 390, 396, and 594; 396 and 594; 456 and 502; 464 and 595; 02; 464 and 17; 17, 235, 464, and 596; 235, 352, 396, 456, and 606; 415, 456, and 502; 456, 502, and 549; 169, 456, 502, and 549; 80, 456, 502, 593, and 606; 1, 42, 80, 456, 502, 593, and 606; 80, 144, 456, 502, 593, and 606; 19, 169, 456, 502 and 549; 43, 415, 456, and 502; 352, 390, 396, and 594; 352, 390, and 396; 283, 349, 39 6, and 594; 11, 55, 120, 362, 584, 600, and 604; 43, 84, 144, 349, and 517; 164 and 165; 164 and 173; 362 and 446; 352, 390, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594 and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526. One or more positions; 43, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 21; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 21, and 67;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 174, 208, 427, 456, and 504; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 139; , 526, 415, 502, 339, and 446;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 19, 460, 569, and 596;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 460, 586, 588, and 608;43, 349, 352, 3 90, 396, 464, 549, 594, 410, 526, and 460; 352, 390, 396, 549, 586, and 594; 63, 158, 352, 390, 396, 549, 586, and 594; 164, 165, 352, 363, 390, 396, 410, 549, 586, and 594; 164, 173, 352, 39 or amino acid sequences having amino acid substitutions at 283, 349, 365, 396, 594, 596, and 131.

[0165] In selected embodiments, the polypeptide comprises an amino acid substitution at positions 352, 390, 396, 594, or any combination thereof, relative to SEQ ID NO:5.

[0166] In selected embodiments, the polypeptide comprises, relative to SEQ ID NO:5, positions F43, Y349, P352, A390, D396, H464, Q549, Q594, and T456; F43, Y349, P352, A390, D396, H464, Q549, Q594, T456, and V526; F43, Y349, P352, A390, D396, H464, Q549, Q594, and P504; F43, Y349, P352, A390, D396, H464, Q549, Q594, and V526; F43, Y349, P352, A390, D396, H464, Q549, Q594, and V526; 6, H464, Q549, Q594, Q410, and V526; F43, Y349, P352, A390, D396, H464, Q549, Q594, A174, and T427; F43, Y349, P352, A390, D396, H464, Q549, Q594, and V208; F43, Y349, P352, A390, D396, H464, Q549, Q594, R63, A145, I182, and V526; F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A415, and T50 2;F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A415, T502, and T21;F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A415, T502, and Q67;F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A415, T502, T21, and Q67;F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A174, V208, T427, T456, and P504; F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A174, V208, T427, T456, and P504; F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A415, T502, and A139; F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, A415, T502, I339, and F446;containing amino acid substitutions at F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, T19, D460, Q569, and H596; F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, D460, S586, E588, and D608; F43, Y349, P352, A390, D396, H464, Q549, Q594, Q410, V526, and D460;

[0167] In some embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of F4L, Y23H, and A590S relative to SEQ ID NO: 5. In some embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of T19P, I169L, and 549 relative to SEQ ID NO: 5. In some embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of F43L and A415V relative to SEQ ID NO: 5. In some embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of G80D and V593M or V593A relative to SEQ ID NO: 5. In some embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of G80D, I144V, V593M or V593A, and D606A or D606V relative to SEQ ID NO: 5. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: M1V, M1I or M1L, T42I or T42A, G80D, V593M or V593A, and D606A or D606V. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: T42I or T42A, G80D, V593M or V593A, and D606A or D606V. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: V156A or V156M, and D604G or D604N. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: A283S or A283T, Y349H, and K365R. In some embodiments, the polypeptide further comprises amino acid substitutions of D396K and Q594K or Q594L relative to SEQ ID NO: 5. In some embodiments, the polypeptide further comprises amino acid substitutions of P352S or P352T, H596Y or H596L, and K131M relative to SEQ ID NO: 5. In some embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of P352S or P352T and A390V relative to SEQ ID NO: 5.In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: A390V, D396K, and Q594K or Q594L. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: D396K and Q594K or Q594L. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: T456P, T456I, or T456A and T502I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: H464N or H464R and T502I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: H464N or H464R and P17T, P17L, or P17S. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: P17T, P17L, or P17S, I235V or I235T, H464N or H464R, and Q569K, Q569L, or Q569R. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: I235V or I235T, P352S or P352T, D396K, T456P, T456I, or T456A, and D606A or D606V. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: A415V, T456P, and T502I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: T456P, T502I, and Q549K. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: I169L, T456P, T502I, and Q549K. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: G80D, T456P, T502I, V593M, and D606A.In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: M1V, T42I, G80D, T456P, T502I, V593M, and D606A. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: G80D, I144V, T456P, T502I, V593M, and D606A. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: T19P, I169L, T456P, T502I, and Q549K. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43L, A415V, T456P, and T502I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: P352T, A390V, D396N, and Q594L. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: P352T, A390V, and D396N. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: A283T, Y349H, D396N, and Q594L. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: A283T, Y349H, P352S, K365R, D396N, Q594L, H596L, and K131M. In selected embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 5: P352T, A390V, D396N, and Q594L.

[0168] In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, and T456I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, T456P, and V526E. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, and P504S. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, and V526E. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, Q410K, and V526E. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, A174S, and T427S. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, and V208M. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N, P352T, A390V, D396N, H464R, Q549R, Q594L, R63G, A145S, I182T, and V526E.

[0169] In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, and T502I. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and T21A. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO: 5: F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, T21A, and Q67K.

[0170] As used herein, a sequence identity of at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO:6, as well as sequences at positions 1, 2, 3, 5, 6, 7, 9, 11, 12, 14, 21, 22, 26, 27, 31, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 12 38, 43, 44, 46, 47, 54, 59, 60, 61, 64, 65, 67, 68, 71, 72, 74, 76, 79, 80, 81, 84, 89, 95, 102, 105, 109, 110, 111, 112, 113, 114, 116, 118, 119, 120, 123, 129, 130, 131, 132, 134, 142, 145, 146, 147, 148, 150, 154, 155, 166, 169, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233 80, 181, 183, 184, 187, 190, 194, 197, 201, 204, 207, 209, 213, 219, 221, 225, 226, 227, 229, 232, 233, 234, 236, 238, 241, 246, 251, 252, 256, 257, 261, 263, 265, 267, 269, 271, 272, 274, 280, 281, 285, 286, 288, 291, 292, 296, 299 , 301, 303, 304, 306, 307, 308, 310, 313, 314, 316, 317, 318, 319, 320, 323, 324, 326, 328, 330, 331, 332, 340, 341, 343, 344, 355, 412, 418, 427, 514, 1198, 1201, 1206, 1212, 1260, and 1282 are provided. In some embodiments, the polypeptide is selected from the group consisting of M1L, M1V, N2S, A3T, T5P, T5A, T5S, E6D, I7S, I7V, I9F, Q11R, L12M, N14D, N14S, M21I, H22P, H22Y, K26N, K26R, T27I, M31I, L35R, N38S, S43P, D44N, D44G, Q46L, C47S, T54I, S59T, H60Y, T61A, H64Y, Y65H, K67N, K67R, R68Q, A71G, T72A, N74D, S76C, S76Y, T79I, M80I, P81S, V84L,R89L, A95D, A95T, A102T, E105D, E105K, S109N, S109R, S110P, Q111R, I112T, K113N, K113E, K114N, K114M, K114E, G116D, K118N, K118R, T1 19I, D120V, K123N, L129M, I130V, K131R, A132S, K134M, K134N, F142V, L145M, I146T, E147K, F148S, S150F, R154K, Q155H, E166D, K169E, P 178S, A180V, A181T, A181S, I183V, A184S, A184T, A184V, P187S, A190T, A190V, V194M, V194A, R197I, Y201N, L204M, D207N, K209N, Q213H, Q213V, A219S, K221N, D225N, V226E, P227T, K229E, S232N, K233N, K233R, N234H, T236A, A238V, A238S, A241S, E246D, K251N, H252Y, H252R, E256D, A257S, A261V, S263I, S263N, N265D, Y267C, E269K, E269D, K271E, K271R, H272Y, I274V, F280L, D281N, D281G, K285G, K286N, K288R , S291F, S291P, K292N, K296R, K296N, I299S, D301G, E303D, I304T, I304V, E306G, V307L, V307G, V307A, V307D, V307G, I308N, N310S, Y313 H, N314K, N316K, N316D, A317D, L318Q, D319N, P320S, P320L, M323I, L324M, D326N, V328M, V328A, A330D, I331V, V332G, S340L, T341A, A343G, S344N, I355V, F412V, V418F, Y427C, R514K, S1198L, A1201V, G1206S, C1212G, F1260L, and V1282M.

[0171] In some embodiments, the polypeptide does not comprise a substitution at one or more positions selected from 6, 9, 28, 44, 64, 76, 80, 95, 110, 113, 114, 116, 118, 130, 132, 142, 155, 187, 190, 194, 221, 233, 234, 238, 261, 272, 280, 281, 299, 303, 304, 307, 308, 313, 316, and 328 relative to SEQ ID NO:6. In some embodiments, the polypeptide is selected from the group consisting of E6D, I9F, I28T, D44N, D44G, H64Y, S76Y, M80I, A95T, S110P, K113N, K113E, K114N, K114E, G116D, K118N, K118R, I130V, A132S, F142V, Q 155H, P187S, A190T, V194A, K221N, K233R, N234H, A238V, A261V, H272Y, F280L, D281G, I299D, E303F, I304T, V307S, I308N, Y313H, N316D, and V328M.

[0172] In some embodiments, the polypeptide is a polypeptide of any of the following sequences, relative to SEQ ID NO:6: positions 2, 67, 95, and 226; 6 and 316; 38, 95, 303; 44 and 118; 67, 95, and 226; 44 and 76; 44, 76, and 118; 130, 234, 303; 118 and 1201; 118, 1201, and 44; 118, 1201, and 76; 130, 234, and 303; 154 and 269; 221 and 44; 44, 76, 130, 234, and 303; 44, 76 , 118, and 1201; 197 and 314; 76, 181, and 194; 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; 59, 76, 306, and 316; or 280 and 340. In selected embodiments, the polypeptide comprises an amino acid sequence with an amino acid substitution at position 76 relative to SEQ ID NO:6.

[0173] In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: N2S, K67N or K67R, A95D, and V226E. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: E6D and N316K or N316D. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: N38S, A95D, E303D. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: K67N or K67R, A95D, and V226E. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: D44G or D44N and S76Y or K118N or K118R. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: D44G or D44N, I130V, N234H, and E303D. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: K118N or K118R and A1201V. In some embodiments, the polypeptide further comprises the following amino acid substitutions: D44G or D44N or S76Y. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: I130V, N234H, and E303D. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: R154K and E269K or E269D. In some embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: K221N and D44G or D44N. In some embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO: 6: F280L and S340L. In some embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO: 6: D44N or D44G, S76Y, I130V, N234H, and E303D.In some embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO:6: D44N or D44G, S76Y, K118R, and A1201V. In selected embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO:6: S76Y. In selected embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO:6: R197I and N314K. In selected embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO:6: S76Y, A181S, and V194M. In selected embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO:6: S76Y, K118R, H252R, and K292N. In selected embodiments, the polypeptide comprises an amino acid sequence with the following amino acid substitutions relative to SEQ ID NO:6: S76Y and I274V. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of S76Y, A102T, K118R, and V307G relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of L12M and S76Y relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of K67N, A95D, and V226E relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of K26N and S76Y relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of H22Y, S76Y, and D319N relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of R154K and E269D relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence with amino acid substitutions of S76Y and A238S relative to SEQ ID NO:6. In selected embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: S76Y and S263N.In selected embodiments, the polypeptide comprises an amino acid sequence having the following amino acid substitutions relative to SEQ ID NO:6: S59T, S76Y, E306G, and N316D.

[0174] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 7 and one or more amino acid substitutions at positions 99, 133, 189, 265, 266, 336, and 343 relative to SEQ ID NO: 7. In some embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 7: M99I, S189N, H265Q, A266V, L336F, and V343A.

[0175] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO:8 and one or more amino acid substitutions at positions 119, 134, 155, 180, 183, 274, 319, 447, 454, 458, 461, 512, 538, and 580 relative to SEQ ID NO:8. In some embodiments, the polypeptide comprises an amino acid sequence having one or more amino acid substitutions of Y119H, N134R, N134Q, D155N, Q180R, D183N, R274L, N319D, V447I, A454S, E458G, D461N, A512T, D538K, and P580Q relative to SEQ ID NO:8.

[0176] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 9 and one or more amino acid substitutions at positions 28, 82, 144, 151, 162, 182, 273, 327, 346 relative to SEQ ID NO: 9. In some embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 9: R28K, A82T, K144E, C151R, N162S, K182E, D273G, A327D, and M346I.

[0177] Provided herein are polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity to SEQ ID NO: 10, and one or more amino acid substitutions at positions 21 and 90 relative to SEQ ID NO: 10. In some embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 10: A21S and V90A.

[0178] As used herein, a sequence identity of at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 11, as well as sequences at positions 2, 3, 7, 9, 11, 12, 14, 16, 20, 26, 29, 32, 34, 35, 40, 43, 45, 46, 54, 61, 64, 65, 70, 77, 101, 103, 105, 106, 108, 109, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 238, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 2 Polypeptides having one or more amino acid substitutions at 111, 119, 120, 123, 126, 127, 130, 131, 148, 149, 151, 157, 159, 164, 166, 185, 194, 196, 203, 211, 217, 218, 219, 236, 242, 257, 267, 279, 283, 286, 288, 291, 293, 296, 303, 306, 313, 314, 316, 326, 331, 336, 347, 352, 361, 374, 377, 395, 396, 398, and 408 are provided.In some embodiments, the polypeptide is selected from the group consisting of A2T, F3S, P7R, A9S, A9G, A11G, F12I, D14N, S16Y, Y20H, S26N, F29S, S32N, E34K, G35V, G35S, G35D, I40S, E43D, H45P, E46K, A54S, R61W, V64M, Y65C, N70S, A77T, D101N, K103E, N105K, N105D, S106G, V108M, A109G, Y111N, L119M, R120S , R123S, A126T, E127G, V130M, D131N, Q148R, S149Y, H151Y, A157D, T159I, A164V, L166M, T 185A, S194G, A196T, T203A, K211R, E217K, R218K, R218S, N219S, A236T, E242D, N257K, N26 7S, M279I, M279V, D283G, N286S, T288I, K291Q, I293V, D296N, S303I, S303G, K306N, S310Y , S310P, I313T, Y314F, A316T, E326G, T331I, A336V, A347T, A347S, T352S, Y361H, M374T, M374I, R377G, T395I, S396T, S396F, G398V, and A408V. In some embodiments, the polypeptide comprises an amino acid sequence having one or more amino acid substitutions at positions 105, 109, 131, 148, 279, and 310; or 9, 105, 109, 131, 148, 279, and 310 relative to SEQ ID NO: 11. In some embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 11: N105K, A109G, D131N, Q148R, M279I, and S310Y or S310P. In some embodiments, the polypeptide comprises an amino acid sequence having one or more of the following amino acid substitutions relative to SEQ ID NO: 11: A9S, N105K, A109G, D131N, A148R, M279I, and S310P.

[0179] In some embodiments, the polypeptide comprises an amino acid sequence having one or more additions relative to SEQ ID NO: 11. In some embodiments, the polypeptide comprises an amino acid sequence having a C-terminal addition of at least one amino acid. In some embodiments, the polypeptide comprises an amino acid sequence having 410L. As used herein, a sequence identity of at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 12, as well as sequences at positions 4, 5, 6, 8, 9, 11, 12, 13, 16, 17, 20, 21, 24, 26, 28, 29, 34, 37, 38, 41, 49, 54, 59, 60, 63, 65, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 128, 129, 230, 231, 232, 233, 234, 235, 236, 23 74, 77, 81, 88, 92, 93, 94, 96, 102, 105, 106, 108, 110, 121, 126, 128, 134, 138, 142, 147, 150, 151, 153, 156, 157, 160, 162, 165, 170, 171, 173, 174, 179, 181, 183, 185, 186, 187, 188, 191, 198, 201, 206, 207, 226, 228, 233, 236, 241, 249, 250, 256, 267, 268, 270, 275, 276, 277, 279, 283, 286, 289, 303, 305, 306, 310, 312, 314, 315, 316, 323, 326, 329, 349, 353, 355, 356, 357, 358, 361, 370, 372, 373, 376, 378, 382, ​​388, 391, 397, 399, 403, 405, 419, 421, 423, 424, 425, 427, 428, 430, 431, 432, 433, 449, 457, 473, 477, 480, 485, 487, 489, 494, 496, 497, Polypeptides are provided having one or more amino acid substitutions at 498, 500, 502, 509, 511, 515, 518, 519, 520, 540, 545, 550, 555, 557, 570, 571, 580, 583, 585, 590, 594, 603, 607, 608, 611, 617, 620, 624, 636, 639, 641, 642, 644, 646, 655, 658, 660, 663, 665, 668, 672, 673, 678, 682, 685, 688, and 695. In some embodiments, the polypeptide is Based on SEQ ID NO: 12, K4N, E5K, L6M, L6I, E8K, E8D, I9T, D11N, T12A, T13I, D16G, R17C, R17S, R20K, R20E, R21E, R21K, S24K, S24Q, S24R, Y26S, Y26H, A28S, A28D, M29I, G34D, A37S, V38M, V38G, I41V, R49L, D54G, K59R, K60N, K63N, A65T, A65V, K67E, K74E, K77E, W81C, K88R, K88E, I92T, R93E, R93K, V94M, K96N, E10 2D, E102G, T105A, L106M, S108P, V110A, G121S, S126P, K128R, L134M, Y138S , Q142H, W147L, K150N, V151M, V151L, A153T, S156R, S156G, D157N, K160R, K 160E, A162T, S165N, S165G, V170E, K171E, F173V, K174N, K174R, T179A, K18 1T, S183N, P185T, E186K, E186D, E187K, A188S, A188V, D191Y, D191E, R198H, R198C, R198S, R201K, D206G, G207D, A226T, I228V, R233K, N236T, R241E, A2 49S, A250S, I256T, S267G, S267N, K268N, H270P, S275N, S275G, R276G, A277 D, A277S, A277T, K279N, G283D, V286G, V289M, G303D, I305T, F306S, A310D, A310T, A312G, A312D, A312T, K314N, Q315R, R316G, N323S, E326A, E326K, N32 9S, G349D, E353D, L355M, L355R, E356G, E356D, S357P, A358V, R361S, P370T , N372K, E373D, S376F, T378I, F382L, M388V, G391S, R397K, A399S, K403N, M 405I, L419P, D421N, K423R, H424N, H424R, V425L, I427V, E428K, D430A, D43 1G, E432D, H433N, A449T, G457D, R473K, E477D, G480D, F485L, S487R, S487G,S489N, N494D, S496N, A497S, V498G, K500N, K502N, Q509R, A511T, A511E, R515S, R518S, P519T, G520D, G520V, Y54 0C, Q545H, K550N, K555E, H557Q, P570S, E571D, C580R, S583R, E585K, E585G, E590D, R594K, M603I, H607N, H607L, K and an amino acid sequence having one or more amino acid substitutions selected from the group consisting of 608R, D611N, L617P, N620S, K624N, T636P, M639V, N641S, V642G, S644N, S644G, E646D, A655V, V658M, K660N, T663A, T665I, R668S, I672V, G673V, S678R, M682L, A685V, A685D, K688N, and V695M.

[0180] In some embodiments, the polypeptide comprises an amino acid sequence having amino acid substitutions at positions 134, 179, 185, 540, 555, 624, and 646; 138, 250, 275, and 421; 303, 405, 520, and 590; 134, 179, 185, 540, 555, and 646, 4 and 49; 4 and 388; 4 and 571; 4, 162, and 480; 4 and 315; 5 and 316; 17 and 156; 38, 108, 497 and 583; 59, 157 and 644; 96, 305, 550 and 642; 106, 160 and 228; 312, 424, 449 and 457; or 376 and 611 relative to SEQ ID NO:12. In some embodiments, the polypeptide comprises an amino acid sequence having ...

Claims

1. A polypeptide comprising one or more amino acid sequences having at least 70% identity, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any of SEQ ID NOs: 1-14, which has one or more amino acid substitutions, deletions, or additions relative to SEQ ID NOs: 1-14.

2. a) at least 70% identity to SEQ ID NO: 1 and one or more amino acid substitutions at positions 2, 3, 5, 28, 57, 77, 80, 107, 110, 116, 122, 142, 155, 161, 166, 173, 177, 185, 211, 216, 227, and 230 relative to SEQ ID NO: 1; b) at least 70% identity to SEQ ID NO:2 and one or more amino acid substitutions at positions 2, 5, 22, 24, 25, 29, 75, 141, 199, 215, 319, 347, 364, 370, 383, 439, 454, 458, 485, 509, 533, 538, 565, 581, 586, 595, 596, 597, and 600 relative to SEQ ID NO:2; c) at least 70% identity to SEQ ID NO:3 and one or more amino acid substitutions at positions 9, 15, 16, 18, 21, 64, 81, 86, 87, 99, 109, 142, 147, 153, 168, 180, 216, 230, 285, and 304 relative to SEQ ID NO:3; d) at least 70% identity to SEQ ID NO:4, and at positions 4, 5, 9, 10, 12, 21, 23, 25, 26, 31, 32, 34, 35, 37, 41, 45, 47, 48, 51, 52, 55, 60, 61, 65, 67, 69, 72, 75, 79, 80, 82, 87, 88, 90, 91, 93, 94, 96, 98, 99, 100, 103, 106, 108, 113, 116 relative to SEQ ID NO:4; one or more amino acid substitutions at 125, 126, 128, 129, 135, 139, 143, 146, 147, 149, 153, 154, 156, 158, 159, 160, 162, 164, 166, 167, 168, 169, 170, 177, 179, 180, 182, 183, 185, 187, 188, 190, 191, 192, 193, 195, 196, 200, 204, 207, and 208; e) at least 70% identity to SEQ ID NO:5, and the amino acid sequence at positions 1, 2, 4, 5, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 49, 52, 55, 56, 58, 60, 62, 63, 67, 71, 74, 76, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 91, 92, 95, 97, 100, 101, 104, 106, 11 0, 112, 113, 115, 117, 119, 120, 124, 125, 127, 129, 130, 131, 134, 139, 142, 144, 145, 146, 147, 149, 150, 155, 156, 157, 158, 159, 163, 164, 165, 167, 169 , 173, 174, 176, 181, 182, 186, 187, 190, 195, 197, 198, 205, 208, 209, 211, 215, 218, 223, 226, 227, 231, 232, 235, 239, 246, 248, 250, 259, 260, 261, 262, 263, 267, 269, 273, 274, 277, 278, 280, 281, 282, 283, 285, 287, 288, 290, 295, 298, 302, 303, 307, 313, 316, 317, 320, 323, 325, 331, 332, 339, 345, 348, 3 49, 352, 353, 354, 356, 361, 362, 363, 364, 365, 366, 367, 369, 370, 371, 372, 373, 375, 376, 380, 383, 385, 386, 389, 390, 392, 396, 397, 399, 402, 403, 40 4, 407, 408, 410, 411, 412, 413, 414, 415, 416, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 434, 435, 437, 440, 443, 445, 446, 448, 450, 452, 456 , 459, 460, 463, 464, 470, 472, 473, 494, 495, 498, 501, 502, 504, 505, 506, 508, 509, 510, 512, 513, 514, 517, 520, 521, 522, 525, 526, 527, 530, 531, 532,533, 535, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 567, 568, 569, 570, 571, 574, 575, 576 one or more amino acid substitutions at 6, 580, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 599, 600, 601, 602, 603, 604, 606, 607, 608, 611, 613, 618, 620, and 656; f) at least 70% identity to SEQ ID NO:6, and relative to SEQ ID NO:6, positions 1, 2, 3, 5, 6, 7, 9, 11, 12, 14, 21, 22, 26, 27, 31, 35, 38, 43, 44, 46, 47, 54, 59, 60, 61, 64, 65, 67, 68, 71, 72, 74, 76, 79, 80, 81, 84, 89, 95, 102, 105, 109, 110, 111, 112, 113, 114, 116, 118, 119, 120, 123, 129, 130, 131, 132, 134, 142, 145, 146, 147, 148, 150, 154, 155, 166, 169, 178, 180, 181, 183, 184, 187, 190, 194, 197, 201, 204, 207, 208 09, 213, 219, 221, 225, 226, 227, 229, 232, 233, 234, 236, 238, 241, 246, 251, 252, 256, 257, 261, 263, 265, 267, 269, 271, 272, 274, 280, 281, 285, 286, 288, 291, 292, 296, 299, 301, 303, 304, 305 one or more amino acid substitutions at: 06, 307, 308, 310, 313, 314, 316, 317, 318, 319, 320, 323, 324, 326, 328, 330, 331, 332, 340, 341, 343, 344, 355, 412, 418, 427, 514, 1198, 1201, 1206, 1212, 1260, and 1282; g) at least 70% identity to SEQ ID NO:7 and one or more amino acid substitutions at positions 99, 133, 189, 265, 266, 336, and 343 relative to SEQ ID NO:7; h) at least 70% identity to SEQ ID NO:8 and one or more amino acid substitutions at positions 119, 134, 155, 180, 183, 274, 319, 447, 454, 458, 461, 512, 538, and 580 relative to SEQ ID NO:8; i) at least 70% identity to SEQ ID NO:9 and one or more amino acid substitutions at positions 28, 82, 144, 151, 162, 182, 273, 327, 346 relative to SEQ ID NO:9; j) at least 70% identity to SEQ ID NO: 10 and one or more amino acid substitutions at positions 21 and 90 relative to SEQ ID NO: 10; k) at least 70% identity to SEQ ID NO: 11 and, relative to SEQ ID NO: 11, at positions 2, 3, 7, 9, 11, 12, 14, 16, 20, 26, 29, 32, 34, 35, 40, 43, 45, 46, 54, 61, 64, 65, 70, 77, 101, 103, 105, 106, 108, 109, 111, 119, 120, 123, 126, 127, 130, 131, 148, 149, 151; one or more amino acid substitutions at 157, 159, 164, 166, 185, 194, 196, 203, 211, 217, 218, 219, 236, 242, 257, 267, 279, 283, 286, 288, 291, 293, 296, 303, 306, 313, 314, 316, 326, 331, 336, 347, 352, 361, 374, 377, 395, 396, 398, and 408; l) at least 70% identity to SEQ ID NO:12, and the amino acids at positions 4, 5, 6, 8, 9, 11, 12, 13, 16, 17, 20, 21, 24, 26, 28, 29, 34, 37, 38, 41, 49, 54, 59, 60, 63, 65, 67, 74, 77, 81, 88, 92, 93, 94, 96, 102, 105, 106, 108, 110, 121, 126, 128, 134, 138, 142, 147, 15 0, 151, 153, 156, 157, 160, 162, 165, 170, 171, 173, 174, 179, 181, 183, 185, 186, 187, 188, 191, 198, 201, 206, 207, 226, 228, 233, 236, 241, 249, 250, 256, 267, 268, 270, 275, 276, 277, 279, 283, 286, 289, 303, 305, 306, 310, 312, 314, 315, 316, 323, 326, 329, 349, 353, 355, 356, 357, 358, 361, 370, 372, 373, 376, 378, 382, ​​388, 391, 397, 399, 403, 405, 419, 421, 423, 424, 425, 427, 428, 430, 431, 432, 433, 449, 457, 473, 477, 480, 485, 487, 489, 494, 496, 497, 498, 500, 502, 50 one or more amino acid substitutions at 9, 511, 515, 518, 519, 520, 540, 545, 550, 555, 557, 570, 571, 580, 583, 585, 590, 594, 603, 607, 608, 611, 617, 620, 624, 636, 639, 641, 642, 644, 646, 655, 658, 660, 663, 665, 668, 672, 673, 678, 682, 685, 688, and 695; m) at least 70% identity to SEQ ID NO: 13, and at least one of the following positions relative to SEQ ID NO: 13: 5, 10, 11, 26, 30, 35, 40, 42, 45, 46, 47, 58, 61, 65, 71, 72, 75, 77, 78, 80, 82, 83, 94, 98, 113, 115, 116, 117, 121, 128, 133, 138, 146, 148, 161, 171, 175, 177, 182, 184, 191, 193, 20 one or more amino acid substitutions at 1, 203, 211, 212, 219, 225, 226, 232, 233, 235, 236, 237, 238, 240, 250, 274, 282, 286, 292, 295, 304, 307, 309, 312, 313, 315, 316, 317, 318, 320, 321, 322, 323, 328, 340, 343, 344, 345, 347, 348, 349, and 350; or n) The polypeptide of claim 1, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 14 and one or more amino acid substitutions at positions 2, 9, 13, 14, 15, 34, 38, 42, 46, 50, 59, 60, 73, 75, 77, 82, 83, 85, 86, 97, 110, 115, 120, 124, 130, 132, 134, 140, 143, 145, 156, 159, 162, 164, 177, 199, 232, and 270 relative to SEQ ID NO:

14.

3. a) at least 70% identity to SEQ ID NO: 1 and one or more amino acid substitutions from SEQ ID NO: 1: A2T, T3I, L5S, T28A, A57T, F77L, Y80D, K107M, K107R, Y110C, Y110D, D116G, E122A, D142E, M155I, K161R, N166D, K173E, Y177N, Y177D, C185R, D211Y, K216E, A227P, G230D and G230S; b) at least 70% identity to SEQ ID NO:2 and one or more of the following amino acid substitutions relative to SEQ ID NO:2: A2T, A2S, G5R, S22P, E24D, L25I, A29S, P75T, I141T, V199I, S215R, D319V, Y347F, S364N, E370K, N383D, V439A, E454D, E454G, S458N, V485F, R509G, D533A, A538V, H565Y, A581T, H586L, N595K, D596N, D597N, D597Y, and I600V; c) at least 70% identity to SEQ ID NO: 3 and one or more of the following amino acid substitutions relative to SEQ ID NO: 3: I9V, A15V, F16Y, S18F, S21N, N64D, H81Y, D86Y, N87K, V99I, E109D, E142K, V147I, N153D, I168M, A180E, A216S, L230F, K285E, and R304R; d) at least 70% identity to SEQ ID NO: 4, and the following sequences relative to SEQ ID NO: 4: R4K, N5K, P9S, A10P, N12D, T21I, V23M, S25N, S25R, V26M, V26G, S31N, S32I, E34A, F35L, A37D, H41L, D45N, I47V, E48G, G51V, S52I, E55K, E55D, E60K, F61L, S65T, S65A, P67T, P67L , P67S, P67H, T69A, A72V, A72D, S75I, S75R, S75T, K79E, T80P, K82E, K87R, P88L, P88T, P88A, S90F, K91N, K91E, A9 3T, A93S, S94N, L96P, R98Q, A99D, A99V, E100K, A103T, A106T, S108A, I113F, V116F, V116I, V125M, V125A, N126T, I 128V, I128L, L129P, L135M, S139N, S139G, G143V, G143C, G146D, G146S, I147V, K149E, K149T, K149R, S153I, S153 R, S153N, F154C, H156R, H156L, S158N, S158R, G159V, V160A, K162R, N164D, I166L, S167I, S168I, S168R, S168N, Q one or more amino acid substitutions among 169R, V170M, V170G, V170L, T177I, T177A, S179R, F180C, F180L, F182C, F182L, G183S, M185I, K187R, G188D, V190I, K191N, A192S, D193N, G195V, G195D, G195S, C196W, T200A, T204I, A207V, A207T, and T208I; e) at least 70% identity to SEQ ID NO: 5, and the following sequences based on SEQ ID NO: 5: M1V, M1I, M1L, T2I, T2A, F4L, F5L, F8L, F8V, F8S, D9N, E10K, E10D, S11I, S11R, S11G, L12P, V13M, V13G, V13E, V13L, P14L, L15Q, K16N, K16R, P17T, P17L, P17S, T19I, T19S, T19A, T19P, P20S, P20L, T21A, Q22R, Y23H, V24M, K25R, L26M, D27A, D27G, D28N, D28Y, A29T, A29V, N30K, I32F, I32S, Q33H, L36M, D37A, D37Y, F39L, S40P, D41E, T4 2I, T42K, T42A, F43L, F43S, F43V, K44N, N45D, N45S, Q49R, K52Q, S55A, T56A, D58E, K60Q, S62T, R63K, R63G, Q67R, Q67H, Q67K, D71Y, K74R, E76K, F78C, K7 9R, G80V, G80D, G81S, G81V, G81D, D82N, V83G, V83M, V83A, V84A, V84G, R85G, R85K, P86L, N87S, R89C, V91G, V91A, A92V, A92T, R95K, K97R, E100D, S101A, D104V, A106D, A106T, D110N, N112H, H113Y, M115R, N117Y, T119A, N120D, N12 0K, N120S, G124V, D125N, D125E, K127R, F129L, D130N, K131M, E134D, E134G , A139S, A139T, P142S, I144V, A145S, A145T, T146A, A147V, Q149R, Y150H, I1 55L, V156A, V156L, V156M, K157V, E158A, N159S, V163G, E164A, E164G, E164 D, G165D, I167V, I169L, I169T, N173S, N173H, N173T, A174S, A174T, N176D, A 181S, I182L, I182V, I182T, A186E, A186T, V187G, V187A, A190T, A190S, F19 5S, A197P, D198G, D198N, A205S, V208M, P209T, T211I, E215D, E218D, P223S,P223H、L226V、I227V、D231N、E232K、I235V、I235T、R239G、I246V、V248E、V2 48M, S250I, S259N, Y260C, K261R, S262N, P263L, S267N, A269V, T273I, T273 N、H274Y、K277N、K277R、P278S、S280T 、L281M、D282E、D282N、A283T、A283S、 N285S、E287D、L288M、N290K、F295S、F298I、F298S、V302I、V303M、A307S、N31 3S、H316R、A317V、S320N、S320R、I323L、I325V、R331K、K332E、I339V、V345L 、㼶㼓㼔㼕㼭、㼥㼓㼔㼘㼫、㼹㼓㼔㼙㼨、㼹㼓㔼9D、㼹㼓㔼�㼤、㼹㼓㔼�㼮、㼹㼓㔼9C、㼰㼓㼕㼒㼳、㼰㼓㼕㼒㼴、㼥㼓㼕㼓㼱、㼥㼓㼕㼓㼤、㼬354M、G356S、N361D、I362V、I362T、L3 63P、L363T、L363M、E364G、K365R、E36 6G、E367G、K369N、K369E、K369M、P370 S、E371K、V372M、D373G、I375V、M376I、 T380P、T380A、E383K、E383D、F385L、H386Y、I389V、A390V、A390I、V392I、D3 96N, D396G, D396K, S397P, S399N, S399G, T402I, R403G, R403I, R403K, R403 S、I404T、I404V、K407R、K407E、R408K 、Q410K、Q410H、Q410R、Q411H、G412V、 F413L、D414N、A415V、A415T、Y416C、M 421I、N422K、E423K、E423D、E424A、E42 5K、E426D、T427A、T427S、R428K、F429 L、S430A、M431L、R434H、R434C、R434S 、I435V、D437G、D437N、T440S、T440I、 R443C、G445S、F446L、F446I、Y448C、E 450D、E450G、M452I、T456P、T456A、T4 56I、A459T、D460N、K463N、H464N、H46 4R、H464S、E470K、V472M、V472A、K473 D、K473N、E494D、E494G、S495A、E498A、E498K, C501Y, T502I, T502S, P504S, P504L, T505A, G506Y, G506D, G506L, G506S, T508A, D509E, D509Y, C510Y, S512N, I513L, I513V, I513 F, Y514H, K517M, K517N, K517Q, K520R, K521N, I522T, I522V, I522F, E525K, V526E, V526M, I527V, S530N, S530R, K531T, D532G, D532Y, S53 3Y, G535D, A537T, K538R, K538N, R540K, R540G, M541L, A542T, I543L, H544R, E545A, R546G, R546K, V547M, K548Q, K548R, Q549K, Q549R, E5 50A, Q551K, E552D, E552K, V553I, F554V, E556K, E556G, S557A, K558R, T559P, T559I, T559A, K560R, A561T, A561G, K562R, K562N, I563L, T 564I, A565S, A565V, K567R, K568N, K568R, Q569K, Q569L, Q569R, A570V, Q571R, D574N, V575M, V575A, S576R, T580I, T580A, T582I, T582S , I583V, K584R, V585M, S586P, S586A, S586F, E587A, E588K, E588G, E588D, S589I, S589R, S589N, A590S, A590T, A591V, P592L, V593M, V593 one or more amino acid substitutions selected from A, Q594L, K595R, K595N, H596Y, H596L, H596P, I597T, I597V, N599H, D600L, D600N, D600G, D600V, N601S, N601K, S602A, S602P, S602Y, D603A, D603V, D604G, D604Y, D604N, D606A, D606V, D606Y, D607Y, D607E, D608N, A611T, E613D, R618I, T620P, and A656V; f) at least 70% identity to SEQ ID NO: 6, and the following sequences, taken as reference to SEQ ID NO: 6: M1L, M1V, N2S, A3T, T5P, T5A, T5S, E6D, I7S, I7V, I9F, Q11R, L12M, N14D, N14S, M21I, H22P, H22Y, K26N, K26R, T27I, M31I, L35R, N38S, S43P, D44N, D44G, Q46L, C47S, T54I, S59T, H60Y, T61A, H64Y, Y65H, K67N, K67R, R68Q, A71G, T72A, N74D, S76C, S76Y, T79I, M80I, P81S, V84L, R89L, A95D, A95T, A102T, E105D, E105K, S109N, S109R, S1 10P, Q111R, I112T, K113N, K113E, K114N, K114M, K114E, G116D, K118N, K118R , T119I, D120V, K123N, L129M, I130V, K131R, A132S, K134M, K134N, F142V, L 145M, I146T, E147K, F148S, S150F, R154K, Q155H, E166D, K169E, P178S, A180 V, A181T, A181S, I183V, A184S, A184T, A184V, P187S, A190T, A190V, V194M, V194A, R197I, Y201N, L204M, D207N, K209N, Q213H, Q213V, A219S, K221N, D22 5N, V226E, P227T, K229E, S232N, K233N, K233R, N234H, T236A, A238V, A238S , A241S, E246D, K251N, H252Y, H252R, E256D, A257S, A261V, S263I, S263N, N2 65D, Y267C, E269K, E269D, K271E, K271R, H272Y, I274V, F280L, D281N, D281 G, K285G, K286N, K288R, S291F, S291P, K292N, K296R, K296N, I299S, D301G, E 303D, I304T, I304V, E306G, V307L, V307G, V307A, V307D, V307G, I308N, N31 0S, Y313H, N314K, N316K, N316D, A317D, L318Q, D319N, P320S, P320L, M323I,one or more amino acid substitutions among L324M, D326N, V328M, V328A, A330D, I331V, V332G, S340L, T341A, A343G, S344N, I355V, F412V, V418F, Y427C, R514K, S1198L, A1201V, G1206S, C1212G, F1260L, and V1282M; g) at least 70% identity to SEQ ID NO: 7 and one or more of the following amino acid substitutions relative to SEQ ID NO: 7: M99I, S189N, H265Q, A266V, L336F, and V343A; h) at least 70% identity to SEQ ID NO: 8 and one or more amino acid substitutions from SEQ ID NO: 8: Y119H, N134R, N134Q, D155N, Q180R, D183N, R274L, N319D, V447I, A454S, E458G, D461N, A512T, D538K, and P580Q; i) at least 70% identity to SEQ ID NO: 9 and one or more of the following amino acid substitutions relative to SEQ ID NO: 9: R28K, A82T, K144E, C151R, N162S, K182E, D273G, A327D, and M346I; j) at least 70% identity to SEQ ID NO: 10 and one or more of the following amino acid substitutions relative to SEQ ID NO: 10: A21S and V90A; k) at least 70% identity to SEQ ID NO: 11, and the following sequences relative to SEQ ID NO: 11: A2T, F3S, P7R, A9S, A9G, A11G, F12I, D14N, S16Y, Y20H, S26N, F29S, S32N, E34K, G35V, G35S, G35D, I40S, E43D, H45P, E46K, A54S, R61W, V64M, Y65C, N70S, A77T, D101N, K103E, N105K, N105D, S106G, V108M, A109G, Y111N, L119M, R120S, R123S, A126T, E127G, V130M, D131N, Q148R, S149Y, H151Y, A157D, T159I, A1 64V, L166M, T185A, S194G, A196T, T203A, K211R, E217K, R218K, R218S, N219S, A236T, E24 2D, N257K, N267S, M279I, M279V, D283G, N286S, T288I, K291Q, I293V, D296N, S303I, S303G , K306N, S310Y, S310P, I313T, Y314F, A316T, E326G, T331I, A336V, A347T, A347S, T352S, Y361H, M374T, M374I, R377G, T395I, S396T, S396F, G398V, and A408V; l) at least 70% identity to SEQ ID NO: 12, and Based on SEQ ID NO: 12, K4N, E5K, L6M, L6I, E8K, E8D, I9T, D11N, T12A, T13I, D16G, R17C, R17S, R20K, R20E, R21E, R21K, S24K, S24Q, S24R, Y26S, Y26H, A28S, A28D, M29I, G34D, A37S, V38M, V38G, I41V, R49L, D54G, K59R, K60N, K63N, A65T, A65V, K67E, K74E, K77E, W81C, K88R, K88E, I92T, R93E, R93K, V94M, K96N, E10 2D, E102G, T105A, L106M, S108P, V110A, G121S, S126P, K128R, L134M, Y138S , Q142H, W147L, K150N, V151M, V151L, A153T, S156R, S156G, D157N, K160R, K 160E, A162T, S165N, S165G, V170E, K171E, F173V, K174N, K174R, T179A, K18 1T, S183N, P185T, E186K, E186D, E187K, A188S, A188V, D191Y, D191E, R198H, R198C, R198S, R201K, D206G, G207D, A226T, I228V, R233K, N236T, R241E, A2 49S, A250S, I256T, S267G, S267N, K268N, H270P, S275N, S275G, R276G, A277 D, A277S, A277T, K279N, G283D, V286G, V289M, G303D, I305T, F306S, A310D, A310T, A312G, A312D, A312T, K314N, Q315R, R316G, N323S, E326A, E326K, N32 9S, G349D, E353D, L355M, L355R, E356G, E356D, S357P, A358V, R361S, P370T , N372K, E373D, S376F, T378I, F382L, M388V, G391S, R397K, A399S, K403N, M 405I, L419P, D421N, K423R, H424N, H424R, V425L, I427V, E428K, D430A, D43 1G, E432D, H433N, A449T, G457D, R473K, E477D, G480D, F485L, S487R, S487G,S489N, N494D, S496N, A497S, V498G, K500N, K502N, Q509R, A511T, A511E, R515S, R518S, P519T, G520D, G520V, Y540C, Q545H, K550N, K555E, H557Q, P570S, E571D, C580R, S583R, E585K, E585G, E590D, R594K, M603I, H607N, H one or more amino acid substitutions among 607L, K608R, D611N, L617P, N620S, K624N, T636P, M639V, N641S, V642G, S644N, S644G, E646D, A655V, V658M, K660N, T663A, T665I, R668S, I672V, G673V, S678R, M682L, A685V, A685D, K688N, and V695M; m) at least 70% identity to SEQ ID NO: 13, and the following sequences based on SEQ ID NO: 13: N5K, N5T, D10N, R11K, D26N, V30E, D35N, R40L, P42A, G45S, G45V, F46V, T47R, T47S, N58T, P61L, T65I, T71I, T71R, T71D, L72M, C75S, V77A, P78L, N80T, E82D, H83Y, H83N, A94S, V98M, E113D, C121F, A128S, A155S, E116D, T117I, R133K, G138V, N146D, G148V, C161R, A171V, A171S, K175T, A177 V, K182E, L184M, I191V, S193A, S193F, F201S, S203N, E211K, A212V, Y219R, N225S, N225T, D226Y, E232K, E232Q, A 233N, A233S, A233K, K235R, Q236R, Q236S, F237L, V238Q, V238M, A240T, A240V, S250A, R274G, A282V, I286N, I28 6T, I286F, P292S, S295N, K304R, E307D, Y309C, A312V, L313M, N315K, N315T, N315S, C316G, I317V, T318A, T318P, one or more amino acid substitutions among K320R, N321D, E322K, K323N, I328T, M340I, K343E, K343R, K344E, K344R, A345T, A345D, A345S, A345Y, A345R, A345K, A345E, A345G, A347K, A347S, A347D, K348N, K349R, A350K, A350D, A350V, and A350T; or n) at least 70% identity to SEQ ID NO: 14; and Based on SEQ ID NO: 14, Q2K, H9L, K13E, Q14K, A15G, K34N, E38K, V42I, A46D, S50I, V59G, Y60H, A73S, A73T, F75L, D77G, G82S, F83L, F83V, F83C, K85E, V86I, E97, I110S, I110L, S115R, K120N 3. The polypeptide of claim 1 or 2, comprising an amino acid sequence having one or more amino acid substitutions selected from the group consisting of K124R, G130D, D132E, N134T, A140T, E143K, D145G, S156I, E159K, I162V, H164Y, H164F, Y177C, S199I, S232L, and L270S.

4. a) at least 70% identity to SEQ ID NO: 1 and amino acid substitutions at positions 2 and 230; 107 and 166; one or both of 107, 166, and 2 and 227; 211 and 110 or 142; 110, 155 and 230; 122 and 155; or 155 and 177 relative to SEQ ID NO: 1; b) at least 70% identity to SEQ ID NO:2 and amino acid substitutions at positions 2 and 597; 24 and 25; 24, 25, 458, 509, 565, and 600; 75 and 597; 141, 454, 533 and 595; 581, 370, and 454; 370 and 581; 370 and 454; 458 and 509; 458, 509 and 565; 458, 509, 565, and 600; 565, 586, and 596; or 565, 509, 458, 600 and 24, 25, 29, 215, 319, 364, 383, and 586 relative to SEQ ID NO:2; c) at least 70% identity to SEQ ID NO:3 and amino acid substitutions at positions 142 and 216 relative to SEQ ID NO:3; d) at least 70% identity to SEQ ID NO:4 and amino acid substitutions at positions 108 and 47 or 208; 170 and 207; 88 and 147; 47, 88 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 88, 128, 147, 170, and 182; or 170, 207, and 108 relative to SEQ ID NO:4; e) At least 70% identity to SEQ ID NO:5 and, relative to SEQ ID NO:5, positions 4, 23 and 590; 19, 169 and 549; 43 and 415; 80 and 593; 80, 144, 593 and 606; 1, 42, 80, 593 and 606; 42, 80, 593 and 606; 156 and 604; 283, 349 and 365; 283, 349, 365, 396 and 594; 283, 349, 365, 396, 594, 596 and 131; 352 and 390; 390, 396 and 594; 396 and 594; 456 and 502; 464 and 502; 464 and 17; 17, 235, 464, and 596; 235, 352, 396, 456, and 606; 415, 456, and 502; 456, 502, and 549; 169, 456, 502, and 549; 80, 456, 502, 593, and 606; 1, 42, 80, 456, 502, 593, and 606; 80, 144, 456, 502, 593, and 606; 19, 169, 456, 502 and 549; 43, 415, 456, and 502; 352, 390, 396, and 594; 352, 390, and 396; 283, 34 352, 390, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594; and 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526. one or more positions; 43, 352, 390, 396, 464, 549, and 594; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, and 502; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 21; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, 21, and 67;43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 174, 208, 427, 456, and 504; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 415, 502, and 139; 410, 526, 415, 502, 339, and 446; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 19, 460, 569, and 596; 43, 349, 352, 390, 396, 464, 549, 594, 410, 526, 460, 586, 588, and 608; 43, 349 , 352, 390, 396, 464, 549, 594, 410, 526, and 460; 352, 390, 396, 549, 586, and 594; 63, 158, 352, 390, 396, 549, 586, and 594; 164, 165, 352, 363, 390, 396, 410, 549, 586, and 594; 164, 1 73, 352, 390, 396, 549, 586, and 594; 83, 352, 390, 396, 549, 586, and 594; 8, 43, 174, 349, 352, 390, 396, 427, 464, 549, and 594; or amino acid substitutions at 283, 349, 365, 396, 594, 596, and 131; f) at least 70% identity to SEQ ID NO:6 and, relative to SEQ ID NO:6, positions 2, 67, 95, and 226; 6 and 316; 38, 95, and 303; 67, 95, and 226; 44 and 76; 44, 76, and 118; 130, 234, and 303; 118 and 1201; 118, 1201, and 44; 118, 1201, and 76; 130, 234, and 303; 154 and 269; 221 and 44; 44, 76, 130, 234, and 30 3; 44, 76, 118, and 1201; 197 and 314; 76, 181, and 194; 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; 59, 76, 306, and 316; or 280 and 340 amino acid substitutions, g) at least 70% identity to SEQ ID NO: 11 and amino acid substitutions at positions 105, 109, 131, 148, 279, and 310; or 9, 105, 109, 131, 148, 279, and 310 relative to SEQ ID NO: 11; h) at least 70% identity to SEQ ID NO: 12 and amino acid substitutions at positions 134, 179, 185, 540, 555, 624, and 646; 138, 250, 275, and 421; 303, 405, 520, and 590; 134, 179, 185, 540, 555, and 646, 4 and 49; 4 and 388; 4 and 571; 4, 162, and 480; 4 and 315; 5 and 316; 17 and 156; 38, 108, 497 and 583; 59, 157 and 644; 96, 305, 550 and 642; 106, 160 and 228; 312, 424, 449 and 457; or 376 and 611 relative to SEQ ID NO: 12; i) at least 70% identity to SEQ ID NO: 13 and amino acid substitutions at positions 30, 46, 240, 304, and 316; 30, 46, 240, and 316; 42 and 318; 184, 240, 315, and 345; 211 and 274; 237 and 237; 286 and 350; 317 and 347; 171, 286, and 315; or 328 and 350 relative to SEQ ID NO: 13; or j) The polypeptide of any one of claims 1 to 3, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 14 and having amino acid substitutions at positions 82, 110, 115, 164, and 199; 82, 110, 115, 124, 164, and 199; 110, 115, and 164; 110, 115, 164, and 199; 110, 115, 164, 199, and 124; or 110, 115, 164, 199, and 82 or 124 relative to SEQ ID NO:

14.

5. a) at least 70% identity to SEQ ID NO: 1 and amino acid substitutions at positions 155; 122 and 155; or 107, 166, and 227 relative to SEQ ID NO: 1; b) at least 70% identity to SEQ ID NO:2 and amino acid substitutions at positions 24, 25, 458, 509, 565, and 600; 22, 347, and 454; or 485 relative to SEQ ID NO:2; c) at least 70% identity to SEQ ID NO:4 and amino acid substitutions at positions 75, 182; 88, 147, and 177; 88 and 147; 88, 116 and 147; 88, 147, 170, and 182; 88, 147, 170, 182, and 51 or 180; 88, 147, and 154; 75, 88, and 147; 47, 88 and 147; 88, 128, 147, 170, and 182; or 88, 93, and 147 relative to SEQ ID NO:4; d) at least 70% identity to SEQ ID NO:5 and, relative to SEQ ID NO:5, the following positions: 352, 390, 396, 594, and 596; 352, 390, 396, 549, and 594; 352, 390, 396, 464, 549, and 594; 289, 352, 390, 396, 549, 594, and 596; 235, 352, 390, 396, 567, and 594; 352, 363, 390, 396, 549, 586, and 594; 352, 390, 396, 549, 580, and 594; 94 and one or more positions selected from 63, 145, 174, 182, 208, 410, 427, 456, 504, and 526; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 67; 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, and 21; or 43, 349, 352, 390, 396, 464, 549, 594, 415, 502, 21, and 67; or e) A polypeptide according to any one of claims 1 to 4, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 6 and having amino acid substitutions at positions 197, 314, and optionally one of 7, 12, or 114; 197 and 314; 76, 181, and 194; 76, 118, 252, and 292; 76 and 274; 76, 102, 118, and 307; 12 and 76; 67, 95, and 226; 26 and 76; 22, 76, 319; 154 and 269; 76 and 238; 76, 238, 296, and 328; 7 and 76; 76 and 263; or 59, 76, 306, and 316 relative to SEQ ID NO:

6.

6. a) at least 70% identity to SEQ ID NO: 1 and, relative to SEQ ID NO: 1, the amino acid substitutions M155I; E122A and M155I; or K107M, N166D, and A227P; b) at least 70% identity to SEQ ID NO:2 and amino acid substitutions at positions E24D, L25I, S458N, R509G, H565Y, and I600V; S22P, Y347F, and E454G; or V485F relative to SEQ ID NO:2; c) at least 70% identity to SEQ ID NO: 4 and the following amino acid substitutions relative to SEQ ID NO: 4: S75I; F182L; P88T, I147V, and T177I; P88T and I147V; P88T, V116I and I147V; P88T, I147V, V170L, and F182L; P88T, I147V, V170L, F180L, and F182L; G51V, P88T, I147V, V170L, and F182L; P88T, I147V, and F154C; S75I, P88T, and I147V; or P88T, A93T, and I147V, d) at least 70% identity to SEQ ID NO: 5 and the following amino acid substitutions relative to SEQ ID NO: 5: P352T, A390V, D396N, Q594L, and H596Y; P352S, A390V, D396N, Q549R, and Q594L; P352T, A390V, D396N, H464R, Q549R, and Q594L; Q289H, P352T, A390V, D396N, Q549R, Q594L, and H596Y; I235T , P352T, A390V, D396N, K567R, and Q594L; P352T, L363P, A390V, D396N, Q549R, S586A, and Q594L; P352T, A390V, D396N, Q549R, and Q594L; P352T, A390V, D396N, Q549R, T580I, and Q594L; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q5 94L and one or more substitutions selected from R63G, A145S, A174S, I182R, V208M, Q410K, T427S, T456I or T456P, P504S, and V526E; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, and T502I; F43S, Y349N or Y349D, P352T, A390V, D396N , H464R, Q549R, Q594L, A415V, T502I, and T21A; F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, and Q67K; or F43S, Y349N or Y349D, P352T, A390V, D396N, H464R, Q549R, Q594L, A415V, T502I, T21A, and Q67K, or e) at least 70% identity to SEQ ID NO: 6 and, relative to SEQ ID NO: 6, at least one of positions R197I, N314K, and optionally I7S, L12M, or K114M; R197I and N314K; S76Y, A181S, and V194M; S76Y, K118R, H252R, and K292N; S76Y and I274V; S76Y, A102T, K118R, and V307G; L12M and S76Y; K 6. The polypeptide of any one of claims 1 to 5, comprising an amino acid sequence having amino acid substitutions at: 67N, A95D, and V226E; K26N and S76Y; H22Y, S76Y, and D319N; R154K and E269D; S76Y and A238S; S76Y, A238S, K296N, and V328M; I7V and S76Y; S76Y and S263N; or S59T, S76Y, E306G, and N316D.

7. 2. The polypeptide of claim 1, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 13 and optionally at least one amino acid substitution with a positively charged amino acid selected from arginine or lysine.

8. 8. The polypeptide of claim 7, wherein the at least one amino acid substitution is at position 2, 5, 6, 7, 8, 9, 10, 12, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 64, 65, 66, 67, 68, 69, 70, 222, 224, 225, 227, 228, 229, 231, 232, 233, 234, 235, 255, 256, 257, 258, 277, 286, 287, 337, 338, 339, 340, 345, 346, 347, 348, 349, 350, or a combination thereof, relative to SEQ ID NO:

13.

9. 9. The polypeptide of claim 7 or 8, wherein the at least one amino acid substitution is at positions 346 and 348; 346, 348 and 349; 346, 348, 349, and 350; 350 and 351; 350, 351, and 352; 350, 351, 352, and 353; 235 and 227; 235 and 345; 235 and 346; 235 and 347; 235 and 348; 235 and 349; 235 and 350; 235, 227, and 349; 5, 235 and 346; 5, 235 and 348; 5, 235 and 349; 227, 235, and 346; or 227, 235, and 348, based on SEQ ID NO:

13.

10. a first amino acid sequence having at least 70% identity to SEQ ID NO: 1, which has one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 1; and a second amino acid sequence having at least 70% identity to SEQ ID NO:2, which has one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO:2; or a first amino acid sequence having at least 70% identity to SEQ ID NO: 4, which has one or more amino acid substitutions, deletions, or additions relative to SEQ ID NO: 4; and A polypeptide according to any one of claims 1 to 6, comprising a second amino acid sequence having at least 70% identity to SEQ ID NO: 5, which has one or more amino acid substitutions, deletions, or additions based on SEQ ID NO:

5.

11. A composition comprising one or more polypeptides of any of claims 1 to 10, or one or more nucleic acids encoding same, and optionally one or more Cas proteins or one or more nucleic acids encoding them, and / or at least one unfoldase protein or at least one nucleic acid encoding same.

12. The CRISPR-Tn system comprises a modified Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated transposon (CRISPR-Tn) system or one or more nucleic acids encoding the modified CRISPR-Tn system, a) one or more Cas proteins selected from Cas5, Cas6, Cas7, Cas8, Cas9, and combinations thereof; and b) one or more transposon-associated proteins selected from TnsA, TnsB, TnsC, TnsD, TniQ, and combinations thereof and at least one or both of 11. A system, wherein at least one of the one or more Cas proteins or at least one of the one or more transposon-associated proteins comprises a polypeptide according to any one of claims 1 to 10.

13. at least one guide RNA (gRNA) complementary to at least a portion of the target nucleic acid, or at least one nucleic acid encoding the same; a donor nucleic acid comprising a cargo nucleic acid sequence flanked by at least one transposon end sequence; at least one unfoldase protein or at least one nucleic acid encoding same; the target nucleic acid, or The system of claim 12 further comprising a combination thereof.

14. A method for nucleic acid modification or incorporation, comprising contacting a target nucleic acid sequence, or a cell containing the target nucleic acid, with a polypeptide according to any one of claims 1 to 10, a composition according to claim 11, or a system or composition thereof according to claim 12 or 13.

15. A cell comprising the polypeptide according to any one of claims 1 to 10, the composition according to claim 11, or the system according to claim 12 or 13.