Circularly Permuted Dehalogenase Variants

JP2025515179A5Pending Publication Date: 2026-05-15PROMEGA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PROMEGA CORP
Filing Date
2023-05-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The existing HALOTAG system is difficult to dynamically control self-labeling activity and cannot monitor protein interactions or changes in metabolite concentration in real time.

Method used

Circular arrangement (CP) dechlorase variants were developed that could be covalently bound with chloroalkyl ligands to form structurally assembled active dechlorase complexes.

Benefits of technology

Dynamic control of the self-labeling activity of the HALOTAG system is realized, and it can monitor protein interactions and metabolite concentration changes in real time, improving the sensitivity and accuracy of cell imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are circularly permuted (cp) dehalogenase variants capable of covalently binding to haloalkyl ligands. Specifically provided are peptides and polypeptides comprising the cp dehalogenase variants, as well as split versions thereof, that are structurally assembled to form active dehalogenase complexes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 338,364, filed May 4, 2022, which is incorporated herein by reference.

[0002] Provided herein are circularly permuted (cp) dehalogenase variants that are capable of covalently binding to haloalkyl ligands. Specifically provided are peptides and polypeptides that include the cp dehalogenase variants, as well as split versions thereof that are structurally assembled to form active dehalogenase complexes. [Background technology]

[0003] The utility of self-labeling protein systems such as HALOTAG and its chloroalkane-based ligands has continually expanded over their lifetime as research tools. Gene fusions to HALOTAG as a general strategy have enabled a wide range of applications including fluorescent labeling for cell biology and imaging, recombinant protein purification, biosensors and diagnostics, energy transfer technologies (BRET, FRET), and therapeutic targeted protein degradation (PROTACs). The development of new fluorophores and fluorogenic dyes (such as Janelia Fluor dyes) as chloroalkane conjugates serves as an example highlighting the renewed interest in HALOTAG for fluorescence detection in cellular imaging applications. The advantages of such dyes in brightness, photostability, sensitivity, and far-red spectral detection over traditional tools such as the widely used fluorescent proteins are particularly evident in challenging or highly sensitive imaging applications in endogenous biology. As chloroalkane conjugates, they can take advantage of the self-labeling activity of HALOTAG to measure protein abundance and localization in a target-specific manner via gene fusion. However, there is a lack of available tools that can take advantage of these improvements in fluorescence detection and measure important functional dynamics by cellular imaging, such as protein interactions or changes in metabolite concentrations. What is needed in the field are tools to dynamically control autolabeling activity in systems such as HALOTAG. Summary of the Invention

[0004] Provided herein are circularly permuted (cp) dehalogenase variants that are capable of covalently binding to haloalkyl ligands. Specifically provided are peptides and polypeptides that include the cp dehalogenase variants, as well as split versions thereof that are structurally assembled to form active dehalogenase complexes.

[0005] In some embodiments, provided herein are compositions comprising a cp variant of a polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity to SEQ ID NO: 1. In some embodiments, provided herein are compositions comprising a first sequence and a second sequence, each comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity to a portion of SEQ ID NO: 1. In some embodiments, the cp variant comprises (i) a first segment comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity with a first portion of SEQ ID NO:1, and (ii) a second segment comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity with a second portion of SEQ ID NO:1. In some embodiments, the first fragment and the second fragment together comprise an amino acid sequence corresponding to at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, 100%) of the length of SEQ ID NO:1. In some embodiments, the polypeptide amino acid corresponding to position 297 of SEQ ID NO:1 is a peptide linked to the polypeptide amino acid corresponding to position 1 of SEQ ID NO:1. In some embodiments, the polypeptide amino acid corresponding to position 297 of SEQ ID NO:1 is connected to the polypeptide amino acid corresponding to position 1 of SEQ ID NO:1 by a linker peptide. In some embodiments, the linker peptide is between 2 and 100 amino acids in length (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or any range therebetween). In some embodiments, the linker peptide comprises a cleavable element (e.g., a protease cleavable site (e.g., TEV protease), a chemically cleavable site, a photocleavable site, etc.). In some embodiments,Circular permutation variants can be any position between 5 and 290 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 5, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146 , 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 2 09, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 24 0, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271,272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290). In some embodiments, the circularly permuted variants are between positions 5 and 13 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, or ranges therebetween), between positions 36 and 51 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, or ranges therebetween), between positions 63 and 72 (e.g., 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, or ranges therebetween) of SEQ ID NO:1. , between 84th and 92nd (e.g., 84, 85, 86, 87, 88, 89, 90, 91, 92, or a range therebetween), between 104th and 130th (e.g., 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, or a range therebetween), between 142nd and 148th (e.g., 14 2, 143, 144, 145, 146, 147, 148, and ranges therebetween), between positions 160 and 174 (e.g., 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, or ranges therebetween), between positions 186 and 189 (e.g., 186, 187, 188, 189, or ranges therebetween), between positions 201 and 203 (e.g., 201, 202, 203, or ranges therebetween), ranges), between positions 221 and 229 (e.g., 221, 222, 223, 224, 225, 226, 227, 228, 229, or ranges therebetween), or between positions 269 and 290 (e.g., 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290, or ranges therebetween).

[0006] In some embodiments, SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 12 2, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 1 49, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 1 80, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 21 1, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242 , 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273,274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, and 289.

[0007] In some embodiments, cp variants are provided that include at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity to one of SEQ ID NOs: 2-289, but include a 1-100 amino acid linker at the cp site (e.g., after the sequence corresponding to the ...ISG and before the sequence corresponding to the MAE... of SEQ ID NOs: 2-289). In some embodiments, the cp variant is selected from the group consisting of (i) SEQ ID NOs: 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 4 7, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, 577, 57 9, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, 601, 603, 605, 607, 609, 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, 649, 651, 653, 655, 657, 659, 661, 663, 665, 667, 669, 671, 673, 675, 677, 679,681, 683, 685, 687, 689, 691, 693, 695, 697, 699, 701, 703, 705, 707, 709, 711, 713, 715, 717, 719, 721, 723, 725, 727, 729, 731, 733, 735, 737, 739, 741, 743, 745, 747, 749, 751, 753, 755, 757, 759, 761, 763, 765, 767, 769, 771, 773, 775, 777, 779, 781, 783, 785, 787, 789, 791, 793, 795, 797, 799, 801, 803, 8 and (ii) a first segment that comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity to one of SEQ ID NOs: 290, 292, 294, 296, 298, 300, 305, 807, 809, 811, 813, 815, 817, 819, 821, 823, 825, 827, 829, 831, 833, 835, 837, 839, 841, 843, 845, 847, 849, 851, 853, 855, 857, 859, 861, 863, and 865; and 02, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 3 64, 366, 368, 370, 372, 374, 376, 378, 380, 382, ​​384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 42 6, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488 , 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550,552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, 576, 578, 580, 582, 584, 586, 588, 590, 592, 594, 596, 598, 600, 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624, 626, 628, 630, 632, 634, 636, 638, 640, 64 2, 644, 646, 648, 650, 652, 654, 656, 658, 660, 662, 664, 666, 668, 670, 672, 674, 676, 678, 680, 682, 684, 686, 688, 690, 692, 694, 696, 698, 700, 702, 704, 706, 708, 710, 712, 714, 716, 718, 720, 722, 724, 726, 728, 730, 732, 7 34, 736, 738, 740, 742, 744, 746, 748, 750, 752, 754, 756, 758, 760, 762, 764, 766, 768, 770, 772, 774, 776, 778, 780, 782, 784, 786, 788, 790, 792, 794, 796, 798, 800, 802, 804, 806, 808, 810, 812, 814, 816, 818, 820, 822, 824, and a second segment comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to one of 826, 828, 830, 832, 834, 836, 838, 840, 842, 844, 846, 848, 850, 852, 854, 856, 858, 860, 862, and 864.

[0008] In some embodiments, the cp variant is selected from the group consisting of (i) SEQ ID NOs: 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 40 1, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463 , 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, 577, 579, 581, 583, 585, 587, 5 89, 591, 593, 595, 597, 599, 601, 603, 605, 607, 609, 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, 649, 65 1, 653, 655, 657, 659, 661, 663, 665, 667, 669, 671, 673, 675, 677, 679, 681, 683, 685, 687, 689, 691, 693, 695, 697, 699, 701, 703, 705, 707, 709, 711, 713 , 715, 717, 719, 721, 723, 725, 727, 729, 731, 733, 735, 737, 739, 741, 743, 745, 747, 749, 751, 753, 755, 757, 759, 761, 763, 765, 767, 769, 771, 773, 775,777, 779, 781, 783, 785, 787, 789, 791, 793, 795, 797, 799, 801, 803, 805, 807, 809, 811, 813, 815, 817, 819, 821, 823, 825, 827, 829, 831, 833, 835, 837, 839, 841, 843, 845, 847, 849, 851, 853, 855, 857, 859, 861, 863, and 865 and at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least (ii) a first segment having a sequence identity (at least 95%) with a 1-100 amino acid linker at the N-terminus (e.g., preceding a sequence corresponding to MAE...); and (iii) a second segment having a sequence identity (at least 95%) with a 1-100 amino acid linker at the N-terminus (e.g., preceding a sequence corresponding to MAE...), , 376, 378, 380, 382, ​​384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 5 00, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 56 2, 564, 566, 568, 570, 572, 574, 576, 578, 580, 582, 584, 586, 588, 590, 592, 594, 596, 598, 600, 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624,626, 628, 630, 632, 634, 636, 638, 640, 642, 644, 646, 648, 650, 652, 654, 656, 658, 660, 662, 664, 666, 668, 670, 672, 674, 676, 678, 680, 682, 684, 686, 688, 690, 692, 694, 696, 698, 700, 702, 7 04, 706, 708, 710, 712, 714, 716, 718, 720, 722, 724, 726, 728, 730, 732, 734, 736, 738, 740, 742, 744, 746, 748, 750, 752, 754, 756, 758, 760, 762, 764, 766, 768, 770, 772, 774, 776, 778, 780, 78 2, 784, 786, 788, 790, 792, 794, 796, 798, 800, 802, 804, 806, 808, 810, 812, 814, 816, 818, 820, 822, 824, 826, 828, 830, 832, 834, 836, 838, 840, 842, 844, 846, 848, 850, 852, 854, 856, 858, 860 , 862, and 864, but with a 1-100 amino acid linker at the C-terminus (e.g., following a sequence corresponding to an ISG). In some embodiments, the cp variant comprises a second segment having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to one of SEQ ID NOs: 290 and 291, 292 and 293, 294 and 295, 296 and 297, 298 and 299, 300 and 301, 302 and 303, 304 and 305, 306 and 307, 308 and 309, 310 and 311, 312 and 313, 314 and 315, 316 and 317, 318 and 319, 320 and 321, 322 and 323, 324 and 325, 326 and 327, 328 and 329, 330 and 331, 332 and 333, 334 and 335, 336 and 337, 338 and 339, 340 and 341, 342 and 343, 344 and 345, 346 and 347, 348 and 349, 350 and 351, 352 and 353, 354 and 355, 356 and 357, 358 and 359, 360 and 361, 362 and 363, 364 and 365,366 and 367, 368 and 369, 370 and 371, 372 and 373, 374 and 375, 376 and 377, 378 and 379, 380 and 381, 382 and 383, 384 and 385, 386 and 387, 388 and 389, 390 and 391, 392 and 393, 394 and 395, 396 and 397, 398 and 399, 400 and 401, 402 and 403, 404 and 405, 406 and 407, 408 and 409, 410 and 411, 412 and 413, 414 and 415, 416 and 417, 418 and 419, 420 and 421 21, 422 and 423, 424 and 425, 426 and 427, 428 and 429, 430 and 431, 432 and 433, 434 and 435, 436 and 437, 438 and 439, 440 and 441, 442 and 443, 444 and 445, 446 and 447, 448 and 449, 450 and 451, 452 and 453, 454 and 455, 456 and 457, 458 and 459, 460 and 461, 462 and 463, 464 and 465, 466 and 467, 468 and 469, 470 and 471, 472 and 473, 474 and 475, 476 and and 477, 478 and 479, 480 and 481, 482 and 483, 484 and 485, 486 and 487, 488 and 489, 490 and 491, 492 and 493, 494 and 495, 496 and 497, 498 and 499, 500 and 501, 502 and 503, 504 and 505, 506 and 507, 508 and 509, 510 and 511, 512 and 513, 514 and 515, 516 and 517, 518 and 519, 520 and 521, 522 and 523, 524 and 525, 526 and 527, 528 and 529, 530 and 531, 53 2 and 533, 534 and 535, 536 and 537, 538 and 539, 540 and 541, 542 and 543, 544 and 545, 546 and 547, 548 and 549, 550 and 551, 552 and 553, 554 and 555, 556 and 557, 558 and 559, 560 and 561, 562 and 563, 564 and 565, 566 and 567, 568 and 569, 570 and 571, 572 and 573, 574 and 575, 576 and 577, 578 and 579, 580 and 581, 582 and 583, 584 and 585, 586 and 587,588 and 589, 590 and 591, 592 and 593, 594 and 595, 596 and 597, 598 and 599, 600 and 601, 602 and 603, 604 and 605, 606 and 607, 608 and 609, 610 and 611, 612 and 613, 614 and 615, 616 and 617, 618 and 619, 620 and 621, 622 and 623, 624 and 625, 626 and 627, 628 and 629, 630 and 631, 632 and 633, 634 and 635, 636 and 637, 638 and 639, 640 and 641, 642 and 6 43, 644 and 645, 646 and 647, 648 and 649, 650 and 651, 652 and 653, 654 and 655, 656 and 657, 658 and 659, 660 and 661, 662 and 663, 664 and 665, 666 and 667, 668 and 669, 670 and 671, 672 and 673, 674 and 675, 676 and 677, 678 and 679, 680 and 681, 682 and 683, 684 and 685, 686 and 687, 688 and 689, 690 and 691, 692 and 693, 694 and 695, 696 and 697, 698 and and 699, 700 and 701, 702 and 703, 704 and 705, 706 and 707, 708 and 709, 710 and 711, 712 and 713, 714 and 715, 716 and 717, 718 and 719, 720 and 721, 722 and 723, 724 and 725, 726 and 727, 728 and 729, 730 and 731, 732 and 733, 734 and 735, 736 and 737, 738 and 739, 740 and 741, 742 and 743, 744 and 745, 746 and 747, 748 and 749, 750 and 751, 752 and 753, 75 4 and 755, 756 and 757, 758 and 759, 760 and 761, 762 and 763, 764 and 765, 766 and 767, 768 and 769, 770 and 771, 772 and 773, 774 and 775, 776 and 777, 778 and 779, 780 and 781, 782 and 783, 784 and 785, 786 and 787, 788 and 789, 790 and 791, 792 and 793, 794 and 795, 796 and 797, 798 and 799, 800 and 801, 802 and 803, 804 and 805, 806 and 807, 808 and 809,810 and 811, 812 and 813, 814 and 815, 816 and 817, 818 and 819, 820 and 821, 822 and 823, 824 and 825, 826 and 827, 828 and 829, 830 and 831, 832 and 833, 834 and 835, 836 and 837, The first and second segments comprise at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity with 838 and 839, 840 and 841, 842 and 843, 844 and 845, 846 and 847, 848 and 849, 850 and 851, 852 and 853, 854 and 855, 856 and 857, 858 and 859, 860 and 861, 862 and 863, or 864 and 865. In some embodiments, one or both members of the aforementioned pairs comprise a 1-100 amino acid linker at the C-terminus or N-terminus. In some embodiments, the pairs are fused (e.g., with a linker) as a single cp polypeptide. In some embodiments, the pairs are split (e.g., with a linker). In some embodiments, other pairs of parent sequences (e.g., sequences having greater than 70% sequence identity to one of SEQ ID NOs: 1-289) are provided (e.g., resulting in duplication or deletion of a segment (e.g., 1-100 amino acids).

[0009] In some embodiments, the first segment and the second segment are fused with a linker. In other embodiments, the first segment and the second segment exist as separate, unlinked peptides / polypeptides. In such embodiments, the first segment and the second segment are typically expressed or synthesized as a single polypeptide and cleaved (e.g., at a cleavable linker element) to generate separate peptides / polypeptides.

[0010] In some embodiments, SEQ ID NOs: 2-865 contain a linker peptide of 0-100 amino acids in length (e.g., at the N-terminus, C-terminus, cp site, etc.). In some embodiments, the linker is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 amino acids in length. The linker may be of any suitable length and may contain amino acids of any suitable properties (e.g., plastic, rigid, hydrophobic, aliphatic, ionic, acidic, basic, bulky, etc., and combinations thereof). In some embodiments, the linker peptide comprises a cleavable element (eg, a protease cleavable site (eg, TEV protease), a chemically cleavable site, a photocleavable site, etc.).

[0011] In some embodiments, the circularly permuted variant is capable of forming a covalent bond with a haloalkane substrate. In some embodiments, the circularly permuted variant comprises 100% sequence identity with SEQ ID NO: 1. In some embodiments, the circularly permuted variant comprises a deletion of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or a range therebetween) at a position corresponding to one or more of the N-terminus of SEQ ID NO: 1, the C-terminus of SEQ ID NO: 1, and on either side of the cp site. In some embodiments, circularly permuted variants comprise duplications of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or a range therebetween) at positions corresponding to one or more of the N-terminus of SEQ ID NO:1, the C-terminus of SEQ ID NO:1, and on either side of the cp site.

[0012] In some embodiments, provided herein are circularly permuted variants that have been cleaved (e.g., at a cleavable element (e.g., a protease site) in the linker) to form two separate peptide / polypeptide fragments. In some embodiments, provided herein are SEQ ID NOs: 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 43 7, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 4 99, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, 601, 603, 605, 607, 609, 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, 649, 651, 653, 655, 657, 659, 661, 663, 665, 667, 669, 671, 673, 675, 677, 679, 681, 683, 685, 687, 689, 691, 693, 695, 697, 699, 701, 703, 705, 707, 709, 711, 713, 715, 717, 719, 721,723, 725, 727, 729, 731, 733, 735, 737, 739, 741, 743, 745, 747, 749, 751, 753, 755, 757, 759, 761, 763, 765, 767, 769, 771, 773, 775, 777, 779, 781, 783, 785, 787, 789, 791, 793, 795, 797, 799, 801, 803, 805, 807, 809, 811, 813, 815, 81 7, 819, 821, 823, 825, 827, 829, 831, 833, 835, 837, 839, 841, 843, 845, 847, 849, 851, 853, 855, 857, 859, 861, 863, and 865. In some embodiments, provided herein are SEQ ID NOs: 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 42 52, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, ​​384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428 , 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 552, 553, 554, 555, 556, 557, 558, 560, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482 06, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, 576, 578, 580, 582,584, 586, 588, 590, 592, 594, 596, 598, 600, 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624, 626, 628, 630, 632, 634, 636, 638, 640, 642, 644, 646, 648, 650, 652, 654, 656, 658, 660, 662, 664, 666, 668, 670, 672, 674, 676, 678, 680, 682, 684, 686, 688, 690, 692, 694, 696, 698, 700, 702, 704, 706, 708, 710, 712, 714, 716, 718, 720, 722, 724, 726, 728, 730, 732, 734, 736, 738, 740, 742, 744, 746, 748, 750, 752, 754, 756, 758, 760, 762, 764, 766, 768, 770, 772, 774, 776, 778, 780, 782, 784, 786, 788, 790, 792, 794, 796, 798, 800, 802, 804, 806, 808, 810, 812, 814, 816, 818, 820, 822, 824, 826, 828, 830, 832, 834, 836, 838, 840, 842, 844, 846, 848, 850, 852, 854, 856, 858, 860, 862, and 864.

[0013] In some embodiments, provided herein are circularly permuted variants that have been cleaved (e.g., at a cleavable element (e.g., a protease site) in the linker) to form two separate peptide / polypeptide fragments. In some embodiments, provided herein are SEQ ID NOs: 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 43 7, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 4 99, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, 577, 579, 581, 583, 585, 587, 589, 591, 593, 595, 597, 599, 601, 603, 605, 607, 609, 611, 613, 615, 617, 619, 621, 623, 625, 627, 629, 631, 633, 635, 637, 639, 641, 643, 645, 647, 649, 651, 653, 655, 657, 659, 661, 663, 665, 667, 669, 671, 673, 675, 677, 679, 681, 683, 685, 687, 689, 691, 693, 695, 697, 699, 701, 703, 705, 707, 709, 711, 713, 715, 717, 719, 721,723, 725, 727, 729, 731, 733, 735, 737, 739, 741, 743, 745, 747, 749, 751, 753, 755, 757, 759, 761, 763, 765, 767, 769, 771, 773, 775, 777, 779, 781, 783, 785, 787, 789, 791, 793, 795, 797, 799, 801, 803, 805, 807, 809, 811, 813, 815, 817, 819, 821, 823, 825, 827, 82 9, 831, 833, 835, 837, 839, 841, 843, 845, 847, 849, 851, 853, 855, 857, 859, 861, 863, and 865, but having a 1-100 amino acid linker at the N-terminus (e.g., preceding a sequence corresponding to MAE...). In some embodiments, provided herein are SEQ ID NOs: 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 41 6, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, ​​384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 4 18, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560,562, 564, 566, 568, 570, 572, 574, 576, 578, 580, 582, 584, 586, 588, 590, 592, 594, 596, 598, 600, 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624, 626, 628, 630, 632, 634, 636, 638, 640, 642, 644, 646, 648, 650, 652, 654 , 656, 658, 660, 662, 664, 666, 668, 670, 672, 674, 676, 678, 680, 682, 684, 686, 688, 690, 692, 694, 696, 698, 700, 702, 704, 706, 708, 710, 712, 714, 716, 718, 720, 722, 724, 726, 728, 730, 732, 734, 736, 738, 740, 742, 744, 746, 748 , 750, 752, 754, 756, 758, 760, 762, 764, 766, 768, 770, 772, 774, 776, 778, 780, 782, 784, 786, 788, 790, 792, 794, 796, 798, 800, 802, 804, 806, 808, 810, 812, 814, 816, 818, 820, 822, 824, 826, 828, 830, 832, 834, 836, 838, 840, 84 2, 844, 846, 848, 850, 852, 854, 856, 858, 860, 862, and 864, but having a 1-100 amino acid linker at the C-terminus (e.g., following a sequence corresponding to an ISG).

[0014] In some embodiments, the cp polypeptide is present as a fusion protein with a first peptide, polypeptide or protein of interest. In some embodiments, the first peptide, polypeptide or protein of interest is selected from the group consisting of an antibody, an antibody fragment, Protein A, an Ig binding domain of Protein A, Protein G, an Ig binding domain of Protein G, Protein A / G, an Ig binding domain of Protein A / G, Protein L, an Ig binding domain of Protein L, Protein M, an Ig binding domain of Protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte binding domain(s) of a protein. In some embodiments, the first peptide, polypeptide or protein of interest is fused to the C-terminus or N-terminus of the cp polypeptide. In some embodiments, the fusion of the cp polypeptide comprises a second peptide, polypeptide or protein of interest. In some embodiments, the second peptide, polypeptide, or protein of interest is selected from the group consisting of an antibody, an antibody fragment, Protein A, an Ig binding domain of Protein A, Protein G, an Ig binding domain of Protein G, Protein A / G, an Ig binding domain of Protein A / G, Protein L, an Ig binding domain of Protein L, Protein M, an Ig binding domain of Protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, an anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte binding domain(s) of a protein. In some embodiments, the first peptide, polypeptide, or protein of interest is fused to the C-terminus or N-terminus of the cp polypeptide. In some embodiments, the first peptide, polypeptide, or protein of interest and the second peptide, polypeptide, or protein of interest are interacting elements capable of forming a complex with each other. In some embodiments, the first peptide, polypeptide, or protein of interest and the second peptide, polypeptide, or protein of interest are colocalizing elements configured to colocalize within a cellular compartment, a cell, a tissue, or an organism.In some embodiments, the cp polypeptide is linked to a molecule of interest.

[0015] In some embodiments, provided herein are polynucleotides that encode the circularly permuted variants described herein. In some embodiments, provided herein are polynucleotides that encode fusion proteins that include the circularly permuted variants described herein.

[0016] In some embodiments, provided herein are expression vectors comprising a polynucleotide encoding a circularly permuted variant, or fusions comprising the circularly permuted variants herein.

[0017] In some embodiments provided herein is a circularly permuted variant described herein, a fusion of a circularly permuted variant described herein, or a cell comprising a polynucleotide or expression vector encoding a circularly permuted variant or a fusion of a circularly permuted variant described herein.

[0018] In some embodiments, provided herein are compositions comprising a split / cp variant of a polypeptide comprising at least 70% sequence identity (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) with SEQ ID NO: 1. In some embodiments, the split / cp variant comprises (i) a first fragment of the cp polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) sequence identity with a portion of SEQ ID NO: 1, and (ii) a second fragment of the cp polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity with a portion of SEQ ID NO: 1. In some embodiments, the first fragment and the second fragment together comprise an amino acid sequence that corresponds to at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, 100%) of SEQ ID NO:1.In some embodiments, the split / cp variant is between positions 5 and 13 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, or a range therebetween), between positions 36 and 51 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, or a range therebetween), between positions 63 and 72 (e.g., 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, or a range therebetween) of SEQ ID NO:1. range), between positions 84 and 92 (e.g., 84, 85, 86, 87, 88, 89, 90, 91, 92, or any range therebetween), between positions 104 and 130 (e.g., 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, or any range therebetween), between positions 142 and 148 (e.g., 142, 143, 144, 145, 146, 147, 148, and ranges therebetween), between 160 and 174 (e.g., 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, or ranges therebetween), between 186 and 189 (e.g., 186, 187, 188, 189, or ranges therebetween), between 201 and 203 (e.g., 201, 202, 203, or ranges therebetween), in the range of), between positions 221 and 229 (e.g., 221, 222, 223, 224, 225, 226, 227, 228, 229, or ranges therebetween), or between positions 269 and 290 (e.g., 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290, or ranges therebetween).In some embodiments, the split / cp variant comprises a deletion of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or any range therebetween) at a position corresponding to one or more of the N-terminus of SEQ ID NO: 1, the C-terminus of SEQ ID NO: 1, and on either side of the cp site. In some embodiments, the split / cp variant comprises a duplicated sequence of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or any range therebetween) at a position corresponding to either side of the cp site. In some embodiments, the split / cp variant is capable of forming a covalent bond with a haloalkane substrate. In some embodiments, the first fragment is present as a fusion protein with a first peptide, polypeptide, or protein of interest. In some embodiments, the first peptide, polypeptide, or protein of interest is selected from the group consisting of an antibody, an antibody fragment, Protein A, an Ig binding domain of Protein A, Protein G, an Ig binding domain of Protein G, Protein A / G, an Ig binding domain of Protein A / G, Protein L, an Ig binding domain of Protein L, Protein M, an Ig binding domain of Protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, an anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte binding domain(s) of a protein. In some embodiments, the second fragment is present as a fusion protein with a second peptide, polypeptide, or protein of interest.In some embodiments, the second peptide, polypeptide, or protein of interest is selected from the group consisting of an antibody, an antibody fragment, Protein A, an Ig binding domain of Protein A, Protein G, an Ig binding domain of Protein G, Protein A / G, an Ig binding domain of Protein A / G, Protein L, an Ig binding domain of Protein L, Protein M, an Ig binding domain of Protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, an anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte binding domain(s) of a protein. In some embodiments, the first peptide, polypeptide, or protein of interest and the second peptide, polypeptide, or protein of interest are interacting elements capable of forming a complex with each other. In some embodiments, the first peptide, polypeptide, or protein of interest and the second peptide, polypeptide, or protein of interest are colocalizing elements configured to colocalize within a cellular compartment, cell, tissue, or organism. In some embodiments, the second fragment is linked to a molecule of interest.

[0019] In some embodiments, provided herein are one or more polynucleotides encoding the cp variants described herein. In some embodiments, provided herein are one or more expression vectors comprising one or more polynucleotides described herein. In some embodiments, provided herein are host cells comprising one or more polynucleotides or one or more expression vectors described herein. [Brief description of the drawings]

[0020] [Figure 1A] 1 is a schematic diagram showing the organization of a circularly permuted (cp) polypeptide, with the N- and C-termini of the construct ("N-terminus" and C-terminus) and the naturally occurring N- and C-termini ("N-terminus" and "C-terminus") positions indicated. [Figure 1B]Schematic diagram showing the organization of split circularly permuted (sp / cp) polypeptides, with the N- and C-termini of the construct ("N-terminus" and C-terminus) and the naturally occurring N- and C-termini ("N-terminus" and "C-terminus") positions indicated. [Figure 2A] Enzyme activity, thermostability, and TEV protease-induced stability changes of cpHT library variants. E. coli lysates containing overexpressed cpHT proteins (cp positions shown along the x-axis) were diluted 5-fold and then mixed 1:1 with CA-AlexaFluor488 ligand to a final concentration of 10 nM. Fluorescence polarization (FP) was monitored for 30 min and initial velocities (ΔmP / s) were calculated. Relative activity was calculated by dividing the cpHT rate by the rate of lysates containing overexpressed 6xHis-HaloTag7 control protein. [Figure 2B] Enzyme activity, thermostability, and TEV protease-induced stability changes of cpHT library variants. The same lysates (undiluted) were heated at 40-90 °C for 30 min, then cooled to room temperature (25 °C) and mixed 1:1 with CA-TMR to a final concentration of 10 nM. FP was measured after 2 h of room temperature incubation. FP intensity is represented by grey shading, with darker shading indicating higher FP values. [Figure 2C] Enzyme activity, thermostability, and TEV protease-induced stability changes of cpHT library variants. The experiment in (B) was repeated using lysates treated with TEV protease. Changes in FP compared to non-TEV-treated lysates are indicated by color shading, with lighter gray indicating more negative changes and darker gray indicating more positive changes. [Figure 2D] Enzymatic activity, thermostability, and TEV protease-induced stability changes of cpHT library variants. Real-time fluorescence polarization assay with HaloTag Alexa488 ligand used to monitor activity of cpHT variants against HALOTAG in E. coli lysates (red dotted line reflects baseline relative activity level of 0.03 (red dotted line in graph), the amount of signal visually separated above background during the real-time assay). [Diagram 3] Fold increase in JF646 signal after addition of rapamycin to non-overlapping split HaloTag fragments. E. coli lysates containing overexpressed spHT protein fragments fused to FRB or FKBP were mixed in the combinations shown on the left side of the table. The lysate mixtures were incubated with 50 nM rapamycin (or no rapamycin as control) for 30 min at room temperature. 100 nM Janelia Fluor 646 ligand was added 1:1 (by volume) to the mixture (50 nM final concentration). Samples were incubated for 24 min at room temperature. Samples were analyzed for fluorescence (excitation: 646 nm, emission: 664 nm) on a Tecan Infinite M1000 microplate reader. The fold signal increase was calculated as Frap+ / Frap- for each combination. [Figure 4] Fold increase in JF646 signal after addition of rapamycin to partially overlapping split HaloTag fragments. Experimental conditions were identical to those in FIG. 3. [Diagram 5] Reactivity of circularly permuted (cp) HaloTag constructs to Janelia Fluor HaloTag ligands, with permutations localized to the fluorophore-interacting lid subdomain of HaloTag. E. coli lysates containing overexpressed cpHT variants were mixed with each of the four JF dye ligands (50 nM final concentration) and incubated for 22 hours at room temperature. LgBiT lysates were included for each ligand as a non-binding negative control. Unsubstituted HaloTag (HT) was included as a positive control. Fluorescence was measured on a Tecan Infinite M1000 microplate reader using the following excitation and emission wavelengths: JF525, 525 nm / 549 nm; JF549, 549 nm / 571 nm; JF585, 585 nm / 609 nm; JF646, 646 nm / 664 nm. [Figure 6]TMR-labeled cpHT lysates visualized by SDS-PAGE and in-gel fluorescence. These gels are provided as a reference for the general reactivity of cpHT variants to high affinity but non-fluorogenic ligands. The samples shown in these gels are from previous experiments and are not the same lysates used in Figure 5. With the exception of cpHT177 and cpHT178, all cpHT variants between 138-180 are fairly well expressed and reactive with TMR. The cpHT variants between 160-178 fail to activate the fluorogenic JF ligand (Figure 5). [Figure 7] Development of fluorogenic signals from cpHT constructs in E. coli lysates corresponding to the current spHT design. 6xHis-HT7 serves as a positive control and FRB-LgBiT serves as a negative control. The red, green, and blue dashed lines allow easy comparison to the positive control fluorescent signal. Measurements were performed at a constant instrument gain of 100 for direct brightness comparison. Top, 45 min incubation; middle, 2 h incubation; bottom, 24 h incubation at room temperature. Fluorescence was measured on a Tecan Infinite M1000 microplate reader using the following excitation and emission wavelengths: JF585, 585 nm / 609 nm; JF635, 635 nm / 652 nm; JF646, 646 nm / 664 nm. [Figure 8] Graph showing changes in fluorescence polarization of cpHT after TEV cleavage. [Figure 9] Gels and graphs showing ligand specificity of exemplary cpHT variants. [Figure 10] Graph showing the thermostability profile of cpHT variants in E. coli lysates using fluorescence polarization after heat treatment to determine the effect of circular permutation and TEV cleavage on stability.

[0021] definition Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the embodiments described herein, some preferred methods, compositions, devices, and materials are described herein. However, before describing the materials and methods, it should be understood that the invention is not limited to the specific molecules, compositions, methodologies, or procedures described herein, as these may vary according to routine experimentation and optimization. It should also be understood that the terminology used in the description is for the purpose of describing the particular versions or embodiments only, and is not intended to limit the scope of the embodiments described herein.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. However, in case of conflict, the present specification, including definitions, shall control. Therefore, in the context of the embodiments described herein, the following definitions apply.

[0023] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "polypeptide" is a reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth.

[0024] As used herein, the term "and / or" includes any and all combinations of the listed items, including any of the individually listed items. For example, "A, B and / or C" includes A, B, C, AB, AC, BC and ABC, each of which should be considered as being individually listed with the explicit reference "A, B and / or C."

[0025] As used herein, the term "comprise" and linguistic variations thereof indicate the presence of the recited feature(s), element(s), method step(s), etc., without excluding the presence of additional feature(s), element(s), method step(s), etc. Conversely, the term "consisting of" and linguistic variations thereof indicate the presence of the recited feature(s), element(s), method step(s), etc., and excludes any unrecited feature(s), element(s), method step(s), etc., except for impurities normally associated therewith. The phrase "consisting essentially of" indicates the recited feature(s), element(s), method step(s), etc., and any additional feature(s), element(s), method step(s), etc. that do not materially affect the basic nature of the composition, system, or method. Many embodiments herein are described using the open term "comprising." Such embodiments encompass the multiple closed "consisting of" and / or "consisting essentially of" embodiments, which may alternatively be claimed or described using such language.

[0026] As used herein, the term "substantially" means that the recited properties, parameters, and / or values ​​need not necessarily be achieved exactly, but deviations or variations may occur, including, for example, tolerances, measurement errors, limits of measurement precision, and other factors known to those of skill in the art, to an extent that does not interfere with the effect intended to be provided by the property. A substantially absent (e.g., substantially non-fluorescent) property or characteristic may be one that is within the noise, below background, below the detection capabilities of the assay being used, or one that is a small percentage (e.g., <1%, <0.1%, <0.01%, <0.001%, <0.00001%, <0.000001%, <0.0000001%) of a notable property (e.g., fluorescence intensity of an active fluorophore).

[0027] As used herein, when referring to an amino acid sequence or a position within an amino acid sequence, the phrase "corresponding to" refers to the relative position of an amino acid residue or amino acid segment, and refers to the sequence, but not necessarily the specific identity of the amino acid at that position. For example, a "peptide corresponding to positions 36-48 of SEQ ID NO:1" may have less than 100% sequence identity (e.g., greater than 70% sequence identity) with positions 36-48 of SEQ ID NO:1, but within the context of the composition or system being described, the peptide is relative to those positions.

[0028] As used herein, the term "system" refers to multiple components (e.g., devices, compositions, etc.) used for a particular purpose. For example, two separate biological molecules can comprise a system if they are useful together for a common purpose, whether or not they are present in the same composition.

[0029] As used herein, the term "complementary" refers to the property of two or more structural elements (e.g., peptides, polypeptides, nucleic acids, small molecules, etc.) that can hybridize, dimerize, or otherwise form a complex with each other. For example, "complementary peptides and polypeptides" can combine to form a complex. Complementary elements may require assistance (facilitation) to form a complex (e.g., from interacting elements), for example, to place the elements in the proper conformation for complementarity, to place the elements in the proper proximity for complementarity, to colocalize complementary elements, to lower the interaction energy for complementary elements, to overcome insufficient affinity for each other.

[0030] As used herein, the term "complex" refers to an assembly or aggregate of molecules (e.g., peptides, polypeptides, etc.) that are in direct and / or indirect contact with each other. In one aspect, "contact," or more specifically, "direct contact," means that two or more molecules are sufficiently close that attractive non-covalent interactions, such as van der Waals forces, hydrogen bonds, ionic and hydrophobic interactions, dominate the interaction of the molecules. In such an aspect, a complex of molecules (e.g., peptides, polypeptides, etc.) is formed under assay conditions such that the complex is thermodynamically favorable (e.g., compared to the unaggregated or uncomplexed states of its constituent molecules). As used herein, the term "complex" refers to an assembly of two or more molecules (e.g., peptides, polypeptides, etc.), unless otherwise specified.

[0031] As used herein, the term "interacting element" refers to a moiety that assists or facilitates two or more structural elements (e.g., peptides, polypeptides, etc.) to come together to form a complex. In some embodiments, a pair of interacting elements (also known as an "interacting pair") is attached to a pair of structural elements (e.g., peptides, polypeptides, etc.), and an attractive interaction between the two interacting elements facilitates the formation of a complex of the structural elements. The interacting element may facilitate the formation of the complex by any suitable mechanism (e.g., bringing the structural elements into close proximity, positioning the structural elements in a suitable conformation for stable interaction, reducing the activation energy for complex formation, combinations thereof, etc.). The interacting element may be a protein, polypeptide, peptide, small molecule, cofactor, nucleic acid, lipid, carbohydrate, antibody, etc. An interacting pair may be made of two of the same interacting elements (i.e., a homopair) or two different interacting elements (i.e., a heteropair). In the case of a heteropair, the interacting elements may be the same type of moiety (e.g., a polypeptide) or two different types of moieties (e.g., a polypeptide and a small molecule). In some embodiments in which complex formation by an interacting pair is studied, the interacting pair may be referred to as a "target pair" or "pair of interest," and the individual interacting elements are referred to as "target elements" (e.g., "target peptide," "target polypeptide," etc.) or "elements of interest" (e.g., "peptide of interest," "polypeptide of interest," etc.).

[0032] As used herein, the term "low affinity" describes an intermolecular interaction between two or more entities that is too weak to result in significant complex formation between the entities, except at concentrations substantially higher (e.g., 2x, 5x, 10x, 100x, 1000x, or more) than physiological or assay conditions, or concentrations that involve promotion of the attached elements (e.g., interacting elements) from forming secondary complexes.

[0033] As used herein, the term "high affinity" describes an intermolecular interaction between two or more (e.g., three) entities that is strong enough to produce detectable complex formation under physiological or assay conditions without promotion from the formation of a secondary complex of the attached elements (e.g., interacting elements).

[0034] As used herein, the term "existing protein" refers to an amino acid sequence that physically existed prior to a particular event or date. A "peptide that is not a fragment of an existing protein" is a short chain of amino acids that is not a fragment or subsequence of a protein (e.g., synthetic or naturally occurring) that physically existed prior to the design and / or synthesis of the peptide.

[0035] As used herein, the term "fragment" refers to a peptide or polypeptide that results from dissociation or "fragmentation" of a larger whole entity (e.g., protein, polypeptide, enzyme, etc.), or that has been prepared to have the same sequence as such. Thus, a fragment is a subsequence of the whole entity (e.g., protein, polypeptide, enzyme, etc.) for which the fragment is made and / or designed. A peptide or polypeptide that is not a subsequence of an existing whole protein is not a fragment (e.g., not a fragment of an existing protein). A peptide or polypeptide that is "not a fragment of an existing protein" is an amino acid chain that is not a subsequence of a protein (e.g., natural or synthetic) that physically existed prior to the design and / or synthesis of the peptide or polypeptide. As used herein, a fragment of a hydrolase or dehalogenase is a sequence that is less than the full-length sequence but is not capable of forming a substrate binding site by itself and / or has substantially reduced or no substrate binding activity, but is closely adjacent to a second fragment of the hydrolase or dehalogenase and exhibits substantially increased substrate binding activity. In one embodiment, a fragment of a hydrolase or dehalogenase comprises at least 5, e.g., at least 10, at least 20, at least 30, at least 40, or at least 50 contiguous residues of a wild-type or mutant hydrolase, or a sequence having at least 70% sequence identity thereto, and may not necessarily include the N- or C-terminal residues or N- or C-terminal sequences of the corresponding full-length protein.

[0036] As used herein, the term "subsequence" refers to a peptide or polypeptide that has 100% sequence identity with a portion of another larger peptide or polypeptide, the subsequence being a perfect sequence match for a portion of the larger amino acid chain.

[0037] The term "amino acid" refers to natural amino acids, unnatural amino acids, and amino acid analogs, all in their D and L stereoisomeric forms, unless otherwise indicated, if their structures allow for such stereoisomeric forms.

[0038] The term "proteinogenic amino acid" refers to the 20 amino acids encoded for the human genetic code, including alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine ​​(Cys or C), glutamine (Gln or Q), glutamic acid (Glu or E), glycine (Gly or G), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V). Selenocysteine ​​and pyrrolysine may also be considered proteinogenic amino acids.

[0039] The term "non-proteinogenic amino acid" refers to an amino acid that is not naturally encoded or found in the genetic code of any organism and is not biosynthetically incorporated into a protein during translation. A non-proteinogenic amino acid may be a "non-natural amino acid" (an amino acid that does not occur in nature) or a "naturally occurring non-proteinogenic amino acid" (e.g., norvaline, ornithine, homocysteine, etc.). Examples of non-proteinogenic amino acids include, but are not limited to, azetidine carboxylic acid, 2-aminoadipic acid, 3-aminoadipic acid, beta-alanine, naphthylalanine, aminopropionic acid, 2-aminobutyric acid, 4-aminobutyric acid, 6-aminocaproic acid, 2-aminoheptanoic acid, 2-aminoisobutyric acid, 3-aminoisobutyric acid, 2-aminopimelic acid, tertiary butylglycine, 2,4-diaminoisobutyric acid, desmosine, 2,2'-diaminopimelic acid, 2,3-diaminopropionic acid, N-ethylglycine, N-ethylasparagine, homoproline, hydroxylysine, allo-hydroxylysine, 3-hydroxyproline, 4-hydroxyproline, isodesmosine, allo-isoleucine, N-methylalanine, N-alkylglycines including N-methylglycine, N-methylisoleucine, N-alkylpentylglycines including N-methylpentylglycine. Included in the non-proteinaceous structures are N-methylvaline, naphthylalanine, norvaline, norleucine ("Norleu"), octylglycine, ornithine, pentylglycine, pipecolic acid, thioproline, homolysine, and homoarginine. Non-proteinaceous structures also include D-amino acid forms of any of the amino acids herein, as well as non-alpha amino acid forms of any of the amino acids herein (such as beta amino acids, gamma amino acids, delta amino acids, etc.), all of which are within the scope of the present invention and may be included in the peptides herein.

[0040] The term "amino acid analog" refers to an amino acid (e.g., natural or non-natural, proteinogenic or non-proteinogenic) in which one or more of the C-terminal carboxy group, the N-terminal amino group, and the side chain bioactive group are chemically blocked, reversibly or irreversibly, or otherwise modified to a bioactive group. For example, aspartic acid-(beta-methyl ester) is an amino acid analog of aspartic acid, N-ethylglycine is an amino acid analog of glycine, or alanine carboxamide is an amino acid analog of alanine. Other amino acid analogs include methionine sulfoxide, methionine sulfone, S-(carboxymethyl)-cysteine, S-(carboxymethyl)-cysteine ​​sulfoxide, and S-(carboxymethyl)-cysteine ​​sulfone.

[0041] As used herein, unless otherwise specified, the terms "peptide" and "polypeptide" refer to a polymeric compound of two or more amino acids joined through the backbone by peptide amide bonds (-C(O)NH-). The term "peptide" typically refers to short amino acid polymers (e.g., chains having fewer than 30 amino acids) and the term "polypeptide" typically refers to longer amino acid polymers (e.g., chains having more than 30 amino acids).

[0042] As used herein, the term "artificial" refers to compositions and systems that are designed or prepared by man and do not occur in nature, for example, an artificial peptide, peptide, or nucleic acid that contains a non-naturally occurring sequence (e.g., a peptide that does not have 100% identity to a naturally occurring protein or fragment thereof).

[0043] As used herein, a "conservative" amino acid substitution refers to the replacement of an amino acid in a peptide or polypeptide with another amino acid that has similar chemical properties, such as size or charge. For purposes of this disclosure, each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) Alanine (A) and Glycine (G); 2) Aspartic acid (D) and glutamic acid (E); 3) Asparagine (N) and Glutamine (Q); 4) arginine (R) and lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), and Valine (V); 6) phenylalanine (F), tyrosine (Y), and tryptophan (W); 7) serine (S) and threonine (T); and 8) Cysteine ​​(C) and methionine (M).

[0044] Naturally occurring residues may be divided into classes based on common side chain properties, e.g., polar positive (or basic) (histidine (H), lysine (K), and arginine (R)), polar negative (or acidic) (aspartic acid (D), glutamic acid (E)), polar neutral (serine (S), threonine (T), asparagine (N), glutamine (Q)), nonpolar fatty acids (alanine (A), valine (V), leucine (L), isoleucine (I), methionine (M)), nonpolar aromatic (phenylalanine (F), tyrosine (Y), tryptophan (W)), proline and glycine, and cysteine. As used herein, a "semi-conservative" amino acid substitution refers to the replacement of an amino acid in a peptide or polypeptide with another amino acid within the same class.

[0045] In some embodiments, unless otherwise specified, conservative or semi-conservative amino acid substitutions may also include non-naturally occurring amino acid residues that have similar chemical properties to the natural residues. These non-natural residues are typically incorporated by chemical peptide synthesis rather than by synthesis in biological systems. These include, but are not limited to, peptidomimetics and other reversed or inverted forms of amino acid moieties. The embodiments herein may, in some embodiments, be limited to natural amino acids, non-natural amino acids, and / or amino acid analogs.

[0046] Non-conservative substitutions may involve exchanging a member of one class for a member of another class.

[0047] As used herein, the term "sequence identity" refers to the degree to which two polymeric sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have the same sequential composition of monomeric subunits. The term "sequence similarity" refers to the degree to which two polymeric sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have similar polymeric sequences. For example, similar amino acids are those that share the same biophysical properties and can be grouped, for example, into acidic (e.g., aspartate, glutamate), basic (e.g., lysine, arginine, histidine), non-polar (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), and uncharged polar (e.g., glycine, asparagine, glutamine, cysteine, serine, threonine, tyrosine). "Percent sequence identity" (or "percent sequence similarity") is calculated by: (1) comparing two optimally aligned sequences over a window of comparison (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); (2) determining the number of positions that contain identical (or similar) monomers (e.g., the same amino acid occurs in both sequences, a similar amino acid occurs in both sequences) to obtain the number of matched positions; (3) dividing the number of matched positions by the total number of positions in the comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); and (4) multiplying the result by 100 to obtain the percent sequence identity or percent sequence similarity. For example, if peptide A and peptide B are both 20 amino acids long and have identical amino acids at all but one position, peptide A and peptide B have 95% sequence identity. If the amino acids at the non-identical positions share the same biophysical properties (e.g., both were acidic), peptide A and peptide B will have 100% sequence similarity.As another example, if peptide C is 20 amino acids long, peptide D is 15 amino acids long, and 14 of the 15 amino acids in peptide D are identical to a portion of peptide C, then peptides C and D have 70% sequence identity, but peptide D has 93.3% sequence identity with the optimal comparison window of peptide C. For purposes of calculating "percent sequence identity" (or "percent sequence similarity") herein, any gap in the aligned sequences is treated as a mismatch at that position.

[0048] Any peptide / polypeptide described herein as having a particular percent sequence identity or similarity (e.g., at least 70%) with a reference sequence ID number may also be expressed as having a maximum number of substitutions (or terminal deletions) relative to that reference sequence. For example, a sequence having at least Y% sequence identity (e.g., 90%) with SEQ ID NO: Z (e.g., 100 amino acids) may have a maximum of X substitutions (e.g., 10) with SEQ ID NO: Z, and thus may also be expressed as "having no more than X (e.g., 10) substitutions with SEQ ID NO: Z."

[0049] As used herein, the term "wild type" refers to a gene or gene product (e.g., a protein, polypeptide, peptide, etc.) that has the characteristics (e.g., sequence) of that gene or gene product isolated from a naturally occurring source and is most frequently observed in a population. In contrast, the term "mutant" or "variant" refers to a gene or gene product that exhibits a modification of the sequence compared to the wild-type gene or gene product. Note that a "naturally occurring variant" is a gene or gene product that occurs in nature but has an altered sequence compared to the wild-type gene or gene product, and is not the most commonly occurring sequence. An "artificial variant" is a gene or gene product that has an altered sequence when compared to the wild-type gene or gene product and does not occur in nature. A variant gene or gene product may be naturally occurring but not the most common variant of the gene or gene product, or "synthetic" produced by human or experimental intervention.

[0050] As used herein, the term "physiological conditions" encompasses any conditions compatible with living cells, e.g., primarily aqueous conditions of temperature, pH, salinity, chemical composition, and the like, that are compatible with living cells.

[0051] As used herein, the term "sample" is used in its broadest sense. In one sense, it is meant to include specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from animals (including humans) and encompass fluids, solids, tissues, and gases. Biological samples include blood products such as plasma, serum, etc. Samples can also refer to cell lysates or purified forms of the enzymes, peptides, and / or polypeptides described herein. Cell lysates can include cells lysed with a lysing agent, or lysates, e.g., rabbit reticulocyte or wheat germ lysates. Samples can also include cell-free expression systems. Environmental samples include environmental materials, such as surface materials, soil, water, crystals, and industrial samples. However, such examples should not be construed as limiting the sample types applicable to the present invention.

[0052] As used herein, the terms "fusion," "fusion polypeptide," and "fusion protein" refer to a chimeric protein that contains a first protein or polypeptide of interest joined to a second, different peptide, polypeptide, or protein (e.g., an interacting element).

[0053] As used herein, the terms "conjugate" and "conjugation" refer to the covalent joining of two molecular entities (e.g., post-synthetic and / or during synthetic production). The chemical (e.g., "chemically" conjugated) or enzymatic attachment of a peptide or small molecule tag to a protein or small molecule is an example of a conjugate.

[0054] As used herein, the terms "polypeptide component" or "peptide component" are used synonymously with the terms "polypeptide component of a [modified dehalogenase] complex" or "peptide component of a [modified dehalogenase] complex." Typically, as used herein, a polypeptide component or peptide component is capable of forming a complex with a second component under appropriate conditions to form a desired complex.

[0055] As used herein, the term "dehalogenase" refers to an enzyme that catalyzes the removal of halogen atoms from a substrate. The term "haloalkane dehalogenase" refers to an enzyme that catalyzes the removal of halogens from a haloalkane substrate to produce alcohols and halides. Dehalogenases and haloalkyl dehalogenases belong to the hydrolase enzyme family and may be referred to as such herein or elsewhere.

[0056] As used herein, the term "modified dehalogenase" refers to a dehalogenase variant (artificial variant) that has a mutation that prevents release of the substrate from the protein after removal of the halogen and results in a covalent bond between the substrate and the modified dehalogenase. The HALOTAG system (Promega) is a commercially available modified dehalogenase and substrate system.

[0057] As used herein, the term "circularly permuted" ("cp") refers to a polypeptide in which the N-terminus and C-terminus are joined together, either directly or through a linker, to produce a circular polypeptide, which is then opened at a location other than between the N-terminus and C-terminus to produce a new linear polypeptide that differs from the ends of the original polypeptide. The location at which the circular polypeptide is opened is referred to herein as the "cp site." Circularly permuted polypeptides include polypeptides that have the same sequence and structure as a circularly permuted and then opened polypeptide. Thus, cp polypeptides may be synthesized de novo as linear molecules, without undergoing the circular permutation and opening steps. The preparation of circularly permuted derivatives is described in International Publication No. WO 95 / 27732, which is incorporated by reference in its entirety.

[0058] As used herein, the term "split" ("sp") refers to a polypeptide that has been divided into two fragments at an internal site in the original polypeptide. The fragments of the sp polypeptide are structurally complementary and can reconstitute the activity of the original polypeptide if they are able to form an active complex.

[0059] As used herein, the term "gapped" refers to a variant of a polypeptide that lacks a segment of the original polypeptide. For example, a "gapped cp polypeptide" or a "gapped sp polypeptide" is one that lacks a segment of the original sequence that occurs at the site of circular permutation or split.

[0060] As used herein, the term "overlapping" refers to a variant of a polypeptide that includes duplication of a segment of an original polypeptide. For example, an "overlapping sp polypeptide" is one in which a segment of the original sequence adjacent to the split site is present (overlapped) at the C-terminus of a first fragment and at the N-terminus of a second fragment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0061] Provided herein are circularly permuted (cp) dehalogenase variants that are capable of covalently binding to haloalkyl ligands. Specifically provided are peptides and polypeptides that include the cp dehalogenase variants, as well as split versions thereof that are structurally assembled to form active dehalogenase complexes.

[0062] In a circularly permuted polypeptide sequence (e.g., SEQ ID NO:X), (1) the final amino acid of the sequence (e.g., corresponding to the final position of SEQ ID NO:X) is peptide-linked (e.g., directly or via a peptide linker) to an initial amino acid of the sequence (e.g., corresponding to the first position of SEQ ID NO:X), and (2) the polypeptide is split at an internal position (cp site) within the sequence, thereby creating a linear polypeptide in which the initial position of the substitute corresponds to the amino acid position immediately following the cp site and the final position of the substitute corresponds to the amino acid position immediately preceding the cp site (Figure 1A).

[0063] In some embodiments, provided herein are modified dehalogenases and circularly permuted hydrolases and dehalogenases such as those derived from commercially available HALOTAG (Promega) and / or mutant hydrolases disclosed in U.S. Published Application No. 2006 / 0024808, the disclosures of which are incorporated herein by reference. In experiments conducted during the development of the embodiments herein, a comprehensive screen of all possible circularly permuted sites in HALOTAG was performed to identify variants that retain activity and stability in the context of a single polypeptide (e.g., cpHT) and / or conditionally separable fragments (e.g., sp / cpHT).

[0064] In some embodiments, provided herein are HALOTAG-based systems tailored for functional biology, such as circularly permuted HATOTAG polypeptides or split versions of cp HATOTAG polypeptides that have similar properties to existing full-length proteins in terms of fragment stability, solubility, and expression, with the added property of being able to reconstitute a significant portion of its activity upon reconstitution of the complete enzyme. HALOTAG ligands of particular importance to certain embodiments herein include fluorogenic ligands. Systems including cpHT and sp / cpHT can be engineered to have a wide range of fragment affinities to enable both facilitated and spontaneous complementation systems. Both circularly permuted and split / cp HALOTAG systems facilitate endogenous tagging of proteins, leading to better fluorogenic ligands or sensors through higher signal, stability, dynamic range, etc. The HALOTAG-based functional biology tools described herein are highly suitable for measuring protein dynamics in live cells using fluorescence imaging, where other techniques lack the utility of HALOTAG's self-labeling activity or the sensitivity of fluorescent chloroalkane ligands.

[0065] As described herein, embodiments are not limited to the HALOTAG sequence. In some embodiments, provided herein are circularly permuted modified dehalogenases that differ in sequence from SEQ ID NO:1. In some embodiments, provided herein are circularly permuted dehalogenases that lack the mutation(s) (e.g., 272 and / or 106) that result in covalent attachment to the haloalkane substrate. Such cp dehalogenases allow for substrate conversion but are otherwise true enzymes that include the sequences and properties of the embodiments described herein.

[0066] During the development of the embodiments herein, experiments were performed to examine the split dehalogenases, their ability to form into active dehalogenase structures, and their ability to activate fluorogenic substrates. A comprehensive screen of all circular permutations of HaloTag (cpHT) revealed that 228 / 296 (77%) reacted with CA-TMR, with 50 variants having at least 10% of the native HT activity on CA-AlexaFluor488. Seventeen cpHT variants had increased thermal stability compared to HT, and 38 variants showed activity recovery after thermal denaturation, likely due to protein refolding. The most active variants due to AlexaFluor488 kinetics clustered in a region distal to the lid domain, although this effect may be specific to this substrate, which is negatively charged and may be sensitive to lid domain perturbation. Indeed, the clustering effect was less pronounced when a neutral TMR ligand was used. With the exception of cpHT around residues 111 and 120, all refolding variants were localized to the lid domain, and all thermostabilized variants were also within the lid domain.

[0067] In some embodiments, provided herein are cpHT polypeptides and systems thereof. Specifically, provided are cp modified dehalogenases that can retain all or a portion of the activity of the parent dehalogenase. In some embodiments, the cp modified dehalogenases exhibit desired functions and properties that are different from or enhanced relative to the parent dehalogenase (e.g., stability, refolding, solubility, etc.).

[0068] In some embodiments, the polypeptides, peptides, fragments and combinations thereof described herein are derived from the modified dehalogenase sequence of SEQ ID NO:1. MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETF QAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG.

[0069] In some embodiments, the peptides and polypeptides herein comprise at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) with all or a portion of SEQ ID NO:1. In some embodiments, the peptides and polypeptides herein comprise 100% sequence identity with all or a portion of SEQ ID NO:1. In some embodiments, the peptides and polypeptides herein comprise at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:1. In some embodiments, the peptides and polypeptides herein comprise 100% sequence similarity to all or a portion of SEQ ID NO:1.

[0070] In some embodiments, a peptide or polypeptide herein comprises an A at a position corresponding to position 2 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a V at a position corresponding to position 47 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a T at a position corresponding to position 58 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a G at a position corresponding to position 78 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an F at a position corresponding to position 88 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an M at a position corresponding to position 89 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an F at a position corresponding to position 128 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a T at a position corresponding to position 155 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a K at a position corresponding to position 160 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a V at a position corresponding to position 167 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a T at a position corresponding to position 172 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an M at a position corresponding to position 175 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a G at a position corresponding to position 176 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an N at a position corresponding to position 195 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an E at a position corresponding to position 224 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a D at a position corresponding to position 227 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a K at a position corresponding to position 257 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an A at a position corresponding to position 264 of SEQ ID NO:1.In some embodiments, a peptide or polypeptide herein comprises an N at a position corresponding to position 272 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an L at a position corresponding to position 273 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an S at a position corresponding to position 291 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a T at a position corresponding to position 292 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an E at a position corresponding to position 294 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an I at a position corresponding to position 295 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises an S at a position corresponding to position 296 of SEQ ID NO:1. In some embodiments, a peptide or polypeptide herein comprises a G at a position corresponding to position 297 of SEQ ID NO:1.

[0071] In some embodiments, a cp dehalogenase (e.g., cpHT) comprises two portions that together comprise at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:1. In some embodiments, the cp dehalogenase (e.g., cpHT) comprises two portions that together comprise at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) with the complete sequence of SEQ ID NO:1. For example, the portion of the cp polypeptide that is N-terminal to the cp site corresponds to a first portion of SEQ ID NO:1 (e.g., at least 70% sequence identity with the first portion) and the portion of the cp polypeptide that is C-terminal to the cp site corresponds to a second portion of SEQ ID NO:1 (e.g., at least 70% sequence identity with the second portion). In some embodiments, the cp dehalogenase (e.g., cpHT) comprises two portions that together comprise 100% sequence identity with all or a portion of SEQ ID NO:1. For example, the portion of the cp polypeptide that is N-terminal to the cp site has 100% sequence identity to a first portion of SEQ ID NO:1, and the portion of the cp polypeptide that is C-terminal to the cp site has 100% sequence identity to a second portion of SEQ ID NO:1.

[0072] In some embodiments, the cp dehalogenase (e.g., cpHT) comprises two portions that together comprise at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:1. For example, the portion of the cp polypeptide that is N-terminal to the cp site corresponds to a first portion of SEQ ID NO:1 (e.g., at least 70% sequence similarity with the first portion) and the portion of the cp polypeptide that is C-terminal to the cp site corresponds to a second portion of SEQ ID NO:1 (e.g., at least 70% sequence similarity with the second portion). In some embodiments, the cp dehalogenase (e.g., cpHT) comprises two portions that together comprise 100% sequence similarity with all or a portion of SEQ ID NO:1. For example, the portion of the cp polypeptide that is N-terminal to the cp site has 100% sequence similarity to a first portion of SEQ ID NO:1, and the portion of the cp polypeptide that is C-terminal to the cp site has 100% sequence similarity to a second portion of SEQ ID NO:1.

[0073] In some embodiments, fragments of a parent sequence (e.g., a dehalogenase (e.g., HALOTAG)) are directly connected (e.g., terminal amino acid to the first amino acid) to form a cp polypeptide (e.g., a cp dehalogenase (e.g., cpHT)). In other embodiments, fragments of a parent sequence are fused together via a peptide linker. In some embodiments, the linker sequence is 1-100 amino acids in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acids, or a range therebetween). A suitable linker may be any sequence of amino acids unless otherwise specified herein.

[0074] In some embodiments, the cp dehalogenase (e.g., cpHT) provides enhanced functionality and / or properties compared to the parent dehalogenase (e.g., HALOTAG). In some embodiments, the cp dehalogenase (e.g., cpHT) differs significantly from the parent dehalogenase.

[0075] For example, in certain embodiments, the cp dehalogenase (e.g., cpHT) retains the ability of the parent dehalogenase (e.g., HALOTAG) to covalently bind to chloroalkane substrates, but does not exhibit the ability to activate fluorogenic cargoes bound to chloroalkanes. Thus, when using constitutively fluorescent ligands such as chloroalkane-tetramethylrhodamine (CA-TMR) or Janelia Fluor 549 (JF549), these cp dehalogenases emit visible fluorescence. However, fluorogenic ligands such as JF646, JF635, and JF585 do not emit fluorescence when bound to this class of cp dehalogenases. An exemplary application of such dehalogenases would be a system using two dehalogenases, one natural (e.g., HALOTAG) and one fluorogenic silent (e.g., cpHT, which cannot activate fluorogenic probes), in a single cell imaging experiment. A constitutively fluorescent substrate (e.g., chloroalkane-CA-TMR) will be visible / detectable if found on both dehalogenases. However, when a fluorogenic substrate (e.g., chloroalkane-JF646) is used, only the substrate bound to the native dehalogenase will be visible / detectable.

[0076] In some embodiments, the cp dehalogenase (e.g., cpHT) exhibits enhanced thermostability compared to the parent dehalogenase (e.g., HALOTAG). While native HALOTAG has a melting temperature of about 70° C., further stabilization increases its value for denaturation-based biochemical applications. In some embodiments, such thermostable cpHT finds use in diagnostic applications that require heating of samples. In some embodiments, the cp dehalogenase (e.g., cpHT) exhibits increased ambient stability or “shelf life” that is desirable for products, particularly rapid or on-demand laboratory or consumer testing. In some embodiments, for example, when fused to a protein of interest, the thermostable cp dehalogenase remains folded during heating of cell lysates in preparation for gel electrophoresis. Under moderate gel conditions, the thermostable cp dehalogenase retains its enzymatic activity and allows for in-gel fluorescent labeling, which may achieve effects similar to Western blotting. Additionally, improved thermostability is desirable for applications to thermophilic organisms.

[0077] In some embodiments, the cp polypeptide (as described above) comprises two fragments of the parent polypeptide sequence connected in reverse order by a linker sequence (Figure 1A). In some embodiments, the linker sequence is a cleavable linker (Figure 1B). In some embodiments, the linker is a sequence that is recognized by an enzyme, e.g., a cleavable sequence or a photocleavable sequence. Exemplary cleavable linker sequences are: [ka] (TEV protease recognition sequence underlined; cleavable peptide bond indicated by slash). Other TEV cleavable linkers (e.g., containing a TEV protease recognition sequence) or other cleavable linkers are within the scope of this specification.

[0078] In some embodiments, cleavage of the linker results in cleavage of the cp polypeptide into peptide or polypeptide fragments (FIG. 1B). However, because the fragments are expressed as a single cp polypeptide and allowed to fold as a single cp polypeptide, the fragments may retain a function and / or structure that would not be achieved by de novo assembly of separate fragments. In some embodiments, provided herein is a cp polypeptide comprising a cleavable linker sequence. In some embodiments, provided herein are peptide and / or polypeptide fragments generated by cleavage of the linker sequence of the cp polypeptide. In some embodiments, provided herein is a cpHT polypeptide comprising a cleavable linker sequence. In some embodiments, provided herein are peptide and / or polypeptide fragments generated by cleavage of the linker sequence of the cpHT polypeptide, referred to herein as sp / cpHT polypeptide.

[0079] sp / cp mutant proteins (e.g., sp / cp dehalogenase, sp / cpHT, etc.) are expressed or synthesized as a single cp polypeptide, but are capable of being cleaved into separate fragments due to the cleavable linker. Depending on the sp / cp protein, cleavage of the single cp polypeptide either (1) resulted in loss of substrate binding activity, (2) retained substrate binding activity as long as the fragments associated with each other but were unable to reassociate the fragments into an active complex, (3) retained the ability to reassociate the fragments into an active complex only when facilitated by a component bound to the fragment, or (4) retained the ability to reassociate the fragments into an active complex.

[0080] The sp / cp proteins are used to reveal and analyze protein interactions within cells, for example, where each portion (e.g., fragment) of the sp / cp protein is fused to a different protein. Provided herein are sp / cp mutant hydrolases, such as those derived from commercially available HALOTAG and / or mutant hydrolases (e.g., modified dehalogenases) disclosed in U.S. Published Application No. 2006 / 0024808, the disclosure of which is incorporated herein by reference. These mutant hydrolases (e.g., modified dehalogenases) are not enzymes (no substrate conversion), but depend on proper protein structure for stable binding of substrate to them. The result of reassociating split cp fragments of mutant hydrolases (e.g., modified dehalogenases) differs from that of traditional split enzyme systems, because the labeling function of the mutant hydrolase (e.g., modified dehalogenase) is retained in one of the fragments even after separation from its partner, whereas the split enzymes are only active while they are together. In effect, the labeling reaction of the split cp mutant hydrolase (e.g., modified dehalogenase) provides a molecular memory of protein interactions.

[0081] As an example of a mutant hydrolase, a mutant dehalogenase (or an intact cp-modified dehalogenase) provides efficient labeling in living cells or their lysates. This labeling is conditional only on the presence or expression of the protein and the presence of a labeled hydrolase substrate. In contrast, labeling of split, modified dehalogenases (e.g., split / cpHT) is dependent on specific protein interactions occurring within the cell and the presence of a labeled hydrolase substrate. For example, a beta-arrestin can be fused to one fragment of a mutant hydrolase (e.g., a modified dehalogenase) and a G-coupling receptor can be fused to the other fragment. Upon stimulation of the receptor in the presence of a labeled substrate, the beta-arrestin binds to the receptor, triggering a labeling reaction of either the receptor or the beta-arrestin (depending on which portion of the mutant hydrolase contains the reactive nucleophilic amino acid).

[0082] In some embodiments, provided herein is a split cp hydrolase (e.g., modified dehalogenase) system comprising a first fragment of a cp hydrolase (e.g., modified dehalogenase) fused to a protein of interest and optionally a second fragment of the cp hydrolase fused to a ligand of the first protein of interest. At least one of the hydrolase fragments has a substitution that, when present in a full-length mutant hydrolase having the sequence of the two fragments, forms a bond with the hydrolase substrate that is more stable than the bond formed between the corresponding full-length wild-type hydrolase and the hydrolase substrate. In one embodiment, each fragment of the cp hydrolase is fused to a protein of interest, and the proteins of interest interact, e.g., bind to each other. In another embodiment, one hydrolase fragment is fused to a protein of interest that interacts with a molecule in a sample. In another embodiment, a complex is formed by binding of the first hydrolase fragment, a second protein fused to the second hydrolase fragment, or a fusion having the second hydrolase fragment and a protein of interest fused to a cellular molecule, in the presence of a drug (or one or more drugs of interest) or under certain conditions.

[0083] Thus, the two fragments of the cp hydrolase (e.g., modified dehalogenases) together provide mutant hydrolases that are structurally related (and contain significant sequence identity / similarity (e.g., greater than 70%) to the full-length hydrolase, but contain at least one amino acid substitution that results in covalent binding of a hydrolase substrate. The full-length mutant hydrolase lacks or has reduced catalytic activity compared to the corresponding full-length wild-type hydrolase, and specifically binds a substrate that can be specifically bound by the corresponding full-length wild-type hydrolase, but no product is formed or substantially less, e.g., 2-fold, 10-fold, 100-fold, or 1000-fold less product is formed from the interaction between the mutant hydrolase and the substrate under conditions that result in product formation from the reaction between the corresponding full-length wild-type hydrolase and the substrate. The lack or reduced amount of product formation by the mutant hydrolase is due to at least one substitution in the full-length mutant hydrolase that results in the mutant hydrolase forming a bond with the substrate that is more stable than the bond formed between the corresponding full-length wild-type hydrolase and the substrate.

[0084] Because reversible protein complementation systems and biosensors have proven to be particularly useful tools for measuring functional dynamics using cellular imaging, such as protein interactions or changes in metabolite concentrations, during development of the embodiments herein, experiments were performed to identify regions within the HALOTAG sequence where strategies could be designed to allow for dynamic control of autolabeling activity.

[0085] In some embodiments, sp / cp dehalogenases (e.g., sp / cpHT) are able to refold after heat denaturation, relying on proteolytic cleavage of a flexible linker. If the linker is left intact, these cp dehalogenases denature and aggregate under a pulse of high heat in a manner similar to native HALOTAG. However, once the linker is cleaved, these sp / cp dehalogenases regain their enzymatic activity by refolding (e.g., reduced amounts). The protease-dependent behavior of these sp / cp dehalogenases makes them an output for screening protease mutants. For example, in a microfluidic protease library screen, using a thermostable oil phase, active proteases cleave the linker on the co-encapsulated cp dehalogenase, and subsequent heating, cooling, and labeling allows for fluorescent selection for the refolded sp / cp dehalogenase. Additionally, enzymes that can withstand rapid and repeated temperature cycling can make useful functional additives for polymerase chain reaction (PCR) applications.

[0086] The sp / cp dehalogenase complementation system offers several technical advantages over intact dehalogenases (including intact cp versions). Covalent labeling of intact dehalogenases with chloroalkane ligands can allow direct readout of protein location and concentration, whereas split dehalogenases (e.g., split / cpHT) direct such labeling to sites of protein-protein interactions. Many important cellular functions, including signal transduction, transcription, translation, and cargo transport, require specific interactions between proteins, membranes, organelles, and subcellular structures. The sp / cp dehalogenase system reports on the location, timing, and frequency of these events, whereas intact dehalogenases can only report on the presence of molecules.

[0087] Like other bimolecular complementation systems, sp / cp dehalogenases are usually inactive until both fragments assemble into an active complex with the help of an interaction partner protein fused to each fragment. Bimolecular fluorescence complementation (BiFC) of green fluorescent protein (GFP) and other fluorescent proteins (FPs) has been used by researchers for many years, but these BiFC systems have several significant drawbacks. Fluorophores take time to mature, proteins tend to assemble irreversibly, and performance is reduced in hypoxic conditions. In contrast, some sp / cp dehalogenases use exogenously supplied cell-permeable fluorescent ligands that assemble reversibly and do not require maturation or oxygen. Chloroalkane ligands feature bright and stable fluorophores that outperform protein-based fluorophores in terms of quantum yield and image resolution, making them ideal for state-of-the-art super-resolution microscopy.

[0088] In contrast to other enzyme complementation-based reporter systems, such as split luciferase, sp / cp dehalogenases form permanent covalent bonds with the substrate, creating durable event marks that can be observed over time. Furthermore, multiple complementation events can result in signal accumulation that does not decrease as the substrate is depleted. This is in contrast to split luciferases, where the signal decreases over time.

[0089] The utility of sp / cp dehalogenases extends beyond fluorescence imaging. Dehalogenases can accept a wide variety of ligands, provided the ligands have a haloalkane functional group. The cargo of the ligands can include, but is not limited to, fluorophores, chromophores, analyte sensing complexes, affinity tags (such as biotin), signals for protein degradation, nucleic acids, or solid supports. Thus, sp / cp dehalogenases can use cellular events as initiating signals for color formation, sensor activation, affinity tagging, protein degradation, DNA / RNA barcoding, crosslinking, or assembly onto supports or molecular scaffolds. The ultimate functional output of the split / cp dehalogenase is determined by the choice of ligand supplied by the user.

[0090] When used in conjunction with a fluorogenic substrate, the bound fluorogenic substrate is retained on one of the fragments upon fragment dissociation but may not be detectable after complex dissociation (because fluorogen-activating contacts with the protein may be disrupted / nonexistent). Thus, the combination of sp / cpHT with a fluorogenic ligand results in a unique situation of labeling but with dynamic (on / off) fluorescence detection of the retained label.

[0091] In some embodiments, an sp / cp dehalogenase (e.g., sp / cpHT) comprises two peptide and / or polypeptide components that together comprise at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:1. For example, a first peptide / polypeptide component of the sp / cp polypeptide corresponds to a first portion of SEQ ID NO:1 (e.g., at least 70% sequence identity to the first portion) and a second peptide / polypeptide component of the sp / cp polypeptide corresponds to a second portion of SEQ ID NO:1 (e.g., at least 70% sequence identity to the second portion). In some embodiments, an sp / cp dehalogenase (e.g., sp / cpHT) comprises two fragments that together comprise 100% sequence identity to all or a portion of SEQ ID NO: 1. For example, a first fragment of the sp / cp polypeptide has 100% sequence identity to a first portion of SEQ ID NO: 1, and a second fragment of the sp polypeptide has 100% sequence identity to a second portion of SEQ ID NO: 1.

[0092] In some embodiments, an sp / cp dehalogenase (e.g., sp / cpHT) comprises two peptide and / or polypeptide components that together comprise at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 1. For example, a first peptide / polypeptide component of an sp / cp polypeptide corresponds to a first portion of SEQ ID NO: 1 (e.g., at least 70% sequence similarity with the first portion) and a first peptide / polypeptide component of an sp / cp polypeptide corresponds to a second portion of SEQ ID NO: 1 (e.g., at least 70% sequence similarity with the second portion). In some embodiments, an sp / cp dehalogenase (e.g., sp / cpHT) comprises two fragments that together comprise 100% sequence similarity to all or a portion of SEQ ID NO: 1. For example, a first fragment of the sp / cp polypeptide has 100% sequence similarity to a first portion of SEQ ID NO: 1, and a second fragment of the sp / cp polypeptide has 100% sequence similarity to a second portion of SEQ ID NO: 1.

[0093] In some embodiments, a cp dehalogenase (e.g., cpHT) comprises a cp site. A cp site is an internal position within a parent sequence that defines the N-terminus and C-terminus of a cp dehalogenase. For example, if a theoretical 100 amino acid polypeptide is circularly permuted with a cp site between residues 33 and 34 of the parent polypeptide (referred to herein as a cp site of 33), the N-terminus of the cp polypeptide corresponds to position 34 of the parent polypeptide, position 100 is the peptide bound to position 1 of the parent polypeptide, and the C-terminus of the cp polypeptide corresponds to position 33 of the parent polypeptide. In some embodiments herein, the cp site in SEQ ID NO:1 can occur anywhere from position 5 of SEQ ID NO:1 to position 290 of SEQ ID NO:1. The following are non-limiting examples of cpHT polypeptides having 100% sequence identity to SEQ ID NO:1 (the segment of cpHT occurring before the cp site of the parent sequence is underlined): cpHT(10) (SEQ ID NO: 7) [ka] cpHT(45) (SEQ ID NO: 43) [ka] cpHT(68) (SEQ ID NO:66) [ka] cpHT(88) (SEQ ID NO: 85) [ka] cpHT(122) (SEQ ID NO: 119) [ka] cpHT(146) (SEQ ID NO: 143) [ka] cpHT(167) (SEQ ID NO: 164) [ka] cpHT(183) (SEQ ID NO: 180) [ka] cpHT(202) (SEQ ID NO: 199) [ka] cpHT(225) (SEQ ID NO:222 [ka] cpHT(275) (SEQ ID NO: 272) [ka]

[0094] Based on the above, other cp sites corresponding to positions between positions 5 and 290 of SEQ ID NO:1 are readily envisioned and are within the scope of the present specification.

[0095] In the above exemplary cpHT, the two parts of the parent sequence are directly fused without a linker sequence. However, as described throughout, some embodiments herein utilize a linker (e.g., a cleavable linker) to fuse the two segments. Examples of cpHT with a cleavable linker sequence (bold) include, but are not limited to, the following (either in the cp site or sequence of the linker): cpHT(10)-TEV linker (SEQ ID NO: 866) [ka] cpHT(45)-TEV linker (SEQ ID NO: 867 [ka] cpHT(68)-TEV linker (SEQ ID NO: 868 [ka] cpHT(88)-TEV linker (SEQ ID NO: 869 [ka] cpHT(122)-TEV linker (SEQ ID NO: 870 [ka] cpHT(146)-TEV linker (SEQ ID NO: 871 [ka] cpHT(167)-TEV linker (SEQ ID NO: 872 [ka] cpHT(183)-TEV linker (SEQ ID NO: 873 [ka] cpHT(202)-TEV linker (SEQ ID NO: 874 [ka] cpHT(225)-TEV linker (SEQ ID NO: 875 [ka] cpHT(275)-TEV linker (SEQ ID NO: 876 [ka]

[0096] In some embodiments, sp / cpHT is provided, in which the cleavable linker of the cpHT has been cleaved by enzymatic, chemical, or light-induced cleavage. Examples of cleaved sp / cpHT having (either in the cp site or sequence of the linker) include, but are not limited to, the following: sp / cpHT(10)-TEV linker [ka] [ka] sp / cpHT(45)-TEV linker [ka] [ka] sp / cpHT(68)-TEV linker [ka] [ka] sp / cpHT(88)-TEV linker [ka] [ka] sp / cpHT(122)-TEV linker [ka] [ka] sp / cpHT(146)-TEV linker [ka] [ka] sp / cpHT(167)-TEV linker [ka] [ka] sp / cpHT(183)-TEV linker [ka] [ka] sp / cpHT(202)-TEV linker [ka] [ka] sp / cpHT(225)-TEV linker [ka] [ka] sp / cpHT(275)-TEV linker [ka]

change

[0097] In some embodiments, the cpHT comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 of SEQ ID NO:1. , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 7, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148 , 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 2 11, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242 2, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273,cp sites corresponding to positions 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290 are provided.

[0098] In some embodiments, cpHT is provided with cp sites corresponding to positions between positions 5 and 13, 36 and 51, 63 and 72, 84 and 92, 104 and 130, 142 and 148, 160 and 174, 186 and 189, 201 and 203, 221 and 229, or 269 and 290 of SEQ ID NO:1.

[0099] In some embodiments, the cp polypeptide of the sp / cp fragment pair lacks one or more portions of the parent sequence. In some embodiments, the deleted portion is 1-50 amino acids in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or ranges therebetween). In some embodiments, the deleted portion is the N-terminal or C-terminal portion of the parent sequence. In some embodiments, the deleted portion is N-terminal to the cp site, C-terminal to the cp site, or overlaps with the cp site in the parent sequence. In such embodiments, the cp polypeptide comprises a first portion and a second portion, each of which comprises sequence identity to a portion of the parent sequence, but the first portion and the second portion of the cp polypeptide do not together comprise the entire sequence of the parent sequence. In some embodiments, a portion of the cpHT fragment or sp / cpHT fragment corresponds to a parent sequence having 70% to 100% sequence identity to SEQ ID NO:1 (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity). In some embodiments, the cpHT fragment or the sp / cpHT fragment portion corresponds to a parent sequence having 70% to 100% sequence similarity to SEQ ID NO:1 (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity).

[0100] In some embodiments, the first portion of the cpHT or fragment of the sp / cpHT complement pair is selected from positions 1-5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73 of SEQ ID NO:1. , 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111 1, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142 , 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 2 05, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 23 6, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267,Corresponding to positions 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, or 292.

[0101] In some embodiments, the second portion of the cpHT or fragment of the sp / cpHT complement pair is selected from the group consisting of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 21, 2 5, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112 , 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 1 75, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 20 6, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237 , 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268,Corresponding to positions 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, or 292 to 297 of SEQ ID NO: 1.

[0102] In some embodiments, the cp polypeptides of a pair of sp / cp fragments comprise portions of the parent sequence that overlap in each portion of the cpHT or fragment of the sp / cpHT. In some embodiments, the overlapping portions are 1-50 amino acids in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any range therebetween). In some embodiments, the overlapping portions are C-terminal to the cp site, N-terminal to the cp site, or overlap with the cp site. In some embodiments, the overlapping portions of the parent sequence are present in both the cp site or the sp / cp fragment.

[0103] The exemplary cpHT and sp / cpHT peptides and polypeptides provided above comprise 100% sequence identity to portions of SEQ ID NO: 1, and there are no portions of the peptides and polypeptides that do not align with 100% sequence identity to SEQ ID NO: 1. However, as described herein, cpHT and sp / cpHT peptides and polypeptides can have less than 100% sequence identity with SEQ ID NO: 1 (e.g., greater than 70%, greater than 75%, greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99% but less than 100% sequence identity).

[0104] In some embodiments, the circularly permuted hydrolases (eg, cpHT) and fragments thereof have enhanced thermostability compared to the parent hydrolase sequence (eg, HALOTAG).

[0105] In some embodiments, sp / cpHT or cpHT can be denatured, renatured, and capable of reconstituting its activity, hi some embodiments, such sp / cpHT and cpHT are used in methods that include exposing samples containing cpHT and sp / cpHT to denaturing conditions (e.g., manufacturing conditions, storage conditions, etc.) prior to substrate binding.

[0106] In some embodiments, provided herein are fusions of a circularly permuted hydrolase (e.g., a dehalogenase, e.g., HALOTAG, etc.) with a protein of interest, an interaction element, a localization element, a heterologous sequence, a peptide tag, a luciferase, or a bioluminescent complex, etc.

[0107] In certain embodiments, the circularly permuted hydrolase (e.g., cpHT) is fused to a heterologous sequence (e.g., a protein of interest). In some embodiments, the cp hydrolase allows for attachment of the heterologous sequence to a functional group bound to a substrate of the hydrolase (e.g., cpHT) or to a solid surface.

[0108] In certain embodiments, both portions of the cp hydrolase (e.g., cpHT) are fused to a heterologous sequence. In some embodiments, the heterologous sequences are substantially identical and, optionally, specifically bind to each other in the absence of one or more exogenous agents, e.g., form a dimer. In another embodiment, the heterologous sequences are different and, optionally, specifically bind to each other in the absence of one or more exogenous agents. In one embodiment, one hydrolase fragment is fused to a heterologous sequence, which interacts with a cellular molecule. In another embodiment, each hydrolase fragment is fused to a heterologous sequence, which interacts in the presence of one or more exogenous agents or under certain conditions. For example, in the presence of rapamycin, a fragment of a hydrolase fused to a rapamycin binding protein (FRB) and another fragment fused to an FK506 binding protein (FKBP) results in a complex of the two fusion proteins. In one embodiment, in the presence of an exogenous agent(s) or under different conditions, the fusion protein complex is not formed. In one embodiment, one heterologous sequence contains a domain, e.g., three or more amino acid residues, which may optionally be covalently modified, e.g., phosphorylated, and which non-covalently interacts with a domain in the other heterologous sequence. Two fragments of a hydrolase, at least one of which is fused to a protein of interest, may be used to detect reversible interactions, e.g., binding of two or more molecules, or other conformational changes or changes in conditions, such as pH, temperature or solvent hydrophobicity, or irreversible interactions.

[0109] Heterologous sequences useful in the present invention include, but are not limited to, those that interact in vitro and / or in vivo. For example, the fusion protein may contain a cp hydrolase or a fragment of a hydrolase and an enzyme of interest, such as a luciferase, RNasin or RNase, and / or a channel protein, a receptor, a membrane protein, a cytoplasmic protein, a nuclear protein, a structural protein, a phosphoprotein, a kinase, a signal protein, a metabolic protein, a mitochondrial protein, a receptor-associated protein, a fluorescent protein, an enzyme substrate, a transcription factor, a transporter protein, and / or a targeting sequence, such as a myristoylation sequence, a mitochondrial localization sequence, or a nuclear localization sequence, to direct the hydrolase fragment, e.g., the fusion protein, to a specific location. The protein of interest fused to the cp hydrolase or hydrolase fragment may be a fragment of a wild-type protein, e.g., a functional or structural domain of a protein, such as a domain of a kinase, a transcription factor, etc. The protein of interest may be fused to the N-terminus or C-terminus of the hydrolase fragment or cp hydrolase. In one embodiment, the fusion protein comprises a protein of interest at the N-terminus and another protein, e.g., a different protein, at the C-terminus of the hydrolase fragment or cp hydrolase. For example, the protein of interest can be an antibody. Optionally, the proteins in the fusion are separated by a linker, e.g., a linker sequence of 1-100 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 acid residues). In some embodiments, the presence of a linker in the fusion protein of the invention does not substantially alter the function of any of the proteins in the fusion compared to the function of each individual protein. For any particular combination of proteins in the fusion, a wide variety of linkers can be used. In one embodiment, the linker is a sequence recognized by an enzyme, e.g., a cleavable sequence or a photocleavable sequence.

[0110] Exemplary heterologous sequences include, but are not limited to, sequences in FRB and FKBP, the regulatory subunit of protein kinase (PKa-R) and the catalytic subunit of protein kinase (PKa-C), src homology regions (SH2) and phosphorylatable sequences, e.g., tyrosine-containing sequences, isoforms of 14-3-3, e.g., 14-3-3t (see Mills et al., 2000), and phosphorylatable sequences, proteins with WW regions (sequences of proteins that bind proline-rich molecules (see Ilsley et al., 2002, and Einbond et al., 1996), and phosphorylatable heterologous sequences, e.g., serine and / or threonine-containing sequences, and sequences in dihydrogenfolate reductase (DHFR) and gyrase B (GyrB).

[0111] As described throughout, the cpHT and sp / cpHT peptides and polypeptides provided herein are used as part of fusion proteins with peptides, polypeptides, antibodies, antibody fragments, and proteins of interest. For example, the invention provides a fusion protein comprising (1) a cpHT or sp / cpHT peptide or polypeptide, and (2) the amino acid sequence of a protein or peptide of interest, e.g., a marker protein, e.g., a selectable marker protein, an enzyme of interest, e.g., luciferase, RNasin, RNase, and / or the sequence of GFP, a nucleic acid binding protein, an extracellular matrix protein, a secreted protein, an antibody or portion thereof, e.g., Fc, a bioluminescent protein, a receptor ligand, a regulatory protein, a serum protein, an immunogenic protein, a fluorescent protein, a protein with a reactive cysteine, a receptor protein, e.g., the NMDA receptor, a channel protein, e.g., the HERG channel protein, Fusion proteins are provided that include an ion channel protein, such as a sodium, potassium, or calcium sensitive channel protein, a membrane protein, a cytoplasmic protein, a nuclear protein, a structural protein, a phosphoprotein, a kinase, a signaling protein, a metabolic protein, a mitochondrial protein, a receptor-associated protein, a fluorescent protein, an enzyme substrate, e.g., a protease substrate, a transcription factor, a protein destabilization sequence, or a transporter protein, e.g., the EAAT1-4 glutamate transporters, as well as a targeting signal that directs the fusion to a particular location, e.g., a plastid targeting signal, such as a mitochondrial localization sequence, a nuclear localization signal, or a myristoylation sequence.

[0112] In some embodiments, the fusion protein comprises (1) a cpHT or sp / cpHT peptide or polypeptide, and (2) a protein that associates with a membrane or a portion thereof, e.g., a targeting protein such as for endoplasmic reticulum targeting, a cell membrane-associated protein, e.g., an integrin protein or domain thereof, e.g., the cytoplasmic, transmembrane and / or extracellular stalk domain of an integrin protein, and / or a protein that links the mutant hydrolase to the cell surface, e.g., a glycosylphosphoinositol signal sequence.

[0113] Fusion partners can include those with enzymatic activity. For example, functional protein sequences can encode a kinase catalytic domain (Hanks and Hunter, 1995), producing a fusion protein capable of enzymatically adding a phosphate moiety to a specific amino acid, or can encode a Src homology 2 (SH2) domain (Sadowski et al., 1986; Mayer and Baltimore, 1993), producing a fusion protein that specifically binds phosphorylated tyrosine.

[0114] In some embodiments, the fusion includes an affinity domain, which includes a peptide sequence that can interact with a binding partner, such as one immobilized on a solid support, useful for identification or purification. A DNA sequence encoding multiple consecutive single amino acids, such as histidine, when fused to an expressed protein, can be used for one-step purification of recombinant proteins by binding with high affinity to a resin column, such as nickel sepharose. Exemplary affinity domains include HisV5 (HHHHH) (SEQ ID NO: 900), HisX6 (HHHHHH) (SEQ ID NO: 901), C-myc (EQKLISEEDL) (SEQ ID NO: 902), Flag (DYKDDDDK) (SEQ ID NO: 903), SteptTag (WSHPQFEK) (SEQ ID NO: 904), hemagglutinin, e.g., HA tag (YPYDVPDYA) (SEQ ID NO: 905), GST, thioredoxin, cellulose binding domain, RYIRS (SEQ ID NO: 906), Phe-His-His-Thr (SEQ ID NO: 907), chitin binding domain, S-peptide, T7 peptide, SH2 domain, C-end RNA tag, WEAAAREACCRECCARA (SEQ ID NO: 908), a metal binding domain, e.g., a zinc binding domain, or a calcium binding domain, such as from a calcium binding protein, e.g., calmodulin, troponin C, calcineurin B, myosin light chain, recoverin, S-modulin, visinin, VILIP, neurocalcin, hippocalcin, fryquenin, caltractin, calpain large subunit, S100 protein, parvalbumin, calbindin, D 9K , Calbindin D 28K and calretinin, intein, biotin, streptavidin, MyoD, Id, leucine zipper sequences, maltose binding protein, and SPYTAG peptides or SPYCATCHER proteins (e.g., SYYHHHHHHDYDIPTTENLYFQGAMVTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEDGRELAGATMELRDSSGKTISTWISDGHVKDFYLYPGKYTFVETAAPDGYEVATPIEFTVNEDGQVTVDGEATEGDAHTGSSGS (SEQ ID NO: 909), SYYHHHHHHDYDIPTTENLYFQGAMVTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEDGRELAGATMELRDCSGKTISTWISDGHVKDFYLY PGKYTFVETAAPDGYEVATPIEFTVNEDGQVTVDGEATEGDAHTGSSGS (SEQ ID NO: 910), GSSHHHHHSSGLVPRGSRGVPHIVMVDAYKRYKGSGESGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSS (SEQ ID NO: 911).

[0115] In some embodiments, the circularly permuted polypeptide or sp / cp fragment described herein (e.g., cpHT or sp / cpHT) is fused to a reporter protein. In some embodiments, the reporter is a bioluminescent reporter (e.g., expressed as a fusion protein with sp / cpHT or cpHT). In certain embodiments, the bioluminescent reporter is a luciferase. In some embodiments, the luciferase is selected from those found in Omphalotus olearius, fireflies (e.g., Photinini), Renilla reniformis, Aequoria, mutants thereof, portions thereof, variants thereof, and any other luciferase enzyme suitable for the systems and methods described herein. In some embodiments, the bioluminescent reporter is a modified, enhanced luciferase enzyme from Oplophorus (e.g., Promega Corporation, NANOLUC enzyme of SEQ ID NO: 3, or a sequence having at least 70% identity thereto (e.g., greater than 70%, greater than 80%, greater than 90%, greater than 95%)). Exemplary bioluminescent reporters are described, for example, in US Patent Application Nos. 2010 / 0281552 and 2012 / 0174242, both of which are incorporated by reference in their entireties.

[0116] In some embodiments, the circularly permuted polypeptide or split fragment thereof (e.g., cpHT or sp / cpHT) is fused to a peptide or polypeptide component of commercially available NanoLuc®-based technologies (e.g., NanoLuc® luciferase, NanoBiT, NanoTrip, etc.). PCT Application No. PCT / US2010 / 033449, U.S. Patent No. 8,557,970, PCT Application No. PCT / 2011 / 059018, and U.S. Patent No. 8,669,103 (each of which is incorporated herein by reference in their entirety for all purposes) describe compositions and methods that include bioluminescent polypeptides used as heterologous sequences in fusions herein. Such polypeptides are used in embodiments herein and may be used in conjunction with the compositions and methods described herein. PCT Application No. PCT / US14 / 26354 and U.S. Patent No. 9,797,889 (each of which is incorporated herein by reference in their entirety for all purposes) describe compositions and methods for the assembly of bioluminescent complexes, and such complexes, as well as peptide and polypeptide components thereof, are used as heterologous sequences in embodiments herein and may be used in conjunction with the compositions and methods described herein. In some embodiments, NanoBiT and other related technologies utilize peptide and polypeptide components in assembly into complexes that are significantly enhanced (e.g., 2-fold, 5-fold, 10 ... 2 Double, 10 3 Double, 10 4In some embodiments, the NanoBiT® peptides and polypeptides are fused to the cpHT and / or sp / cpHT fragments herein. U.S. Patent Publication No. 2020 / 0270586 and International Application No. PCT / US19 / 36844 (incorporated herein by reference in their entireties for all purposes) describe multipartite luciferase complexes (e.g., NanoTrip) that may be used as heterologous sequences in embodiments herein and in conjunction with the compositions and methods described herein.

[0117] As described herein, the cpHT and sp / cpHT systems herein utilize a haloalkane substrate. In some embodiments, the substrate is of formula (I): R-linker-AX, where R is a solid surface, one or more functional groups, or is absent, and the linker is a multi-atom straight or branched chain containing C, N, S, or O, or a group containing one or more rings, e.g., saturated or unsaturated rings, such as one or more aryl rings, heteroaryl rings, or any combination thereof, and where AX is a substrate for a dehalogenase, hydrolase, HALOTAG, cpHT, or sp / cpHT system herein (e.g., A is (CH2) 4~20 and X is a halide (e.g., Cl or Br). Suitable substrates are described, for example, in U.S. Pat. Nos. 11,072,812, 11,028,424, 10,618,907, and 10,101,332, which are incorporated by reference in their entireties.

[0118] In some embodiments, R is one or more functional groups, such as a fluorophore, biotin, a luminophore, or a fluorescent or luminescent molecule. Exemplary functional groups for use in the present invention include, but are not limited to, amino acids, proteins, such as enzymes, antibodies or other immunogenic proteins, radionuclides, nucleic acid molecules, drugs, lipids, biotin, avidin, streptavidin, magnetic beads, solid supports, electron opaque molecules, chromophores, MRI contrast agents, dyes, such as xanthene dyes, calcium sensitive dyes, such as 1-[2-amino-5-(2,7-dichloro-6-hydroxy-3-oxy-9-xanthenyl)-phenoxy]-2 -(2'-amino-5'-methylphenoxy)ethane-N,N,N',N'-tetraacetic acid (Fluo-3), sodium sensitive dyes such as 1,3-benzenedicarboxylic acid, 4,4'-[1,4,10,13-tetraoxa-7,16-diazacyclooctadecane-7,16-diylbis(5-methoxy-6,2-benzofurandiyl)bis (PBFI), NO sensitive dyes such as 4-amino-5-methylamino-2',7'-difluorescein, and other fluorophores. In one embodiment, the functional group is an immunogenic molecule, i.e., one that is bound by an antibody specific for that molecule.In some embodiments, the functional group is an E3 ubiquitin ligase ligand or targeting chimera (TAC) system, such as the phosphorylation-targeting chimera (PhosTAC; Chen et al. ACS Chem. Biol. 3121, 16, 12, 2808-2815; incorporated by reference in its entirety) system, the deubiquitinase-targeting chimera (DUBTAC; Henning et al. Deubiquitinase-Targeting Chimeras for Targeted Protein Stabilization. bioRxiv; 2021. DOI: 10.1101 / 2021.04.30.441959; incorporated by reference in its entirety) system, the lysosome-targeting chimera (LyTAC; Banik et al. Nature 584, 291-297 (2020); incorporated by reference in its entirety) system, the autophagy-targeting chimera (AUTAC; Takahashi et al. Mol. Cell. 2019 Dec 5;76(5):797-810.e10; incorporated by reference in its entirety) system, the Autophagy Linking Compound (ATTEC; Fu et al. Cell Research volume 31, pages 965-979 (2021); incorporated by reference in its entirety) system, and other functional groups used in mobilizing components of oligo-based TACs.

[0119] In some embodiments, a substrate of the invention is permeable to the plasma membrane of a cell.

[0120] In some embodiments, the substrates herein include a cleavable linker, such as those described in US Pat. No. 10,618,907, which is incorporated by reference in its entirety.

[0121] In some embodiments, the substrate comprises a fluorescent functional group (R). Suitable fluorescent functional groups include xanthene derivatives (e.g., fluorescein, rhodamine, Oregon Green, eosin, Texas Red, etc.), cyanine derivatives (e.g., cyanine, indocarbocyanine, oxacarbocyanine, thiacarbocyanine, merocyanine, etc.), naphthalene derivatives (e.g., dansyl and prodan derivatives), oxadiazole derivatives (e.g., pyridyloxazole, nitrobenzoxadiazole, benzoxadiazole, etc.), pyrene derivatives (e.g., cascade blue), oxazine derivatives (e.g., Nile red, Nile blue, cresyl violet, oxazine 170, etc.), acridine derivatives (e.g., proflavine, acridine orange, acridine yellow, etc.), arylmethine derivatives (e.g., auramine, crystal violet, malachite green, etc.), tetrapyrrole derivatives (e.g., porphine, phthalocyanine, bilirubin, etc.), CF dyes (Biotium), BODIPY (Invitrogen), ALEXA Examples of dyes that may be used include, but are not limited to, FLUOR (Invitrogen), DYLIGHT FLUOR (Thermo Scientific, Pierce), ATTO and TRACY (Sigma Aldrich), FluoProbes (Interchim), DY and MEGASTOKES (Dyomics), SULFO CY dyes (CYANDYE, LLC), SETAU and SQUARE dyes (SETA BioMedicals), QUASAR and CAL FLUOR dyes (Biosearch Technologies), SURELIGHT dyes (APC, RPE, PerCP, phycobilisomes) (Columbia Biosciences), APC, APCXL, RPE, BPE (Phyco-Biotech), autofluorescent proteins (e.g., YFP, RFP, mCherry, mKate), quantum dot nanocrystals, and the like.

[0122] In some embodiments, the substrate comprises a fluorogenic functional group (R). The fluorogenic functional group is a functional group that generates and enhances a fluorescent signal upon binding of the substrate to a target (e.g., binding of a haloalkane to a modified dehalogenase). By generating significantly increased fluorescence (e.g., 10-fold, 20-fold, 50-fold, 100-fold, 200-fold, 500-fold, 100-fold, or more) upon target engagement, problems with background signal are mitigated. Exemplary fluorogenic dyes for use in embodiments herein include the JANELIA FLUOR family of fluorophores, such as: [ka] [ka] (See, e.g., U.S. Pat. Nos. 9,933,417, 10,018,624, 10,161,932, and 10,495,632, each of which is incorporated by reference in its entirety.) In some embodiments, exemplary conjugates of JANELIA FLUOR 549 and JANELIA FLUOR 646 with haloalkane substrates of modified dehalogenases (e.g., HALOTAG) are commercially available (Promega Corp.). The use and design of fluorogenic functional groups, dyes, probes, and substrates are described, for example, in Grimm et al. Nat Methods. 2017 Oct;14(10):987-994., Wang et al. Nat Chem. 2020 Feb;12(2):165-172, each of which is incorporated by reference in its entirety.

[0123] In some embodiments, provided herein are isolated nucleic acid molecules (polynucleotides) comprising a nucleic acid sequence encoding a circularly permuted hydrolase (e.g., cpHT) as described herein. Also provided are isolated nucleic acid molecules comprising a nucleic acid sequence encoding a fusion protein comprising a cp hydrolase fragment (e.g., cpHT, etc.) and one or more amino acid residues at the N-terminus (N-terminal fusion partner) and / or C-terminus (C-terminal fusion partner). In one embodiment, the fusion protein comprises at least two different fusion partners (e.g., as described herein), one at the N-terminus and one at the C-terminus, where one of the fusions can be a sequence used for purification, e.g., glutathione S-transferase (GST) or polyHis sequence, a sequence intended to alter the properties of the remainder of the fusion protein, e.g., a protein destabilizing sequence, or a sequence with distinguishable properties. In one embodiment, the isolated nucleic acid molecule comprises a nucleic acid sequence that is optimized for expression in at least one selected host. Optimized sequences include codon-optimized sequences, i.e., codons that are more frequently used in one organism relative to another, e.g., distantly related organisms, as well as modifications to add or modify Kozak sequences and / or introns, and / or modifications to remove undesirable sequences, e.g., potential transcription factor binding sites. In one embodiment, the polynucleotide comprises a nucleic acid sequence encoding a fragment of a dehalogenase, which nucleic acid sequence is optimized for expression in a selected host cell. In one embodiment, the optimized polynucleotide no longer hybridizes to a corresponding non-optimized sequence, e.g., does not hybridize to a non-optimized sequence under moderate or high stringency conditions. In another embodiment, the polynucleotide encodes a polypeptide having less than 90%, e.g., less than 80%, nucleic acid sequence identity to a corresponding non-optimized sequence, and optionally at least 80%, e.g., at least 85%, 90% or more amino acid sequence identity to a polypeptide encoded by the non-optimized sequence.

[0124] Constructs, e.g., vectors comprising the expression cassettes and isolated nucleic acid molecules, as well as host cells having one or more of the constructs, and kits comprising the isolated nucleic acid molecule(s), or one or more of the constructs or vectors, are also provided. Host cells include prokaryotic or eukaryotic cells, e.g., plant cells or vertebrate cells, e.g., mammalian cells, including, but not limited to, human, non-human primate, canine, feline, bovine, equine, ovine, sheep, or rodent (e.g., rabbit, rat, ferret, or mouse) cells. In some embodiments, the expression cassette comprises a promoter, e.g., a constitutive promoter or a regulatable promoter, operably linked to the nucleic acid molecule. In some embodiments, the expression cassette comprises an inducible promoter. In certain embodiments, the invention comprises a vector comprising a nucleic acid sequence encoding a fusion protein comprising a fragment of a dehalogenase. In some embodiments, an optimized nucleic acid sequence, e.g., a human codon-optimized sequence, encoding at least a fragment of a hydrolase, preferably a fusion protein comprising a fragment of a hydrolase, is used in the nucleic acid molecules of the invention. Optimization of nucleic acid sequences is known in the art, see, for example, WO02 / 16944, which is incorporated by reference in its entirety.

[0125] Also provided herein are cells comprising a circularly permuted hydrolase (e.g., cpHT), a split / circularly permuted hydrolase fragment(s) (e.g., sp / cpHT), a polynucleotide, an expression vector, and the like. In some embodiments, the components described herein are expressed in a cell. In some embodiments, the components described herein are introduced into a cell, for example, via transfection, electroporation, infection, cell fusion, or any other means.

[0126] In some embodiments, the systems herein (e.g., including cp hydrolases (e.g., cpHT, sp / cpHT, etc.)) can be used to measure or detect various states and / or molecules of interest. For example, protein-protein interactions are essential to virtually all aspects of cell biology, from gene transcription, protein translation, signal transduction, and cell division and differentiation. Protein complementation assays (PCA) are one of several methods used to monitor protein-protein interactions. In PCA, protein-protein interactions bring the two non-functional halves of an enzyme into physical proximity with each other, which allows for refolding into a functional enzyme. Thus, the interaction is monitored by enzyme activity. In protein complementation labeling (PCL), the detection enzyme is mutated to capture the substrate, for example, via an acyl-mutated enzyme intermediate. Thus, a covalent bond is created between the substrate and the reconstituted mutant enzyme, allowing for cumulative labeling over time, thus increasing the sensitivity for detecting weak protein-protein interactions. In one embodiment, a vector encoding a cp-modified dehalogenase having a cleavable linker (e.g., cpHT) is expressed intracellularly as a fusion with at least one protein of interest or introduced into a cell, cell lysate, in vitro transcription / translation mixture, or supernatant, and a hydrolase substrate (e.g., a haloalkane) labeled with a functional group is added thereto. The functional group is then detected or determined, for example, relative to a control sample at one or more time points.

[0127] In some embodiments, provided herein is a method for detecting an interaction between two proteins in a sample. The method includes providing a sample having a cell, a lysate of a cell, or an in vitro transcription / translation reaction having a plurality of expression vectors of the invention, and a hydrolase substrate (e.g., a haloalkane) having at least one functional group under conditions effective to allow association of the first fusion protein and the second fusion protein. The presence, amount, or location of the at least one functional group in the sample is detected.

[0128] In some embodiments, the present invention provides a method for detecting a molecule of interest in a sample. The method includes providing a cell having a plurality of expression vectors of the present invention, a lysate thereof, an in vitro transcription / translation reaction having a plurality of expression vectors of the present invention, and a sample having a hydrolase substrate (e.g., a haloalkane) having at least one functional group under conditions effective to allow a first heterologous amino acid sequence to interact with the molecule of interest in the sample. The presence, amount, or location of the at least one functional group in the sample is detected, thereby detecting the presence, amount, or location of the molecule of interest.

[0129] Also provided herein is a method for detecting an agent that alters the interaction of two proteins, which comprises providing a cell comprising a plurality of expression vectors of the invention, a lysate thereof, or an in vitro transcription / translation reaction comprising a plurality of expression vectors of the invention, a hydrolase substrate (e.g., a haloalkane) having at least one functional group, and a sample having an agent under conditions effective to allow association of the first fusion protein and the second fusion protein. The agent is suspected of altering the interaction between the first heterologous amino acid sequence and the second heterologous amino acid sequence. The presence or amount of the at least one functional group in the sample is detected compared to a sample without the agent.

[0130] In another embodiment, the present invention provides a method for detecting an agent that alters the interaction of a molecule of interest with a protein. The method includes providing a cell, a lysate thereof, or an in vitro transcription / translation reaction with a plurality of expression vectors of the present invention, a hydrolase substrate (e.g., a haloalkane) having at least one functional group, and an agent suspected of altering the interaction between a heterologous amino acid sequence and a molecule of interest in a sample. The presence or amount of the functional group in the sample compared to a sample with the agent.

[0131] In some embodiments, provided herein is a method for detecting the presence of a molecule of interest. For example, a cell is contacted with a vector comprising a promoter, e.g., a regulatable promoter, and a nucleic acid sequence encoding two complementary fragments of a mutant hydrolase, at least one of which is fused to a protein that interacts with the molecule of interest. In one embodiment, the transfected cell is cultured under conditions in which the promoter induces transient expression of the fragment or regulated expression of one of the fragments, and activity associated with a labeled substrate is detected.

[0132] In some embodiments, the systems herein (e.g., including cp hydrolases (e.g., cpHT, sp / cpHT, etc.)) can be used as biosensors to detect the presence / amount of a molecule of interest or a particular condition (e.g., pH or temperature). Upon interacting with a molecule of interest or being subjected to a certain condition, the biosensor undergoes a conformational change or is chemically altered to cause a change in activity. In some embodiments, the cp hydrolases herein comprise an interaction domain of a molecule of interest. For example, biosensors can be generated to detect proteases (e.g., for detecting the presence of a particular viral protease, which is an indication of the presence of a virus), kinases (e.g., by inserting a kinase site into a reporter protein), RNAi (e.g., by inserting a sequence suspected to be recognized by RNAi into the coding sequence of a reporter protein and then monitoring reporter activity after addition of the RNAi), ligands, binding proteins such as antibodies, cyclic nucleotides such as cAMP or cGMP, or metals such as calcium by inserting suitable sensor regions into the cp hydrolases (e.g., cpHT, sp / cpHT, etc.). One or more sensor regions may be inserted at the C-terminus, N-terminus, and / or at one or more suitable positions within the cp hydrolase sequence, the sensor region comprising one or more amino acids. One or all of the inserted sensor regions may include linker amino acids for coupling the sensor to the remainder of the polypeptide. Examples of biosensors are disclosed in U.S. Patent Application Publication Nos. 2005 / 0153310 and 2009 / 0305280, and PCT Publication No. WO2007 / 120522A2, each of which is incorporated herein by reference.

[0133] experiment Example 1 Comprehensive screening of circularly permuted dehalogenase substitutes. Plasmids encoding all possible circularly permuted versions of HaloTag, as well as two linker control versions of the unpermuted HaloTag with linkers simply added to the N- or C-terminus, were constructed by PCR for a total of 298 gene constructs. The linkers connecting the native N- and C-termini were: [ka] (TEV protease recognition sequence underlined, cleavable peptide bond indicated by slash). Expression was performed in E. coli and cell lysates were prepared by adding chemical lysis reagents. Lysates were treated with TEV protease (or water as a negative control) and subjected to a battery of biochemical tests.

[0134] Lysates were assayed for protein solubility by centrifugation followed by conjugation with 10 μM CA-TMR ligand and gel electrophoresis. To determine the thermal stability of each cpHT, lysates were heated from 40 to 90 °C for 30 min and cooled to room temperature before being mixed with 10 nM CA-TMR and subjected to fluorescence polarization (FP) measurements. Enzyme activity was quantitatively measured by mixing lysates with 10 nM CA-AlexaFluor488 and monitoring their FP changes over a 30 min period.

[0135] This screen revealed that 228 / 296 (77%) cpHT variants reacted with CA-TMR, the majority of these were soluble, and 50 variants had at least 10% of the native HT activity on CA-AlexaFluor488 (Figure 2). 17 cpHT variants had increased thermostability compared to HT, and 38 variants showed activity recovery after thermal denaturation, presumably due to protein refolding. The most active variants by AlexaFluor488 kinetics clustered in a region distal to the lid domain, but this effect may be specific to this substrate, which is negatively charged and may be sensitive to lid domain perturbation. Indeed, the clustering effect was less pronounced when a neutral TMR ligand was used in the solubility and stability assays. With the exception of cpHT near residues 111 and 120, all refolded variants were localized to the lid domain, and all thermostabilized variants were also within the lid domain.

[0136] A real-time fluorescence polarization assay with HaloTag Alexa488 ligand was used to monitor the activity of cpHT variants in E. coli lysates (Figure 2D). The Alexa488 ligand reacts slowly enough with HaloTag to allow calculation of initial velocities and comparison of enzyme activity to full-length HaloTag. Activity is not normalized to concentration and is therefore, in this case, a qualitative measure of enzyme activity after circular permutation. Using a baseline relative activity level of 0.03 (red dotted line in Figure 2D), an amount that visually separates the signal above background during the real-time assay, a total of 118 / 297 cpHT variants were observed to retain measurable activity. With only a few exceptions, activity measurements in this assay did not change significantly after TEV cleavage, indicating that the constructs remained in an intact functional state after the linker between the fragments was cleaved. Using this comprehensive map of functional cpHT, approximately 10 common regions of sequence were identified that retained relatively high expression and / or activity after circular permutation.

[0137] Example 2 Testing of split dehalogenase variants After completing a screen of all 298 possible circular permutations of HaloTag (cpHT) (see Example 1), 22 split sites were selected for testing as split HaloTag fragment pairs (spHT). The spHT designs were selected based on the properties of their cpHT counterparts, including thermostability, expression, enzymatic activity, and changes in biophysical properties upon cleavage of the TEV protease recognition sequence in the linker connecting the native N- and C-termini. Particular interest was paid to variants (e.g., circular permutations within the sequence region proximal to residue 120) that showed the ability to reform or refold after thermal denaturation upon TEV protease cleavage of the cpHT form.

[0138] An initial set of spHT N- and C-terminal fragments (spHT 80, 97, and 121) were expressed in E. coli as fusions to several different domains, including maltose binding protein (MBP), a 6x polyhistidine tag (His-tag), the large and small components of the bimolecular NanoLuc system (LgBiT and SmBiT), and full-length NanoLuc variants. Moderate expression was noted for some of these fusions, but all suffered from low solubility. The low solubility was usually attributed to exposure of core hydrophobic residues buried in the intact HT structure that form an aggregation-prone surface on the spHT fragments. Estimates based on NanoLuc activity suggest that the solubility of these fragments is less than 5% in E. coli lysates.

[0139] Despite their low solubility, all 22 exemplary spHT designs were generated as fusions to the FRB and FKBP domains. FRB and FKBP undergo chemically induced high affinity heterodimerization in the presence of rapamycin. Thus, spHT fragments fused to these domains can be brought into close proximity with each other by the addition of rapamycin, providing an assay for functional reconstitution of HaloTag enzymatic activity. Each of the spHT fragments was fused to FRB or FKBP at either the N- or C-terminus (n=44), generating a total of 176 unique fusion proteins that were expressed in E. coli. Because the best orientation of FRB and FKBP relative to the spHT fragment domains cannot be predicted ab initio, all possible orientations and combinations were assayed (8 per spHT site). Fusion combinations were assayed using fluorogenic Janelia Fluor 646 (JF646) ligand in the presence of 50 nM rapamycin. JF646 is available through the regular Promega catalog and was chosen because it has low background fluorescence (allowing direct fluorescence measurements in 96-well plates) and provides a higher stringency test than non-fluorogenic ligands (such as TMR).

[0140] Six of the 22 spHT FKBP / FKBP designs showed a 2-fold or greater increase in fluorescence signal in the presence of rapamycin (spHT 80, 133, 145, 157, 180, and 195), with a maximum 4.7-fold induction for the [1-195]-FKBP+[196-297]-FRB combination (Figure 3). The corresponding cpHT 195 showed a T of approximately 7°C higher than HT. m spHT hits (at 80) were the most thermostable of all variants in the circular permutation screen. All but one spHT hit (spHT 80) were located within the lid subdomain of HT, including a region of sequence covering residues 133-216. Among the spHT hits, there were multiple orientations of FRB and FKBP that allowed reconstitution of activity. In general, fusion combinations in which FKBP was at the C-terminus of either fragment performed best.

[0141] In addition to the combination of the "blunt" spHT fragments (where all HT residues were present exactly once), several "gapped" and "overlapping" combinations were tested (where certain residues were missing from both fragments or present in both fragments, respectively). Residues missing or doubly represented in these combinations were restricted to the lid subdomain, specifically, helix 6, helix 7, helix 8, and / or helix 9. Gapped combinations failed to reconstitute detectable ligand-binding activity. However, overlapping combinations showed up to 3-fold reconstitution over background (Figure 4). These results indicate that (a) the lid helices are important for ligand binding in HT and (b) the lid subdomain tolerates sequence overlap and can participate in secondary structure swapping, a useful feature for designing biosensors and conformationally dynamic protein switches.

[0142] Example 3 Testing of circularly permuted dehalogenase variants To further understand the effects of perturbations in the lid subdomain, we retested a series of cpHT variants containing breaks in this region. Specifically, we assayed their ability to activate fluorescence with three fluorogenic ligands (JF525, JF585, and JF646) and compared these reactivities (as well as total TMR labeling) to non-cpHT. We found that cpHT 160-178 retained virtually zero fluorogenic activation ability, while other cpHT variants in the 138-180 region generally retained this ability, albeit to a lesser extent than non-cpHT (Figure 5). cpHT 160-178 is labeled by TMR chloroalkane ligands as efficiently or more efficiently than other cpHT variants in the lid region (Figure 6). Taken together, this evidence indicates that perturbation of helix 8, which encompasses most of the 160-178 region of HT sequence space, nearly eliminates the fluorogenic activation properties of HT without interfering with chloroalkane catalysis.

[0143] A similar analysis was applied to cpHT variants corresponding to all 22 of the spHT designs, assaying them on a small panel of fluorogenic substrates over three time points (Figure 7). Ligands accumulated signal at different rates, in the order JF646>JF585>JF635. Four cpHT variants stood out for not having significant fluorogenic activation capacity: cpHT44, 164, 272, and 274 (cpHT164 was previously noted for this property). All four of these cpHT variants are labeled by TMR, although cpHT44, 272, and 274 are labeled with low efficiency. The lack of fluorogenic activation observed in cpHT44, 272, and 274 may have a different cause than cpHT164, since the lid domain is predicted not to be disrupted by substitutions at these distal sequence sites.

[0144] Example 4 TEV cleavage of cpHT Treatment of cpHT variants with TEV protease provides an opportunity to assess function after the resulting fragments have had a chance to physically separate, providing insight into their function, for example, as sp / cpHT (Figure 8). The majority of variants in the cpHT library showed little or no response to TEV treatment and retained non-cleavage activity. However, some sites, for example, regions around positions 25, 88, 244, and 272, showed a marked reduction in activity as measured by fluorescence polarization with the TMR-HaloTag ligand. The reduced activity of these variants indicates that circular permutation at these sites results in fragments that allow spontaneous dissociation, making them candidates for designing low affinity biosensors that require enhanced complementation.

[0145] Example 5 Ligand specificity of cpHT It was observed that many cpHT variants have different ligand specificities or activities. Many examples occurred throughout the sequence of HaloTag, including circular permutations at 10, 27, 42, 68, 72, 124, 130, 145, 153, 162, 173, 181, 205, 219, 244, and 257. All of these constructs had detectable activity in gel-based detection assays with TMR ligands, but little or no measurable activity in fluorescence polarization kinetic assays with Alexa488 ligands. Figure 9 illustrates this with an example at positions 66-68. The opposite specificity was also observed, with positions such as 239 showing measurable activity with FP with Alexa488 ligands, but no activity with gel with TMR ligands.

[0146] Example 6 Changes in thermostability of cpHT During development of embodiments herein, experiments were performed to measure the activity of cpHT variants in E. coli lysates using fluorescence polarization after heat treatment and to determine the effect of circular permutation and TEV cleavage on stability (Figure 10). Many constructs were destabilized by TEV cleavage as indicated by a left shift in their melting profiles. Others showed increased stability after treatment at elevated temperatures, up to 90°C, due to refolding of these constructs into active enzymes after returning to ambient temperature. Finally, a combination of the two phenotypes was observed that showed not only destabilization of the melting profiles but also refolding activity at higher temperatures.

Claims

[Claim 1] A composition comprising a circular permutation variant of a polypeptide, comprising a first sequence and a second sequence, each having at least 70% sequence identity with the portion of sequence number 1.