Split-modified dehalogenase variant
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PROMEGA CORP
- Filing Date
- 2023-05-04
- Publication Date
- 2026-05-15
AI Technical Summary
Current tools lack the ability to dynamically measure important functional dynamics such as changes in protein interactions or metabolite concentrations using cell imaging, despite advancements in fluorescence detection with HALOTAG and its chloroalkane-based ligands.
The development of peptide and polypeptide sequences that structurally assemble into a modified dehalogenase structure capable of binding to a haloalkyl ligand, specifically through split dehalogenase variants that assemble via structural complementation into an active dehalogenase complex.
This approach enables dynamic control of self-labeling activity, allowing for the measurement of protein interactions and metabolite concentrations with enhanced sensitivity and specificity, thereby overcoming the limitations of existing technologies.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 338,323, filed May 4, 2022, which is incorporated herein by reference.
[0002] Provided herein are peptide and polypeptide sequences that structurally assemble to form an active modified dehalogenase structure capable of binding to a haloalkyl ligand. In particular, provided herein are split dehalogenase variants that assemble via structural complementation into an active dehalogenase complex, as well as systems and methods of using the same.
Background Art
[0003] The utility of self-labeling protein systems such as HALOTAG and its chloroalkane-based ligands has continued to expand over time as research tools. Genetic fusions to HALOTAG as a general strategy have enabled a wide range of applications including fluorescent labeling for cell biology and imaging, recombinant protein purification, biosensors and diagnostics, energy transfer technologies (BRET, FRET), and therapeutic target proteolysis (PROTAC). The development of new fluorophores and fluorogenic dyes (such as JANELIA FLUOR dyes) as chloroalkane conjugates serves as an example highlighting the renewed interest in HALOTAG for fluorescence detection in cell imaging applications. The advantages of such dyes over conventional tools such as widely used fluorescent proteins in terms of brightness, photostability, sensitivity, and far-red spectral detection are particularly prominent in difficult or sensitive imaging applications in endogenous biology. As chloroalkane conjugates, they can utilize the self-labeling activity of HALOTAG to measure protein abundance and localization in a target-specific manner via genetic fusion. However, there is a lack of available tools that can measure important functional dynamics by cell imaging such as changes in protein interactions or metabolite concentrations that can take advantage of these improvements in fluorescence detection. What is needed in this field are tools for dynamically controlling the self-labeling activity in systems such as HALOTAG.
Summary of the Invention
[0004] Provided herein are peptide and polypeptide sequences that structurally assemble to form a modified dehalogenase structure capable of binding to a haloalkyl ligand. In particular, provided herein are split dehalogenase variants that assemble via structural complementation into an active dehalogenase complex, as well as systems and methods of use thereof.
[0005] In some embodiments, provided herein is a composition comprising a split variant of a polypeptide having at least 70% sequence similarity (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) to SEQ ID NO:1. In some embodiments, the split variant comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to SEQ ID NO:1.
[0006] In some embodiments, the split variant is a binary system comprising a first fragment and a second fragment. In some embodiments, the split variant comprises (i) a first fragment of a polypeptide having at least 70% sequence similarity (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%) to a first portion of SEQ ID NO:1, and (ii) a second fragment of a polypeptide having at least 70% sequence similarity (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) to a second portion of SEQ ID NO:1. In some embodiments, the first fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to the first portion of SEQ ID NO:1. In some embodiments, the second fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to the second portion of SEQ ID NO:1. In some embodiments, the first fragment and the second fragment together comprise an amino acid sequence corresponding to at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, 100%) of the length of SEQ ID NO:1.
[0007] In some embodiments, each of the first fragment and the second fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity with one of SEQ ID NOs: 2-577. In some embodiments, each of the first fragment and the second fragment comprises 100% sequence similarity with one of SEQ ID NOs: 2-577. In some embodiments, each of the first fragment and the second fragment comprises 100% sequence identity with one of SEQ ID NOs: 2-577.
[0008] In some embodiments, the first fragment is SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514,One of 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, and 576, and includes at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity. The second fragment is SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413,One of 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, and 577, and includes at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity.
[0009] In some embodiments, the first fragment is SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514,Comprising a first reference sequence selected from one of 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, and 576 and having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity.
[0010] In some embodiments, the second fragment is SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515,It comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity with a second reference sequence selected from one of 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, and 577.
[0011] In some embodiments, the first fragment and the second fragment exhibit enhancement of one or more traits as compared to the first reference sequence and the second reference sequence, and the traits are selected from affinity for each other, expression, intracellular solubility, intracellular stability, and activity when combined.
[0012] In some embodiments, the split variant includes a split (“sp”) site at a position corresponding to any position between position 5 and position 290 (e.g., positions 19 to 34). In some embodiments, the split variant is between positions 5 and 13 of SEQ ID NO: 1 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, or a range therebetween), between positions 36 and 51 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, or a range therebetween), between positions 63 and 72 (e.g., 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, or a range therebetween), between positions 84 and 92 (e.g., 84, 85, 86, 87, 88, 89, 90, 91, 92, or a range therebetween), between positions 104 and 130 (e.g., 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, or a range therebetween), between 142 and 148 (e.g., 142, 143, 144, 145, 146, 147, 148, and a range therebetween), between positions 160 and 174 (e.g., 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, or a range therebetween), between positions 186 and 189 (e.g., 186, 187, 188, 189, or a range therebetween), between positions 201 and 203 (e.g., 201, 202, 203, or a range therebetween), between positions 221 and 229 (e.g., 221, 222, 223, 224, 225, 226, 227, 228, 229, or a range therebetween), or between positions 269 and 290 (e.g., 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290, or a range therebetween), and includes an sp site at a position corresponding thereto.
[0013] In some embodiments, the split variant can form a covalent bond with a haloalkane substrate.
[0014] In some embodiments, the split variant comprises 100% sequence identity to SEQ ID NO: 1.
[0015] In some embodiments, the split variant contains deletions of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or a range between them) at positions corresponding to one or more of the N-terminus of SEQ ID NO: 1, the C-terminus of SEQ ID NO: 1, and either side of the sp site. In some embodiments, the split variant contains duplicate sequences of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or a range between them) at positions corresponding to either side of the sp site.
[0016] In some embodiments, provided herein is a composition comprising (i) a peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to one or more of SEQ ID NOS: 578 - 1187, and (ii) a polypeptide having at least 70% sequence similarity to one or more of SEQ ID NOS: 1188 - 3033, wherein the complex of the peptide and the polypeptide can form a covalent bond with a haloalkane substrate. In some embodiments, the peptide has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to one of SEQ ID NOS: 578 - 1187. In some embodiments, the peptide has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to one of SEQ ID NOS: 1188 - 3033.
[0017] In some embodiments, provided herein is a peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence similarity to one or more of SEQ ID NOs: 578-1187. In some embodiments, provided herein is a peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to one of SEQ ID NOs: 578-1187. In some embodiments, the peptide can form a complex (e.g., facilitated or non-facilitated) with the polypeptide of SEQ ID NO: 1188, and this complex can form a covalent bond with a haloalkane substrate.
[0018] In some embodiments, provided herein are SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512,A peptide or polypeptide that has at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity with one of 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, and 576, and this peptide or polypeptide has the sequences of SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399,One selected from 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, and 577 can interact with a peptide or polypeptide to form a modified dehalogenase complex and can form a covalent bond with a haloalkane substrate. In some embodiments, the peptide or polypeptide has the sequence numbers 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292,One of 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, and 576, and includes at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity.
[0019] In some embodiments, provided herein are SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511,A peptide or polypeptide that has at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity with one of 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, and 577, and this peptide or polypeptide has the sequences set forth in SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396,A peptide or polypeptide selected from one of 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504, 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, and 576 can interact to form a modified dehalogenase complex, and the modified dehalogenase complex can form a covalent bond with a haloalkane substrate. In some embodiments, the peptide or polypeptide has the sequence numbers 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283,One of 285, 287, 289, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, and 577, and includes at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity.
[0020] In some embodiments, provided herein is a peptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to one of SEQ ID NOs: 578-1187, which peptide can interact with a polypeptide selected from one of SEQ ID NOs: 1188-3033 to form a modified dehalogenase complex, and this modified dehalogenase complex can form a covalent bond with a haloalkane substrate. In some embodiments, the peptide has at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to one of SEQ ID NOs: 578-1187.
[0021] In some embodiments, provided herein is a peptide having 100% sequence identity to SEQ ID NO: 3034 or 3035.
[0022] In some embodiments, provided herein is a peptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to one of SEQ ID NOs: 1188-3033, which polypeptide can interact with a peptide selected from one of SEQ ID NOs: 578-1187, 3034, or 3035 to form a modified dehalogenase complex, and this modified dehalogenase complex can form a covalent bond with a haloalkane substrate. In some embodiments, the polypeptide has at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to one of SEQ ID NOs: 1188-3033.
[0023] In some embodiments, the first fragment, peptide, or polypeptide component of the sp-modified dehalogenase herein exists as a fusion protein with a first peptide, polypeptide, or protein of interest. In some embodiments, the first peptide, polypeptide, or protein of interest is selected from the group consisting of an antibody, antibody fragment, protein A, the Ig-binding domain of protein A, protein G, the Ig-binding domain of protein G, protein A / G, the Ig-binding domain of protein A / G, protein L, the Ig-binding domain of protein L, protein M, the Ig-binding domain of protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, an anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte-binding domain(s) of a protein. In some embodiments, the second fragment, peptide, or polypeptide component of the sp-modified dehalogenase herein exists as a fusion protein with a second peptide, polypeptide, or protein of interest. In some embodiments, the second peptide, polypeptide, or protein of interest is selected from the group consisting of an antibody, antibody fragment, protein A, the Ig-binding domain of protein A, protein G, the Ig-binding domain of protein G, protein A / G, the Ig-binding domain of protein A / G, protein L, the Ig-binding domain of protein L, protein M, the Ig-binding domain of protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, an anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte-binding domain(s) of a protein. In some embodiments, the first peptide, polypeptide, or protein of interest and the second peptide, polypeptide, or protein of interest are interacting elements that can form a complex with each other. In some embodiments, the first peptide, polypeptide, or protein of interest and the second peptide, polypeptide, or protein of interest are co-localization elements configured to co-localize within a cell compartment, cell, tissue, or organism. In some embodiments, the second fragment is linked to a molecule of interest.
[0024] In some embodiments, the first fragment, peptide, or polypeptide component and the second fragment, peptide, or polypeptide component of the sp-modified dehalogenase are fused to an antibody or other binding protein such that their proximity is promoted by the presence of an analyte of the antibody or other binding protein (e.g., in a diagnostic assay).
[0025] In some embodiments, the first fragment, peptide, or polypeptide component of the sp-modified dehalogenase herein and / or the second fragment, peptide, or polypeptide component of the sp-modified dehalogenase herein are linked (either directly or via a linker) to a small molecule. In some embodiments, the small molecule linked to the fragment can interact (e.g., bind) with a small molecule or other element (e.g., a peptide or polypeptide, see above) linked or fused to the other fragment.
[0026] In some embodiments, each fragment of the dehalogenase is linked (e.g., fused, conjugated, etc.) to a complementary interaction or dimerization element. In some embodiments, the interaction or dimerization element promotes the formation of an active dehalogenase complex. For example, the first fragment of the dehalogenase is linked to FRB and the second fragment of the dehalogenase is linked to FKBP. In such embodiments, the presence of rapamycin induces dimerization of FRB and FKBP and promotes the formation of the dehalogenase complex. In some embodiments, the sp dehalogenase is used in such systems that are not capable of independent active complex formation but form an active complex upon promotion.
[0027] In some embodiments, provided herein are one or more polynucleotides encoding the split variants described herein. In some embodiments, provided herein are one or more expression vectors comprising one or more polynucleotides described herein. In some embodiments, provided herein are host cells comprising one or more polynucleotides or one or more expression vectors described herein. In some embodiments, provided are cells whose genomes are edited to incorporate sequences encoding the split variants described herein.
[0028] The split dehalogenase complementation system offers several technical advantages over intact or circularly permuted dehalogenases. Covalent labeling of intact dehalogenases with chloroalkane ligands can enable direct readout of protein location and concentration, whereas split dehalogenases direct such labeling to sites of molecular interactions (e.g., protein–protein interactions). Many important cellular functions, including signal transduction, transcription, translation, and cargo transport, require specific interactions between proteins, membranes, organelles, and subcellular structures. The split dehalogenase system reports on the location, timing, and frequency of these events, while intact dehalogenases can only report on the presence of molecules.
[0029] In some embodiments, the split dehalogenase systems, compositions, and methods herein are used in fluorescence microscopy and / or imaging applications. For example, the split modified dehalogenase enables monitoring of functional / molecular events (e.g., protein:protein interactions) with fluorescent ligands beyond cell culture, e.g., in live animals, tissues, organoid model systems, etc. The split dehalogenase is used to measure the localization and occurrence of molecular events in intracellular structures, cell:cell interactions or interfaces, and deep tissues of living organisms. These uses can further be configured in high-throughput formats for screening or diagnostic applications.
[0030] The components of the split dehalogenase exhibit individually reduced activity compared to the active complex assembled therefrom. In some embodiments, assembly of the active complex is performed using interaction partner proteins fused to each fragment. Bimolecular fluorescence complementation (BiFC) of green fluorescent protein (GFP) and other FPs has been used by researchers for many years, but these BiFC systems have several significant drawbacks. The fluorophore takes time to mature, the proteins tend to assemble irreversibly, and performance decreases under hypoxic conditions. In contrast, experiments conducted during the development of the embodiments herein demonstrate that some split dehalogenases assemble reversibly and, when coupled with a fluorescently tagged ligand, use an exogenously supplied cell-permeable fluorescent ligand that does not require maturation or oxygen. In some embodiments, provided herein is a chloroalkane ligand characterized by a bright and stable fluorophore that outperforms protein-based fluorophores in terms of signal intensity (e.g., quantum yield and extinction coefficient) as well as temporal and spatial resolution (e.g., image resolution), and is ideal for advanced imaging applications such as super-resolution microscopy and light-sheet microscopy.
[0031] In contrast to other enzyme complementation-based reporter systems such as split luciferase, split dehalogenase forms a permanent covalent bond with the substrate, creating a durable event mark that can be observed for hours, days, or longer. In the absence of complementation of the split dehalogenase fragments, ligand binding cannot form, but the covalent bond remains even after the dehalogenase complex has dissociated. Furthermore, multiple complementation events can result in signal accumulation that does not decrease as the substrate becomes depleted, which is in contrast to split luciferase where the signal decreases over time.
[0032] The utility of split dehalogenase extends beyond fluorescence imaging. Dehalogenase can accept a wide variety of ligands when the ligand has a haloalkane functional group. The cargo of the ligand can include, but is not limited to, fluorophores, chromophores, analyte sensing complexes, affinity tags (such as biotin), signals for proteolysis or post-translational modification, nucleic acids, peptides, polypeptides, chemical derivatives for dimerization, or solid supports. Thus, in certain embodiments, split dehalogenase is utilized as an initiation signal for cell events for chromogenesis, sensor activation, affinity tagging, proteolysis, DNA / RNA barcoding, crosslinking, dimerization, or assembly onto a support or molecular scaffold. The ultimate functional output of split dehalogenase is determined by the choice of ligand supplied by the user. The flexibility of the split dehalogenase system described herein is used in a variety of methods and applications.
[0033] In some embodiments, for the utility of certain split-modified dehalogenases having fluorescence and for the detection of protein:protein interactions, the embodiments herein find use in a variety of cell sorting applications. For example: ● Screen for the presence of complementary LgHT:SmHT or “dual” tags (SmHT-HiBiT) during CRISPR cell line generation. This helps solve the problem of isolating clonal cell lines edited with tags without “blind” screening, which adds significant labor and time to isolate cell lines with tags. With a selectable tag that enables fluorescence detection, the user can immediately screen edited cells against cells with that edit. ● Screen for cells expressing a complementary spHaloTag, typically when expressed from a plasmid fused to another protein. ● Screen cells for the presence (or absence) of a specific PPI. This provides enrichment of cells containing the interacting proteins to enable downstream assays, diagnostics, or purification of cells (such as engineered T cells). ● Screen cells that have undergone promoted molecular interactions or molecular proximity through stimuli such as small molecules or hormones. Specific examples are screening cells that have formed ternary complexes through treatment with PROTACs, molecular glues, or other “TACs”. Other examples are screening cells for molecular interactions via BRET and screening cells with differences in fluorescence signals resulting from target engagement detected by split HaloTag (e.g., for drug screening). ● Screen cells for virus-infected cells through viral delivery of nucleic acid sequences encoding split HaloTag fragments, or more directly when the viral protein itself that infects the cells consists of a fusion to an spHaloTag component (such as a viral coat protein).
[0034] A method of combining cell imaging with flow cytometry or sorting to simultaneously measure morphological cell characteristics and the localization of a reporter or dye, evaluate a cell population (e.g., for diagnosis), and identify or isolate rare or difficult-to-culture cell types, or complex phenotypes. The use of split dehalogenases with these methods enables, for example, among other things, cell cycle analysis, apoptosis detection, immunophenotyping, detection and quantification of intracellular signaling, drug screening, microbial population analysis, and stem cell analysis.
[0035] In some embodiments, provided herein is a method of detecting a protein-protein interaction in a sample, comprising contacting (a) (i) a first complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the first portion of SEQ ID NO: 1, and (ii) a first fusion comprising a first protein of interest, (b) (i) a second complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the second portion of SEQ ID NO: 1, and (ii) a second fusion comprising a second protein of interest, and (c) a substrate comprising R-linker-A-X, wherein R is a functional group or a solid support, X is a halogen, and A-X is a substrate for a dehalogenase enzyme, such that binding of the first protein of interest and the second protein of interest results in the formation of a complex between the first complementary fragment and the second complementary fragment that can form a covalent bond with the substrate.
[0036] In certain embodiments, provided herein is a method for detecting an interaction between two proteins in a sample. The method herein includes providing a sample having a cell comprising a fusion of the split dehalogenase or expression vector(s) of the invention, a first heterologous protein sequence and a second heterologous protein sequence, and a first complementary fragment and a second complementary fragment (e.g., encoding complementary fragments of the split dehalogenase), its lysate, or an in vitro transcription / translation reaction comprising such components, and a hydrolase substrate having at least one functional group (e.g., a haloalkane) under conditions effective to permit association of the first fusion protein and the second fusion protein. The presence, amount, or position of at least one functional group in the sample is detected.
[0037] In some embodiments, provided herein is a method for detecting an interaction between two proteins in a sample, comprising: (a) expressing in the sample (i) a first complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1, and (ii) a first fusion comprising a first protein of interest; (b) expressing in the sample (i) a second complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1, and (ii) a second fusion comprising a second protein of interest; (c) contacting the sample with a substrate comprising R-linker-A-X, where R is a functional group or a solid support, X is a halogen, and A-X is a substrate for a dehalogenase enzyme; and (d) detecting the presence, amount, and / or position of at least one functional group.
[0038] In another embodiment, provided herein is a method for detecting a molecule of interest in a sample. The method includes a cell containing an in vitro transcription / translation reaction comprising a molecule of interest, a lysate thereof, or components thereof that are bound to a fusion of a first complementary fragment of split dehalogenase and a second complementary fragment of split dehalogenase with a heterologous protein (or an expression vector encoding the fusion), and a hydrolase substrate having at least one functional group (e.g., a haloalkane) under conditions effective to allow the heterologous protein to interact with the molecule of interest in the sample. The presence, amount, or location of at least one functional group in the sample is detected, thereby detecting the presence, amount, or location of the molecule of interest.
[0039] In some embodiments, provided herein is a method for detecting a molecule of interest in a sample, comprising: (a) contacting the sample with a first complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1 linked to the molecule of interest; (b) in the sample, expressing a fusion comprising (i) a second complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1, and (ii) a protein capable of binding to the molecule of interest, or contacting the sample with the fusion; (c) contacting the sample with a substrate comprising R-linker-A-X (wherein R is a functional group or a solid support, X is a halogen, and A-X is a substrate for a dehalogenase enzyme); and (d) detecting the presence, amount, and / or location of at least one functional group.
[0040] In some embodiments, provided herein is a method for detecting the effect of an agent on the interaction of two proteins, the method comprising: (a) in a sample, expressing (i) a first complementary fragment of a split variant of a polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to a portion of SEQ ID NO: 1, and (ii) a first fusion comprising a first protein sequence, or contacting the sample with the first fusion; (b) in the sample, expressing (i) a second complementary fragment of a split variant of a polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to a portion of SEQ ID NO: 1, and (ii) a fusion comprising a second protein sequence capable of binding to the first protein sequence, or contacting the sample with the fusion; (c) contacting the sample with a substrate comprising R - linker - A - X (wherein R is a functional group or a solid support, X is a halogen, and A - X is a substrate for a dehalogenase enzyme); (d) contacting the sample with an agent; and (e) detecting the presence, amount, and / or location of at least one functional group.
[0041] In some embodiments, provided herein is a method for detecting the effect of an agent on the interaction between a target protein and a protein ligand, the method comprising: (a) in a sample, (i) expressing a fusion comprising a first complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1 and (ii) the target protein, or contacting the sample with the fusion; (b) contacting the sample with a second complementary fragment of a split variant of a polypeptide having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1 linked to the ligand; (c) contacting the sample with a substrate comprising R-linker-A-X (wherein R is a functional group or a solid support, X is a halogen, and A-X is a substrate for a dehalogenase enzyme); (d) contacting the sample with the agent; and (e) detecting the presence, amount, and / or location of at least one functional group.
[0042] In some embodiments, provided herein is a method of controllable targeted proteolysis, comprising: (a) providing, or expressing in a sample, a first fusion comprising (i) a first complementary fragment of a split variant of a polypeptide comprising a portion of SEQ ID NO:1 and having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity thereto, and (ii) a target protein; (b) contacting the sample with a haloalkane proteolysis targeting chimera (PROTAC) and a ligand capable of engaging an E3 ubiquitin ligase; and (c) contacting the sample with a second complementary fragment of a split variant of a polypeptide comprising a portion of SEQ ID NO:1 and having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity thereto, wherein formation of the split variant complex results in binding of the haloalkane by the split variant complex, bringing the ligand capable of engaging the E3 ubiquitin ligase into proximity with the target protein, resulting in ubiquitination of the target protein and directing the target protein for proteasomal degradation. In some embodiments, the first fusion further comprises luciferase or a first component of a bioluminescence complex, one of the complementary fragments is linked to a fluorophore, and luminescence from the luciferase or bioluminescence complex is capable of exciting the fluorophore.
[0043] In some embodiments, provided herein is a method of controllable target protein modification, comprising: (a) (i) providing, or expressing in a sample, a first complementary fragment of a split variant of a polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1, and (ii) a first fusion comprising a target protein; (b) contacting the sample with a ligand capable of engaging a halocarbon chimera and a protein modification enzyme; (c) contacting the sample with a second complementary fragment of a split variant of a polypeptide comprising at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity to the portion of SEQ ID NO: 1, wherein formation of the split variant complex results in binding of the halocarbon by the split variant complex, bringing a ligand capable of engaging a protein modification enzyme into proximity to the target protein and resulting in modification of the target protein. In some embodiments, the chimera is a PhosTAC and the protein modification enzyme is a phosphatase.
[0044] In some embodiments of any of the methods herein, the first complementary fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to the first portion of SEQ ID NO: 1, and the second complementary fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity to the second portion of SEQ ID NO: 1.
[0045] In some embodiments of any of the methods of this specification, the first portion of SEQ ID NO: 1 is SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, 126, 128, 130, 132, 134, 136, 138, 140, 142, 144, 146, 148, 150, 152, 154, 156, 158, 160, 162, 164, 166, 168, 170, 172, 174, 176, 178, 180, 182, 184, 186, 188, 190, 192, 194, 196, 198, 200, 202, 204, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, 228, 230, 232, 234, 236, 238, 240, 242, 244, 246, 248, 250, 252, 254, 256, 258, 260, 262, 264, 266, 268, 270, 272, 274, 276, 278, 280, 282, 284, 286, 288, 290, 292, 294, 296, 298, 300, 302, 304, 306, 308, 310, 312, 314, 316, 318, 320, 322, 324, 326, 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 380, 382, 384, 386, 388, 390, 392, 394, 396, 398, 400, 402, 404, 406, 408, 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 476, 478, 480, 482, 484, 486, 488, 490, 492, 494, 496, 498, 500, 502, 504,Selected from 506, 508, 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, 556, 558, 560, 562, 564, 566, 568, 570, 572, 574, and 576, the second part of SEQ ID NO: 1 is SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, 125, 127, 129, 131, 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, 163, 165, 167, 169, 171, 173, 175, 177, 179, 181, 183, 185, 187, 189, 191, 193, 195, 197, 199, 201, 203, 205, 207, 209, 211, 213, 215, 217, 219, 221, 223, 225, 227, 229, 231, 233, 235, 237, 239, 241, 243, 245, 247, 249, 251, 253, 255, 257, 259, 261, 263, 265, 267, 269, 271, 273, 275, 277, 279, 281, 283, 285, 287, 289, 291, 293, 295, 297, 299, 301, 303, 305, 307, 309, 311, 313, 315, 317, 319, 321, 323, 325, 327, 329, 331, 333, 335, 337, 339, 341, 343, 345, 347, 349, 351, 353, 355, 357, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, 411, 413, 415, 417, 419, 421, 423, 425, 427, 429, 431, 433, 435, 437, 439, 441,Selected from 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, 477, 479, 481, 483, 485, 487, 489, 491, 493, 495, 497, 499, 501, 503, 505, 507, 509, 511, 513, 515, 517, 519, 521, 523, 525, 527, 529, 531, 533, 535, 537, 539, 541, 543, 545, 547, 549, 551, 553, 555, 557, 559, 561, 563, 565, 567, 569, 571, 573, 575, and 577.,
[0046] In some embodiments of any of the methods herein, the first complementary fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity with one of SEQ ID NOs: 578-1187 (or 100% identity with SEQ ID NO: 3034 or 3035), and the second complementary fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence similarity with one of 1188-3033.
[0047] In some embodiments of any of the methods herein, the first complementary fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity with one of SEQ ID NOs: 578-1187 (or 100% identity with SEQ ID NO: 3034 or 3035), and the second complementary fragment comprises at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%) sequence identity with one of 1188-3033.
[0048] In certain embodiments, cells, beads, nanoparticles, liposomes, or other structures are provided that display the first and / or second complementary fragments of a split dehalogenase (e.g., spHT). In some embodiments, cell surface-displayed split dehalogenases are used in bacterial display, yeast display, mammalian display, phage display, and the like. In some embodiments, surface-displayed split dehalogenases are free to interact with an impermeable substrate and can be used to detect analytes in solution or, if both cells display complementary split protein fragments, to detect cell-cell interactions.
[0049] Also provided herein is a method for detecting an agent that alters the interaction of two proteins, comprising providing a sample having cells that contain fusions of the first and second complementary fragments of a split dehalogenase with a first heterologous protein and a second heterologous protein (or expression vector(s) encoding the fusions), a lysate thereof, or an in vitro transcription / translation reaction containing such components, a hydrolase substrate having at least one functional group (e.g., a haloalkane), and providing the sample under conditions effective to permit association of the first fusion protein and the second fusion protein. The agent is suspected of altering the interaction of the first heterologous protein and the second heterologous protein. The presence or amount of at least one functional group in the sample is detected and compared to a sample that does not contain the agent. In some embodiments, multiple concentrations of the agent are assayed to determine the effect of the agent on the protein-protein interaction. In some embodiments, screening is provided using the systems herein to screen libraries (e.g., 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 100,000, or more) of agents and / or heterologous protein sequences.
[0050] In another embodiment, a method is provided for detecting an agent that changes the interaction between a target molecule and a protein. This method includes a target molecule bound to a first complementary fragment of a split dehalogenase, a fusion of a second complementary fragment of the split dehalogenase and a heterologous protein (or an expression vector encoding such a fusion), a lysate thereof, or a cell containing an in vitro transcription / translation reaction containing such components, a hydrolase substrate having at least one functional group (e.g., a haloalkane), and a sample containing an agent suspected of changing the interaction between the heterologous amino acid sequence and the target molecule in the sample, provided under conditions effective to allow the heterologous protein to interact with the target molecule in the sample. The presence or amount of the functional group in the sample as compared to a sample having the agent. In some embodiments, multiple concentrations of the agent are assayed to determine the effect of the agent on the protein-protein interaction. In some embodiments, provided is a screening in which a library of agents (e.g., 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 100,000, or more), a target molecule, and / or a heterologous protein sequence are screened using the systems herein.
[0051] In some embodiments, provided herein is a method for detecting the presence of a target molecule. For example, a cell is contacted with a vector (s) comprising a nucleic acid sequence encoding a promoter, e.g., a regulatable promoter, and two complementary fragments of a mutant hydrolase, at least one of which is fused to a protein that interacts with the target molecule. In one embodiment, the transfected cell is cultured under conditions in which the promoter induces transient expression of the fragment or regulated expression of one of the fragments and an activity associated with the labeled substrate is detected.
[0052] In some embodiments, methods are provided for expressing one or both complementary fragments of a split dehalogenase (e.g., spHT) intracellularly. In some embodiments, the split dehalogenase, or a fragment thereof (or a fusion thereof), is transiently expressed by the cell. In some embodiments, the nucleic acid encoding the split dehalogenase or a fragment thereof (or a fusion thereof) is stably integrated into the cell (or its genome). In some embodiments, provided herein are cells or cell lines that encode and are capable of expressing one or both complementary fragments of a split dehalogenase (e.g., spHT) or a fusion thereof. In some embodiments, methods are provided for generating such cells, for example, by transfecting a nucleic acid vector into the cell and / or by CRISPR insertion of a split dehalogenase (e.g., spHT) construct into the genome of the cell.
[0053] Other methods described herein or otherwise practicable with split dehalogenases are within the scope of the technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0054]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25A
Figure 25B
Figure 26A
Figure 26B
Figure 26C
Figure 27A
Figure 27B
Figure 28A
Figure 28B
Figure 29A
Figure 29B
Figure 29C
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Figure 64
Figure 65
Figure 66
Figure 67
Figure 68
Figure 69
Figure 70
Figure 71
Figure 72
Figure 73
Figure 74
Figure 75
Figure 76
Figure 77
Figure 78
Figure 79
Figure 80
Figure 81
Figure 82
Figure 83
Figure 84
Figure 85
Figure 86
Figure 87
Figure 88
Figure 89
Figure 90
Figure 91
Figure 92
Figure 93
Figure 94
Figure 95
Figure 96
Figure 97
Figure 98
Figure 99
Figure 100
Figure 101
Figure 102
Figure 103
Figure 104
Figure 105
Figure 106
Figure 107
Figure 108
Figure 109
Figure 110
Figure 111
Figure 112
Figure 113
Figure 114
Figure 115
Figure 116
Figure 117
Figure 118
Figure 119
Figure 120
Figure 121
Figure 122
Figure 123
Figure 124
Figure 125
Figure 126
Figure 127
Figure 128
Figure 129
Figure 130
Figure 131
Figure 132
Figure 133
Figure 134
Figure 135
Figure 136
Figure 137
Figure 138
Figure 139
Figure 140
Figure 141
Figure 142
Figure 143
Figure 144
Figure 145
Figure 146
Figure 147
Figure 148
Figure 149
Figure 150
Figure 151
Figure 152
Figure 153
Figure 154
Figure 155
Figure 156
[0055] Definitions When practicing or testing the embodiments described herein, it is possible to use any methods and materials similar or equivalent to those described herein, but some preferred methods, compositions, devices, and materials are described herein. However, before describing these materials and methods, it should be understood that the present invention is not limited to the specific molecules, compositions, methodologies, or procedures described herein, as certain molecules, compositions, methodologies, or procedures may vary according to routine experimentation and optimization. It should also be understood that the terms used in the description are for the purpose of describing only a particular version or embodiment and are not intended to limit the scope of the embodiments described herein.
[0056] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. However, in case of conflict, the present specification, including definitions, will control. Accordingly, the following definitions apply in the context of the embodiments described herein.
[0057] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polypeptide" includes reference to one or more polypeptides and their equivalents known to those skilled in the art.
[0058] As used herein, the term "and / or" includes any and all combinations of the listed items, including any one of the individually listed items. For example, "A, B, and / or C" includes A, B, C, AB, AC, BC, and ABC, and each of these should be considered to be individually recited with the specification of "A, B, and / or C".
[0059] As used herein, the term "comprise" and its linguistic variations indicate the presence of the recited features (plural), elements (plural), method steps (plural), etc., without excluding the presence of additional features (plural), elements (plural), method steps (plural), etc. Conversely, the term "consisting of" and its linguistic variations indicate the presence of the recited features (plural), elements (plural), method steps (plural), etc., and exclude any unrecited features (plural), elements (plural), method steps (plural), etc., except for impurities normally associated therewith. The phrase "consisting essentially of" indicates the recited features (plural), elements (plural), method steps (plural), etc., and any additional features (plural), elements (plural), method steps (plural), etc., that do not substantially affect the basic nature of the composition, system, or method. Many embodiments of this specification are described using the open word "comprise". Such embodiments include multiple closed "consisting of" and / or "consisting essentially of" embodiments, which could alternatively be claimed or described using such words.
[0060] As used herein, the term "substantially" means that the recited characteristics, parameters, and / or values need not be achieved exactly, but to the extent that they do not preclude the intended effects provided by the characteristics, e.g., tolerances, measurement errors, limits of measurement precision, and other factors known to those of skill in the art, such that deviations or variations can occur. A characteristic or feature that is substantially absent (e.g., substantially non-fluorescent) is one that is within noise, below background, below the detection capabilities of the assay being used, or a minor percentage (e.g., <1%, <0.1%, <0.01%, <0.001%, <0.001%, <0.00001%, <0.000001%, <0.0000001%) of a prominent characteristic (e.g., the fluorescence intensity of an active fluorophore).
[0061] As used herein, when referring to an amino acid sequence or a position within an amino acid sequence, the phrase "corresponding to" refers to the relative position of an amino acid residue or amino acid segment, and although a sequence is recited, it is not necessarily the specific identity of the amino acid at that position. For example, a "peptide corresponding to positions 36 - 48 of SEQ ID NO:1" can include less than 100% sequence identity (e.g., greater than 70% sequence identity) with positions 36 - 48 of SEQ ID NO:1, but within the context of the described composition or system, the peptide is relevant to those positions.
[0062] As used herein, the term "system" refers to a plurality of components (e.g., devices, compositions, etc.) used for a particular purpose. For example, two distinct biological molecules can be included in a system if they are useful together for a common purpose, regardless of whether they are present in the same composition.
[0063] As used herein, the term "complementary" refers to the property of two or more structural elements (e.g., peptides, polypeptides, nucleic acids, small molecules, etc.) that can hybridize, dimerize, or otherwise form a complex with each other. For example, "complementary peptides and polypeptides" can together form one complex. Complementary elements may require assistance (facilitation), for example, to form a complex (e.g., from interacting elements), to arrange the elements in an appropriate conformation for complementarity, to place the elements in appropriate proximity for complementarity, to co-localize complementary elements, to reduce the interaction energy for complementary elements, or to overcome insufficient affinity for each other.
[0064] As used herein, the term "complex" refers to an aggregate or assembly of molecules (e.g., peptides, polypeptides, etc.) that contact each other directly and / or indirectly. In one aspect, "contact", or more specifically "direct contact", means that two or more molecules are sufficiently close such that attractive non-covalent interaction forces, such as van der Waals forces, hydrogen bonds, ionic and hydrophobic interactions, dominate the interaction of the molecules. In such an aspect, a complex of molecules (e.g., peptides, polypeptides, etc.) is formed under assay conditions such that the complex is thermodynamically favorable (e.g., as compared to the non-aggregated or non-complexed state of its constituent molecules). As used herein, the term "complex" refers to an aggregate of two or more molecules (e.g., peptides, polypeptides, etc.) unless otherwise stated.
[0065] As used herein, the term "interaction element" refers to a portion that aids or facilitates two or more structural elements (e.g., peptides, polypeptides, etc.) coming together to form a complex. In some embodiments, a pair of interaction elements (aka "interaction pair") are attached to a pair of structural elements (e.g., peptides, polypeptides, etc.), and the attractive interaction between the two interaction elements promotes the formation of a complex of the structural elements. The interaction elements can facilitate the formation of a complex by any suitable mechanism (e.g., bringing the structural elements into proximity, arranging the structural elements in a suitable conformation for stable interaction, reducing the activation energy for complex formation, combinations thereof, etc.). The interaction elements can be proteins, polypeptides, peptides, small molecules, cofactors, nucleic acids, lipids, carbohydrates, antibodies, etc. The interaction pair can be made from two of the same interaction element (i.e., homopair) or two different interaction elements (i.e., heteropair). In the case of a heteropair, the interaction elements can be of the same type of moiety (e.g., polypeptide) or two different types of moieties (e.g., polypeptide and small molecule). In some embodiments where complex formation by an interaction pair is being studied, the interaction pair can be referred to as a "target pair" or "pair of interest", and the individual interaction elements can be referred to as "target elements" (e.g., "target peptide", "target polypeptide", etc.) or "elements of interest" (e.g., "peptide of interest", "polypeptide of interest", etc.).
[0066] As used herein, the term "low affinity" describes an intermolecular interaction between two or more entities that is too weak to result in significant complex formation between the entities, except at concentrations substantially higher (e.g., 2-fold, 5-fold, 10-fold, 100-fold, 1000-fold, or more) than physiological or assay conditions, or at concentrations that involve facilitation from the formation of a second complex of the attached element (e.g., interaction element).
[0067] As used herein, the term "high affinity" describes an intermolecular interaction between two or more (e.g., three) entities having sufficient strength to produce detectable complex formation under physiological or assay conditions without facilitation from the formation of a second complex of the attached element (e.g., interacting element).
[0068] As used herein, the term "existing protein" refers to an amino acid sequence that physically existed prior to a particular event or date. A "peptide that is not a fragment of an existing protein" is a short amino acid chain that is not a fragment or subsequence of a protein (e.g., synthetic or naturally occurring) that physically existed prior to the design and / or synthesis of the peptide.
[0069] As used herein, the term "fragment" refers to a peptide or polypeptide resulting from the dissociation or "fragmentation" of a larger whole entity (e.g., a protein, polypeptide, enzyme, etc.), or a peptide or polypeptide prepared to have the same sequence as such. Thus, a fragment is a partial sequence of the entire entity (e.g., a protein, polypeptide, enzyme, etc.) from which the fragment is made and / or designed. A peptide or polypeptide that is not a partial sequence of an existing whole protein is not a fragment (e.g., not a fragment of an existing protein). A peptide or polypeptide that is "not a fragment of an existing protein" is an amino acid chain that is not a partial sequence of a protein (e.g., natural or synthetic) that physically existed prior to the design and / or synthesis of the peptide or polypeptide. As used herein, a fragment of a hydrolase or dehalogenase has fewer residues than the full-length sequence, cannot form a substrate-binding site alone, and / or has substantially reduced or no substrate-binding activity, but shows substantially increased substrate-binding activity when in proximity to a second fragment of the hydrolase or dehalogenase. In one embodiment, a fragment of a hydrolase or dehalogenase is at least 5, e.g., at least 10, at least 20, at least 30, at least 40, or at least 50 consecutive residues of a wild-type or mutant hydrolase, or a sequence having at least 70% sequence identity thereto, and does not necessarily include the N-terminal or C-terminal residue, or N-terminal or C-terminal sequence, of the corresponding full-length protein.
[0070] As used herein, the term "subsequence" refers to a peptide or polypeptide having 100% sequence identity to a portion of another larger peptide or polypeptide. This subsequence is a perfect sequence match for the portion of the larger amino acid chain.
[0071] The term "amino acid" refers to natural amino acids, unnatural amino acids, and amino acid analogs, and unless otherwise indicated, all of their D and L stereoisomers when their structures permit such stereoisomeric forms.
[0072] The term "proteinogenic amino acid" refers to the 20 amino acids encoded by the human genetic code, including alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine (Cys or C), glutamine (Gln or Q), glutamic acid (Glu or E), glycine (Gly or G), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V). Selenocysteine and pyrrolysine may also be considered proteinogenic amino acids.
[0073] The term "non-proteinogenic amino acid" refers to an amino acid that is not naturally encoded or found in the genetic code of any organism and is not biosynthetically incorporated into proteins during translation. Non-proteinogenic amino acids can be "unnatural amino acids" (amino acids that do not occur naturally) or "naturally occurring non-proteinogenic amino acids" (e.g., norvaline, ornithine, homocysteine, etc.). Examples of non-proteinogenic amino acids include, but are not limited to, azetidinecarboxylic acid, 2-aminoadipic acid, 3-aminoadipic acid, beta-alanine, naphthylalanine, aminopropionic acid, 2-aminobutyric acid, 4-aminobutyric acid, 6-aminocaproic acid, 2-aminoheptanoic acid, 2-aminoisobutyric acid, 3-aminoisobutyric acid, 2-aminopimelic acid, tert-butylglycine, 2,4-diaminoisobutyric acid, desmosine, 2,2'-diaminopimelic acid, 2,3-diaminopropionic acid, N-ethylglycine, N-ethylasparagine, homoproline, hydroxylysine, allo-hydroxylysine, 3-hydroxyproline, 4-hydroxyproline, isodesmosine, allo-isoleucine, N-methylalanine, N-alkylglycine including N-methylglycine, N-methylisoleucine, N-alkylpentylglycine including N-methylpentylglycine. Included are N-methylvaline, naphthylalanine, norvaline, norleucine ("Norleu"), octylglycine, ornithine, pentylglycine, pipecolic acid, thioproline, homolysine, and homoarginine. Non-proteinogenic forms include D-amino acid forms of any of the amino acids herein, as well as non-alpha amino acid forms (beta amino acids, gamma amino acids, delta amino acids, etc.) of any of the amino acids herein, all of which are within the scope hereof and may be included in the peptides herein.
[0074] The term "amino acid analog" refers to an amino acid (e.g., natural or non-natural, proteinaceous or non-proteinaceous) in which one or more of the C-terminal carboxy group, N-terminal amino group, and side-chain bioactive group are chemically blocked or modified to the bioactive group, either reversibly or irreversibly, or by other means. For example, aspartic acid-(beta-methyl ester) is an amino acid analog of aspartic acid, N-ethylglycine is an amino acid analog of glycine, or alanine carboxamide is an amino acid analog of alanine. Other amino acid analogs include methionine sulfoxide, methionine sulfone, S-(carboxymethyl)-cysteine, S-(carboxymethyl)-cysteine sulfoxide, and S-(carboxymethyl)-cysteine sulfone.
[0075] As used herein, unless otherwise specified, the terms "peptide" and "polypeptide" refer to polymeric compounds of two or more amino acids joined through the backbone by peptide amide bonds (-C(O)NH-). The term "peptide" typically refers to short amino acid polymers (e.g., chains having less than 30 amino acids), and the term "polypeptide" typically refers to longer amino acid polymers (e.g., chains having more than 30 amino acids).
[0076] As used herein, the terms "artificial" or "synthetic" refer to compositions and systems that do not occur in nature. For example, an artificial or synthetic peptide, polypeptide, or nucleic acid contains a non-natural sequence (e.g., a peptide having less than 100% identity to a naturally occurring protein or a fragment thereof).
[0077] As used herein with respect to the production of peptides and polypeptides, the term "synthesis" and its linguistic variants can refer to chemical peptide synthesis techniques, as well as the gene expression of peptides and polypeptides.
[0078] As used herein, "conservative" amino acid substitutions refer to the substitution of an amino acid in a peptide or polypeptide with another amino acid having similar chemical properties, such as size or charge. For the purposes of the present disclosure, each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) Alanine (A) and glycine (G); 2) Aspartic acid (D) and glutamic acid (E); 3) Asparagine (N) and glutamine (Q); 4) Arginine (R) and lysine (K); 5) Isoleucine (I), leucine (L), methionine (M), and valine (V); 6) Phenylalanine (F), tyrosine (Y), and tryptophan (W); 7) Serine (S) and threonine (T); and 8) Cysteine (C) and methionine (M).
[0079] Amino acid residues can be grouped into classes based on common side chain properties, such as polar positive (or basic) (e.g., histidine (H), lysine (K), and arginine (R)); polar negative (or acidic) (e.g., aspartic acid (D), glutamic acid (E)); polar neutral (e.g., serine (S), threonine (T), asparagine (N), glutamine (Q)); nonpolar aliphatic (e.g., alanine (A), valine (V), leucine (L), isoleucine (I), methionine (M)); nonpolar aromatic (e.g., phenylalanine (F), tyrosine (Y), tryptophan (W)); proline and glycine; and cysteine. As used herein, "semi-conservative" amino acid substitutions refer to the substitution of an amino acid in a peptide or polypeptide with another amino acid within the same class.
[0080] In some embodiments, unless otherwise specified, conservative or semi-conservative amino acid substitutions may also include non-naturally occurring amino acid residues that have chemical properties similar to natural residues. These non-natural residues are typically incorporated by chemical peptide synthesis rather than by synthesis in biological systems. These include, but are not limited to, peptidomimetics and other reverse or inverted forms of amino acid moieties. Embodiments herein may, in some embodiments, be limited to natural amino acids, non-natural amino acids, and / or amino acid analogs.
[0081] Non-conservative substitutions may involve exchanging a member of one class for a member of another class.
[0082] As used herein, the term "sequence identity" refers to the degree to which two polymer sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have a continuous composition of the same monomer subunits. The term "sequence similarity" refers to the degree to which two polymer sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have similar polymer sequences. For example, similar amino acids are those that share the same biophysical properties, and can be grouped, for example, as acidic (e.g., aspartate, glutamate), basic (e.g., lysine, arginine, histidine), non-polar (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), and uncharged polar (e.g., glycine, asparagine, glutamine, cysteine, serine, threonine, tyrosine). "Percent sequence identity" (or "percent sequence similarity") is calculated by: (1) comparing two sequences optimally aligned over a window of comparison (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); (2) determining the number of positions that contain identical (or similar) monomers (e.g., where the same amino acid occurs in both sequences, where similar amino acids occur in both sequences) to obtain the number of matched positions; (3) dividing the number of matched positions by the total number of positions in the comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); and (4) multiplying the result by 100 to obtain the percent sequence identity or percent sequence similarity. For example, if Peptide A and Peptide B are both 20 amino acids in length and have the same amino acids at all positions except one, then Peptide A and Peptide B have 95% sequence identity. If the non-identical amino acids at the non-identical position share the same biophysical characteristics (e.g., both are acidic), then Peptide A and Peptide B would have 100% sequence similarity.As another example, if peptide C is 20 amino acids in length, peptide D is 15 amino acids in length, and 14 out of the 15 amino acids in peptide D are identical to a portion of peptide C, then peptides C and D have 70% sequence identity, while peptide D has 93.3% sequence identity with the optimal comparison window of peptide C. As used herein, for the purposes of calculating "percent sequence identity" (or "percent sequence similarity"), any gaps in the aligned sequences are treated as mismatches at that position.
[0083] Any peptide / polypeptide described herein as having a particular reference sequence ID number and a specific percent sequence identity or similarity (e.g., at least 70%) may also be represented as having the maximum number of substitutions (or terminal deletions) relative to that reference sequence. For example, a sequence having at least Y% sequence identity (e.g., 90%) with SEQ ID NO: Z (e.g., 100 amino acids) may have a maximum of X substitutions (e.g., 10) relative to SEQ ID NO: Z, and thus may also be represented as "having X (e.g., 10) or fewer substitutions relative to SEQ ID NO: Z".
[0084] As used herein, the term "physiological conditions" encompasses any conditions that are compatible with living cells, such as aqueous conditions including temperature, pH, salinity, chemical composition, etc., that are primarily compatible with living cells.
[0085] As used herein, the term "sample" is used in its broadest sense. In one sense, it is meant to include specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from animals (including humans) and include fluids, solids, tissues, and gases. Biological samples include blood preparations such as plasma and serum. A sample can also refer to a cell lysate or purified form of an enzyme, peptide, and / or polypeptide described herein. A cell lysate can include cells lysed with a lysing agent, or lysates such as rabbit reticulocyte or wheat germ lysates. A sample can also include a cell-free expression system. Environmental samples include environmental materials such as surface substances, soil, water, crystals, and industrial samples. However, such examples should not be construed as limiting the types of samples applicable to the present invention.
[0086] As used herein, the terms "fusion", "fusion polypeptide", and "fusion protein" refer to chimeric proteins containing a first protein or polypeptide of interest conjugated to a second, different peptide, polypeptide, or protein (e.g., an interaction element).
[0087] As used herein, the terms "conjugate" and "conjugation" refer to a covalent bond between two molecular entities (e.g., after synthesis and / or during synthetic production). The attachment of a peptide or small molecule tag to a protein or small molecule, either chemically (e.g., "chemically" conjugated) or enzymatically, is an example of a conjugate.
[0088] As used herein, the terms "polypeptide component" or "peptide component" are used synonymously with the terms "polypeptide component of the [mutant dehalogenase] complex" or "peptide component of the [mutant dehalogenase] complex". Typically, as used herein, a polypeptide component or peptide component can form a complex with a second component under appropriate conditions to form a desired complex.
[0089] As used herein, the term "dehalogenase" refers to an enzyme that catalyzes the removal of halogen atoms from a substrate. The term "haloalkane dehalogenase" refers to an enzyme that catalyzes the removal of halogen from a haloalkane substrate to produce an alcohol and a halide. Dehalogenases and haloalkyl dehalogenases belong to the hydrolase enzyme family and may be so referred to herein or elsewhere.
[0090] As used herein, the term "modified dehalogenase" refers to a dehalogenase variant (artificial variant) that has a mutation that prevents the release of the substrate from the protein after halogen removal and results in a covalent bond between the substrate and the modified dehalogenase. Since the modified dehalogenase does not release the substrate, it cannot turnover and is not a classical enzyme. The HALOTAG system (Promega) is a commercially available modified dehalogenase and substrate system.
[0091] As used herein, the term "circularly permuted" ("cp") refers to a polypeptide in which the N- and C-termini are joined together, either directly or via a linker, to produce a circularly permuted polypeptide, which is then opened at a position other than between the N- and C-termini to produce a new linear polypeptide that is different from the termini of the original polypeptide. The position at which the circularly permuted polypeptide is opened is referred to herein as the "cp site". Circular permutants include polypeptides having the same sequence and structure as the polypeptide that has been circularly permuted and then opened. Thus, a cp polypeptide may be newly synthesized as a linear molecule and may not undergo the circularization and opening steps. The preparation of circularly permuted derivatives is described in International Publication No. WO95 / 27732, which is incorporated herein by reference in its entirety.
[0092] As used herein, the term "split" ("sp") refers to a polypeptide that has been split into two fragments at an internal site of the original polypeptide. The fragments of the sp polypeptide are structurally complementary and, if capable of forming an active complex, can reconstitute the activity of the original polypeptide. The nomenclature herein for referring to the split components of a polypeptide enumerates the position numbers from the full polypeptide corresponding to the last residue in the N-terminal component of the split polypeptide. For example, if a polypeptide is 100 residues in length, the sp52 version of that polypeptide comprises a first fragment corresponding to positions 1-52 of the parent polypeptide and a second fragment corresponding to positions 53-100 of the parent polypeptide. As another example, spHT(45) refers to a split variant of the commercially available HALOTAG protein, where the first fragment comprises residues 1-45 of the HALOTAG polypeptide sequence and the second fragment comprises residues 46-297 of the HALOTAG polypeptide sequence.
[0093] Alternatively, the components of a split polypeptide may be expressed herein by referring to the name of the polypeptide from which it is derived, the residues in the source polypeptide that are present within the component (in parentheses), followed by any substitutions within the component relative to the source polypeptide (in parentheses). For example, a split component of the commercially available HALOTAG protein corresponding to positions 22-297 of the HALOTAG sequence can be described as HaloTag[22-297]. If the second position of the component includes an M-F substitution, the component can be designated HaloTag[22-297](M2F). The component may include an N-terminal methionine residue that is not present in the source sequence, and such a residue indicates the position of the substitution but is not included in the numbering of the fragment within the source polypeptide.
[0094] As used herein, the term "gapped" refers to a split variant of a polypeptide that lacks a segment of the original polypeptide. For example, a "gapped sp polypeptide" is one that lacks the segment of the original sequence that occurs at the site of the split.
[0095] As used herein, the term "overlapped" refers to a split variant of a polypeptide that includes an overlap of segments of the original polypeptide. For example, an "overlapped sp polypeptide" is one in which segments of the original sequence adjacent to the split site are present (overlapped) at the C-terminus of the first fragment and the N-terminus of the second fragment.
Best Mode for Carrying Out the Invention
[0096] Provided herein are peptide and polypeptide sequences that structurally assemble to form an active modified dehalogenase structure that can bind (e.g., covalently) to a haloalkyl ligand. In particular, provided herein are split dehalogenase variants that assemble via structural complementation into an active dehalogenase complex, as well as systems and methods of use thereof.
[0097] Split variant proteins, i.e., enzymes mutated to inhibit or eliminate catalytic activity, are used to reveal and analyze protein interactions within cells, e.g., protein interactions in which each part (fragment) of a split protein is fused to a different protein. Provided herein are split variant hydrolases such as those derived from the commercially available HALOTAG protein (Promega) disclosed in U.S. Patent Application Publication No. 20060024808, the disclosure of which is incorporated herein by reference, and / or mutant hydrolases.
[0098] Although these mutant hydrolases are not technically enzymes (there is no substrate conversion), the stable binding of the substrate to them depends on the appropriate protein structure. The result of reassociating the split fragments of the mutant hydrolase is that the labeling function of the mutant hydrolase is retained in one of the fragments even after separation from its partner, whereas split enzymes are active only while they are together and do not have artifacts of their previous activity after they are separated, which is different from that of the split enzyme system. In fact, the labeling reaction of the split mutant hydrolase provides a molecular memory of protein interaction. In the case of a fluorogenic ligand, the label is retained in one of the fragments but may not be detectable after complex dissociation (due to the possible disruption / absence of the fluorogen-activating contact with the protein). Therefore, the combination of split dehalogenase and a fluorogenic ligand results in a unique situation of permanent labeling but with dynamic (on / off) fluorescence detection of the retained label.
[0099] As an example of a mutant hydrolase, mutant dehalogenase provides efficient labeling within living cells or their lysates. This labeling is conditional only on the presence or expression of the protein and the presence of the labeled hydrolase substrate. In contrast, the labeling of split mutant dehalogenase depends on specific protein interactions occurring within the cell and the presence of the labeled hydrolase substrate. For example, beta-arrestin can be fused to one fragment of the mutant hydrolase, and a G-coupled receptor can be fused to the other fragment. Stimulating the receptor in the presence of the labeled substrate causes beta-arrestin to bind to the receptor and trigger a labeling reaction either at the receptor fusion or the beta-arrestin fusion (depending on which part of the mutant hydrolase contains the reactive nucleophilic amino acid).
[0100] In some embodiments, provided herein is a split variant hydrolase (e.g., split modified dehalogenase) system comprising a first fragment of a hydrolase fused to a protein of interest and optionally a second fragment of a hydrolase fused to a ligand of the first protein of interest. At least one of the hydrolase fragments has a substitution that forms a bond with a hydrolase substrate that is more stable than the bond formed between the corresponding full-length wild-type hydrolase and the hydrolase substrate when present in a full-length variant hydrolase (e.g., modified dehalogenase) having the sequences of the two fragments. In one embodiment, each fragment of the hydrolase is fused to a protein of interest, and the proteins of interest interact, e.g., bind to each other. In another embodiment, one hydrolase fragment is fused to a protein of interest that interacts with a molecule in a sample. In another embodiment, a complex is formed by the binding of a first hydrolase fragment, a second protein fused to a second hydrolase fragment, or a fusion having a second hydrolase fragment and a protein of interest fused to a cellular molecule, in the presence of an agent (one or more agents of interest) or under certain conditions.
[0101] Accordingly, both fragments of the hydrolase (e.g., modified dehalogenase) are structurally related to the full-length hydrolase (and include significant sequence identity / similarity, e.g., greater than 70%), and provide a mutant hydrolase that includes at least one amino acid substitution that results in a covalent bond of the hydrolase substrate. The full-length mutant hydrolase lacks or has reduced catalytic activity compared to the corresponding full-length wild-type hydrolase, specifically binds to a substrate that can be specifically bound by the corresponding full-length wild-type hydrolase, but under conditions, no product or substantially less product is formed from the interaction between the mutant hydrolase and the substrate, e.g., 2-fold, 10-fold, 100-fold, or 1000-fold less product is formed, which results in product formation from the reaction between the corresponding full-length wild-type hydrolase and the substrate. The lack or reduction in product formation by the mutant hydrolase is due to at least one substitution in the full-length mutant hydrolase, and the substitution results in the mutant hydrolase forming a bond with the substrate, which is more stable than the bond formed between the corresponding full-length wild-type hydrolase and the substrate.
[0102] HALOTAG is a 297-residue self-labeling polypeptide (33 kDa) derived from a bacterial hydrolase (dehalogenase) enzyme that has been modified to covalently bind to its ligand, a haloalkane moiety. The HALOTAG ligand can be linked to a solid surface (e.g., beads) or a functional group (e.g., fluorophore), and the HALOTAG polypeptide can be fused to various proteins of interest, enabling covalent attachment of the protein of interest to a solid surface or functional group.
[0103] A HALOTAG polypeptide is a hydrolase (e.g., a modified dehalogenase) with a genetically modified active site that specifically binds to a haloalkane ligand chloroalkane linker, and the rate of ligand binding is improved and increased (Pries et al. The Journal of Biological Chemistry. 270(18):10405-11; incorporated in its entirety by reference). The reaction that forms the bond between the protein tag and the chloroalkane linker is fast and essentially irreversible under physiological conditions (Waugh DS (June 2005). Trends in Biotechnology. 23(6):316-20; incorporated in its entirety by reference). In the native hydrolase enzyme, nucleophilic attack by the chloroalkane reactive linker causes displacement of the halogen by an amino acid residue, resulting in the formation of a covalent alkyl enzyme intermediate. This intermediate is then hydrolyzed by an amino acid residue within the wild-type hydrolase (Chen et al. (February 2005) Current Opinion in Biotechnology. 16(1):35-40; incorporated in its entirety by reference). This would result in the regeneration of the enzyme after the reaction. However, in HALOTAG, a modified haloalkane dehalogenase, the enzyme mutation renders it non-hydrolyzable, so the reaction intermediate cannot proceed with a second reaction. As a result, the intermediate persists as a stable covalent adduct with no associated reverse reaction (Marks et al. (August 2006) Nature Methods. 3(8):591-6; incorporated in its entirety by reference).
[0104] The HaloTag fusion protein can be expressed using standard recombinant protein expression techniques (Adams et al. (May 2002) Journal of the American Chemical Society. 124(21):6063-76; which is incorporated by reference in its entirety). The HaloTag polypeptide is a relatively small protein, and since the reaction is foreign to mammalian cells, there is no interference by endogenous mammalian metabolic reactions (Naested et al. The Plant Journal. 18(5):571-6; which is incorporated by reference in its entirety). When the fusion protein is expressed, there is a wide range of potential experimental fields including enzyme assays, cell imaging, protein arrays, determination of intracellular localization, and many additional possibilities (Janssen DB (April 2004). Current Opinion in Chemical Biology. 8(2):150-9; which is incorporated by reference in its entirety).
[0105] Various HaloTag ligands, functional groups, fusions, assays, modifications, uses, etc. are described in U.S. Patent No. 8,748,148, U.S. Patent No. 9,593,316, U.S. Patent No. 10,246,690, U.S. Patent No. 8,742,086, U.S. Patent No. 9,873,866, U.S. Patent No. 10,604,745, U.S. Patent Application No. 2009 / 0253131, U.S. Patent Application No. 2010 / 0273186, 20130337539, U.S. Patent Application No. 2012 / 0258470, U.S. Patent Application No. 2012 / 0252048, U.S. Patent Application No. 2011 / 0201024, U.S. 2014 / 0322794, each of which is incorporated by reference in its entirety.
[0106] Reversible protein complementation systems and biosensors have proven to be particularly useful tools for measuring functional dynamics using cell imaging, such as protein interactions or changes in metabolite concentrations. Therefore, during the development of the embodiments herein, experiments were conducted to identify regions within the HALOTAG sequence that would allow for the design of strategies to dynamically control self-labeling activity. An exhaustive screening was first performed to identify all possible circular permutation sites in the HALOTAG protein that retain activity and stability in relation to single polypeptides and / or conditionally separable fragments. The information obtained from this screening was used to design and test split HALOTAG pairs.
[0107] In some embodiments, provided herein is a HALOTAG-based system engineered for functional biology, such as split HATOTAG polypeptides, having properties similar to existing full-length proteins in terms of fragment stability, solubility, and expression, and having the additional feature that a significant portion of its activity can be reconstituted upon reconstitution of the complete enzyme. Particularly important HALOTAG ligands for certain embodiments herein include fluorescent generating ligands. Systems that combine spHT can be engineered to have a broad range of fragment affinities to enable both facilitated and spontaneous complementation systems. The split HALOTAG system facilitates endogenous tagging of proteins and improves fluorescent ligands or sensors through higher signals, stability, dynamic range, etc. The HALOTAG-based functional biology tools described herein are well-suited for measuring protein dynamics in live cells using fluorescence imaging, where other technologies lack the utility of the self-labeling activity of HALOTAG or the sensitivity of fluorescent chloroalkane ligands.
[0108] As described herein, embodiments are not limited to HALOTAG sequences. In some embodiments, provided herein are split modified dehalogenases that differ in sequence from SEQ ID NO: 1. In some embodiments, provided herein are split dehalogenases that lack mutations (e.g., 272 and / or 106) that result in a covalent bond to a haloalkane substrate. Such sp dehalogenases enable substrate conversion but are otherwise true enzymes that include the sequences and properties of the embodiments described herein.
[0109] During the development of the embodiments of this specification, experiments were conducted to examine split dehalogenases, their ability to assemble into an active dehalogenase structure, and their ability to activate a fluorogenic substrate. Initially, an exhaustive screening of all circular permutants of HaloTag (cpHT) revealed that 228 / 296 (77%) reacted with CA-TMR and 50 variants had at least 10% of the native HT activity on CA-AlexaFluor488. 17 cpHT variants had increased thermal stability compared to HT, and 38 variants showed recovery of activity after heat denaturation, presumably due to protein refolding. The most active variants in terms of Alexa Fluor488 kinetics were clustered in the region distal to the lid domain (residues 133 - 215), but this effect may be specific to this substrate, which is negatively charged and sensitive to lid domain perturbation. Indeed, when using a neutral TMR ligand, the clustering effect was not as prominent. With the exception of cpHT near residues 111 and 120, all refolding variants were localized to the lid domain, and all thermostabilized variants were also within the lid domain. From these results, 22 candidates identified in the cpHT screening were pursued for further testing as true split proteins (spHT). As fusions to FRB and FKBP, a series of spHT variants showing rapamycin-inducible complementation were identified and demonstrated by activation of a fluorogenic HT ligand (e.g., spHT(133), spHT(145), spHT(157), spHT(180), and spHT(195), etc.). This functionality also extends to pairs of spHT fragments containing varying degrees of sequence overlap localized to the lid subdomain of HT. Further investigation of the perturbations in the lid subdomain revealed an important function of Helix 8 in activating the bound fluorogenic ligand. The spHT complexes showed diverse behaviors from the perspective of reversibility, with three fully reversible complexes and one irreversible complex identified in rapamycin / FK506 competition experiments, and an overall stabilization effect noted for the JF646 binding state of all complexes.It was noted that the spHT-FRB / FKBP fragments were co-expressed in mammalian cells and that the complex was likely formed spontaneously by co-translational folding. Collectively, this study demonstrates a broad functional utility for spHT designs, some of which exhibit unique properties.
[0110] In some embodiments, provided herein are spHT polypeptides and their systems. Specifically, an sp-modified dehalogenase is provided that can reconstitute all or part of the activity of a parental dehalogenase.
[0111] In some embodiments, the polypeptides, peptides, fragments, and combinations thereof described herein are derived from the modified dehalogenase sequence of SEQ ID NO: 1. MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG.
[0112] In some embodiments, the peptides and polypeptides herein include all or part of SEQ ID NO: 1 and at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity). In some embodiments, the peptides and polypeptides herein include all or part of SEQ ID NO: 1 and 100% sequence identity. In some embodiments, the peptides and polypeptides herein include all or part of SEQ ID NO: 1 and at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity). In some embodiments, the peptides and polypeptides herein include all or part of SEQ ID NO: 1 and 100% sequence similarity.
[0113] In some embodiments, the peptide or polypeptide herein includes A at the position corresponding to position 2 of SEQ ID NO: 1. In other embodiments, the peptide or polypeptide herein includes S at the position corresponding to position 2 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes V at the position corresponding to position 47 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes T at the position corresponding to position 58 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes G at the position corresponding to position 78 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes F at the position corresponding to position 88 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes M at the position corresponding to position 89 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes F at the position corresponding to position 128 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes T at the position corresponding to position 155 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes K at the position corresponding to position 160 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes V at the position corresponding to position 167 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes T at the position corresponding to position 172 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes M at the position corresponding to position 175 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes G at the position corresponding to position 176 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes N at the position corresponding to position 195 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes E at the position corresponding to position 224 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes D at the position corresponding to position 227 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein includes K at the position corresponding to position 257 of SEQ ID NO: 1.In some embodiments, the peptide or polypeptide herein contains A at the position corresponding to position 264 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains N at the position corresponding to position 272 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains L at the position corresponding to position 273 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains S at the position corresponding to position 291 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains T at the position corresponding to position 292 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains E at the position corresponding to position 294 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains I at the position corresponding to position 295 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains S at the position corresponding to position 296 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein contains G at the position corresponding to position 297 of SEQ ID NO: 1.
[0114] In some embodiments, the peptide or polypeptide herein does not have an S at the position corresponding to position 2 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an L at the position corresponding to position 47 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an S at the position corresponding to position 58 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a D at the position corresponding to position 78 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a Y at the position corresponding to position 88 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an L at the position corresponding to position 89 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a C at the position corresponding to position 128 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an A at the position corresponding to position 155 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an E at the position corresponding to position 160 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an A at the position corresponding to position 167 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an A at the position corresponding to position 172 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a K at the position corresponding to position 175 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a C at the position corresponding to position 176 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a K at the position corresponding to position 195 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an A at the position corresponding to position 224 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an N at the position corresponding to position 227 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an E at the position corresponding to position 257 of SEQ ID NO: 1.In some embodiments, the peptide or polypeptide herein does not have a T at the position corresponding to position 264 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an H at the position corresponding to position 272 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a Y at the position corresponding to position 273 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have a P at the position corresponding to position 291 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an A at the position corresponding to position 292 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an amino acid at the position corresponding to position 294 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an amino acid at the position corresponding to position 295 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an amino acid at the position corresponding to position 296 of SEQ ID NO: 1. In some embodiments, the peptide or polypeptide herein does not have an amino acid at the position corresponding to position 297 of SEQ ID NO: 1.
[0115] As described herein, embodiments are not limited to the HALOTAG sequence. In some embodiments, provided herein is a split-modified dehalogenase that differs in sequence from SEQ ID NO: 1. In some embodiments, provided herein is a split dehalogenase lacking a mutation(s) (e.g., 272 and / or 106) that results in a covalent bond to a haloalkane substrate. Such split dehalogenases enable substrate conversion but are otherwise true enzymes that include the sequences and properties of the embodiments described herein.
[0116] In some embodiments, the sp dehalogenase (e.g., spHT) comprises two peptide and / or polypeptide components having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) combined with all or part of SEQ ID NO: 1. For example, the first peptide / polypeptide component of the sp polypeptide corresponds to the first part of SEQ ID NO: 1 (e.g., having at least 70% sequence identity to the first part), and the first peptide / polypeptide component of the sp polypeptide corresponds to the second part of SEQ ID NO: 1 (e.g., having at least 70% sequence identity to the second part). In some embodiments, the sp dehalogenase (e.g., spHT) comprises two fragments having 100% sequence identity combined with all or part of SEQ ID NO: 1. For example, the first fragment of the sp polypeptide has 100% sequence identity to the first part of SEQ ID NO: 1, and the second fragment of the sp polypeptide has 100% sequence identity to the second part of SEQ ID NO: 1.
[0117] In some embodiments, the sp dehalogenase (e.g., spHT) comprises two peptide and / or polypeptide components that together include at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 1. For example, the first peptide / polypeptide component of the sp polypeptide corresponds to the first portion of SEQ ID NO: 1 (e.g., at least 70% sequence similarity to the first portion), and the second peptide / polypeptide component of the sp polypeptide corresponds to the second portion of SEQ ID NO: 1 (e.g., at least 70% sequence similarity to the second portion). In some embodiments, the sp dehalogenase (e.g., spHT) comprises two fragments that together include 100% sequence similarity to all or a part of SEQ ID NO: 1. For example, the first fragment of the sp polypeptide has 100% sequence similarity to the first portion of SEQ ID NO: 1, and the second fragment of the sp polypeptide has 100% sequence similarity to the second portion of SEQ ID NO: 1.
[0118] In some embodiments, the sp dehalogenase (e.g., spHT) includes an sp site. The sp site is an internal position within the parent sequence that defines the C-terminus of the first component or fragment and the N-terminus of the second component or fragment of the sp dehalogenase. For example, if a theoretical 100-amino acid polypeptide is split at the sp site (referred to herein as the sp site at 57) between residues 57 and 58 of the parent polypeptide, the first component polypeptide would correspond to positions 1 to 57 of SEQ ID NO: 1, and the second component polypeptide would correspond to positions 58 to 100 of SEQ ID NO: 1. In some embodiments herein, the sp site within SEQ ID NO: 1 can occur at any position from position 5 to position 290 of SEQ ID NO: 1. In some embodiments, SEQ ID NOs: 2 to 577 are exemplary components of an spHT polypeptide having 100% sequence identity to SEQ ID NO: 1. In some embodiments, the active spHT complex is formed between two fragments that together contain the amino acids corresponding to each position of SEQ ID NO: 1. For example, a polypeptide having the sequence of SEQ ID NO: 26 and a peptide having the sequence of SEQ ID NO: 27 together contain the amino acids corresponding to each position of SEQ ID NO: 1. Any pair of a peptide and a polypeptide (or two polypeptides) corresponding to two of SEQ ID NOs: 2 to 577 and together containing the amino acids corresponding to each position of SEQ ID NO: 1 are used in the embodiments herein. In some embodiments, the spHT dehalogenase includes any of the following fragment pairs: SEQ ID NOs: 2 and 3, 4 and 5, 6 and 7, 8 and 9, 10 and 11, 12 and 13, 14 and 15, 16 and 17, 18 and 19, 20 and 21, 22 and 23, 24 and 25, 26 and 27, 28 and 29, 30 and 31, 32 and 33, 34 and 35, 36 and 37, 38 and 39, 40 and 41, 42 and 43, 44 and 45, 46 and 47, 48 and 49, 50 and 51, 52 and 53, 54 and 55, 56 and 57, 58 and 59, 60 and 61, 62 and 63, 64 and 65, 66 and 67, 68 and 69, 70 and 71, 72 and 73, 74 and 75, 76 and 77, 78 and 79, 80 and 81, 82 and 83, 84 and 85, 86 and 87, 88 and 89, 90 and 91, 92 and 93, 94 and 95, 96 and 97, 98 and 99, 100 and 101,102 and 103, 104 and 105, 106 and 107, 108 and 109, 110 and 111, 112 and 113, 114 and 115, 116 and 117, 118 and 119, 120 and 121, 121, 122 and 123, 124 and 125, 126 and 127, 128 and 129, 130 and 131, 132 and 133, 134 and 135, 136 and 137, 138 and 139, 140 and 141, 142 and 143, 144 and 145, 146 and 147, 148 and 149, 150 and 151, 152 and 153, 154 and 155, 156 and 157, 158 and 159, 160 and 161, 172 and 173, 174 and 175, 176 and 177, 178 and 179, 180 and 181, 182 and 183, 184 and 185, 186 and 187, 188 and 189, 190 and 191, 192 and 193, 194 and 195, 196 and 197, 198 and 199, 200 and 201, 202 and 203, 204 and 205, 206 and 207, 208 and 209, 190 and 211, 212 and 213, 214 and 215, 216 and 217, 218 and 219, 220 and 221, 222 and 223, 224 and 225, 226 and 227, 228 and 229, 300 and 301, 302 and 303, 304 and 305, 306 and 307, 308 and 309, 310 and 311, 312 and 313, 314 and 315, 316 and 317, 318 and 319, 320 and 321, 322 and 323, 324 and 325, 326 and 327, 328 and 329, 330 and 331, 332 and 333, 334 and 335, 336 and 337, 338 and 339, 340 and 341, 342 and 343, 344 and 345, 346 and 347, 348 and 349, 350 and 351, 352 and 353, 354 and 355, 356 and 357, 358 and 359, 360 and 361, 362 and 363, 364 and 365, 366 and 367, 368 and 369, 370 and 371, 372 and 373, 374 and 375, 376 and 377, 378 and 379, 380 and 381, 382 and 383, 384 and 385, 386 and 387, 388 and 389, 390 and 391, 392 and 393, 394 and 395, 396 and 397, 398 and 399, 400 and 401,402 and 403, 404 and 405, 406 and 407, 408 and 409, 410 and 411, 412 and 413, 414 and 415, 416 and 417, 418 and 419, 420 and 421, 422 and 423, 424 and 425, 426 and 427, 428 and 429, 430 and 431, 432 and 433, 434 and 435, 436 and 437, 438 and 439, 440 and 441, 442 and 443, 444 and 445, 446 and 447, 448 and 449, 450 and 451, 452 and 453, 454 and 455, 456 and 457, 458 and 459, 460 and 461, 462 and 463, 464 and 465, 466 and 467, 468 and 469, 470 and 471, 472 and 473, 474 and 475, 476 and 477, 478 and 479, 480 and 481, 482 and 483, 484 and 485, 486 and 487, 488 and 489, 490 and 491, 492 and 493, 494 and 495, 496 and 497, 498 and 499, 500 and 501, 502 and 503, 504 and 505, 506 and 507, 508 and 509, 510 and 511, 512 and 513, 514 and 515, 516 and 517, 518 and 519, 520 and 521, 522 and 523, 524 and 525, 526 and 527, 528 and 529, 530 and 531, 532 and 533, 534 and 535, 536 and 537, 538 and 539, 540 and 541, 542 and 543, 544 and 545, 546 and 547, 548 and 549, 550 and 551, 552 and 553, 554 and 555, 556 and 557, 558 and 559, 560 and 561, 562 and 563, 564 and 565, 566 and 567, 568 and 569, 570 and 571, 572 and 573, 574 and 575, and 576 and 577.
[0119] In some embodiments, spHT includes both a peptide and a polypeptide (or two polypeptides) pair corresponding to two of SEQ ID NOs: 2 to 577, and includes the amino acids corresponding to each position of SEQ ID NO: 1, but has a deletion at the C-terminus or N-terminus of one or both of the fragments, with a length of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, or a range therebetween). For example, the pair corresponding to SEQ ID NOs: 7 and 28 both correspond to the positions of SEQ ID NO: 1 but have an 11-residue deletion. In some embodiments, any pair of SEQ ID NOs: 2 to 577 that both correspond to the sequence of SEQ ID NO: 1 but have a deletion of up to 40 amino acids is within the scope of spHT herein. In some embodiments, the deletion is adjacent to the split site. In some embodiments, the deletion corresponds to the N-terminus or C-terminus of SEQ ID NO: 1.
[0120] In some embodiments, spHT includes both a peptide and a polypeptide (or two polypeptides) pair corresponding to two of SEQ ID NOs: 2 to 577, and includes the amino acids corresponding to each position of SEQ ID NO: 1, but has a duplication at the C-terminus or N-terminus of one or both of the fragments, with a length of up to 40 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, or a range therebetween). For example, the pair corresponding to SEQ ID NOs: 6 and 29 both correspond to the positions of SEQ ID NO: 1 but have an 11-residue duplication. In some embodiments, any pair of SEQ ID NOs: 2 to 577 that both correspond to the sequence of SEQ ID NO: 1 but have a duplication of up to 40 amino acids is within the scope of spHT herein. In some embodiments, the duplication is adjacent to the split site. In some embodiments, the duplication corresponds to the N-terminus or C-terminus of SEQ ID NO: 1.
[0121] For example, a fragment that utilizes any sp site corresponding to a position between position 5 and position 290 of SEQ ID NO: 1 is readily envisioned and is within the scope of this specification.
[0122] In some embodiments, spHT includes the 5th, 6th, 7th, 8th, 9th, 10th, 11th, 12th, 13th, 14th, 15th, 16th, 17th, 18th, 19th, 31st, 21st, 22nd, 23rd, 24th, 25th, 26th, 27th, 28th, 29th, 30th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, 39th, 40th, 41st, 42nd, 43rd, 44th, 45th, 46th, 47th, 48th, 49th, 50th, 51st, 52nd, 53rd, 54th, 55th, 56th, 57th, 58th, 59th, 60th, 61st, 62nd, 63rd, 64th, 65th, 66th, 67th, 68th, 69th, 70th, 71st, 72nd, 73rd, 74th, 75th, 76th, 77th, 78th, 79th, 80th, 81st, 82nd, 83rd, 84th, 85th, 86th, 87th, 88th, 89th, 90th, 91st, 92nd, 93rd, 94th, 95th, 96th, 97th, 98th, 99th, 100th, 101st, 102nd, 313th, 104th, 105th, 106th, 107th, 108th, 109th, 110th, 111th, 112th, 113th, 114th, 115th, 116th, 117th, 118th, 119th, 120th, 121st, 122nd, 123rd, 124th, 125th, 126th, 127th, 128th, 129th, 130th, 131st, 132nd, 133rd, 134th, 135th, 136th, 137th, 138th, 139th, 140th, 141st, 142nd, 143rd, 144th, 145th, 146th, 147th, 148th, 149th, 150th, 151st, 152nd, 153rd, 154th, 155th, 156th, 157th, 158th, 159th, 160th, 161st, 162nd, 163rd, 164th, 165th, 166th, 167th, 168th, 169th, 170th, 171st, 172nd, 173rd, 174th, 175th, 176th, 177th, 178th, 179th, 180th, 181st, 182nd, 183rd, 184th, 185th, 186th, 187th, 188th, 189th, 190th, 191st, 192nd, 193rd, 194th, 195th, 196th, 197th, 198th, 199th, 310th, 311th, 312th, 313th, 314th, 315th, 316th, 317th, 318th, 319th, 210th, 211th, 212th, 213th, 214th, 215th, 216th, 217th, 218th, 219th, 220th, 221st, 222nd, 223rd, 224th, 225th, 226th, 227th, 228th, 229th, 230th, 231st, 232nd, 233rd, 234th, 235th, 236th, 237th, 238th, 239th, 240th, 241st, 242nd, 243rd, 244th, 245th, 246th, 247th, 248th, 249th, 250th, 251st, 252nd, 253rd, 254th, 255th, 256th, 257th, 258th, 259th, 260th, 261st, 262nd, 263rd, 264th, 265th, 266th, 267th, 268th, 269th, 270th, 271st, 272nd, 273rd of SEQ ID NO: 1An sp site corresponding to position 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290 is provided.
[0123] In some embodiments, spHT has an sp site corresponding to a position between position 5 and 13, position 36 and 51, position 63 and 72, position 84 and 92, position 104 and 130, position 142 and 148, position 160 and 174, position 186 and 189, position 311 and 313, position 221 and 229, or position 269 and 290 of SEQ ID NO: 1.
[0124] In some embodiments, sp peptides and polypeptides having 70% - 100% sequence identity (e.g., more than 70% sequence identity, more than 75% sequence identity, more than 80% sequence identity, more than 85% sequence identity, more than 90% sequence identity, more than 95% sequence identity, more than 96% sequence identity, more than 97% sequence identity, more than 98% sequence identity, more than 99% sequence identity) to one of SEQ ID NOs: 2 - 557 are provided. In some embodiments, sp peptides and polypeptides having 70% - 100% sequence similarity (e.g., more than 70% sequence similarity, more than 75% sequence similarity, more than 80% sequence similarity, more than 85% sequence similarity, more than 90% sequence similarity, more than 95% sequence similarity, more than 96% sequence similarity, more than 97% sequence similarity, more than 98% sequence similarity, more than 99% sequence similarity) to one of SEQ ID NOs: 2 - 557 are provided.
[0125] In some embodiments, pairs of sp peptides and / or polypeptides that can form an active sp dehalogenase complex (active spHT complex) are provided. Such pairs include at least 70% sequence identity or similarity to two of SEQ ID NOs: 2 - 557, both contain residues corresponding to 100% of the positions of SEQ ID NO: 1, and allow for up to 40 deletions or duplications at the C - terminus or N - terminus of the peptide / polypeptide.
[0126] In some embodiments, the first fragment of the spHT complementary pair is positions 1 to 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 31, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 313, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270 of SEQ ID NO: 1correspond to the 271st, 272nd, 273rd, 274th, 275th, 276th, 277th, 278th, 279th, 280th, 281st, 282nd, 283rd, 284th, 285th, 286th, 287th, 288th, 289th, or 290th position.,
[0127] In some embodiments, the second fragment of the spHT complementary pair is set forth in SEQ ID NO: 1 at positions 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 31, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 313, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272It corresponds to positions 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or positions 290 to 294 of SEQ ID NO: 1.,
[0128] In some embodiments, the overlapping portion of the spHT complementary pair is 1 to 40 amino acids in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 31, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, or in the range between them).
[0129] In some embodiments, the deleted portion of the spHT complementary pair is 1 to 40 amino acids in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 31, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, or in the range between them).
[0130] The exemplary spHT fragment sequences of SEQ ID NOs: 2 to 577 contain 100% sequence identity to the portion of SEQ ID NO: 1, and there is no portion of these sequences that does not align with 100% sequence identity to SEQ ID NO: 1. However, as described herein, spHT peptides and polypeptides can have less than 100% sequence identity to SEQ ID NO: 1 (e.g., greater than 70%, greater than 75%, greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, but less than 100% sequence identity). Accordingly, peptides and polypeptides having less than 100% sequence identity to one of SEQ ID NOs: 2 to 577 (e.g., greater than 70%, greater than 75%, greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, but less than 100% sequence identity) are provided herein and are used in the complementary pairs and complexes herein.
[0131] In some embodiments, the spHT complementary pair of the present specification includes a peptide corresponding to SEQ ID NO: 578 and a polypeptide corresponding to SEQ ID NO: 1188. SEQ ID NOs: 578 and 1188 are fragments of SEQ ID NO: 1 and have 100% sequence identity with a portion of SEQ ID NO: 1. In some embodiments, the spHT complementary pair includes a peptide having 100% sequence identity with SEQ ID NO: 578, and such a peptide is referred to herein as "SmHT". In some embodiments, the spHT complementary pair includes a polypeptide having 100% sequence identity with SEQ ID NO: 1188, and such a polypeptide is referred to herein as "LgHT". During the development of the embodiments of the present specification, a wide range of experiments were conducted to analyze variants of SmHT and LgHT. SEQ ID NOs: 579-1187 correspond to peptide variants in which at least one and up to all positions of SEQ ID NO: 588 are substituted. Each of the peptides of SEQ ID NOs: 578-1187 was synthesized and tested for various properties, including the ability to form an active complex with a complementary LgHT variant polypeptide. SEQ ID NOs: 1189-3033 correspond to polypeptide variants having one or more substitutions relative to SEQ ID NO: 1188. Each of the polypeptides of SEQ ID NOs: 1188-3033 was synthesized and tested for various properties, including the ability to form an active complex with a complementary SmHT variant peptide.
[0132] In some embodiments, provided herein is a SmHT peptide or SmHT variant peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence similarity (e.g., conservative or semi-conservative similarity) to one of SEQ ID NOs: 578 - 1187. In some embodiments, the peptide corresponds to SmHT (SEQ ID NO: 578), but has one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or a range therebetween) of one or more substitutions of SEQ ID NOs: 588 - 1187 as compared to SEQ ID NO: 578. In some embodiments, the SmHT variant has 1 - 8 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, or a range therebetween) non-conservative substitutions as compared to one of SEQ ID NOs: 578 - 1187.
[0133] In some embodiments, provided herein is X 1 X 2 X 3 X 4 X 5 (F / W / Y / M / H)X 7 (F / W / Y / D / R)X 9 X 10 X 11 (F / W / Y / M / H / R)(V / I / L / M / A / C)X 14 (V / I / L / A / C / MI / L / F / W)X 16 X 17 (SEQ ID NO: 3034), and / or X 1 X 2 X 3 X 4 X 5 (F / W / Y)X 7 (F / W / Y)X 9 X 10 X 11 (F / W / Y)(V / I / L / M)X 14 (V / I / L)X 16 X 17 a SmHT peptide or SmHT variant peptide comprising (SEQ ID NO: 3035), Here, each X is an arbitrary amino acid (e.g., a protein - constituting amino acid).
[0134] In some embodiments, provided herein is an LgHT peptide or an LgHT variant peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence similarity (e.g., conservative or semi - conservative similarity) to one of SEQ ID NOs: 1188 - 3033. In some embodiments, the polypeptide corresponds to LgHT (SEQ ID NO: 1188), but has one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more, or a range therebetween) of the substitutions of one or more of SEQ ID NOs: 1189 - 3033 compared to SEQ ID NO: 1188. In some embodiments, the LgHT variant has at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 1188 - 3033.
[0135] In some embodiments, provided herein is a spHT complementary pair comprising: (a) a SmHT peptide or SmHT variant peptide having (1) at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence similarity (e.g., conserved or semi-conserved similarity) to one of SEQ ID NOs: 578-1187, (2) one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or a range therebetween) substitutions relative to SEQ ID NO: 578, and / or (3) 1-8 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, or a range therebetween) non-conserved substitutions relative to one of SEQ ID NOs: 578-1187; and (b) an LgHT polypeptide or LgHT variant polypeptide having (1) at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence similarity (e.g., conserved or semi-conserved similarity) to one of SEQ ID NOs: 1188-3033, (2) one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more, or a range therebetween) substitutions relative to SEQ ID NO: 1188, and / or (3) at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 1188-3033.
[0136] In some embodiments, split hydrolases (e.g., spHT) and fragments thereof have enhanced thermal stability compared to the parent hydrolase sequence (e.g., HALOTAG).
[0137] The formation of the spHT complex from two complementary fragments can be reversible or irreversible. In some embodiments, the spHT complex can be denatured, renatured, and its activity reconstituted. In some embodiments, such spHT is used in a method that includes exposing a sample containing spHT to denaturing conditions (e.g., manufacturing conditions, storage conditions, etc.) prior to substrate binding.
[0138] In some embodiments, provided herein are fusions of split hydrolases (e.g., dehalogenases (e.g., HALOTAG, etc.)) with proteins of interest, interaction elements, localization elements, heterologous sequences, peptide tags, luciferases, or bioluminescent complexes, etc.
[0139] In certain embodiments, both fragments of a split hydrolase (e.g., spHT) are fused to heterologous sequences. In some embodiments, the heterologous sequences are substantially identical and optionally specifically bind to each other in the absence of one or more exogenous agents, e.g., to form a dimer. In another embodiment, the heterologous sequences are different and optionally specifically bind to each other in the absence of one or more exogenous agents. In one embodiment, one hydrolase fragment is fused to a heterologous sequence that interacts with a cellular molecule. In another embodiment, each hydrolase fragment is fused to a heterologous sequence and the heterologous sequences interact in the presence of one or more exogenous agents or under certain conditions. For example, in the presence of rapamycin, a fragment of a hydrolase fused to the rapamycin-binding protein (FRB) and another fragment fused to the FK506-binding protein (FKBP) result in a complex of the two fusion proteins. In one embodiment, no complex of the fusion proteins is formed in the presence of an exogenous agent(s) or under different conditions. In one embodiment, one heterologous sequence contains a domain, e.g., three or more amino acid residues, which may optionally be covalently modified, e.g., phosphorylated, and non-covalently interacts with a domain in another heterologous sequence. Two fragments of the hydrolase, at least one of which is fused to a protein of interest, can be used to detect reversible interactions, e.g., binding of two or more molecules, or other conformational or condition changes such as changes in pH, temperature or solvent hydrophobicity, or irreversible interactions.
[0140] The rapamycin / FRB / FKBP system provides an example of a small molecule that induces a protein-protein interaction that can be detected / monitored by the spHT system of the present specification. However, other systems that induce the formation of the spHT complex are within the scope of the present specification. Other small molecule-induced protein interactions are used in the embodiments of the present specification. Additionally, proteins interact (i.e., associate or dissociate) as a result of other events within the cell that affect their local concentrations, such as additive / subtractive abundance caused by direct physical association, co-localization, stabilization or degradation stimuli, additive / subtractive abundance controlled at the gene level (i.e., upregulation, downregulation). Embodiments of the present specification can be used in monitoring such effects in vitro and in vivo.
[0141] Heterologous sequences useful in the present invention include, but are not limited to, those that interact in vitro and / or in vivo. For example, a fusion protein may include (1) a hydrolase fragment (e.g., a portion of spHT) and (2) an enzyme of interest, such as luciferase, RNasin or RNase, and / or a channel protein, receptor, membrane protein, cytoplasmic protein, nuclear protein, structural protein, phosphoprotein, kinase, signaling protein, metabolic protein, mitochondrial protein, receptor-related protein, fluorescent protein, enzyme substrate, transcription factor, transporter protein, and / or a target sequence, such as a myristoylation sequence, a mitochondrial localization sequence, or a nuclear localization sequence, and the hydrolase fragment, e.g., the fusion protein, directs to a specific location. The protein of interest fused to the hydrolase fragment can be a fragment of a wild-type protein, e.g., a functional or structural domain of a protein such as a kinase domain, a transcription factor, etc. The protein of interest can be fused to the N-terminus or C-terminus of the fragment (e.g., a portion of spHT). In one embodiment, the fusion protein includes the protein of interest at the N-terminus and another protein, e.g., a different protein, at the C-terminus of the fragment (e.g., a portion of spHT). For example, the protein of interest can be an antibody. Optionally, the proteins in the fusion can be separated by a linker, e.g., a linker sequence of 1 to 20 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 acid residues). In some embodiments, the presence of the linker in the fusion proteins of the present invention does not substantially alter the function of any of the proteins in the fusion as compared to the function of each individual protein. A wide variety of linkers can be used for any particular combination of proteins in the fusion. In one embodiment, the linker is a sequence recognized by an enzyme, e.g., a cleavable sequence, or a photocleavable sequence.
[0142] Exemplary heterologous arrays include, but are not limited to, the arrays in FRB and FKBP, the regulatory subunit of protein kinase (PKa-R) and the catalytic subunit of protein kinase (PKa-C), the src homology region (SH2) and phosphorylatable arrays, for example, tyrosine-containing arrays, 14-3-3, for example, the isoform of 14-3-3t (see Mils et al., 3100), and phosphorylatable arrays, proteins having a WW region (sequences of proteins that bind to proline-rich molecules (see Ilsley et al., 3102, and Einbond et al., 1996), and phosphorylatable heterologous arrays, for example, serine and / or threonine-containing arrays, as well as the arrays in dihydrofolate reductase (DHFR) and gyrase B (GyrB).
[0143] As described throughout, the spHT peptides and polypeptides provided herein are used as part of fusion proteins with a peptide, polypeptide, antibody, antibody fragment, and protein of interest. For example, the present invention provides a fusion protein comprising (1) an spHT peptide or polypeptide, and (2) the amino acid sequence of a protein or peptide of interest, such as a marker protein, such as a selectable marker protein, an enzyme of interest, such as luciferase, RNasin, RNase, and / or the sequence of GFP, a nucleic acid binding protein, an extracellular matrix protein, a secreted protein, an antibody or portion thereof, such as Fc, a bioluminescent protein, a receptor ligand, a regulatory protein, a serum protein, an immunogenic protein, a fluorescent protein, a protein having a reactive cysteine, a receptor protein, such as an NMDA receptor, an ion channel protein, such as a sodium, potassium, or calcium-sensitive channel protein, including a channel protein such as the HERG channel protein, a membrane protein, a cytoplasmic protein, a nuclear protein, a structural protein, a phosphoprotein, a kinase, a signaling protein, a metabolic protein, a mitochondrial protein, a receptor-related protein, a fluorescent protein, an enzyme substrate, such as a protease substrate, a transcription factor, a protein destabilization sequence, or a transporter protein, such as an EAAT1-4 glutamate transporter, and a targeting signal that directs the fusion to a specific location, such as a mitochondrial localization sequence, a nuclear localization signal, or a plastid targeting signal such as a myristoylation sequence.
[0144] In some embodiments, the fusion protein comprises (1) an spHT peptide or polypeptide, and (2) a protein that associates with a membrane or portion thereof, such as a targeting protein, such as one for endoplasmic reticulum targeting, a cell membrane binding protein, such as an integrin protein or domain thereof, such as the cytoplasmic, transmembrane and / or extracellular stem domains of an integrin protein, and / or a protein that links a mutant hydrolase to the cell surface, such as a glycosylphosphatidylinositol signal sequence.
[0145] The fusion partner can include those having enzyme activity. For example, a functional protein sequence can encode a kinase catalytic domain (Hanks and Hunter, 1995), can produce a fusion protein capable of enzymatically adding a phosphate moiety to a specific amino acid, or can encode an Src homology 2 (SH2) domain (Sadowski et al., 1986, Mayer and Baltimore, 1993) and can produce a fusion protein that specifically binds to phosphorylated tyrosine.
[0146] In some embodiments, the fusion includes an affinity domain that includes a peptide sequence that can interact with a binding partner, such as one immobilized on a solid support useful for identification or purification. A DNA sequence encoding a plurality of consecutive single amino acids, such as histidine, when fused to an expressed protein, can be used for one-step purification of a recombinant protein by binding with high affinity to a resin column such as nickel sepharose. Exemplary affinity domains include HisV5 (HHHHH) (SEQ ID NO: 13), HisX6 (HHHHHH) (SEQ ID NO: 3), C-myc (EQKLISEEDL) (SEQ ID NO: 4), Flag (DYKDDDDK) (SEQ ID NO: 5), SteptTag (WSHPQFEK) (SEQ ID NO: 6), hemagglutinin, such as the HA tag (YPYDVPDYA) (SEQ ID NO: 7), GST, thioredoxin, cellulose binding domain, RYIRS (SEQ ID NO: 8), Phe-His-His-Thr (SEQ ID NO: 9), chitin binding domain, S-peptide, T7 peptide, SH2 domain, C-end RNA tag, WEAAAREACCRECCARA (SEQ ID NO: 10), metal binding domain, such as a zinc binding domain, or a calcium binding domain from a calcium binding protein, such as calmodulin, troponin C, calcineurin B, myosin light chain, recoverin, S-modulin, visinin, VILIP, neurocalcin, hippocalcin, frequenin, caltractin, calpain large subunit, S100 protein, parvalbumin, calbindin, D 9K, Calbindin D 28K , and calretinin, intein, biotin, streptavidin, MyoD, Id, leucine zipper sequence, and maltose binding protein.
[0147] In some embodiments, the split hydrolase fragment (e.g., spHT) described herein is fused to a reporter protein. In some embodiments, the reporter is a bioluminescent reporter (e.g., expressed as a fusion protein with spHT). In certain embodiments, the bioluminescent reporter is luciferase. In some embodiments, the luciferase is selected from those found in Omphalotus olearius, fireflies (e.g., Photinini), Renilla reniformis, Aequoria, variants thereof, portions thereof, variants thereof, and any other luciferase enzyme suitable for the systems and methods described herein. In some embodiments, the bioluminescent reporter is a modified and enhanced luciferase enzyme from Oplophorus (e.g., the NANOLUC enzyme of Promega Corporation, SEQ ID NO: 3, or a sequence having at least 70% identity thereto (e.g., greater than 70%, greater than 80%, greater than 90%, greater than 95%)). Exemplary bioluminescent reporters are described, for example, in U.S. Patent Application No. 2010 / 0281552 and U.S. Patent Application No. 2012 / 0174242, both of which are incorporated herein by reference in their entirety.
[0148] In some embodiments, the split hydrolase fragments described herein (e.g., spHT) are fused to peptide or polypeptide components of commercially available NanoLuc®-based technologies (e.g., NanoLuc® luciferase, NanoBiT, NanoTrip, NanoBRET, etc.). PCT Application No. PCT / US2010 / 033449, U.S. Patent No. 8,557,970, PCT Application No. PCT / 2011 / 059018, and U.S. Patent No. 8,669,103 (each of which is hereby incorporated by reference in its entirety for all purposes) describe compositions and methods comprising bioluminescent polypeptides for use as heterologous sequences in the fusions herein. Such polypeptides are used in the embodiments herein and can be used in conjunction with the compositions and methods described herein. PCT Application No. PCT / US14 / 26354 and U.S. Patent No. 9,797,889 (each of which is hereby incorporated by reference in its entirety for all purposes) describe compositions and methods for the assembly of bioluminescent complexes, and such complexes, as well as their peptide and polypeptide components, are used as heterologous sequences in the embodiments herein and can be used in conjunction with the compositions and methods described herein. In some embodiments, NanoBiT and other related technologies utilize peptide and polypeptide components upon assembly into a complex and exhibit significantly enhanced (e.g., 2-fold, 5-fold, 10-fold, 10 2 -fold, 10 3 -fold, 10 4 -fold, or more) luminescence in the presence of an appropriate substrate (e.g., coelenterazine or a coelenterazine analog) as compared to the peptide and polypeptide components alone. In some embodiments, the NanoBiT peptides and polypeptides are fused to the spHT fragments herein. U.S. Patent Publication No. 2020 / 0270586 and International Application No. PCT / US19 / 36844 (each of which is hereby incorporated by reference in its entirety for all purposes) describe multi-part luciferase complexes (e.g., NanoTrip) for use as heterologous sequences in the embodiments herein and can be used in conjunction with the compositions and methods described herein.
[0149] In some embodiments, the sp dehalogenase is used with a split reporter. In some embodiments, a fragment of the sp dehalogenase is linked (e.g., fused, conjugated, etc.) to a fragment of the split reporter. When the two entities are combined, an active dehalogenase and an active reporter are formed. Examples of split fluorescent protein reporters include split GFP and split mCherry. In other embodiments, a first fragment of the split reporter (e.g., split fluorescent protein, split luciferase, etc.) is fused to a first fragment of the sp dehalogenase, and a second fragment of the split reporter is conjugated to a haloalkane substrate. In such embodiments, when an active dehalogenase complex is formed, the complex binds to the haloalkane substrate and an active reporter complex is assembled. In some embodiments, a fragment of the sp dehalogenase and / or the haloalkane is fused to another split protein such as a split TEV protease or other enzyme.
[0150] Also, assume that split HaloTag fragments are used in a "dual tag" configuration, where the split fragments of HaloTag are combined with split fragments of luciferase, fluorescent protein, or other labels / reporters (including SpyCatcher). For example, HiBiT-spHaloTag fragment tags, or GFP11-spHaloTag fragment tags. More generally, split versions of other enzyme classes such as split TEV protease exist, and these can likewise be created in these "dual tag" configurations.
[0151] As described herein, the spHT system of the present specification utilizes haloalkane substrates. In some embodiments, the substrate has the formula (I): R - linker - A - X, where R is a solid surface, one or more functional groups, or absent, the linker is a polyatomic straight or branched chain containing C, N, S, or O, or a group containing one or more rings, such as one or more aryl rings, heteroaryl rings, or any combination thereof, saturated or unsaturated rings, where A - X is a substrate of the dehalogenase, hydrolase, HALOTAG, or spHT system of the present specification (e.g., A is (CH 2 ) 4~ 20 and X is a halide (e.g., Cl or Br)). Suitable substrates are described, for example, in U.S. Patent Nos. 11,072,812, 11,028,424, 10,618,907, and 10,101,332 (which are incorporated herein by reference in their entirety).
[0152] In some embodiments, R is one or more functional groups (such as fluorophores, biotin, chromophores, or fluorescence-generating or luminescent molecules). Exemplary functional groups for use in the present invention include, but are not limited to, amino acids, proteins such as enzymes, antibodies or other immunogenic proteins, radionuclides, nucleic acid molecules, drugs, lipids, biotin, avidin, streptavidin, magnetic beads, solid supports, electron-opaque molecules, chromophores, MRI contrast agents, dyes such as xanthene dyes, calcium-sensitive dyes such as 1-[2-amino-5-(2,7-dichloro-6-hydroxy-3-oxo-9-xanthenyl)-phenoxy]-2-(2'-amino-5'-methylphenoxy)ethane-N,N,N',N'-tetraacetic acid (Fluo-3), sodium-sensitive dyes such as 1,3-benzenedicarboxylic acid, 4,4'-[1,4,10,13-tetraoxa-7,16-diazacyclooctadecane-7,16-diybis(5-methoxy-6,2-benzofurandil)bis(PBFI), NO-sensitive dyes such as 4-amino-5-methylamino-2',7'-difluorescein, or other fluorophores. In one embodiment, the functional group is an immunogenic molecule, i.e., one that is bound by an antibody specific for that molecule.
[0153] In some embodiments, the substrate of the present invention is permeable to the plasma membrane of cells (i.e., it can pass from the outside of the cell (e.g., eukaryotic, prokaryotic) into the cell interior without chemical, enzymatic, or mechanical disruption of the cell membrane).
[0154] In some embodiments, the substrate herein includes a cleavable linker, such as those described in U.S. Patent No. 10,618,907, which is incorporated herein by reference in its entirety.
[0155] In some embodiments, the substrate comprises a fluorescent functional group (R).Suitable fluorescent functional groups include, but are not limited to, stilbazolidium derivatives (Marquesa et al. Mechanism-Based Strategy for Optimizing HaloTag Protein Labeling. ChemRxiv. Cambridge: Cambridge Open Engage; 2021; incorporated herein by reference in its entirety), xanthene derivatives (e.g., fluorescein, rhodamine, Oregon Green, eosin, Texas Red, etc.), cyanine derivatives (e.g., cyanine, indocarbocyanine, oxacarbocyanine, thiacarbocyanine, merocyanine, etc.), naphthalene derivatives (e.g., dansyl and prodan derivatives), oxadiazole derivatives (e.g., pyridyloxazole, nitrobenzoxadiazole, benzoxadiazole, etc.), pyrene derivatives (e.g., cascade blue), oxazine derivatives (e.g., Nile Red, Nile Blue, cresyl violet, oxazine 170, etc.), acridine derivatives (e.g., proflavine, acridine orange, acridine yellow, etc.), arylmethine derivatives (e.g., auramine, crystal violet, malachite green, etc.), tetrapyrrole derivatives (e.g., porphyrin, phthalocyanine, bilirubin, etc.), CF dyes (Biotium), BODIPY (Invitrogen), ALEXA FLOUR (Invitrogen), DYLIGHT FLUOR (Thermo Scientific, Pierce), ATTO and TRACY (Sigma Aldrich), FluoProbes (Interchim), DY and MEGASTOKES (Dyomics), SULFO CY dyes (CYANDYE, LLC), SETAU and SQUARE dyes (SETA BioMedicals), QUASAR and CAL FLUOR dyes (Biosearch Technologies), SURELIGHT dyes (APC, RPE, PerCP, phycobilisome) (Columbia Biosciences), APC, APCXL, RPE, BPE (Phyco-Biotech), autofluorescent proteins (e.g., YFP, RFP, mCherry, mKate), quantum dot nanocrystals, and the like.
[0156] In some embodiments, the substrate comprises a fluorogenic functional group (R). The fluorogenic functional group is a functional group that generates and enhances a fluorescent signal upon binding of the substrate to a target (e.g., binding of a haloalkane to a modified dehalogenase). By generating a significantly increased fluorescence (e.g., 10-fold, 31-fold, 50-fold, 100-fold, 310-fold, 500-fold, 1000-fold, or more) upon target engagement, the problem of background signal is reduced. Exemplary fluorogenic dyes for use in the embodiments herein include fluorophores of the JANELIA FLUOR family such as: JANELIA FLUOR 549:
Chem.
Chem.
Chem.
Chem.
Chem.
[0157] In some embodiments, a “dual warhead” substrate is provided that includes a haloalkane moiety (e.g., a substrate for a modified dehalogenase such as HALOTAG) and a dimerization moiety that is a ligand (or capture element) for a second binding protein (capture element). For example, certain embodiments herein include a haloalkane linked to a SNAP tag ligand (Figure 15A; Cermakova & Hodges. Molecules 2018, 23(8), 1958, which is incorporated by reference in its entirety), a haloalkane linked to cTMP (Figure 15B; Cermakova & Hodges. Molecules 2018, 23(8), 1958, which is incorporated by reference in its entirety), a haloalkane linked to a rapamycin-like moiety that can bind to FKBP or FRB (Chen et al. ACS Chem. Biol. 2021, 16, 12, 2808-2815, which is incorporated by reference in its entirety), or other haloalkane “dual warhead” ligands that can bind to a modified dehalogenase (such as HALOTAG) and a second capture agent. In such embodiments, a system is provided that includes a split modified dehalogenase (spHT), a dual warhead substrate, and a capture agent (such as FKBP, FRB, SNAP tag, eDHFR, etc.) that can bind to the dimerization moiety. In some embodiments, one or both fragments of the capture agent and / or the split modified dehalogenase (spHT) are provided as fusions with a protein of interest. In some embodiments, the dual warhead ligand triggers dimerization of (1) the split modified dehalogenase (spHT) and (2) any element bound or fused to the capture agent. By adding another protein-binding small molecule moiety to the haloalkane, the dual warhead triggers proximity of the fusion partners to the split modified dehalogenase (spHT) and the capture agent. In some embodiments, the cell contains two proteins of interest, one tagged with a fragment of the split modified dehalogenase (spHT) and the other tagged with the capture agent, and in the presence of a dual warhead ligand that includes a haloalkane and a capture element of the capture agent, the tags dimerize and position the fusion proteins of interest in proximity.Such embodiments provide for the forced proximity of the fusion partners to these proteins within the cell and control the timing and extent of their co-localization via the addition of dual warhead ligands. Any suitable linker can be used in the assembly of the dual warhead substrate. The linker can include various combinations of such groups to provide linkers having ester (-C(O)O-), amide (-C(O)NH-), carbamate (-NHC(O)O-), urea (-NHC(O)NH-), phenylene (e.g., 1,4-phenylene), straight or branched alkylene, and / or oligo and polyethylene glycol (-(CH 2 CH 2 O) x -) linkages. In some embodiments, the linker may include two or more atoms (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 or 200 atoms, or any range therebetween (e.g., 2 - 20, 5 - 10, 15 - 35, 25 - 100, etc.) of 2 - 200 atoms). In some embodiments, the linker includes a combination of oligoethylene glycol linkages and carbamate linkages. In some embodiments, the linker has the formula -O(CH 2 CH 2 O) z1 -C(O)NH-(CH 2 CH 2 O) z2 -C(O)NH-(CH 2 ) z3 -(OCH 2 CH 2 ) z4 O-, wherein z1, z2, z3, and z4 are each independently selected in the form of 0, 1, 2, 3, 4, 5, and 6. For example, in some embodiments, the linker has a formula selected from:
Chemical formula
[0158] In some embodiments, the dual warheads used in the embodiments herein are haloalkanes linked to a ligand that can engage an E3 ubiquitin ligase (e.g., thalidomide, cereblon E3 ubiquitin ligase, von Hippel-Lindau (VHL) E3 ligase, or any other E3 ubiquitin ligase), also known as proteolysis targeting chimeras (PROTACs). Haloalkane PROTACs can bind to a modified dehalogenase or modified dehalogenase complex and an E3 ubiquitin ligase, and recruitment of the E3 ligase results in ubiquitination via the proteasome and subsequent degradation of the modified dehalogenase (complex) and any protein components fused thereto (e.g., the target protein). In some embodiments, the split dehalogenase system herein is used in assays / systems for measuring the kinetics of target protein ubiquitination, or in applications such as measuring the dose-response curve of a compound in an endpoint format. For example, in some embodiments, the target protein is expressed / provided in a sample as a fusion with a first component fragment of a split modified dehalogenase (e.g., spHT), the sample is contacted with a PROTAC of a haloalkane and a ligand that can engage an E3 ubiquitin ligase (e.g., thalidomide, cereblon E3 ubiquitin ligase, von Hippel-Lindau (VHL) E3 ligase or any other E3 ubiquitin ligase), and when a second component fragment of the split modified dehalogenase (e.g., spHT) is added, an active modified dehalogenase complex is formed, the haloalkane is bound by the complex that brings the ligand into proximity to the target protein, resulting in ubiquitination and inducing the fusion target to the proteasome for degradation.
[0159] In some embodiments, the components of the split dehalogenase have a high affinity for each other, such that the split dehalogenase complex is formed when the two components are in proximity to each other. The high affinity for the components of the split modified dehalogenase promotes the formation of the split dehalogenase complex and the degradation of the target protein. In such embodiments, the second component can be added to the system at a specific time to induce degradation and can be localized to a specific location or compartment (e.g., cell type, organelle, tissue, etc.) where degradation occurs or can be conditionally expressed. In other embodiments, the components of the split dehalogenase have a low affinity for each other and a second interaction is required to induce the formation of the split dehalogenase complex. For example, the second component of the split dehalogenase is fused to a protein that binds to the target protein or is linked to a ligand of the target protein. The binding of this component to the target protein enables the formation of the split dehalogenase complex, which can then bind to the haloalkane of the PROTAC and induce degradation.
[0160] In related embodiments, the target protein is expressed / provided in a sample as a fusion with (i) a first component fragment of a split modified dehalogenase (e.g., spHT) and (ii) a first interacting protein. The sample is contacted with a ligand capable of engaging a proteolysis targeting chimera (PROTAC) of a haloalkane and an E3 ubiquitin ligase (e.g., thalidomide), and when a fusion of a second component fragment of the split modified dehalogenase (e.g., spHT) and a second interacting protein is added, an active modified dehalogenase complex is formed (facilitated by the binding of the first and second interacting proteins), the haloalkane is bound by a complex that brings the ligase into proximity to the target protein, resulting in ubiquitination and directing the fusion target to the proteasome for degradation. In other embodiments, complex formation and subsequent degradation are monitored by fluorescence, bioluminescence, and / or BRET. For example, in certain embodiments, the target protein is expressed / provided in a sample as a fusion with luciferase (e.g., NANOLUC) or a component of a bioluminescence complex (e.g., a component of the NANOBIT system), the first component fragment of the split modified dehalogenase (e.g., spHT) is expressed / provided as a fusion with ubiquitin or an E3 ubiquitin ligase (e.g., thalidomide, cereblon E3 ubiquitin ligase, von Hippel-Lindau (VHL) E3 ligase, or any other E3 ubiquitin ligase), the sample is contacted with a bifunctional ligand that includes a molecule capable of binding to the haloalkane and the target protein, and when a second component fragment of the split modified dehalogenase (e.g., spHT) having a high affinity for the first component fragment is added, an active modified dehalogenase complex is formed, the haloalkane is bound by a complex that brings ubiquitin into proximity to the target protein, resulting in ubiquitination and directing the target to the proteasome for degradation and eliminating the signal from the luciferase. In similar embodiments, the components of the split modified dehalogenase are linked to fluorophores so that the system can be monitored using BRET between the target fusion and the split modified dehalogenase.
[0161] In other embodiments, the targeted chimeric (TAC) system can monitor the system by utilizing a haloalkane linked to a detectable moiety rather than as a functional component of the system. For example, a first component of a modified dehalogenase is fused to ubiquitin, a second component of the modified dehalogenase (e.g., having low affinity for the first component) is fused to a target protein, and the haloalkane is linked to a fluorophore or other detectable moiety. When ubiquitin is in proximity to the target protein, a modified dehalogenase complex is formed, the haloalkane is bound, and the complex is labeled at the detectable moiety.
[0162] In some embodiments, the split dehalogenase system herein is used in various other targeted chimeric (TAC) systems, such as a phosphorylation-targeted chimera (PhosTAC; Chen et al. ACS Chem. Biol. 3121, 16, 12, 2808-2815; incorporated herein by reference in its entirety) system, a deubiquitinase-targeted chimera (DUBTAC; Henning et al. Deubiquitinase-Targeting Chimeras for Targeted Protein Stabilization. bioRxiv; 2021. DOI: 10.1101 / 2021.04.30.441959; incorporated herein by reference in its entirety) system, a lysosome-targeted chimera (LyTAC; Banik et al. Nature 584, 291-297 (2020); incorporated herein by reference in its entirety) system, an autophagy-targeted chimera (AUTAC; Takahashi et al. Mol Cell. 2019 Dec 5;76(5):797-810.e10; incorporated herein by reference in its entirety) system, an autophagy-linked compound (ATTEC; Fu et al. Cell Research volume 31, pages 965-979 (2021); incorporated herein by reference in its entirety) system, and oligo-based TACs. Dual warheads comprising a haloalkane and a ligand for any of the above TAC systems can be used in the embodiments herein. For example, PhosTAC is similar to the well-described PROTACs in its ability to induce a ternary complex, and PhosTAC focuses on mobilizing a Ser / Thr phosphatase to a phospho-substrate to mediate its dephosphorylation. PhosTAC extends the use of PROTAC technology beyond proteolysis via ubiquitination to other post-translational modifications of proteins.For example, in some embodiments, the target protein is expressed / provided in a sample as a fusion with a first component fragment of a split modified dehalogenase (e.g., spHT), the sample is contacted with a phosphorylated targeting chimera (PhosTAC) of a haloalkane and a ligand capable of engaging a phosphatase enzyme, and when a second component fragment of a split modified dehalogenase (e.g., spHT) having a high affinity for the first component fragment is added, an active modified dehalogenase complex is formed, and the haloalkane is bound by the complex that brings the ligand into proximity to the target protein, resulting in phosphorylation of the target protein.
[0163] In some embodiments, the split dehalogenase system herein is used in other targeting chimera systems where a dual - function ligand comprising a haloalkane and a ligand for a mobilizable enzyme is combined with a fusion of a target protein and a fragment of spHT, and introduction of a second high - affinity spHT fragment into the system induces the enzymatic activity of the mobilizable enzyme towards the target protein.
[0164] Systems and methods that include any combination of the above TAC systems / assays are within the scope of this specification.
[0165] In some embodiments, provided herein is an isolated nucleic acid molecule (polynucleotide) comprising a nucleic acid sequence encoding a split hydrolase (e.g., spHT) fragment described herein. In some embodiments, such a polynucleotide comprises an open reading frame encoding spHT or a fragment thereof. In some embodiments, such a polynucleotide is within an expression vector or integrated into the genomic material of a cell. In some embodiments, such a polynucleotide further comprises regulatory elements such as a promoter. Further provided is an isolated nucleic acid molecule comprising a nucleic acid sequence encoding a fusion protein comprising an sp hydrolase fragment (e.g., spHT, etc.) and one or more amino acid residues at the N-terminus (N-terminal fusion partner) and C-terminus (C-terminal fusion partner). In one embodiment, the fusion protein comprises at least two different fusion partners (e.g., as described herein), one at the N-terminus and the other at the C-terminus, and one of the fusions is a sequence used for purification, such as a glutathione S-transferase (GST) or polyHis sequence, a sequence intended to alter the properties of the remainder of the fusion protein, such as a protein destabilizing sequence, or a sequence having distinguishable properties. In one embodiment, the isolated nucleic acid molecule comprises a nucleic acid sequence optimized for expression in at least one selected host. Optimized sequences include codon-optimized sequences, i.e., codons more frequently used in one organism relative to another organism, e.g., a distantly related organism, as well as modifications for adding or modifying Kozak sequences and / or introns, and / or modifications for removing undesirable sequences, e.g., potential transcription factor binding sites. In one embodiment, the polynucleotide comprises a nucleic acid sequence encoding a fragment of a dehalogenase, and this nucleic acid sequence is optimized for expression in a selected host cell. In one embodiment, the optimized polynucleotide no longer hybridizes to the corresponding non-optimized sequence, e.g., does not hybridize to the non-optimized sequence under moderate or high stringency conditions.In another embodiment, the polynucleotide has less than 90%, for example less than 80%, nucleic acid sequence identity to the corresponding non-optimized sequence and optionally encodes a polypeptide having at least 80%, for example at least 85%, 90% or more amino acid sequence identity to the polypeptide encoded by the non-optimized sequence.
[0166] Also provided are constructs, such as expression cassettes, and vectors containing isolated nucleic acid molecules, and host cells having one or more of the constructs, and kits containing isolated nucleic acid molecules, one or more constructs or vectors. Host cells include prokaryotic or eukaryotic cells, such as plant cells or vertebrate cells, such as mammalian cells including, but not limited to, human, non-human primate, dog, cat, bovine, equine, ovine, caprine, or rodent (e.g., rabbit, rat, ferret, or mouse) cells. In some embodiments, the expression cassette includes a promoter operably linked to the nucleic acid molecule, such as a constitutive promoter or a regulatable promoter. In some embodiments, the expression cassette includes an inducible promoter. In certain embodiments, the invention includes a vector containing a nucleic acid sequence encoding a fusion protein comprising a fragment of a dehalogenase. In some embodiments, an optimized nucleic acid sequence, such as a human codon-optimized sequence, encoding at least a fragment of a hydrolase, preferably a fusion protein comprising a fragment of a hydrolase, is used in the nucleic acid molecules of the invention. Optimization of nucleic acid sequences is known in the art; see, for example, WO02 / 16944, which is incorporated herein by reference in its entirety.
[0167] Also provided herein are cells containing split hydrolase fragment(s) (e.g., spHT), polynucleotides, expression vectors, etc. In some embodiments, the components described herein are expressed intracellularly. In some embodiments, the components of the present specification are introduced into cells via, for example, transfection, electroporation, infection, cell fusion, or any other means.
[0168] In some embodiments, the systems herein (e.g., including sp hydrolase (e.g., spHT, etc.)) can be used to measure or detect various target states and / or molecules. For example, protein-protein interactions are essential for substantially all aspects of cell biology, from gene transcription, protein translation, signal transduction, to cell division and differentiation. Protein complementation assay (PCA) is one of several methods used to monitor protein-protein interactions. In PCA, protein-protein interactions physically bring two non-functional halves of an enzyme into proximity with each other, which enables refolding into a functional enzyme. Thus, the interaction is monitored by enzyme activity. In protein complementation labeling (PCL), a covalent bond is created between the substrate and the complex, resulting in a cumulative label over time, thus increasing the sensitivity for detecting weak and / or rare protein-protein interactions. In a typical split enzyme system, when complementation is disrupted, signal generation is lost due to the absence or reduction of substrate conversion. However, in a split-label protein system (e.g., spHaloTag), due to the covalent nature of the label, it remains on the split protein even after complementation is disrupted. The latter demonstrated advantage is that for molecular events that occur periodically but are present in very low abundance (such as hormones that bind to neurotransmitters or receptors), the signal accumulates (covalently) over time and ultimately provides a signal sufficient to detect the event. This is difficult with split enzyme systems due to the rarity of events leading to low turnover of the enzyme signal.
[0169] In one embodiment, a vector encoding two complementary fragments of a mutant dehalogenase (e.g., spHT) fused to a protein of interest, or a vector encoding two complementary fragments of a mutant dehalogenase fused to a protein of interest, is introduced into a cell, cell lysate, in vitro transcription / translation mixture, or supernatant, and a hydrolase substrate labeled with a functional group (e.g., a haloalkane) is added thereto. The functional group is then detected or determined, for example, at one or more time points relative to a control sample.
[0170] In some embodiments, provided herein is a method for detecting an interaction between two proteins in a sample. The method includes providing a cell, a lysate of the cell, or an in vitro transcription / translation reaction having a plurality of expression vectors of the invention, and a sample having a hydrolase substrate (e.g., a haloalkane) having at least one functional group, under conditions effective to permit association of a first fusion protein and a second fusion protein. The presence, amount, or location of at least one functional group in the sample is detected.
[0171] In some embodiments, the invention provides a method for detecting a molecule of interest in a sample. The method includes providing a cell having a plurality of expression vectors of the invention, a lysate thereof, an in vitro transcription / translation reaction having a plurality of expression vectors of the invention, and a sample having a hydrolase substrate (e.g., a haloalkane) having at least one functional group, under conditions effective to permit interaction of a first heterologous amino acid sequence with the molecule of interest in the sample. The presence, amount, or location of at least one functional group in the sample is detected, whereby the presence, amount, or location of the molecule of interest is detected.
[0172] Also provided herein is a method for detecting an agent that alters the interaction between two proteins, which comprises providing a cell comprising a plurality of expression vectors of the invention, a lysate thereof, or an in vitro transcription / translation reaction having a plurality of expression vectors of the invention, a hydrolase substrate having at least one functional group (e.g., haloalkane), and a sample having the agent, under conditions effective to permit association of a first fusion protein and a second fusion protein. The agent is suspected of altering the interaction between a first heterologous amino acid sequence and a second heterologous amino acid sequence. The presence or amount of at least one functional group in the sample is detected as compared to a sample not containing the agent.
[0173] In another embodiment, the invention provides a method for detecting an agent that alters the interaction between a molecule of interest and a protein. The method comprises providing a cell comprising a plurality of expression vectors of the invention, a lysate thereof, or an in vitro transcription / translation reaction having a plurality of expression vectors of the invention, a hydrolase substrate having at least one functional group (e.g., haloalkane), and an agent suspected of altering the interaction between a heterologous amino acid sequence and a molecule of interest in the sample. The presence or amount of the functional group in the sample as compared to a sample having the agent.
[0174] In some embodiments, provided herein is a method for detecting the presence of a molecule of interest. For example, a cell is contacted with a vector comprising a nucleic acid sequence encoding a promoter, e.g., a regulatable promoter, and two complementary fragments of a mutant hydrolase, at least one of which is fused to a protein that interacts with the molecule of interest. In one embodiment, the transfected cells are cultured under conditions where the promoter induces transient expression of the fragment or regulated expression of one of the fragments and an activity associated with the labeled substrate is detected.
[0175] In some embodiments, the systems herein (e.g., including sp hydrolase (e.g., spHT, etc.)) can be used as biosensors for detecting the presence / amount of a target molecule or a particular condition (e.g., pH or temperature). When interacting with a target molecule or being subjected to a particular condition, the biosensor undergoes a conformational change or chemically changes to cause a change in activity. In some embodiments, the sp hydrolase herein includes an interaction domain for the target molecule. For example, the biosensor can be a protease (e.g., for detecting the presence of a specific viral protease and being an indicator of the presence of a virus), a kinase (e.g., by inserting a kinase site into a reporter protein), RNAi (e.g., by inserting a suspected sequence recognized by RNAi into the coding sequence of a reporter protein and then monitoring the reporter activity after the addition of RNAi), a binding protein such as a ligand, an antibody, a cyclic nucleotide such as cAMP or cGMP, or a metal such as calcium, which can be generated by inserting a suitable sensor region into the sp hydrolase (e.g., spHT, etc.). One or more sensor regions can be inserted at one or more suitable positions within the C-terminus, N-terminus, and / or the sp hydrolase sequence, and the sensor region includes one or more amino acids. One or all of the inserted sensor regions may include linker amino acids for coupling the sensor to the remainder of the polypeptide. Examples of biosensors are disclosed in U.S. Patent Application Publication Nos. 2005 / 0153310 and 2009 / 0305280, and PCT Publication No. WO2007 / 120522A2, each of which is incorporated herein by reference.
[0176] Experiment Example 1 Comprehensive Screening of Circular Permutated Dehalogenase Substitutions A plasmid encoding all possible circular permuted versions of HaloTag and two linker control versions of unsubstituted HaloTag with simple N- or C-terminal appended linkers were constructed by PCR of a total of 298 gene constructs. The linker connecting the native N- and C-termini was JPEG2025516354000008.jpg18165. Expression was carried out in E. coli, and cell lysates were prepared by adding chemical lysis reagents. The lysates were treated with TEV protease (or water as a negative control) and subjected to a series of biochemical tests.
[0177] Lysates were assayed for protein solubility by centrifugation, followed by conjugation with 10 μM CA-TMR ligand, and gel electrophoresis. To determine the thermal stability of each cpHT, lysates were heated to 40 - 90 °C for 30 min, cooled to room temperature, then mixed with 10 nM CA-TMR and subjected to fluorescence polarization (FP) measurements. Enzyme activity was quantitatively measured by mixing lysates with 10 nM CA-AlexaFluor488 and monitoring their FP changes over 30 min.
[0178] This screening revealed that 228 / 296 (77%) of the cpHT variants reacted with CA-TMR, most of which were soluble, and 50 variants had at least 10% of the native HT activity on CA-AlexaFluor488 (Figure 1). Seventeen cpHT variants had increased thermal stability compared to HT, and 38 variants showed recovery of activity after heat denaturation, presumably by protein refolding. The most active variants, as determined by the Alexa Fluor488 rate, were clustered in regions distal to the lid domain, but this effect may be specific to this substrate, which is negatively charged and sensitive to lid domain perturbation. Indeed, when a neutral TMR ligand was used in the solubility and stability assays, the clustering effect was less prominent. With the exception of cpHT near residues 111 and 120, all refolding variants were localized to the lid domain, and all thermostabilized variants were also within the lid domain.
[0179] Example 2 Testing of split HaloTag variants Selection of spHT screening candidates from an inclusive cpHT screening After completing the screening of all 298 possible circular permutants of HaloTag (cpHT) (see Example 1), 22 split sites were selected for testing as split HaloTag fragment pairs (spHT). The candidate spHT designs were selected based on the properties of their cpHT counterparts, including thermal stability, expression, enzyme activity, and changes in biophysical properties upon cleavage of the TEV protease recognition sequence in the linker connecting the native N- and C-termini. Particular interest was paid to variants that showed the ability to reform or refold after heat denaturation upon cleavage of the TEV protease in the cpHT form (e.g., circular permutants within the sequence region near residue 120).
[0180] Expression of spHT fragments as insoluble fusions to various tags Initial sets of spHT N-terminal and C-terminal fragments (spHT 80, 97, and 121) were expressed in E. coli as fusions to several different domains containing the maltose-binding protein (MBP), 6x polyhistidine tag (His-tag), and the large and small components of the bimolecular NanoBiT system (LgBiT and SmBiT). Moderate expression was noted for some of these fusions, but all suffered from low solubility. The low solubility was due to the exposure of core hydrophobic residues buried in the complete HT structure, which usually forms a surface prone to aggregation on the spHT fragments. Based on estimates of NanoLuc activity, the solubility of these fragments was less than 5% in the E. coli lysate.
[0181] Characterization of spHT Variants by Chemically Induced Dimerization of FRB / FKBP Fusions Despite the low solubility, all 22 exemplary spHT designs were generated as fusions to the FRB domain and the FKBP domain. FRB and FKBP undergo high-affinity heterodimerization chemically induced in the presence of rapamycin. Thus, spHT fragments fused to these domains can be brought into proximity with each other by the addition of rapamycin, providing an assay for the functional reconstitution of HaloTag enzyme activity. Each of the spHT fragments was fused to FRB or FKBP at either the N-terminus or the C-terminus (n = 44), generating a total of 176 unique fusion proteins, which were expressed in E. coli. Since the best orientations of FRB and FKBP with respect to the spHT fragment domain could not be predicted a priori, all possible orientations and combinations were assayed (8 per spHT site). The fusion combinations were assayed using the fluorogenic JaneliaFluor 646 (JF646) ligand in the presence of 50 nM rapamycin. JF646 is available through the normal Promega catalog, has low background fluorescence (enabling direct fluorescence measurements in 96-well plates), and was selected to provide a higher stringency test than non-fluorogenic ligands (such as TMR).
[0182] Six of the 22 spHT FRB / FKBP designs showed more than a two-fold increase in fluorescence signal in the presence of rapamycin (spHT 80, 133, 145, 157, 180, and 195), with up to a 4.7-fold induction shown for the [1-195]-FKBP+[196-297]-FRB combination (Figure 2). The corresponding cpHT 195 was the most thermostable among all variants in the circular permutation screen (at a T about 7 °C higher than HT). All but one of the spHT hits (spHT 80) were located within the lid subdomain of HT, which contains the region of the sequence covering residues 133–216. Among the spHT hits, there were multiple orientations of FRB and FKBP that allowed for reconstitution of activity. Generally, the fusion combinations where FKBP was at the C-terminus of either fragment were best. m In addition to combinations of blunt spHT fragments (where all HT residues are present exactly once), several “gapped” combinations (e.g., having deletions relative to the parental sequence) and “overlapped” combinations (e.g., having duplications relative to the parental sequence) were tested (where a particular residue was missing from both fragments or present in both fragments). Residues missing or doubly represented in these combinations were limited to the lid subdomain, specifically helices 6, 7, 8, and / or 9. The gapped combinations were unable to reconstitute detectable ligand-binding activity. However, the overlapped combinations showed up to a three-fold reconstitution relative to the background (Figure 3). These results indicate that (a) the lid helices are important for ligand binding in HT and (b) the lid subdomain can tolerate sequence duplications and may be involved in secondary structure swapping, a useful feature for designing biosensors and conformationally dynamic protein switches.
[0183] Reversibility of spHT complementation
[0184] Reversibility is an important property of the bimolecular reporter system, and during the development of the embodiments herein, experiments were conducted to determine whether the formation of the spHT complex could be reversible and how this could affect ligand binding and signal dynamics. First, the spHT FRB / FKBP pair was retested with higher concentrations of rapamycin. Some combinations of spHT showed a rapid increase in the fluorescence of JF646 (up to 5-fold the background) only when 100 nM or 500 nM rapamycin was added (Figure 4), indicating that the rapamycin used in previous tests may have been insufficient for some spHT designs. spHT 19, 145, 157, 195, and 233 stood out with the largest fold increase in the fluorescence of JF646. 500 nM rapamycin was used in subsequent reversibility experiments.
[0185] Combinations of the spHT FRB / FKBP fusion were incubated with 500 nM rapamycin for 24 hours, then a 31-fold molar excess (10 μM) of the competing ligand FK506 was added. After 24 hours, JF646 was added and allowed to bind for an additional 24 hours (total of 72 hours elapsed). spHT 19 had slightly less fluorescence compared to the no-FK506 control, and spHT 157, 195, and 233 had only background levels of fluorescence compared to the no-FK506 control (Figure 5). However, the fluorescence of spHT 145 did not decrease compared to its no-FK506 control. That is, rapamycin formed an irreversible complex with spHT 145, a semi-reversible complex with spHT 19, and reversible complexes with spHT 157, 195, and 233.
[0186] The combination of the spHT FRB / FKBP fusion was incubated with 500 nM rapamycin for 24 hours, JF646 for 48 hours, and then a 10-fold molar excess of FK506 for 48 hours (Figure 6). In this case, FK506 was unable to reverse the fluorescence generation of spHT 19, 145, 157, 195, and 233. That is, when rapamycin was competed from the FKBP fusion, the fluorescence of JF646 did not decrease and the induced dimerization signal was removed. These observations were confirmed using TMR labeling and SDS-PAGE analysis (Figure 7).
[0187] Collectively, these results demonstrate that some spHT fragments (e.g., split sites 145, 157, 195, etc.) may require long-term proximity to form complexes, perhaps because as spatially separated entities, they form non-complementary non-native structures and require time to sample many conformations in the presence of their stabilizing partners. However, experiments conducted during the development of the embodiments herein have demonstrated that other spHT fragments can form detectable complexes in less than 30 minutes using N-terminal split sites (e.g., split at 19 or 30). Some spHT variants have high affinity and form irreversible (FK506-resistant) complexes such as spHT 145. Other complexes are sensitive to FK506 due to low affinity such as spHT 157, 195, and 233. Complexes that bind ligands can benefit from further stabilization that makes them FK506-resistant spHT complexes and can be reversible in the ligand-free state but may become irreversible in the ligand-bound state.
[0188] Quantitative Reconstitution of spHT 19 by Titration with Short N-Terminal Fragments Based on the sensitivity of spHT to rapamycin concentration, fluorescence was predicted to also be sensitive to the ratio of spHT fragment concentrations. spHT 19 was used as a test case because the larger C-terminal fragment had measurable background activity and the smaller N-terminal fragment was attractive as a potential peptide tag.
[0189] The large C-terminal fragment was kept constant in all eight combinations of spHT 19 FRB / FKBP fusions, while the concentration of the small N-terminal fragment varied. It was found that by increasing the N:C ratio from 1.25 to 10, the TMR labeling efficiency could be increased by more than 100% depending on the orientation of FRB and FKBP in the fusion (Figures 8 and 12). Similarly, the JF646 fluorescence signal could be increased by up to approximately 25% at an N:C ratio of 10 (Figure 10). The higher responsiveness observed with TMR labeling is likely because under the TMR labeling conditions (10 μM substrate), the labeling is limited by the spHT complex concentration, while under the JF646 labeling conditions (0.1 μM substrate), the substrate concentration is the limiting factor.
[0190] These results indicate that the larger C-terminal fragment of spHT can function as a quantitative integrator of the smaller N-terminal fragment in the presence of sufficient ligand and high-affinity partner interactions such as FRB-rapamycin-FKBP interactions.
[0191] Expression and Activity of spHT Variants in Mammalian Cells A small subset of spHT variants (spHT 145, 157, and 195) was selected for expression in mammalian cells (HeLa cells). Cells were co-transfected with the pF4Ag shuttle vector encoding the spHT fragments as fusions to FKBP and FRB, adding FKBP to the C-terminus of the first fragment and FRB to the C-terminus of the second fragment in each case. For all spHT co-transfectants, HT activity was observed in both lysates (using a non-fluorescent TMR ligand, Figure 11) and live cells (using fluorescent JF646 and JF585 ligands, Figure 12). The activity was visible even in the absence of rapamycin. Addition of 50 nM rapamycin increased the fluorescence yields of JF646 and JF585 in live cells by up to about 2-fold maximally, but had no apparent effect on the band intensity in TMR-labeled SDS-PAGE. That is, TMR labeling occurred equally (and significantly) regardless of the presence or absence of rapamycin.
[0192] Similar results were obtained using overlapping combinations of spHT in HeLa cells (Figures 13 and 14). These results show that (a) co-expression of spHT fragments enables the fragments to spontaneously complement in a way not observed during independent expression and subsequent mixing, indicating the possibility of co-translational folding; (b) in such pre-assembled spHT complexes, rapamycin can assist in the formation of properly folded lid subdomains by stabilizing the lid helix interactions or by reducing their conformational freedom; (c) since the total concentration of the spHT complex does not change significantly, such an effect of rapamycin can increase the fluorescent ligand (e.g., JF646 and JF585) signal without increasing the total ligand binding (as measured by saturation with the TMR ligand).
[0193] Example 3 Internal split site Several split sites located in the lid domain were further examined for complementation as circular permutations that internally split the peptide from the lid domain (i.e., remove the internal helix and test for complementation). This included constructs that contained gapped and overlapping residues in the peptide to be complemented. Since several circularly permuted HaloTag variants lacking residues within the 146 - 195 region of the lid domain were expressed and previously observed to be soluble in E. coli, we tested whether complementation could be observed when the residues corresponding to the deleted fragment were reintroduced. When the cpHaloTag lacking residues 146 - 157 (cpHT(Δ146 - 157)), which is fused to FRB or FKBP, was paired with the cognate 146 - 157 peptide (HT146 - 157) as a fusion to FRB or FKBP, rapamycin - dependent complementation activity was found to be observable (Figure 16). Complementation due to a larger lid domain deletion in cpHaloTag (cpHTΔ146 - 180) was not observed when attempting to pair with the cognate deleted peptide fragment.
[0194] Since the cpHT(Δ146 - 157) internal deletion fragment was functional when the complementing missing residues were reintroduced as a separate peptide, we tested whether other peptides containing various other lid domain residues could exhibit rapamycin - dependent complementation activity. Figure 17 shows the enhancement of rapamycin - induced activity when the cpHT(Δ146 - 157) fragment was paired with the HT(158 - 180) peptide as a fusion to FRB or FKBP. This pair of constructs shared an overlap in the 158 - 180 residue region of the complex and a gap in the 146 - 157 region, yet still functioned for activity in the assay and responded to rapamycin.
[0195] Following experiments showing that split cpHaloTag variants lacking fragments of the lid domain (specifically, helix 6; residues 146 - 157) can be complemented with peptides containing smaller lid domain fragments that either correspond to the missing residues or have overlapping and gapped configurations, we tested whether complementation could also occur by donation of lid domain residues via domain swapping. Domain swapping is the phenomenon where two polypeptides exchange similar folded domains, and when it occurs between similar (often identical) proteins, it can reproduce each monomer's structure. In the case of HaloTag, it has been shown that its lid subdomain can "swap" between monomers, creating a dimeric structure where each monomer consists of its own core a / b - hydrolase domain and the lid domain of its partner. Since the function of HaloTag depends on the proper folding of the lid domain for binding to chloroalkane substrates, it was inferred that cpHaloTag variants lacking fragments of the lid domain could have their activity restored if another cpHaloTag construct could exchange or donate the missing residues to form a complete HaloTag structure. To detect activity only when domain swapping occurred, a D106A mutation was created in the domain "donor" construct in the pairs shown in Figure 18. The D106A mutation eliminates covalent binding of chloroalkane (and thus is not detected on gels). However, since those mutant cpHaloTag variants still retain their lid domain residues, they can exchange them for split cpHaloTag variants to complement their missing residues, restore their activity, and enable labeling with a TMR HaloTag ligand that is detectable after SDS - PAGE. Figure 18 shows the success in identifying constructs that can domain - swap and complement split HaloTag fragments under these conditions, facilitated by their fusion to FRB or FKBP and the inclusion of rapamycin.This experiment shows that the cpHaloTag variant lacking the lid domain fragment (Δ146 - 157) can form a complementation system through domain swapping with other cpHaloTag variants that have the terminus in the lid domain and can donate the necessary residues.
[0196] Using the LgBiT and SmBiT tags on fragments of split HaloTag fused to FRB / FKBP, the complementation and reversibility of each complex were measured in a fluorescence-independent manner. The NanoBiT detection of fragment complementation closely matched the pattern of activity associated with JF646 HaloTag ligand labeling. In the absence of rapamycin, low luminescence and JF646 labeling were detected, but upon addition of rapamycin, both signals increased significantly, indicating that complex formation and restoration of enzyme activity are dependent on promotion via the FRB:FKBP interaction. Addition of FK506, an inhibitor of the FRB:FKBP interaction, to the reaction showed a decrease in both the luminescence signal and the fluorescence signal of all constructs after their rapamycin-dependent complementation, demonstrating that these split HaloTag fragments are physically and functionally reversible.
[0197] To evaluate their usefulness for diagnostic or clinical applications, the robustness of the internal split HaloTag complementation system against human body fluid matrices was tested. Both spHT145 and spHT195 showed resistance to each matrix up to the 10% limit tested and retained their rapamycin-dependent complementation and activity. This experiment demonstrates that these split HaloTag fragments are resistant to human body fluid matrices and can be envisioned as a technology for detecting molecular proximity or binding in diagnostic or clinical assays.
[0198] Example 4 N-terminal split site During the development of the embodiments of this specification, experiments were conducted to test combinations of N-terminal split HaloTag fragments to determine whether they could be induced to complement as FRB or FKBP fusions. The role of sequence overlap in determining performance was investigated. A series of N-terminal fragments of small peptide sizes can be observed to exhibit a rapamycin-dependent response in activity with the JF646 HaloTag ligand. Since the larger fragments were composed of residues 22 - 297 or 23 - 297, many of the small fragments tested had either a gap or an overlap in their sequences. This demonstrated complementation with these N-terminal split fragments over a variety of sequence variabilities and lengths.
[0199] The N-terminal split HaloTag system could be optimized by systematic evaluation of the cleavage of the smaller HT(1 - 19) fragment. Figure 22 shows that cleavage of the first 2 - 3 N-terminal residues enhances the fold response of the system to rapamycin, in particular. Complementation activity was demonstrated with small fragments as small as about 11 amino acids (HT(8 - 19)).
[0200] The ability to detect complementation of N-terminal split HaloTag fragments independent of their activity was tested by fusion to NanoBiT components. A 100 - 250-fold increase in luminescence was observed upon addition of rapamycin, which corresponded to an increase in labeling with JF646 in a separate assay, indicating that both fragment complementation and enzyme activity through physical proximity could be detected for asymmetric N-terminal split HaloTag fragments.
[0201] Other N-terminal split HaloTag fragments were functional in a dual-tag configuration with HiBiT. When HiBiT was added to multiple different small N-terminal HaloTag fragments, both HaloTag activity upon binding of the JF646 ligand and NanoBiT activity with the HiBiT tag could be detected simultaneously (in different reactions). This demonstrated that these tags could be used in parallel to perform multiple measurements from a single system, such that the user could add this dual-tag for multiple uses in both luminescence and fluorescence.
[0202] Considering the difference in the large fragment between HT(22-297) and HT(23-297), the role of the Met22 position was investigated through site saturation and mutation to all other amino acids. Amino acids at position 22 that improved brightness (e.g., M22I or M22L), and amino acids that improved the amplification response of the system (M22F) were identified. It was also observed that the introduction of the mutation of "HaloTag9" (Q165H+P174R) significantly improved the labeling rate of the system when added to the HT(22-297) large fragment. These experiments demonstrated that mutations could be introduced into both the large and small fragments of these split HaloTag variants to improve system performance.
[0203] Considering that mutations in each small and large fragment of the N-terminal HaloTag split resulted in improved brightness and amplification response with the JF646 ligand, they were tested with other Janelia Fluor HaloTag ligands (Figure 26). This experiment demonstrated that mutations in each of the N-terminal split fragments could regulate the kinetics of labeling with different HaloTag ligands. Furthermore, when sufficient time up to 66 hours was given, some constructs were shown to continue to increase in brightness and amplification response, resulting in a response of up to 60-fold. It was also shown that some constructs achieved brightness equivalent to that of full-length HaloTag7, with a relative brightness of 1.0 or more of the HaloTag7 control.
[0204] The experiment of FIG. 27 is the detection of the rapamycin-dependent response in split HaloTag activity in living mammalian cells. The response can be detected for multiple small HaloTag fragments (HT(3-19), HT(4-19)) with different large HaloTag fragments. The system performance was optimized, as shown, by introducing mutations to the large HT(22~297) fragment of either the Q165H+P174R double mutant or the M22F mutation. The fold response (labeling rate after pre-incubation with rapamycin) is detectable within 15 minutes, indicating that these split HaloTag constructs are suitable for live cell-based assays for detecting many biological processes in real time.
[0205] The experiment also demonstrated the functionality and utility of the split HaloTag system in live cell fluorescence microscopy. This is a desired modality for detection, especially considering the development and use of Janelia Fluor HaloTag ligands for advanced cell imaging applications such as STED or STORM. The experiment demonstrated that the split HaloTag constructs have brightness equivalent to that of the full-length HaloTag7 (FIG. 28), and that the fluorescence signal in live cells depends on the addition of rapamycin to promote the interaction of the split HaloTag fragments (FIG. 28), highlighting its utility as a reporter for molecular interactions in live cells.
[0206] Example 5 Use of synthetic peptides for measuring split HaloTag complementation During the development of the embodiments of this specification, experiments were conducted that demonstrated that the synthetic peptide version of the HaloTag[3-19] fragment can be used and complementation with the HaloTag[22-297](M2F) fragment was observed (Figure 30). At the highest concentration of the peptides tested, 1 mM, saturation of the fluorescence signal was not observed, indicating that the affinity between the split HaloTag fragments is likely to be greater than 100 micromolar. Their low affinity is expected to be in a suitable range for in vivo protein:protein interaction studies because it requires fusion to partners that interact to form the complemented complex.
[0207] The experiments additionally demonstrated that the synthetic peptide version of the HaloTag[3-19] fragment can be used and complementation was observed, but with a different variant of LgHT, HaloTag[22-297](Q145H+P154R). This LgHT variant is more stable, expresses better, and gives a higher complementation signal, but due to its higher "background" or uncomplemented signal, it gives a lower fold response (Figure 31). This LgHT variant appears to be approaching saturation of the fluorescence signal, indicating that the affinity between the split HaloTag fragments is approximately 1 micromolar, which can be considered a higher affinity than ideal for in vivo protein:protein interaction studies, but complementation in mammalian cells using this LgHT variant has been observed, indicating that it can be used for the purpose.
[0208] Complementation with synthetic peptides can be measured using different HaloTag ligands (TMR) and a fluorescence polarization assay format (Figure 32). The relative kinetic rates of labeling of HaloTag[22-297](M2F) at different levels of complementation by the peptides are demonstrated. At high peptide concentrations, complementation by the peptides results in a higher labeling rate.
[0209] The HaloTag[22-297](Q145H+P154R), which is an LgHT variant, was also tested using the HaloTag ligand (TMR) in a fluorescence polarization assay format to measure complementation by the synthetic peptide (Figure 33). Due to its greater intrinsic stability and expression in E. coli lysates, the relative kinetic rate of labeling of this peptide-free variant is higher. In the presence of the peptide, a faster labeling rate can be achieved compared to the HaloTag[22-297](M2F) variant.
[0210] The use of a fully purified split HaloTag system (6xHis-HaloTag[22-297](M2F) and synthetic HaloTag[3-19] peptide) can complement each other in vitro, resulting in complementation by the JF646 ligand and an increase in fluorescence intensity after labeling (Figure 34). The fold response to various concentrations of synthetic HaloTag[3-19] peptide for purified 6xHis-HaloTag[22-297](M2F) was calculated compared to a non-complemented reaction lacking the peptide (Figure 35).
[0211] A fully purified split HaloTag system with a synthetic peptide modified by adding two consecutive arginine residues to the N-terminus of the HaloTag[3-19] sequence was used in an attempt to improve peptide solubility. Although an increase in peptide solubility was not observed, it was demonstrated that the N-terminus of the peptide can be modified with additional residues while maintaining the function of the split HaloTag system (Figure 36).
[0212] During the development of the embodiments herein, experiments were conducted to determine the fold response to various concentrations of variant synthetic HaloTag[3-19] peptides for purified 6xHis-HaloTag[22-297](M2F) compared to a non-complemented reaction lacking the peptide (Figure 37).
[0213] A fully purified split HaloTag system that can complement each other in vitro using variants of LgHT, 6xHis-HaloTag[22-297](Q145H+P154R), and synthetic HaloTag[3-19] peptide, with complementation by the JF646 ligand and an increase in fluorescence intensity after labeling (Figure 38). This variant of LgHT saturates its fluorescence intensity at higher peptide concentrations, showing a higher apparent affinity for the peptide.
[0214] During the development of the embodiments herein, experiments were performed to demonstrate the fold response of purified 6xHis-HaloTag[22-297](Q145H+P154R) to various concentrations of synthetic HaloTag[3-19] peptide compared to a non-complemented reaction lacking the peptide (Figure 39). Due to its high affinity, this variant LgHT can show a measurable fold response to low concentrations of peptide up to the mid-nanomolar range, suggesting that it may be suitable as a stand-alone SmHT detection system assuming the assay conditions are sufficient for complementation.
[0215] Experiments were performed to demonstrate the use of a purified split HaloTag system with a modified synthetic peptide that adds two consecutive arginine residues to the N-terminus of the HaloTag[3-19] sequence (Figure 40). The N-terminus of the peptide can be modified with additional residues, and in this case, the function of the split HaloTag system can be retained using the HaloTag[22-297](Q145H+P154R) variant.
[0216] Fold response of purified 6xHis-HaloTag[22-297](Q145H+P154R) to various concentrations of variant synthetic HaloTag[3-19] peptide compared to a non-complemented reaction lacking the peptide (Figure 41). Compared to the HaloTag[22-297](M2F) comparator, this LgHT variant shows a higher fold response at lower peptide concentrations and is potentially useful for the detection of variant SmHT peptides as a detection reagent.
[0217] This purification system demonstrated successful detection of shorter peptides based on the residue HaloTag[8-19] by N-terminal or C-terminal arginine addition (Figure 42). The shorter peptides showed lower affinity than the HaloTag[3-19] peptide. This demonstrates that shorter peptides can be used for complementation in the split HaloTag system, that sequence addition to shorter sequences is tolerated, and potentially further optimizes the system. This purification system also demonstrated successful detection of shorter peptides based on the residue HaloTag[8-19] by N-terminal or C-terminal arginine addition, in this case using the variant LgHT, HaloTag[22-297](Q145H*P154R) (Figure 43). However, this LgHT variant has a higher affinity for the full length and has shorter peptides, so the shorter peptides show lower affinity than the HaloTag[3-19] peptide.
[0218] Example 6 Mutagenesis of HaloTag[3-19] or "small HaloTag" sequence During the development of the embodiments of this specification, experiments were carried out to test all possible single mutations within the HaloTag[3-19] fragment. Expression and activity data for the HaloTag[3-19] variants are provided in Table 1 (all substitutions relative to positions 3-19 of SEQ ID NO: 1). Figures 44-60 represent 19 amino acid changes at individual positions (numbers 1-17). HaloTag[3-19] variants were expressed as fusions to the C-terminus of FKBP and tested for complementation promoted by the HaloTag[22-297](M2F)-FRB fusion by adding rapamycin to E. coli lysates. These experiments helped identify positions within the HaloTag[3-19] fragment that are resistant to mutation and made it possible to construct multiple mutations within the fragment. The effect of combining favorable tolerance mutations at each position is shown, starting with double / triple / quadruple, etc. mutations and ultimately reaching a design where all 17 positions can be mutated simultaneously in a functional variant with 0% identity to the starting HaloTag[3-19] sequence. Most interactions with large fragments occur via main-chain, beta-strand interactions and are independent of side chains. For the few side chains involved in the interaction, conservative mutations occur that change the residue but retain the properties of the side-chain interaction, for example, mutations from phenylalanine to tyrosine, where the hydrophobicity of the side chain can be retained. Thus, a functional small HaloTag fragment with all 17 positions mutated can be achieved.
[0219] Mutations of residue E1 of the HaloTag[3-19] sequence showed no deleterious mutations and showed some potential beneficial mutations (E1A, E1H, E1K) (Figure 44).
[0220] Mutations of residue I2 of the HaloTag[3-19] sequence showed beneficial mutations with larger hydrophobic side chains such as I2F, I2W, and I2Y (Figure 45). The I2T mutation was the only mutation that showed a significant loss of performance in the assay.
[0221] Mutations of residue G3 of the HaloTag[3-19] sequence showed improvement with G3N mutation and several deleterious mutations (G3I, G3T, and G3W), and had a significant loss of performance (Figure 46).
[0222] Mutations of residue T4 of the HaloTag[3-19] sequence showed some improvement with T4S and T4E mutations, as well as some deleterious effects from T4L and T4W mutations (Figure 47).
[0223] Mutations of residue G5 of the HaloTag[3-19] sequence generally showed good tolerance across all amino acid changes (Figure 48).
[0224] Mutations of residue F6 of the HaloTag[3-19] sequence showed that this residue is very sensitive to mutations (Figure 49). Most mutations were deleterious, except for F6M, F6W, and F6Y, which showed performance similar to the non-mutated protein. This is an example showing that maintaining the hydrophobic nature of the side chain (see also Ile and Val mutations), especially the ring structure (see also histidine mutations), is beneficial for interaction with large fragments.
[0225] Mutations of residue P7 of the HaloTag[3-19] sequence showed improvement with P7N, and were less improved than P7H or P7A (Figure 50). The mutation to P7K seemed to be significantly harmful to activity. As shown in later figures, mutations such as P7N are well tolerated as single mutations, but generally show a decrease in activity when combined with other mutations in HaloTag[3-19].[[]]END
[0226] Mutations of residue F8 of the HaloTag[3-19] sequence, similar to F6, showed that this is a very sensitive residue to mutations (Figure 51). Only mutations F8W and F8Y were well tolerated, indicating that it is necessary to preserve a large hydrophobic side chain at this position for complementation. Some slight complementation activity was detectable with F8D and F8R, but these were lower than the non-mutated HaloTag[3-19].[[]]END
[0227] Mutations at residue D9 of the HaloTag[3-19] array showed that all mutations were well tolerated by some beneficial mutations from D9A, D9P, D9Q, and D9R (Figure 52).
[0228] Mutations at residue P10 of the HaloTag[3-19] array showed moderate sensitivity to many mutations (Figure 53). Mutations P10A, P10E, P10S, and P10H had the best tolerance. Mutations P10I and P10K were still functional but were the most detrimental. Similar to residue P7, it should be noted that mutations that are well tolerated as single mutations at P10 are mostly detrimental when combined with other mutations in HaloTag[3-19]. Thus, mutations at P7 and P10 can be tolerated, but they appear to be in a unique category of positions that do not combine well with other mutations.
[0229] Mutations at residue H11 of the HaloTag[3-19] array showed improvement through H11N, but H11D, H11G, and H11P were shown to be very detrimental mutations (Figure 54). Similar to F6, F8, and Y12, this is part of a set of residues that contact large fragments through their hydrophobic loop side chains, but H11 appears to tolerate more mutations than the other three residues in the set.
[0230] Mutations at residue Y12 of the HaloTag[3-19] array showed very high sensitivity to mutations, with tolerance to the conservative mutations Y12F and Y12W that preserve the side chain of its hydrophobic loop (Figure 55). Some charged residues were also tolerated, but with low activity such as Y12H and Y12R. All other mutations appear to be very detrimental at this position. Along with F6, F8, and Y12, this is a large hydrophobic side chain that tolerates only a few amino acid changes that maintain its properties.
[0231] Mutations at residue V13 of the HaloTag[3-19] sequence showed tolerance to other similar hydrophobic amino acids such as V13A, V13L, and V13I (Figure 56). There were also lower but active mutations with charged residues and some notable non-functional mutations as shown.
[0232] Mutations at residue E14 of the HaloTag[3-19] sequence showed tolerance across all mutations except for E14L, which was lower than the non-mutated sequence (Figure 57).
[0233] Mutations at residue V15 of the HaloTag[3-19] sequence showed tolerance to other hydrophobic amino acids that were conservative mutations, particularly V15I and V15L (Figure 58). Many still had detectable activity but did not seem to tolerate charged residues well.
[0234] Mutations at residue L16 of the HaloTag[3-19] sequence showed tolerance to all amino acid changes (Figure 59).
[0235] Mutations at residue G17 of the HaloTag[3-19] sequen...
Claims
[Claim 1] A composition comprising a split variant of a polypeptide having at least 70% sequence similarity to Sequence ID No. 1.