Modified dehalogenases having extended surface loop regions

JP2025515181A5Pending Publication Date: 2026-05-15PROMEGA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PROMEGA CORP
Filing Date
2023-05-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The existing HALOTAG protein has shortcomings in substrate interactions, molecular proximity, or molecular geometry, and it is difficult to meet the optimization needs of some applications.

Method used

By extending the circulating region on the surface of the HALOTAG protein, internal fusion insertion locations are provided, and the ability to bind interactions and activate environmentally sensitive chemicals is regulated through these extended regions.

Benefits of technology

It realizes better substrate interactions, molecular proximity, and molecular geometry, which enhances the activation and binding ability of environmentally sensitive chemicals and expands the applicability of HALOTAG protein in various applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are modified dehalogenases with extended surface loop regions that provide locations for internal fusion insertions and modulate binding interactions and activation of environmentally sensitive chemicals.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 338,369, filed May 4, 2022, which is incorporated herein by reference.

[0002] Mega Table This specification contains long tables. Table 1 has been submitted via EFS-Web in the following electronic format: Filename: TABLE_1_Loop_HTs.txt, Created on: May 4, 2023, 2023, File Size: 117,291 bytes. The contents of Table 1 are incorporated herein by reference in their entirety.

[0003] Provided herein are modified dehalogenases with extended surface loop regions that provide locations for internal fusion insertions and modulate binding interactions and activation of environmentally sensitive chemicals. [Background technology]

[0004] The utility of self-labeling protein systems such as HALOTAG and its chloroalkane-based ligands has continually expanded over the lifetime of this research tool. Gene fusions to HALOTAG as a general strategy have enabled a wide range of applications including fluorescent labeling for cell biology and imaging, recombinant protein purification, biosensors and diagnostics, energy transfer technologies (BRET, FRET), and therapeutic targeted protein degradation (PROTACs).

[0005] What is needed is a modified HALOTAG protein that provides substrate interaction, optimal molecular proximity, or optimal molecular geometry. Summary of the Invention

[0006] Provided herein are modified dehalogenases with extended surface loop regions that provide locations for internal fusion insertions and modulate binding interactions and activation of environmentally sensitive chemicals.

[0007] In some embodiments, provided herein is a composition comprising a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to SEQ ID NO:2, wherein X1 to X 25 Each of X1 to X2 is independently selected from any amino acid or is absent; 25 At least five of X1 to X2 are not absent, and the polypeptide has less than 100% sequence identity to SEQ ID NO:1. 25 At least 10 of them are not non-existent.

[0008] In some embodiments, provided herein is a composition comprising a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to SEQ ID NO:3, wherein X1 to X 25 Each of X1 to X2 is independently selected from any amino acid or is absent; 25 At least five of X1 to X2 are not absent, and the polypeptide has less than 100% sequence identity to SEQ ID NO:1. 25 At least 10 of them are not non-existent.

[0009] In some embodiments, provided herein is a composition comprising a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to SEQ ID NO:4, wherein X1 to X 25 Each of X1 to X2 is independently selected from any amino acid or is absent; 25At least five of X1 to X2 are not absent, and the polypeptide has less than 100% sequence identity to SEQ ID NO:1. 25 At least 10 of them are not non-existent.

[0010] In some embodiments, provided herein is a composition comprising a polypeptide having at least 70% sequence identity to SEQ ID NO:5, wherein X1 to X 25 Each of X1 to X2 is independently selected from any amino acid or is absent; 25 at least five of which are not absent and the polypeptide has less than 100% sequence identity to SEQ ID NO:1.

[0011] In some embodiments, X1 to X 25 At least 10 of them are not non-existent.

[0012] In some embodiments, provided herein are compositions comprising a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 6-9, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NOs: 10-13, and an internal segment connecting the N-terminal segment and the C-terminal segment, wherein the internal segment is greater than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:6, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:10, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:7, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:11, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:8, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:12, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:9, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:13, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.

[0013] In some embodiments, provided herein are compositions comprising a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 14-20, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NOs: 21-27, and an internal segment connecting the N-terminal segment and the C-terminal segment, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 14, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 21, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 15, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 22, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.In some embodiments, the polypeptide comprises a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 16, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 23, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 17, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 24, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 18, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 25, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.In some embodiments, the polypeptide comprises a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 19, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 26, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:20, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:27, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.

[0014] In some embodiments, provided herein are compositions comprising a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 81-85, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NOs: 86-90, and an internal segment connecting the N-terminal segment and the C-terminal segment, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:81, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:86, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:82, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:87, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:83, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:88, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:84, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:89, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:85, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:90, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.In some embodiments, the polypeptide comprises a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 19, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO: 26, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length. In some embodiments, the polypeptide comprises an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:20, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:27, and an internal segment connecting the N-terminal and C-terminal segments, wherein the internal segment is more than 25 amino acids in length.

[0015] In some embodiments, the internal segment is less than 1000 amino acids in length (e.g., 900 amino acids, 800 amino acids, 700 amino acids, 600 amino acids, 500 amino acids, 400 amino acids, 300 amino acids, 200 amino acids, 100 amino acids, or less, or ranges therebetween). In some embodiments, the internal segment is a fluorescent or bioluminescent polypeptide capable of emitting energy at a first wavelength. In some embodiments, the internal segment is a component of a bioluminescent complex capable of emitting energy at a first wavelength when contacted with one or more complementary components of the bioluminescent complex and a luminophore. In some embodiments, the internal segment is a binding protein, an enzyme, or an epitope capable of being recognized by a binding protein. In some embodiments, the internal segment comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to one of SEQ ID NOs: 28-32 or a circularly permuted variant thereof. In some embodiments, the internal segment comprises one of SEQ ID NOs: 28-32 or a circularly permuted variant thereof.

[0016] In some embodiments, provided herein are N-terminal segments that include at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to one of SEQ ID NOs: 6-9, 14-20, and 81-85, at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to one of SEQ ID NOs: 28-32 ... a C-terminal segment having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or a range therebetween) sequence identity to one of SEQ ID NOs: 10-13, 21-27, and 86-90, a first internal segment linking the N-terminal segment and the central segment, and a second internal segment linking the central segment and the C-terminal segment. In some embodiments, provided herein are compositions comprising a polypeptide having an N-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:6, a central segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:18, a C-terminal segment that comprises at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to SEQ ID NO:11, a first internal segment linking the N-terminal segment and the central segment, and a second internal segment linking the central segment and the C-terminal segment. In some embodiments, the first internal segment is selected from the group consisting of X1 to X 25 Including X1 to X 25 Each of X1 to X2 is independently selected from any amino acid or is absent; 25 At least five of the X are not absent, and the second internal segment is X 26 ~X 50 Contains X 26 ~X 50 is independently selected from any amino acid or is absent;26 ~X 50 In some embodiments, the first internal segment is selected from the group consisting of X1 to X 25 Including X1 to X 25 Each of X1 to X2 is independently selected from any amino acid or is absent; 25 at least five of the first and second internal segments are not absent and the second internal segment is greater than 25 amino acids in length. In some embodiments, the second internal segment is a binding protein, a fluorescent protein, a bioluminescent protein, a component of a bioluminescent complex, or an enzyme. In some embodiments, the first internal segment and the second internal segment are each greater than 25 amino acids in length. In some embodiments, the first internal segment and the second internal segment are independently selected from a binding protein, a fluorescent protein, a bioluminescent protein, a component of a bioluminescent complex, and an enzyme.

[0017] In some embodiments, provided herein is a composition comprising a peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 33-80. In some embodiments, provided herein is a method comprising contacting a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 33-80 with a luminophore substrate that emits luminescence when contacted with a portion of the polypeptide. In some embodiments, the luminophore substrate is a coelenterazine substrate or a derivative thereof (e.g., furimazine). In some embodiments, the method further comprises contacting a composition herein with a substrate of formula (I): R-Linker-AX where R is a solid surface or a functional moiety, the linker is a polyatomic straight or branched chain containing C, N, S, or O, optionally containing one or more rings, AX is a substrate for a dehalogenase, and A is (CH2) 4~20and X is a halide. In some embodiments, provided herein is a system that includes (a) a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 33-80, and (b) (i) a luminophore substrate that emits luminescence when contacted with a portion of the polypeptide, and / or (ii) a modified dehalogenase substrate of formula (I), R-Linker-AX where R is a solid surface or a functional moiety, the linker is a polyatomic straight or branched chain containing C, N, S, or O, optionally containing one or more rings, AX is a substrate for a dehalogenase, and A is (CH2) 4~20 and X is a halide. In some embodiments, R is a functional moiety selected from the group consisting of a nucleic acid molecule, an amino acid, a peptide, a receptor protein, a glycoprotein, an antibody, a lipid, a hapten, a receptor ligand, a fluorophore, a photocatalyst, and a toxin.

[0018] In some embodiments, provided herein is a composition comprising a peptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 91-120. In some embodiments, provided herein is a method comprising contacting a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or any range therebetween) sequence identity to one of SEQ ID NOs: 91-120 with a peptide having at least 70% sequence identity to SEQ ID NO: 30, and a luminophore substrate that emits luminescence when contacted with a conjugate of the peptide and a portion of the polypeptide. In some embodiments, the luminophore substrate is a coelenterazine substrate or a derivative thereof (e.g., furimazine). In some embodiments, the method further comprises contacting the composition with a substrate of formula (I): R-Linker-AX where R is a solid surface or a functional moiety, the linker is a polyatomic straight or branched chain containing C, N, S, or O, optionally containing one or more rings, AX is a substrate for a dehalogenase, and A is (CH2) 4~20 and X is a halide. In some embodiments, provided herein is a system that includes (a) a polypeptide having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%, or ranges therebetween) sequence identity to one of SEQ ID NOs: 91-120, (b) a peptide having at least 70% sequence identity to SEQ ID NO: 30, and (c) a luminophore substrate that emits luminescence when contacted with a portion of the polypeptide, and / or (ii) a modified dehalogenase substrate of formula (I), R-Linker-AX where R is a solid surface or a functional moiety, the linker is a polyatomic straight or branched chain containing C, N, S, or O, optionally containing one or more rings, AX is a substrate for a dehalogenase, and A is (CH2) 4~20 and X is a halide. In some embodiments, R is a functional moiety selected from the group consisting of a nucleic acid molecule, an amino acid, a peptide, a receptor protein, a glycoprotein, an antibody, a lipid, a hapten, a receptor ligand, a fluorophore, a photocatalyst, and a toxin.

[0019] In some embodiments, provided herein is a system that includes a modified dehalogenase described herein and a substrate of formula (I): R-linker-AX, where AX is a substrate for the dehalogenase and A is (CH) 4~20 wherein X is a halide, the linker is a polyatomic straight or branched chain comprising C, N, S, or O, optionally containing one or more rings, R is a fluorophore, and X1-X25 are capable of interacting with a substrate to enhance one or more of substrate binding to the modified dehalogenase, fluorescence intensity of the fluorophore, activation of the fluorophore, and resonance energy transfer to the fluorophore. In some embodiments, the fluorophore is fluorogenic.

[0020] In some embodiments, provided herein is a method comprising contacting a modified dehalogenase described herein with a substrate of formula (I): R-Linker-AX where R is a solid surface or a functional moiety, the linker is a polyatomic straight or branched chain containing C, N, S, or O, optionally containing one or more rings, AX is a substrate for a dehalogenase, and A is (CH2) 4~20 and X is a halide. [Brief description of the drawings]

[0021] [Figure 1] 3D structure of a HALOTAG-modified dehalogenase bound to a chloroalkane ligand, highlighting loops 165 and 180. [Diagram 2] TMR ligand labeling activity of loop HaloTag constructs. Each loop received an insertion of 2, 5, or 10 amino acids consisting of glycine-serine (Gly-Ser). Constructs were expressed in E. coli and tested in cell lysates to measure TMR ligand labeling activity in the total (T) or soluble (S) fractions of the lysates. Measurements were performed by running samples through SDS-PAGE and scanning the gel for fluorescence. [Diagram 3] JF646 ligand labeling activity and thermostability of loop HaloTag constructs. Each loop received an insertion of 2, 5, or 10 amino acids consisting of glycine-serine (Gly-Ser). Constructs were expressed in E. coli and tested in cell lysates after heating for 30 min at the indicated temperatures by measuring JF646 ligand labeling activity in the lysates. Measurements were performed in a plate-based format measuring the fluorescence of each sample. [Figure 4]Constructs tested to explore optimal loop extension design for loop HALOTAG constructs. (A) Design and positioning of 10X-Gly-Ser sequences inserted into loop 165 or loop 180. (B) TMR ligand labeling activity of loop HaloTag constructs. Each loop received an insertion of 10 amino acids composed of glycine-serine (Gly-Ser). Constructs were expressed in E. coli and tested in cell lysates, and TMR ligand labeling activity was measured in the total (T) or soluble (S) fractions of the lysates. Measurements were performed by running samples through SDS-PAGE and scanning the gel for fluorescence. [Diagram 5] TMR and JF646 ligand labeling activity of loop HaloTag library designs. Each loop library design consisted of insertions at loop-165 or loop-180 with no adjacent "noF" residues in the loops commonly used for CDR3 loops in antibodies. The randomized loop sequences tested were 7, 11, or 15 amino acids in length. Constructs were expressed in E. coli and tested in cell lysates by measuring (A) TMR ligand labeling activity using a fluorescence polarization assay, or (B) JF646 ligand activity in a fluorescence assay. A 6xHis-HaloTag (ATG2733) control is included for comparison. [Figure 6] Comparison of Loop HaloTag library clones with JF646 ligand labeling activity with TMR. Each clone was plotted as a single data point of its fluorescence intensity with JF646 ligand compared to its fluorescence polarization with TMR ligand. (A) Highlighted clones for libraries of 11 or 15 randomized residues in loop 165. (B) Highlighted clones for libraries of 11 or 15 randomized residues in loop 180. A 6xHis-HaloTag (ATG2733) control is included for comparison. Some Loop HaloTag variants show HaloTag-like levels of activity with both ligands, while others are only active for TMR ligand labeling and not for JF646 ligand fluorescence activation. [Figure 7A] Comparison of Loop HaloTag library clones and Alexa488 ligand labeling activity with JF646 ligand. Individual clones with different loop sequences were tested in E. coli lysates for their activity with multiple ligands. (A) Fluorescence intensity of JF646 ligand with Loop HaloTag clones shows that a range of activities is detected. (B) Binding kinetics to Alexa488 ligand for Loop HaloTag clones shows different activity patterns, with some clones showing high activity with JF646 but little detectable activity with Alexa488 and vice versa. (C) Comparison of Loop HaloTag clone activity across multiple ligands. Clones in different quadrants of the graph represent those with more selective substrate specificity. [Figure 7B] Comparison of Loop HaloTag library clones and Alexa488 ligand labeling activity with JF646 ligand. Individual clones with different loop sequences were tested in E. coli lysates for their activity with multiple ligands. (A) Fluorescence intensity of JF646 ligand with Loop HaloTag clones shows that a range of activities is detected. (B) Binding kinetics to Alexa488 ligand for Loop HaloTag clones shows different activity patterns, with some clones showing high activity with JF646 but little detectable activity with Alexa488 and vice versa. (C) Comparison of Loop HaloTag clone activity across multiple ligands. Clones in different quadrants of the graph represent those with more selective substrate specificity. [Figure 7C]Comparison of Loop HaloTag library clones and Alexa488 ligand labeling activity with JF646 ligand. Individual clones with different loop sequences were tested in E. coli lysates for their activity with multiple ligands. (A) Fluorescence intensity of JF646 ligand with Loop HaloTag clones shows that a range of activities is detected. (B) Binding kinetics to Alexa488 ligand for Loop HaloTag clones shows different activity patterns, with some clones showing high activity with JF646 but little detectable activity with Alexa488 and vice versa. (C) Comparison of Loop HaloTag clone activity across multiple ligands. Clones in different quadrants of the graph represent those with more selective substrate specificity. [Figure 8] The stable sequence allows for a dual loop HaloTag configuration. Individual clones with different loop sequences at both positions 165 and 180 were tested in E. coli lysates for activity with TMR. (A) Tested combinations of previously identified sequences at each loop position that resulted in active loop HaloTag clones. (B) Gel electrophoresis of loop HaloTag clones labeled with TMR ligand in E. coli lysates. Protein staining shows consistent amounts of expression across all loop HaloTag clones. Fluorescence detection in the gel indicates detectable TMR labeling activity specific to the loop HaloTag clone being tested. [Figure 9]Characterization of circularly permuted NanoLuc (cpNLuc), circularly permuted thermostable NanoLuc (cptsNLuc), and circularly permuted thermostable NanoLuc with a point mutation, F164C (cptsNLuc(F164C)), generated HaloTag-NLuc fusions and chimeras by inserting the point mutation, F164C (cptsNLuc(F164C)) into loops 165 and 180. Fusions and chimeras were expressed in E. coli, purified, and compared for chloroalkane-TMR ligand binding kinetics, luminescence brightness, and intramolecular BRET efficiency to the bound TMR ligand. (A) Chimera structure, (B) Binding kinetics of 2.5 nM chloroalkane-TMR to 20 nM fusion and chimera monitored by fluorescence polarization, (C) total luminescence for 6 nM fusion and chimera treated with 20 μM fluorofurimazine, (D) intramolecular BRET efficiency for 6 nM fusion and chimera labeled with 5-fold molar excess of chloroalkane-TMR and treated with 20 μM fluorofurimazine. [Figure 10] Fluorescence emission intensities due to BRET to bound fluorophores showing extensive overlap between their excitation spectra and bioluminescent reporter emission. Fusions and chimeras expressed in E. coli lysates were labeled with 1 uM fluorescent chloroalkane ligands. Upon treatment with 20 μM fluorofurimazine, emission spectra (ex=800 nm) were monitored on a SPRK plate reader. [Figure 11] BRET efficiency for NLuc-HaloTag versus HT-178-cpNLuc67 / 68-179 for bound fluorophores showing extensive overlap between their excitation spectra and bioluminescent reporter emission. 6 nM purified fusions and chimeras were labeled with 10-fold molar excess of fluorescent chloroalkane ligands. Upon treatment with 20 μM fluorofurimazine, emission spectra (ex=800 nm) were monitored on a SPRK plate reader. [Figure 12]Binding properties of HaloTag-LgBiT fusions and chimeras generated by insertion of LgBiT, cpLgBiT and cpLgBiT+4 into loop 180. Fusions and chimeras were expressed in E. coli, purified and compared for chloroalkane-TMR ligand binding and complementation affinity with VS-HiBiT. (A) Chimera constructs, (B) Equal concentrations of chimeras were labeled overnight with a 5-fold molar excess of TMR ligand, resolved by SDS-PAGE and scanned for fluorescence. (C) Binding kinetics of 2.5 nM chloroalkane-TMR to 20 nM or 160 nM fusions and chimeras monitored via fluorescence polarization. (D) Binding kinetics of 2.5 nM chloroalkane-TMR to 20 nM or 160 nM chimeras monitored via fluorescence polarization after complementation with a 10-fold molar excess of VS-HiBiT. [Figure 13] Luminescence and BRET efficiencies of HaloTag-LgBiT fusions and chimeras generated by insertion of LgBiT and circularly permuted LgBiT (cpLgBiT) into loop 180. Fusions and chimeras were expressed in E. coli, purified, and compared to bound TMR ligand for their brightness and intramolecular BRET efficiency. (A) Total luminescence of 6 nM fusions and chimeras complemented with 60 nM VS-HiBiT and treated with 20 μM fluoro-furimazine, (B) Intramolecular BRET efficiency of 6 nM fusions and chimeras complemented with 60 nM VS-HiBiT, labeled with a 5-fold molar excess of chloroalkane-TMR, and treated with 20 μM fluoro-furimazine. [Figure 14]Circular permutation of NanoLuc improves donor, acceptor, and BRET when inserted into HaloTag. The circular permutation site as shown in NanoLuc was inserted into loop 180 of HaloTag and expressed in E. coli. Cell lysates containing each construct were labeled with TMR-CA and tested for luminescence and BRET activity upon addition of fluorofurimazine. (A) Donor and (B) Acceptor luminescence was measured 60 seconds after addition of NanoLuc substrate. (C) MilliBRET (mBRET) was calculated by multiplying the donor to acceptor (BRET) signal ratio by 1,000. The activity of NanoLuc inserted into loop 180 of HaloTag without circular permutation is shown in black on the far right. [Figure 15] Variation of linker length connecting circularly permuted NanoLuc inserted into HaloTag. Circularly permuted NanoLuc at position 67 was inserted into loop 180 of HaloTag using different glycine-serine (GS) linker variations and expressed in E. coli. Cell lysates containing each construct were labeled with TMR-CA and tested for luminescence and BRET activity upon addition of fluorofurimazine. (A) Schematic showing the location of the linker inserted into the HaloTag-cpNanoLuc67 chimera. (B) Donor and (C) acceptor luminescence was measured 60 seconds after addition of NanoLuc substrate. (D) MilliBRET (mBRET) was calculated by multiplying the donor to acceptor (BRET) signal ratio by 1,000. The construct is labeled as "HTi_cpNl67" representing the insertion of cpNanoLuc67 into HaloTag loop 180. The linker sites are abbreviated as "L1", "L2", and "L3" according to their position in (A), and the length of the GS-linker is indicated as a suffix to their name (i.e., "3" represents a GS-linker sequence of three amino acids). The activity of NanoLuc inserted without circular permutation into loop 180 of HaloTag is shown in black at the far right. [Figure 16]Biochemical characterization of lead HALOTAG-cpNANOLUC chimeras (i.e., circularly permuted NanoLuc inserted into the surface loop of HaloTag) emerging from a screen for alternative circularly permuted sites in NanoLuc and flexible linkers that can be incorporated between the components of the chimera. Chimeras were expressed in E. coli, purified, and compared for HaloTag-TMR ligand binding kinetics, brightness, and intramolecular BRET efficiency to bound TMR ligand. (A) Structure of HALOTAG-cpNANOLUC chimera. (B) Binding kinetics of 2.5 nM HaloTag-TMR ligand to 20 nM chimera monitored via fluorescence polarization. (C) Total luminescence of 6 nM chimera treated with 20 μM fluorofurimazine. (D) Intramolecular BRET efficiency of 6 nM chimera covalently labeled with HaloTag-TMR ligand and treated with 20 μM fluorofurimazine. [Figure 17] Characterization of transiently expressed lead HALOTAG-cpNANOLUC chimeras emerging from screening for alternative circularly permuted sites in NanoLuc and flexible linkers that can be incorporated between components of the chimera. Constructs encoding NanoLuc-HaloTag fusions and chimeras were transiently expressed in HeLa cells and assessed for expression, brightness, and efficiency of intramolecular BRET to bound TMR ligand. (A) Structure of HALOTAG-cpNANOLUC chimera. (B) Expression levels. Lysates from cells labeled with 1 μM HaloTag-TMR ligand were resolved by SDS-PAGE and scanned with a fluorescent imager. Expression levels quantified using Image J software and normalized to expression of NanoLuc-HaloTag fusion. (C) Total luminescence from cells treated with 20 μM fluorofurimazine and also normalized to expression. (D) Intramolecular BRET efficiency of cells treated with 500 nM HaloTag-TMR ligand. [Figure 18]BRET imaging of cells transiently expressing either the NanoLuc-HaloTag fusion or the lead HALOTAG-cpNANOLUC chimera emerging from screening for alternative circularly permuted sites of NanoLuc. (A) Images of cells in the presence and absence of bound HaloTag TMR ligand taken on an Olympus LV200 bioluminescence microscope after treatment with 20 μM fluoro-furimazine. Images of donor and acceptor emission were acquired sequentially using a 460 / 80 bandpass filter and a 590 nm longpass filter, respectively. (B) BRET ratios of individual cells. [Figure 19] Biochemical characterization of chimeras generated by inserting circularly permuted NanoLuc into loops 180 and 194 / 195 of HaloTag. Chimeras were expressed in E. coli, purified, and compared for HaloTag-TMR ligand binding kinetics, brightness, and intramolecular BRET efficiency to bound TMR ligand. (A) HaloTag structure with loops and insertion site annotated. (B) Structure of HALOTAG-cpNANOLUC chimera. (C) Binding kinetics of 2.5 nM HaloTag-TMR ligand to 20 nM chimera monitored via fluorescence polarization. (C) Total luminescence of 6 nM chimera treated with 20 μM fluoro-furimazine. (D) Intramolecular BRET efficiency of 6 nM chimera covalently labeled with HaloTag-TMR ligand and treated with 20 μM fluoro-furimazine. [Figure 20]Biochemical characterization of chimeras genetically fused to dCas12g1 and incorporating additional mutations in the HaloTag domain. Annotation of the additional mutations is based on the full-length, non-disrupted HaloTag protein. Fusions were expressed in E. coli, purified, and compared for chloroalkane-TMR ligand binding kinetics, brightness, and intramolecular BRET efficiency to bound TMR ligand. (A) Fusion constructs. (B) Binding kinetics of 2.5 nM chloroalkane-TMR to 20 nM fusion monitored via fluorescence polarization. (C) Total luminescence of 6 nM fusion treated with 20 μM fluoro-furimazine. (D) Intramolecular BRET efficiency of 6 nM chimera covalently labeled with HaloTag-TMR ligand and treated with 20 μM fluoro-furimazine. (E) Effect of additional mutations in the HaloTag domain on chlHaloTag-TMR ligand binding kinetics. (F) Effect of additional mutations in the HaloTag domain on brightness and efficiency of intramolecular BRET to bound TMR ligand. [Figure 21] Biochemical characterization of constructs incorporating circularly permuted NanoLuc either as an insertion into loop 180 of HaloTag or as a fusion to the circularly permuted HaloTag. (A) Total luminescence of 6 nM purified protein treated with 20 μM fluoro-furimazine. (B) Intramolecular BRET efficiency of 6 nM protein covalently labeled with HaloTag-TMR ligand and treated with 20 μM fluoro-furimazine. [Figure 22]Biochemical characterization of complementation-based chimeras incorporating a flexible linker and circularly permuted LgBiT+4 at two alternative sites (i.e., 67 / 68 or 49 / 50). (A) Structure of HALOTAG-cpLGBIT chimera. (B-C) Effect of flexible linker on binding kinetics of 2.5 nM HaloTag-TMR ligand to 200 nM of 20 nM chimera complemented with VS-HiBiT. (C-D) Effect of flexible linker on binding affinity to VS-HiBiT peptide. (E-F) Total luminescence of 6 nM chimera complemented with 60 nM VS-HiBiT and treated with 20 μM fluorofurimazine. (G-I) Intramolecular BRET efficiency for 6 nM chimera complemented with 60 nM VS-HiBiT and covalently labeled with HaloTag-TMR ligand. [Figure 23] Characterization of transiently expressed complementation-based chimeras incorporating LgBiT+4 circularly permuted at a flexible linker and two alternative sites (i.e., 67 / 68 or 49 / 50). Constructs encoding the chimeras were transfected into genome-edited HeLa cells expressing HiBiT-tagged GAPDH. Cells were assessed for expression, brightness, and efficiency of intramolecular BRET to bound TMR ligand. (A) Structure of HALOTAG-cpLGBIT chimeras. (B) Expression levels. Lysates from cells labeled with 1 μM HaloTag-TMR ligand were resolved by SDS-PAGE and scanned with a fluorescent imager. Expression levels were quantified using Image J software and normalized to expression of HaloTag178-cpLgBIT+4 67 / 68-179. (C) Total luminescence from cells treated with 20 μM fluorofurimazine. (D) Intramolecular BRET efficiency of cells treated with 500 nM HaloTag-TMR ligand. [Figure 24]Biochemical characterization of complementation-based chimeras incorporating a flexible linker and circularly permuted LgTrip at two alternative sites (i.e., 67 / 68 or 49 / 50). (A) Structure of HALOTAG-cpLGTRIP chimera. (B-C) Effect of flexible linker on binding rate of 2.5 nM chloroalkane-TMR to 20 nM chimera complemented with 200 nM dipeptide (i.e., VS-HiBiT-Trip9). (C-D) Effect of flexible linker on binding affinity to dipeptide. (E-F) Total luminescence of 6 nM chimera complemented with 60 nM dipeptide and treated with 20 μM fluorofurimazine. (G-I) Intramolecular BRET efficiency for 6 nM chimera complemented with 60 nM dipeptide and covalently labeled with HaloTag-TMR ligand. [Diagram 25] Impact of additional LgTrip mutations on the biochemical properties of the lead complementation-based chimera HaloTag178(L1-3)-cpLgBiT+4-179. Annotation of the additional mutations is based on the full-length non-disrupted NanoLuc protein. (A) Structure of the HALOTAG-cpLGBIT chimera, (B) Effect of the mutations on binding affinity to the VS-HiBiT peptide. (C) Effect of the mutations on brightness and efficiency of intramolecular BRET towards bound TMR ligand for 6 nM chimera complemented with 60 nM VS-HiBiT. (D-E) Binding kinetics of 2.5 nM HaloTag-TMR ligand to 20 nM or 80 nM chimera complemented with 200 nM or 800 nM VS-HiBiT, respectively. [Figure 26] Impact of additional mutations in the LgBiT domain on the biochemical properties of the lead complementation-based chimera HaloTag178(L1-3)-cpLgBiT+4-179. Annotation of the additional mutations is based on the full-length non-disrupted NanoLuc protein. (A) Structure of the HALOTAG-cpLGBIT chimera. (B) Effect of the mutations on binding affinity to the VS-HiBiT peptide. (C) Effect of the mutations on brightness and efficiency of intramolecular BRET towards bound TMR ligand for 6 nM chimera complemented with 60 nM VS-HiBiT. [Figure 27]Effect of different L1 linker configurations on the biochemical properties of the lead complementation-based chimera HaloTag178(L1-3)-cpLgBiT+4-179. (A) Structure of the HALOTAG-cpLGBIT chimera. (B) Effect of mutations on binding affinity to the VS-HiBiT peptide. (C) Effect of mutations on brightness for 6 nM chimera complemented with 60 nM VS-HiBiT and efficiency of intramolecular BRET towards bound TMR ligand. (D-E) Binding kinetics of 2.5 nM HaloTag-TMR ligand to 20 nM or 80 nM chimera complemented with 200 nM or 800 nM VS-HiBiT, respectively.

[0022] definition Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the embodiments described herein, some preferred methods, compositions, devices, and materials are described herein. However, before describing the materials and methods, it should be understood that the invention is not limited to the specific molecules, compositions, methodologies, or procedures described herein, as these may vary according to routine experimentation and optimization. It should also be understood that the terminology used in the description is for the purpose of describing the particular versions or embodiments only, and is not intended to limit the scope of the embodiments described herein.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. However, in case of conflict, the present specification, including definitions, shall control. Therefore, in the context of the embodiments described herein, the following definitions apply.

[0024] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "polypeptide" is a reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth.

[0025] As used herein, the term "and / or" includes any and all combinations of the listed items, including any of the individually listed items. For example, "A, B and / or C" includes A, B, C, AB, AC, BC and ABC, each of which should be considered as being individually listed with the explicit reference "A, B and / or C."

[0026] As used herein, the term "comprise" and linguistic variations thereof indicate the presence of the recited feature(s), element(s), method step(s), etc., without excluding the presence of additional feature(s), element(s), method step(s), etc. Conversely, the term "consisting of" and linguistic variations thereof indicate the presence of the recited feature(s), element(s), method step(s), etc., and excludes any unrecited feature(s), element(s), method step(s), etc., except for impurities normally associated therewith. The phrase "consisting essentially of" indicates the recited feature(s), element(s), method step(s), etc., and any additional feature(s), element(s), method step(s), etc. that do not materially affect the basic nature of the composition, system, or method. Many embodiments herein are described using the open term "comprising." Such embodiments encompass the multiple closed "consisting of" and / or "consisting essentially of" embodiments, which may alternatively be claimed or described using such language.

[0027] As used herein, the term "substantially" means that the recited properties, parameters, and / or values ​​need not necessarily be achieved exactly, but deviations or variations may occur, including, for example, tolerances, measurement errors, limits of measurement precision, and other factors known to those of skill in the art, to an extent that does not interfere with the effect intended to be provided by the property. A substantially absent (e.g., substantially non-fluorescent) property or characteristic may be one that is within the noise, below background, below the detection capabilities of the assay being used, or one that is a small percentage (e.g., <1%, <0.1%, <0.01%, <0.001%, <0.00001%, <0.000001%, <0.0000001%) of a notable property (e.g., fluorescence intensity of an active fluorophore).

[0028] As used herein, when referring to an amino acid sequence or a position within an amino acid sequence, the phrase "corresponding to" refers to the relative position of an amino acid residue or amino acid segment, and refers to the sequence, but not necessarily the specific identity of the amino acid at that position. For example, a "peptide corresponding to positions 36-48 of SEQ ID NO:1" may have less than 100% sequence identity (e.g., greater than 70% sequence identity) with positions 36-48 of SEQ ID NO:1, but within the context of the composition or system being described, the peptide is relative to those positions.

[0029] As used herein, the term "system" refers to multiple components (e.g., devices, compositions, etc.) used for a particular purpose. For example, two separate biological molecules can comprise a system if they are useful together for a common purpose, whether or not they are present in the same composition.

[0030] As used herein, the term "complementary" refers to the property of two or more structural elements (e.g., peptides, polypeptides, nucleic acids, small molecules, etc.) that can hybridize, dimerize, or otherwise form a complex with each other. For example, "complementary peptides and polypeptides" can combine to form a complex. Complementary elements may require assistance (facilitation) to form a complex (e.g., from interacting elements), for example, to place the elements in the proper conformation for complementarity, to place the elements in the proper proximity for complementarity, to colocalize complementary elements, to lower the interaction energy for complementary elements, to overcome insufficient affinity for each other.

[0031] As used herein, the term "complex" refers to an assembly or aggregate of molecules (e.g., peptides, polypeptides, etc.) that are in direct and / or indirect contact with each other. In one aspect, "contact," or more specifically, "direct contact," means that two or more molecules are sufficiently close that attractive non-covalent interactions, such as van der Waals forces, hydrogen bonds, ionic and hydrophobic interactions, dominate the interaction of the molecules. In such an aspect, a complex of molecules (e.g., peptides, polypeptides, etc.) is formed under assay conditions such that the complex is thermodynamically favorable (e.g., compared to the unaggregated or uncomplexed states of its constituent molecules). As used herein, the term "complex" refers to an assembly of two or more molecules (e.g., peptides, polypeptides, etc.), unless otherwise specified.

[0032] As used herein, the term "fragment" refers to a peptide or polypeptide that results from dissociation or "fragmentation" of a larger whole entity (e.g., a protein, polypeptide, enzyme, etc.), or that has been prepared to have the same sequence as such. Thus, a fragment is a subsequence of the whole entity (e.g., protein, polypeptide, enzyme, etc.) for which the fragment is made and / or designed. A peptide or polypeptide that is not a subsequence of an existing whole protein is not a fragment (e.g., not a fragment of an existing protein). A peptide or polypeptide that is "not a fragment of an existing protein" is an amino acid chain that is not a subsequence of a protein (e.g., natural or synthetic) that physically existed prior to the design and / or synthesis of the peptide or polypeptide. As used herein, a fragment of a hydrolase or dehalogenase is a sequence that is less than the full-length sequence but is not capable of forming a substrate binding site by itself and / or has substantially reduced or no substrate binding activity, but which in close proximity to a second fragment of the hydrolase or dehalogenase exhibits substantially increased substrate binding activity. In one embodiment, a fragment of a hydrolase or dehalogenase comprises at least 5, e.g., at least 10, at least 20, at least 30, at least 40, or at least 50 contiguous residues of a wild-type or mutant hydrolase, or a sequence having at least 70% sequence identity thereto, and may not necessarily include the N- or C-terminal residues or N- or C-terminal sequences of the corresponding full-length protein.

[0033] As used herein, the term "subsequence" refers to a peptide or polypeptide that has 100% sequence identity with a portion of another larger peptide or polypeptide, the subsequence being a perfect sequence match for a portion of the larger amino acid chain.

[0034] The term "amino acid" refers to natural amino acids, unnatural amino acids, and amino acid analogs, all in their D and L stereoisomeric forms, unless otherwise indicated, if their structures allow for such stereoisomeric forms.

[0035] The term "proteinogenic amino acid" refers to the 20 amino acids encoded for the human genetic code, including alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine ​​(Cys or C), glutamine (Gln or Q), glutamic acid (Glu or E), glycine (Gly or G), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V). Selenocysteine ​​and pyrrolysine may also be considered proteinogenic amino acids.

[0036] The term "non-proteinogenic amino acid" refers to an amino acid that is not naturally encoded or found in the genetic code of any organism and is not biosynthetically incorporated into a protein during translation. A non-proteinogenic amino acid may be a "non-natural amino acid" (an amino acid that does not occur in nature) or a "naturally occurring non-proteinogenic amino acid" (e.g., norvaline, ornithine, homocysteine, etc.). Examples of non-proteinogenic amino acids include, but are not limited to, azetidine carboxylic acid, 2-aminoadipic acid, 3-aminoadipic acid, beta-alanine, naphthylalanine, aminopropionic acid, 2-aminobutyric acid, 4-aminobutyric acid, 6-aminocaproic acid, 2-aminoheptanoic acid, 2-aminoisobutyric acid, 3-aminoisobutyric acid, 2-aminopimelic acid, tertiary butylglycine, 2,4-diaminoisobutyric acid, desmosine, 2,2'-diaminopimelic acid, 2,3-diaminopropionic acid, N-ethylglycine, N-ethylasparagine, homoproline, hydroxylysine, allo-hydroxylysine, 3-hydroxyproline, 4-hydroxyproline, isodesmosine, allo-isoleucine, N-methylalanine, N-alkylglycines including N-methylglycine, N-methylisoleucine, N-alkylpentylglycines including N-methylpentylglycine. Included in the non-proteinaceous structures are N-methylvaline, naphthylalanine, norvaline, norleucine ("Norleu"), octylglycine, ornithine, pentylglycine, pipecolic acid, thioproline, homolysine, and homoarginine. Non-proteinaceous structures also include D-amino acid forms of any of the amino acids herein, as well as non-alpha amino acid forms of any of the amino acids herein (such as beta amino acids, gamma amino acids, delta amino acids, etc.), all of which are within the scope of the present invention and may be included in the peptides herein.

[0037] The term "amino acid analog" refers to an amino acid (e.g., natural or non-natural, proteinogenic or non-proteinogenic) in which one or more of the C-terminal carboxy group, the N-terminal amino group, and the side chain bioactive group are chemically blocked, reversibly or irreversibly, or otherwise modified to a bioactive group. For example, aspartic acid-(beta-methyl ester) is an amino acid analog of aspartic acid, N-ethylglycine is an amino acid analog of glycine, or alanine carboxamide is an amino acid analog of alanine. Other amino acid analogs include methionine sulfoxide, methionine sulfone, S-(carboxymethyl)-cysteine, S-(carboxymethyl)-cysteine ​​sulfoxide, and S-(carboxymethyl)-cysteine ​​sulfone.

[0038] As used herein, unless otherwise specified, the terms "peptide" and "polypeptide" refer to a polymeric compound of two or more amino acids joined through the backbone by peptide amide bonds (-C(O)NH-). The term "peptide" typically refers to short amino acid polymers (e.g., chains having fewer than 30 amino acids) and the term "polypeptide" typically refers to longer amino acid polymers (e.g., chains having more than 30 amino acids).

[0039] As used herein, the term "artificial" refers to compositions and systems that are designed or prepared by man and do not occur in nature, for example, an artificial peptide, peptide, or nucleic acid that contains a non-naturally occurring sequence (e.g., a peptide that does not have 100% identity to a naturally occurring protein or fragment thereof).

[0040] As used herein, a "conservative" amino acid substitution refers to the replacement of an amino acid in a peptide or polypeptide with another amino acid that has similar chemical properties, such as size or charge. For purposes of this disclosure, each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) Alanine (A) and Glycine (G); 2) Aspartic acid (D) and glutamic acid (E); 3) Asparagine (N) and Glutamine (Q); 4) arginine (R) and lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), and Valine (V); 6) phenylalanine (F), tyrosine (Y), and tryptophan (W); 7) serine (S) and threonine (T); and 8) Cysteine ​​(C) and methionine (M).

[0041] Naturally occurring residues may be divided into classes based on common side chain properties, e.g., polar positive (or basic) (histidine (H), lysine (K), and arginine (R)), polar negative (or acidic) (aspartic acid (D), glutamic acid (E)), polar neutral (serine (S), threonine (T), asparagine (N), glutamine (Q)), nonpolar fatty acids (alanine (A), valine (V), leucine (L), isoleucine (I), methionine (M)), nonpolar aromatic (phenylalanine (F), tyrosine (Y), tryptophan (W)), proline and glycine, and cysteine. As used herein, a "semi-conservative" amino acid substitution refers to the replacement of an amino acid in a peptide or polypeptide with another amino acid within the same class.

[0042] In some embodiments, unless otherwise specified, conservative or semi-conservative amino acid substitutions may also include non-naturally occurring amino acid residues that have similar chemical properties to the natural residues. These non-natural residues are typically incorporated by chemical peptide synthesis rather than by synthesis in biological systems. These include, but are not limited to, peptidomimetics and other reversed or inverted forms of amino acid moieties. The embodiments herein may, in some embodiments, be limited to natural amino acids, non-natural amino acids, and / or amino acid analogs.

[0043] Non-conservative substitutions may involve exchanging a member of one class for a member of another class.

[0044] As used herein, the term "sequence identity" refers to the degree to which two polymeric sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have the same sequential composition of monomeric subunits. The term "sequence similarity" refers to the degree to which two polymeric sequences (e.g., peptides, polypeptides, nucleic acids, etc.) have similar polymeric sequences. For example, similar amino acids are those that share the same biophysical properties and can be grouped, for example, into acidic (e.g., aspartate, glutamate), basic (e.g., lysine, arginine, histidine), non-polar (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), and uncharged polar (e.g., glycine, asparagine, glutamine, cysteine, serine, threonine, tyrosine). "Percent sequence identity" (or "percent sequence similarity") is calculated by: (1) comparing two optimally aligned sequences over a window of comparison (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); (2) determining the number of positions that contain identical (or similar) monomers (e.g., the same amino acid occurs in both sequences, a similar amino acid occurs in both sequences) to obtain the number of matched positions; (3) dividing the number of matched positions by the total number of positions in the comparison window (e.g., the length of the longer sequence, the length of the shorter sequence, a specified window); and (4) multiplying the result by 100 to obtain the percent sequence identity or percent sequence similarity. For example, if peptide A and peptide B are both 20 amino acids long and have identical amino acids at all but one position, peptide A and peptide B have 95% sequence identity. If the amino acids at the non-identical positions share the same biophysical characteristics (e.g., both were acidic), peptide A and peptide B will have 100% sequence similarity.As another example, if peptide C is 20 amino acids long, peptide D is 15 amino acids long, and 14 of the 15 amino acids in peptide D are identical to a portion of peptide C, then peptides C and D have 70% sequence identity, but peptide D has 93.3% sequence identity with the optimal comparison window of peptide C. For purposes of calculating "percent sequence identity" (or "percent sequence similarity") herein, any gap in the aligned sequences is treated as a mismatch at that position.

[0045] Any peptide / polypeptide described herein as having a particular percent sequence identity or similarity (e.g., at least 70%) with a reference sequence ID number may also be expressed as having a maximum number of substitutions (or terminal deletions) relative to that reference sequence. For example, a sequence having at least Y% sequence identity (e.g., 90%) with SEQ ID NO: Z (e.g., 100 amino acids) may have a maximum of X substitutions (e.g., 10) with SEQ ID NO: Z, and thus may also be expressed as "having no more than X (e.g., 10) substitutions with SEQ ID NO: Z."

[0046] As used herein, the term "wild type" refers to a gene or gene product (e.g., a protein, polypeptide, peptide, etc.) that has the characteristics (e.g., sequence) of that gene or gene product isolated from a naturally occurring source and is most frequently observed in a population. In contrast, the term "mutant" or "variant" refers to a gene or gene product that exhibits a modification of the sequence compared to the wild-type gene or gene product. Note that a "naturally occurring variant" is a gene or gene product that occurs in nature but has an altered sequence compared to the wild-type gene or gene product, and is not the most commonly occurring sequence. An "artificial variant" is a gene or gene product that has an altered sequence when compared to the wild-type gene or gene product and does not occur in nature. A variant gene or gene product may be naturally occurring but not the most common variant of the gene or gene product, or "synthetic" produced by human or experimental intervention.

[0047] As used herein, the term "physiological conditions" encompasses any conditions compatible with living cells, e.g., primarily aqueous conditions of temperature, pH, salinity, chemical composition, and the like, that are compatible with living cells.

[0048] As used herein, the term "sample" is used in its broadest sense. In one sense, it is meant to include specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from animals (including humans) and encompass fluids, solids, tissues, and gases. Biological samples include blood products such as plasma, serum, etc. Samples can also refer to cell lysates or purified forms of the enzymes, peptides, and / or polypeptides described herein. Cell lysates can include cells lysed with a lysing agent, or lysates, e.g., rabbit reticulocyte or wheat germ lysates. Samples can also include cell-free expression systems. Environmental samples include environmental materials, such as surface materials, soil, water, crystals, and industrial samples. However, such examples should not be construed as limiting the sample types applicable to the present invention.

[0049] As used herein, the terms "fusion," "fusion polypeptide," and "fusion protein" refer to a chimeric protein that contains a first protein or polypeptide of interest joined to a second, different peptide, polypeptide, or protein (e.g., an interacting element).

[0050] As used herein, the terms "conjugate" and "conjugation" refer to the covalent joining of two molecular entities (e.g., post-synthetic and / or during synthetic production). The chemical (e.g., "chemically" conjugated) or enzymatic attachment of a peptide or small molecule tag to a protein or small molecule is an example of a conjugate.

[0051] As used herein, the term "dehalogenase" refers to an enzyme that catalyzes the removal of halogen atoms from a substrate. The term "haloalkane dehalogenase" refers to an enzyme that catalyzes the removal of halogens from a haloalkane substrate to produce alcohols and halides. Dehalogenases and haloalkyl dehalogenases belong to the hydrolase enzyme family and may be referred to as such herein or elsewhere.

[0052] As used herein, the term "modified dehalogenase" refers to a dehalogenase variant (artificial variant) that has a mutation that prevents release of the substrate from the protein after removal of the halogen and results in a covalent bond between the substrate and the modified dehalogenase. The HALOTAG system (Promega) is a commercially available modified dehalogenase and substrate system.

[0053] As used herein, the term "circularly permuted" ("cp") refers to a polypeptide in which the N-terminus and C-terminus are joined together, either directly or through a linker, to produce a circular polypeptide, which is then opened at a location other than between the N-terminus and C-terminus to produce a new linear polypeptide that differs from the ends of the original polypeptide. The location at which the circular polypeptide is opened is referred to herein as the "cp site." Circularly permuted polypeptides include polypeptides that have the same sequence and structure as a circularly permuted and then opened polypeptide. Thus, cp polypeptides may be synthesized de novo as linear molecules, without undergoing the circular permutation and opening steps. The preparation of circularly permuted derivatives is described in International Publication No. WO 95 / 27732, which is incorporated by reference in its entirety.

[0054] As used herein, the term "luminescence" refers to the emission of light by a substance as the result of a chemical reaction ("chemiluminescence") or an enzymatic reaction ("bioluminescence").

[0055] As used herein, the term "bioluminescence" refers to the production and emission of light by a reaction catalyzed or enabled by an enzyme, protein, protein complex, or other biological molecule (e.g., a bioluminescent complex). In a typical embodiment, a substrate of a bioluminescent entity (e.g., a bioluminescent protein or bioluminescent complex) is converted by the bioluminescent entity to an unstable form, which subsequently emits light.

[0056] As used herein, the term "luminophore" refers to a chemical moiety or compound that can be placed in an excited electronic state (e.g., by a chemical or enzymatic reaction) and emit light upon return to the ground electronic state.

[0057] As used herein, the term “imidazopyrazine luminophore” refers to “natural coelenterazine,” as well as synthetic (e.g., derivatives or variants) and natural analogs thereof (e.g., furimazines), in addition to those disclosed in WO 2003 / 040100, U.S. Application No. 12 / 056,073 (paragraph

[0086] ), U.S. Patent No. 8,669,103, and U.S. Provisional Application No. 63 / 379,573, the disclosures of which are incorporated herein by reference in their entireties. Coelenterazine refers to a genus of luminophores that include furimazine analogs (e.g., fluorofurimazine), including coelenterazine-N, coelenterazine-F, coelenterazine-H, coelenterazine-HCP, coelenterazine-CP, coelenterazine-C, coelenterazine-E, coelenterazine-FCP, bis-deoxycoelenterazine ("coelenterazine-HH"), coelenterazine-I, coelenterazine-ICP, coelenterazine-V, and 2-methylcoelenterazine.

[0058] As used herein, the term "coelenterazine" refers to a naturally occurring ("natural") imidazopyrazine of the following structure: [ka]

[0059] As used herein, the term "furimazine" refers to a coelenterazine derivative of the following structure: [ka]

[0060] As used herein, the term "fluorofurimazine" refers to a furimazine derivative of the following structure: [ka] (U.S. Application Serial No. 16 / 548,214, incorporated by reference in its entirety).

[0061] As used herein, the term "bioluminescence resonance energy transfer" ("BRET") refers to a distance-dependent interaction in which energy is transferred from a donor bioluminescent protein / complex and substrate to an acceptor molecule without the emission of a photon. The efficiency of BRET depends on the inverse sixth power of the intermolecular separation, making it useful over distances comparable to the dimensions of biological macromolecules (e.g., within 30-80 Å, depending on the degree of spectral overlap).

[0062] As used herein, the term "Oplophorus luciferase" ("OgLuc") refers to a light-emitting polypeptide having significant sequence identity, structural conservation, and / or functional activity of luciferase produced by and derived from the deep-sea shrimp Oplophorus gracilirostris. Specifically, OgLuc polypeptide refers to a light-emitting polypeptide having significant sequence identity, structural conservation, and / or functional activity of the mature 19 kDa subunit of the Oplophorus luciferase protein complex, such as SEQ ID NO:28 (NANOLUC), which contains 10 β-strands (β1, β2, β3, β4, β5, β6, β7, β8, β9, β10) and utilizes a substrate, such as coelenterazine or a coelenterazine derivative or analog, to generate light. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0063] Provided herein are modified dehalogenases with extended surface loop regions that provide locations for internal fusion insertions and modulate binding interactions, energy transfer and activation of environmentally sensitive chemicals.

[0064] The development of new fluorophores and fluorogenic dyes (such as JANELIA FLUOR dyes) for use with chloroalkanes (CAs) highlights renewed interest in HALOTAG for fluorescence detection in cellular imaging applications. The advantages of such dyes in brightness, photostability, sensitivity, and far-red spectral detection over traditional tools such as the widely used fluorescent proteins are particularly evident in challenging or highly sensitive imaging applications in endogenous biology. As chloroalkane conjugates, they can take advantage of the self-labeling activity of HALOTAG to measure protein abundance and localization in a target-specific manner via gene fusion.

[0065] Recent evidence supports a model of activation of rhodamine-based fluorogenic dyes bound to chloroalkanes through physical interactions with the surface of HALOTAG after binding. This interaction shifts the equilibrium of the dye from its non-fluorescent lactone state toward its fluorescent zwitterionic state. Much effort has been devoted to chemical modification of the dye scaffold itself to enhance this effect through changes that modulate the lactone-zwitterionic structural equilibrium. However, chemical modifications of the dye structure that push the equilibrium toward the zwitterionic state to enhance fluorescence also tend to reduce the cell permeability of the ligand, and similarly, those that favor the lactone state enhance permeability at the expense of fluorescence yield. Compared to chemical modifications of fluorogenic dyes, modifications to HALOTAG itself have been less well studied. Point mutations in HALOTAG have been shown to enhance fluorogenicity by making protein:dye interactions more favorable for fluorescence (Frei et al. Engineered HaloTag variants for fluorescence lifetime multiplexing. Nature Methods volume 19, pages 65-70 (2022), incorporated by reference in its entirety).

[0066] During the development of the embodiments herein, experiments were performed to improve the activation of chemical dyes, such as modified dehalogenases, fluorogenic dyes on the surface of HALOTAG. It was reasoned that one ideal solution would involve an engineered protein surface for optimal interactions that would improve fluorescence activation upon binding. Point mutations at surface residues have been shown to be one such solution, but they are inherently limited in positioning and sequence availability of the native HALOTAG protein scaffold. For example, efforts were made to engineer extended loops into the HALOTAG structure in regions proximal to the dye interaction site to provide additional interaction surface area and greatly increase the available sequence space for optimization. In addition to the increased interaction surface, this solution also provides a new binding mechanism between the dye and the protein that is only achievable through the conformation of the extended loop, thereby providing an entirely new chemical activation scheme. Thus, the range of activatable chemicals is significantly increased in proportion to the vastly new protein sequence space and structures available in the extended loop region. The utility of the extended loop is not limited to improved dye activation and / or interaction with the substrate, and such activation / interaction is not necessary to practice the present invention.

[0067] The extended HALOTAG loop is used for activation of fluorogenic dyes, but can also be extended to a wide range of environmentally sensitive CA-conjugated chemicals that are activated by the optimized binding surface or pocket formed through the engineered loop sequence on the surface of HALOTAG. Thus, engineered "loop HALOTAG" variants may be tailored for activation of environmentally sensitive chemicals in a robust and orthogonal manner after binding. For example, the extended loop is used to enhance dye / chemical activation via BRET, and the extended loop is utilized to further engineer chimeras of HALOTAG with bioluminescent reporters to improve the efficiency of BRET-based activation through more favorable proximity / geometry for BRET between the bioluminescent reporter and the bound ligand. This is particularly important when the spectral overlap between the bioluminescent reporter emission and the ligand excitation is significantly limited. One downstream application of this improved efficiency is the use of a bioluminescent light source as an activator of downstream chemicals.

[0068] The embodiments herein are not limited to enhancing the interaction between the loop and the ligand or interaction partner. In some embodiments, the regions identified herein (e.g., loop 165, loop 180, loop 194 / 195) are used as positions for inserting peptides or polypeptides into the HALOTAG sequence. For example, the extended loop also provides a position for inserting larger polypeptides, such as proteins or enzymes, into HALOTAG for optimal positioning or geometry close to the bound ligand. In some embodiments, chimeras formed at the internal loop sites increase the efficiency of energy transfer between the inserted protein and the HALOTAG ligand via BRET or FRET, especially when the spectral overlap between the emission of the inserted reporter and the excitation of the HALOTAG ligand is significantly limited. For example, it has been demonstrated that circularly permuted NANOLUC luciferase (cpNL) increases the efficiency of BRET with fluorescent HALOTAG ligands when inserted into a position within the HALOTAG lid domain (Hiblot, J., et al. (2017) Angew Chem Int Ed Engl 56(46):14556-14560, incorporated by reference in its entirety). In some embodiments, this strategy provides a solution for similarly increasing FRET efficiency, for example, when a fluorescent protein (e.g., GFP, RFP, etc.) is inserted into the loop region disclosed herein proximal to the fluorescent HALOTAG ligand.

[0069] Evidence from structural analysis of HALOTAG bound to fluorescent and fluorogenic ligands in parallel with mutational studies supports a model of fluorescence activation of rhodamine-based dyes via surface contacts between HALOTAG and the dye moiety. Specifically, Helix 8 of HALOTAG (residues approximately 167-176) is positioned to make direct contact with the dye in some structures. Experiments performed during the development of embodiments herein demonstrated that deletions, circular permutations, and / or splits in or proximal to this region of HALOTAG eliminate fluorogenic activation of the ligand. It was therefore deemed reasonable that point mutations at the surface of HALOTAG might make these interactions more favorable, ultimately resulting in increased fluorogenic activity. However, introducing point mutations at the binding surface of HALOTAG, while a successful strategy thus far, is ultimately limited by the location of existing residues in the protein scaffold, risking perturbing the folding of the protein, which also contribute to its native structure. In addition, due to their close proximity to the dye moiety of the ligand, only a small number of residues may have the potential for optimization, limiting the utility of this approach.

[0070] To increase the binding surface and configuration available for interaction optimization, experiments were performed during development of the embodiments herein to introduce additional protein sequences into this critical region of HALOTAG. The loop regions adjacent to Helix 8 of HALOTAG were targeted for modification because (1) loops are generally more tolerant to insertions, modifications, or deletions without significantly disrupting protein folding or function, and (2) they are in close proximity to the bound dye in the crystal structure, allowing newly inserted residues to be positioned within a distance where interactions can be formed. The two loops adjacent to Helix 8 are designated herein as Loop 165 (residues 164-166) and Loop 180 (residues 177-182), both of which are within the lid subdomain of HALOTAG that constitutes the majority of the ligand binding tunnel and the surface-exposed tunnel opening (Figure 1). At these locations, empirical steps were taken to engineer extended loop regions into HALOTAG. Optimal sites for insertion of residues within Loop 165 or Loop 180 were identified. Preliminary screening was performed to identify several sequence insertions of 7-15 residues in length that resulted in loop HALOTAG variants with unique activity profiles, demonstrating the utility of this concept. To further demonstrate the feasibility of introducing relatively large sequences (e.g., bioluminescent reporters), during the development of the embodiments herein, experiments were performed to insert bioluminescent reporters into the extended loop. The resulting chimeras were used for activation of a bound dye via a unique intramolecular bioluminescent resonance energy transfer mechanism.

[0071] The extended surface loops offer various advantages that are expected to improve and / or extend the capabilities and applications of HALOTAG. First, similar to the complementarity determining region (CDR) loops of antibodies, the extended surface loops can adopt diverse conformations composed of different amino acid sequences, making them suitable for highly divergent but specific binding modes. There are examples of antibodies and other binding scaffolds (e.g., DARPINS, scFVs, and nanobodies) that have been engineered to bind small molecules, including fluorogenic dyes, in a manner that increases fluorescence. However, specific recognition of small molecules by antibodies is not straightforward to engineer, and structural and biophysical analyses have revealed that binding is commonly achieved by dimerization of antibodies around the small molecule target, essentially creating a binding pocket between the monomers. In some embodiments, the advantage of molecular recognition through the extended loops in HALOTAG overcomes this challenge, since binding is already achieved by robust interactions and autolabeling activity with CA in the monomeric complex. In this scenario, the covalent attachment of CA to HALOTAG positions the conjugated small molecule cargo on its surface and allows residues in the proximal extended loop region to interact, thereby reducing the engineering burden required for activation by also eliminating the need to engineer robust and specific ligand affinities.

[0072] Molecular recognition by the extended surface loop in HaloTag is not limited to the purpose of activating CA conjugates. In some embodiments, the extended loop interacts with intermolecular binding partners, such as other proteins, similar to antibody recognition, targeting HALOTAG (and its bound CA ligand) to a specific target, for example, intracellularly or as part of a diagnostic assay. These configurations of extended loop HALOTAG retain many of the advantages of antibodies, but also include the ability to genetically encode the construct and deliver the ligand of interest as a CA conjugate in close proximity to the protein target. Beyond molecular recognition, the utility offered by the extended HALOTAG loop allows for new conformations and geometries of the chimeric protein inserted within the loop. For example, larger polypeptides can be engineered into favorable distances and geometries, allowing for more efficient energy transfer between the inserted polypeptide (such as a bioluminescent enzyme) and the bound HALOTAG ligand. This is particularly important when there is limited spectral overlap between the emission of the bioluminescent reporter and the excitation of the HaloTag ligand, and the distance and geometry within the chimera are important for energy transfer.

[0073] The additional capabilities of the extended loop HALOTAG design impart capacity for molecular interactions that expand the useful applications of HALOTAG. For example:

[0074] 1. The extension loops allow for an increase in the fluorescence (or range of fluorescence activation) of HALOTAG fluorogenic ligands, such as those currently commercially available (i.e., CA-Janelia Fluor dyes, Promega corp., Madison, WI). The increase in fluorescence is realized as either signal intensity or fluorescence lifetime in the presence of engineered extension loops in HALOTAG. The difference in fluorescence lifetime has been shown to be valuable for HALOTAG-9 / 10 / 11 multiplexing in fluorescence imaging (Frei, M. et al (2022). Nature Methods. (19) 65-70., incorporated by reference in its entirety). Additional categories of novel applications of existing HaloTag fluorescent / fluorogenic ligands include BRET and FRET-based applications, where chimeras are created using these extension loops as insertion sites to create chimeras with bioluminescent or fluorescent proteins. For BRET, some applications include a) BRET as a means to tune the emission of NANOLUC-based bioluminescent reporters for cell / animal imaging, b) sorting HIBIT-edited cells where labeling is dependent on complementation with LGBIT, c) BRET-induced activation of light-sensitive molecules including catalysts, and d) BRET-induced bioluminescent lysis.

[0075] 2. HALOTAG Fluorogenic Ligand System. Provided herein are extended loop HALOTAG variants with CA fluorogenic dyes that allow for greater fluorescence yield or signal-to-background upon activation. In some embodiments, the CA fluorogenic dyes do not have significant activation in unmodified HALOTAG. For example, certain Janelia Fluor dyes, for example, have a stronger natural preference for the non-fluorescent lactone state (which is more cell permeable), but are more difficult to transition to the fluorescent zwitterionic state without the additional stabilizing molecular interactions provided by the optimized extended surface in the extended loop modified dehalogenases herein. Such improved systems are used, for example, in cell imaging, where the simultaneous reduction of background signal of the non-fluorescent free ligand and better potential activation of the bound ligand, in addition to the better cell / tissue permeability of the lactone state dye, creates an overall better signal-to-background ratio for imaging.

[0076] 3. Chemicals specifically compatible / activatable with engineered loop-modified dehalogenase variants. Beyond fluorogenic dyes, there are numerous commercially valuable ligands, such as catalysts, biosensors, and proximity labels, that are used as CA conjugates and undergo stabilization of their structural transitions by interaction with extended loop-modified dehalogenases. Such systems are configured to allow a tunable range of responses. For example, BAPTA-CA ligands have been shown to be intracellular indicators that undergo conformational changes and increase their fluorescence upon chelation of Ca2+ ions, making them synthetic biosensors sensitive to Ca2+ flux in live cells. The Ca2+ response of BAPTA-CA ligands can be chemically tuned over a range of affinities, but typically at the expense of quantum yield. Optimized extended loop-modified dehalogenases provide BAPTA-CA responses to physiologically relevant Ca2+ levels with higher quantum yields in a manner that cannot be achieved by synthetic chemical modification of the ligand alone. As another example, calcium-indicating / chelating moieties alter the fluorogenicity of CA dyes in an affinity- and color-tunable manner, which has been particularly useful for providing calcium indicators in the red (approximately 650 nm) range of detection (Mertes et al. J. Am. Chem. Soc. 2022, 144, 15, 6928-6935, incorporated by reference in its entirety).

[0077] 4. Affinity Reagents Based on Extended Loop-Modified Dehalogenases. Extended loop-modified dehalogenases that recognize other molecular targets such as proteins offer a wide range of utilities such as stand-alone affinity reagents, purification / enrichment systems, diagnostics, imaging tools, or genetically encodable intracellular bioassays. All of these systems would benefit from localization of the CA ligand upon binding of the extended loop-modified dehalogenase to its target.

[0078] The modified dehalogenases, systems, and methods herein are not limited by the particular utilities and uses described herein, and an understanding of the utility or use of the modified dehalogenase is not necessary to practice the invention. Any embodiment that includes a modified dehalogenase having an amino acid sequence inserted internally at one of the positions described herein is within the scope of the present invention. Enhanced ability to activate a substrate or effect an interaction is not required for modified dehalogenases with internal insertions to be within the scope of the present invention.

[0079] I. Modified Dehalogenases In some embodiments, provided herein are modified dehalogenases with internal insertions. In some embodiments, the modified dehalogenase is the commercially available HALOTAG protein (SEQ ID NO: 1), or a variant thereof (e.g., greater than 70% sequence identity). HALOTAG is a 297-residue self-labeling polypeptide (33 kDa) derived from a bacterial hydrolase (dehalogenase) enzyme modified to covalently bind to its ligand, a haloalkane moiety. The HALOTAG ligand can be linked to a solid surface (e.g., a bead) or a functional group (e.g., a fluorophore), and the HALOTAG polypeptide can be fused to various proteins of interest, allowing for covalent binding of the protein of interest to a solid surface or functional group.

[0080] HALOTAG polypeptides are hydrolases with genetically modified active sites that specifically bind haloalkane ligand chloroalkane linkers, enhancing and increasing the rate of ligand binding (Pries et al. The Journal of Biological Chemistry. 270(18):10405-11, incorporated by reference in its entirety). The reaction that forms the bond between the protein tag and the chloroalkane linker is fast and essentially irreversible under physiological conditions (Waugh DS (June 2005). Trends in Biotechnology. 23(6):316-20; incorporated by reference in its entirety). In native hydrolase enzymes, nucleophilic attack of the chloroalkane reactive linker causes displacement of the halogen by an amino acid residue, resulting in the formation of a covalent alkyl enzyme intermediate. This intermediate is then hydrolyzed by amino acid residues in the wild-type hydrolase (Chen et al. (February 2005) Current Opinion in Biotechnology. 16(1):35-40; incorporated by reference in its entirety). This would result in regeneration of the enzyme after the reaction. However, in the modified haloalkane dehalogenase, HALOTAG, the reaction intermediate cannot proceed to a second reaction because a mutation in the enzyme prevents it from being hydrolyzed. This causes the intermediate to persist as a stable covalent adduct with no associated back reaction (Marks et al. (August 2006) Nature Methods. 3(8):591-6; incorporated by reference in its entirety).

[0081] HALOTAG fusion proteins can be expressed using standard recombinant protein expression techniques (Adams et al. (May 2002) Journal of the American Chemical Society. 124(21):6063-76; incorporated by reference in its entirety). Because the HALOTAG polypeptide is a relatively small protein and the reaction is foreign to the mammalian cell, there is no interference from endogenous mammalian metabolic reactions (Naested et al. The Plant Journal. 18(5):571-6; incorporated by reference in its entirety). Once the fusion protein is expressed, there is a wide range of potential experimental fields including enzyme assays, cell imaging, protein arrays, determination of subcellular localization, and many additional possibilities (Janssen DB (April 2004). Current Opinion in Chemical Biology. 8(2):150-9; incorporated by reference in its entirety).

[0082] Various HALOTAG ligands, functional groups, fusions, assays, modifications, uses, and the like are described in U.S. Patent No. 8,748,148, U.S. Patent No. 9,593,316, U.S. Patent No. 10,246,690, U.S. Patent No. 8,742,086, U.S. Patent No. 9,873,866, U.S. Patent No. 10,604,745, U.S. Patent Application No. 2009 / 0253131, U.S. Patent Application No. 2010 / 0273186, 20130337539, U.S. Patent Application No. 2012 / 0258470, U.S. Patent Application No. 2012 / 0252048, U.S. Patent Application No. 2011 / 0201024, U.S. Patent Application No. 2014 / 0322794, each of which is incorporated by reference in its entirety.

[0083] As described herein, embodiments are not limited to the HALOTAG sequence. In some embodiments, provided herein are split modified dehalogenases that differ in sequence from SEQ ID NO:1. In some embodiments, provided herein are split dehalogenases that lack the mutation(s) (e.g., 272 and / or 106) that result in covalent attachment to the haloalkane substrate. Such sp dehalogenases allow for substrate conversion but are otherwise true enzymes that include the sequences and properties of the embodiments described herein.

[0084] In some embodiments, provided herein are polypeptides and fusions derived from the modified dehalogenase sequence of SEQ ID NO:1. MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETF QAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG.

[0085] In some embodiments, the modified dehalogenase polypeptides herein comprise at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:1. In some embodiments, the polypeptides herein comprise 100% sequence identity to all or a portion of SEQ ID NO:1. In some embodiments, the polypeptides herein comprise at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:1. In some embodiments, the polypeptides herein comprise 100% sequence similarity with all or a portion of SEQ ID NO:1.

[0086] In some embodiments, provided herein are modified dehalogenase polypeptides that have at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to SEQ ID NO:1, but include an extended loop sequence (e.g., 1-25 amino acids in length) or a peptide or polypeptide insertion at a position or sequence within SEQ ID NO:1 (e.g., replacing loop 165, replacing loop 180, replacing loop 194 / 195, following position 165, following position 180, following position 194, etc.).

[0087] In some embodiments, provided herein are modified dehalogenase polypeptides comprising an insertion of up to 25 amino acids in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids, or ranges therebetween) within loop 165 of SEQ ID NO: 1. In some embodiments, provided herein are polypeptides that comprise at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:2. In some embodiments, the polypeptides herein comprise 100% sequence identity to all or a portion of SEQ ID NO: 2. In some embodiments, the polypeptides herein comprise at least 70% sequence similarity to all or a portion of SEQ ID NO: 2 (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity). In some embodiments, the polypeptides herein comprise 100% sequence similarity to all or a portion of SEQ ID NO: 2.

[0088] In some embodiments, provided herein are modified dehalogenase polypeptides that include an insertion of up to 25 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids, or any range therebetween) in length at a position corresponding to the position following position 165 of SEQ ID NO: 1. In some embodiments, provided herein are polypeptides that include at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:3. In some embodiments, the polypeptides herein comprise 100% sequence identity to all or a portion of SEQ ID NO: 3. In some embodiments, the polypeptides herein comprise at least 70% sequence similarity to all or a portion of SEQ ID NO: 3 (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity). In some embodiments, the polypeptides herein comprise 100% sequence similarity to all or a portion of SEQ ID NO: 3.

[0089] In some embodiments, provided herein are modified dehalogenase polypeptides that include an insertion of up to 25 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids, or ranges therebetween) in length within loop 180 of SEQ ID NO: 1. In some embodiments, provided herein are polypeptides that include at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:4. In some embodiments, the polypeptides herein comprise 100% sequence identity to all or a portion of SEQ ID NO: 4. In some embodiments, the polypeptides herein comprise at least 70% sequence similarity to all or a portion of SEQ ID NO: 4 (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity). In some embodiments, the polypeptides herein comprise 100% sequence similarity to all or a portion of SEQ ID NO: 4.

[0090] In some embodiments, provided herein are modified dehalogenase polypeptides that include an insertion of up to 25 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids, or any range therebetween) in length at a position corresponding to the position following position 180 of SEQ ID NO: 1. In some embodiments, provided herein are polypeptides that include at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:5. In some embodiments, the polypeptides herein comprise 100% sequence identity with all or a portion of SEQ ID NO: 5. In some embodiments, the polypeptides herein comprise at least 70% sequence similarity with all or a portion of SEQ ID NO: 5 (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity). In some embodiments, the polypeptides herein comprise 100% sequence similarity with all or a portion of SEQ ID NO: 5.

[0091] In some embodiments, provided herein are modified dehalogenase polypeptides that include a peptide or polypeptide (e.g., a protein) inserted at an internal position (e.g., replacing loop 165, replacing loop 180, replacing loop 194 / 195, after position 165, after position 180, after position 194, etc.). In some embodiments, the inserted sequence is 1, 2, 5, 10, 20, 50, 100, 150, 200, 250, 300, 400, 500, or more amino acids in length. In some embodiments, the inserted sequence and modified dehalogenase each retain all or a portion (e.g., greater than 10%, greater than 25%, greater than 50%, greater than 75%, greater than 90%) of their activity and / or function (e.g., substrate binding ability).

[0092] In some embodiments, provided herein are modified dehalogenase polypeptides that include a peptide or polypeptide insertion within a loop corresponding to loop 165 of SEQ ID NO: 1. In some embodiments, provided herein are modified dehalogenase polypeptides that include a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of one of SEQ ID NOs: 6-9 fused to the C-terminus of the peptide or polypeptide insert sequence. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of one of SEQ ID NOs: 10-13 fused to the N-terminus of the peptide or polypeptide insert sequence.

[0093] In some embodiments, provided herein is a peptide or polypeptide insert comprising a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:6 fused to the C-terminus of the peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 10 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:6. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:6. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:6. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:10. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 10. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 10.

[0094] In some embodiments, provided herein is a peptide or polypeptide insert comprising a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:7 fused to the C-terminus of the peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 11 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:7. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:7. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:7. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:11. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:11. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:11.

[0095] In some embodiments, provided herein is a peptide or polypeptide insert comprising a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:8 fused to the C-terminus of the peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 12 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:8. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:8. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:8. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:12. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 12. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 12.

[0096] In some embodiments, provided herein is a peptide or polypeptide insert comprising a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:9 fused to the C-terminus of the peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 13 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:9. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:9. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:9. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:13. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 13. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 13.

[0097] In some embodiments, provided herein are modified dehalogenase polypeptides that include a peptide or polypeptide insertion within a loop corresponding to loop 180 of SEQ ID NO: 1. In some embodiments, provided herein are a first dehalogenase polypeptide having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of one of SEQ ID NOs: 14-20 fused to the C-terminus of the peptide or polypeptide insert sequence. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:21-27 fused to the N-terminus of the peptide or polypeptide insert.

[0098] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 14 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 21 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 14. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 14. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO: 14. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:21. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:21. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:21.

[0099] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 15 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 22 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 15. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 15. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO: 15. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 22. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:22. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:22.

[0100] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 16 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 23 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 16. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 16. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO: 16. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:23. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:23. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:23.

[0101] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 17 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 24 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 17. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 17. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO: 17. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:24. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:24. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:24.

[0102] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 18 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 25 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 18. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO: 18. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO: 18. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:25. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:25. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:25.

[0103] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:19 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 26 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 19. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:19. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:19. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:26. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:26. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:26.

[0104] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:20 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:27 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:20. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:20. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:20. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:27. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO:27. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO:27.

[0105] In some embodiments, provided herein are modified dehalogenase polypeptides that include a peptide or polypeptide insertion within a loop corresponding to loop 194 / 195 of SEQ ID NO: 1. In some embodiments, provided herein are a first dehalogenase polypeptide having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of one of SEQ ID NOs: 81-85 fused to the C-terminus of the peptide or polypeptide insert sequence. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:86-90 fused to the N-terminus of the peptide or polypeptide insert.

[0106] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:81 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 86 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 81. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:81. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:81. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:86. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 86. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 86.

[0107] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:82 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 87 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 82. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:82. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:82. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:87. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 87. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 87.

[0108] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:83 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 88 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 83. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:83. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:83. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:88. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 88. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 88.

[0109] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:84 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 89 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO: 84. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:84. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:84. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:89. In some embodiments, the second sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) with all or a portion of SEQ ID NO: 89. In some embodiments, the second sequence comprises 100% sequence similarity with all or a portion of SEQ ID NO: 89.

[0110] In some embodiments, provided herein is a first sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO:85 fused to the C-terminus of a peptide or polypeptide insert. and a second sequence having at least 70% sequence identity (e.g., greater than 70% sequence identity, greater than 75% sequence identity, greater than 80% sequence identity, greater than 85% sequence identity, greater than 90% sequence identity, greater than 95% sequence identity, greater than 96% sequence identity, greater than 97% sequence identity, greater than 98% sequence identity, greater than 99% sequence identity) to all or a portion of SEQ ID NO: 90 fused to the N-terminus of the peptide or polypeptide insert. In some embodiments, the first sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:85. In some embodiments, the first sequence comprises at least 70% sequence similarity (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity) to all or a portion of SEQ ID NO:85. In some embodiments, the first sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:85. In some embodiments, the second sequence comprises 100% sequence identity to all or a portion of SEQ ID NO:90. In some embodiments, the second sequence comprises at least 70% sequence similarity to all or a portion of SEQ ID NO:90 (e.g., greater than 70% sequence similarity, greater than 75% sequence similarity, greater than 80% sequence similarity, greater than 85% sequence similarity, greater than 90% sequence similarity, greater than 95% sequence similarity, greater than 96% sequence similarity, greater than 97% sequence similarity, greater than 98% sequence similarity, greater than 99% sequence similarity). In some embodiments, the second sequence comprises 100% sequence similarity to all or a portion of SEQ ID NO:90. In some embodiments, provided herein are circular permutations of modified dehalogenases described herein (e.g.,In some embodiments, the circularly permuted variants have sequences inserted into the 165 loop and / or the 180 loop. In some embodiments, the circularly permuted variants have sequences inserted into any position between 5 and 290 of SEQ ID NO:1 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 1, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 19 9, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259,260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290). In some embodiments, circularly permuted variants are between positions 5 and 13 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, or a range therebetween), between positions 36 and 51 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, or a range therebetween), between positions 63 and 72 (e.g., 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, or a range therebetween) of SEQ ID NO:1. ), between 84 and 92 (e.g., 84, 85, 86, 87, 88, 89, 90, 91, 92, or any range therebetween), between 104 and 130 (e.g., 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, or any range therebetween), between 142 and 148 (e.g., 14 2, 143, 144, 145, 146, 147, 148, and ranges therebetween), between 160 and 174 (e.g., 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, and ranges therebetween), between 186 and 189 (e.g., 186, 187, 188, 189, and ranges therebetween), between 201 and 203 (e.g., 201, 202, 203, and ranges therebetween), cp sites at positions corresponding to positions 221 and 229 (e.g., 221, 222, 223, 224, 225, 226, 227, 228, 229, or ranges therebetween), or between positions 269 and 290 (e.g., 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, or 290, or ranges therebetween).

[0111] In some embodiments, the cp-modified dehalogenase includes a first segment having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to a first portion of one of SEQ ID NOs:2-5, and a second segment having at least 70% (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 100%) sequence identity to the first portion of one of SEQ ID NOs:2-5.

[0112] In some embodiments, the polypeptides herein retain the ability of the modified dehalogenase to form a stable bond (eg, a covalent bond) with a haloalkane substrate.

[0113] Circularly permuted modified dehalogenase variants (e.g., cpHT) are described in U.S. Provisional Application No. 63 / 338,364 and U.S. Application No. 18 / 311,977, which are incorporated herein by reference in their entireties. In some embodiments, the circularly permuted modified dehalogenase comprises an extended surface loop and / or an insertion of loops 165, 180, and / or 194 / 195. For example, any of the modified dehalogenase sequences provided herein may be provided as a circularly permuted version thereof (e.g., with any suitable cp site described therein). Similarly, any of the cp modified dehalogenases (e.g., cpHT) described in U.S. Provisional Application No. 63 / 338,364 and / or U.S. Application No. 18 / 311,977 may be provided with an extended surface loop and / or an insertion of loops 165, 180, and / or 194 / 195.

[0114] Split modified dehalogenase variants (e.g., spHT) are described in U.S. Provisional Application No. 63 / 338,323 and U.S. Application No. 18 / 312,117, which are incorporated herein by reference in their entireties. In some embodiments, split modified dehalogenases are provided that include an extended surface loop and / or an insertion of loops 165, 180, and / or 194 / 195. For example, any of the modified dehalogenase sequences provided herein may be provided as a split version thereof (e.g., with any suitable sp site described therein). Similarly, any sp modified dehalogenase (e.g., spHT) described in U.S. Provisional Application No. 63 / 338,323 and / or U.S. Application No. 18 / 312,117 may be provided with an extended surface loop and / or an insertion of loops 165, 180, and / or 194 / 195.

[0115] II. Insert The invention includes an amino acid sequence (eg, a peptide or polypeptide) (eg, SEQ ID NO:1 or a sequence derived therefrom (eg, greater than 70% sequence identity)) inserted into a position having a modified dehalogenase.

[0116] In some embodiments, the insertion is an extended loop sequence, e.g., to enhance / modify the interaction between the modified dehalogenase and a substrate (e.g., a functional portion of a substrate). In some embodiments, the extended loop sequence is the sequence X1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 X 15 X 16 X 17 X 18 X 19 X 20 X 21 X 22 X 23 X 24 X 25 and X1~X 25Each of X1 through X2 can be independently selected from any amino acid (e.g., proteinogenic amino acids, natural amino acids, unnatural amino acids, amino acid analogs, etc.) or absent. 25 At least one of X1 to X2 is not absent. 25 is 1 amino acid long, 2 amino acids long, 3 amino acids long, 4 amino acids long, 5 amino acids long, 6 amino acids long, 7 amino acids long, 8 amino acids long, 9 amino acids long, 10 amino acids long, 15 amino acids long, 20 amino acids long, 25 amino acids long, or a range therebetween.

[0117] In some embodiments, the insertion is a peptide or polypeptide having a desired functionality. In such embodiments, the peptide or polypeptide can be of any length (e.g., 10 amino acids, 20 amino acids, 30 amino acids, 40 amino acids, 50 amino acids, 75 amino acids, 100 amino acids, 150 amino acids, 200 amino acids, 300 amino acids, 400 amino acids, 500 amino acids, 600 amino acids, 700 amino acids, 800 amino acids, 900 amino acids, 1000 amino acids, or more, or any range therebetween). In some embodiments, the insertion location is a loop, such that the substrate binding ability of the modified dehalogenase is maintained despite the presence of the insertion.

[0118] In some embodiments, the insert is a heterologous sequence, hi some embodiments, the heterologous sequence interacts (e.g., through contact and / or resonance / energy transfer) with a functional moiety of the substrate.

[0119] Heterologous sequences useful as inserts in the modified dehalogenase include, but are not limited to, an enzyme of interest, such as a luciferase, an RNasin or an RNase, and / or a channel protein, a receptor, a membrane protein, a cytoplasmic protein, a nuclear protein, a structural protein, a phosphoprotein, a kinase, a signal protein, a metabolic protein, a mitochondrial protein, a receptor-associated protein, a fluorescent protein, an enzyme substrate, a transcription factor, a transporter protein, and / or a targeting sequence, such as a myristoylation sequence, a mitochondrial localization sequence, or a nuclear localization sequence that directs the modified dehalogenase to a particular location. The heterologous sequence fused into the loop of the modified dehalogenase can be a fragment of an entire protein, for example, a functional or structural domain of a protein, such as a domain of a kinase, a transcription factor, etc. The heterologous sequence can be a fragment of a protein that interacts with a second fragment of the protein to form an active complex by protein complementation.

[0120] In some embodiments, the heterologous sequence inserted into the loop of the modified dehalogenase interacts with another element to form a complex, for example, FRB or FKBP can be inserted into the 165 loop or the 180 loop and can interact with the other when in close proximity. Exemplary heterologous sequences include, but are not limited to, sequences in FRB and FKBP, the regulatory subunit of protein kinase (PKa-R) and the catalytic subunit of protein kinase (PKa-C), src homology regions (SH2) and phosphorylatable sequences, e.g., tyrosine-containing sequences, isoforms of 14-3-3, e.g., 14-3-3t (see Mills et al., 3100), and phosphorylatable sequences, proteins with WW regions (sequences of proteins that bind proline-rich molecules (see Ilsley et al., 3102, and Einbond et al., 1996), and phosphorylatable heterologous sequences, e.g., serine and / or threonine-containing sequences, and sequences in dihydrogenfolate reductase (DHFR) and gyrase B (GyrB).

[0121] In some embodiments, the heterologous sequence for insertion into the loop of the modified dehalogenase is selected from the group consisting of an antibody, an antibody fragment, Protein A, an Ig binding domain of Protein A, Protein G, an Ig binding domain of Protein G, Protein A / G, an Ig binding domain of Protein A / G, Protein L, an Ig binding domain of Protein L, Protein M, an Ig binding domain of Protein M, an oligonucleotide probe, a peptide nucleic acid, a DARPin, an anticalin, a nanobody, an aptamer, an affimer, a purified protein, and an analyte binding domain(s) of a protein.

[0122] As described throughout, any of a variety of peptides, polypeptides, antibodies, enzymes, reporters, and proteins of interest may be inserted into the 165 loop and 180 loop of the modified dehalogenases herein. For example, the present invention provides a method for the preparation of modified dehalogenases comprising: (1) a modified dehalogenase; (2) an amino acid sequence of a protein or peptide of interest inserted within the 165 loop or the 180 loop, such as a marker protein, e.g., a selectable marker protein, an enzyme of interest, e.g., a luciferase, an RNasin, an RNase, and / or a sequence of a GFP, a nucleic acid binding protein, an extracellular matrix protein, a secreted protein, an antibody or portion thereof, e.g., an Fc, a bioluminescent protein, a receptor ligand, a regulatory protein, a serum protein, an immunogenic protein, a fluorescent protein, a protein with a reactive cysteine, a receptor protein, e.g., an NMDA receptor, a channel protein, e.g., a HERG channel protein, or a nucleic acid binding protein, including ... a selectable marker protein, an enzyme of interest, e.g., a luciferase, an RNasin, an RNase, and / or a GFP, or a nucleic acid binding protein, including a marker protein, a selectable marker protein, an enzyme of interest, In one embodiment, the fusions include ion channel proteins such as ribosomal, potassium, or calcium sensitive channel proteins, membrane proteins, cytoplasmic proteins, nuclear proteins, structural proteins, phosphoproteins, kinases, signaling proteins, metabolic proteins, mitochondrial proteins, receptor-associated proteins, fluorescent proteins, enzyme substrates, e.g., protease substrates, transcription factors, protein destabilization sequences, or transporter proteins, e.g., EAAT1-4 glutamate transporters, as well as targeting signals that direct the fusion to a particular location, e.g., a mitochondrial localization sequence, a nuclear localization signal, or a plastid targeting signal such as a myristoylation sequence.

[0123] In some embodiments, the heterologous sequence is associated with a membrane or portion thereof, e.g., a targeting protein such as for endoplasmic reticulum targeting, a cell membrane-associated protein, e.g., an integrin protein or domain thereof, e.g., the cytoplasmic, transmembrane and / or extracellular stalk domain of an integrin protein, and / or a protein linking the mutant hydrolase to the cell surface, e.g., a glycosylphosphoinositol signal sequence.

[0124] Heterologous sequences for insertion into the modified dehalogenase loop can include sequences with enzymatic activity. For example, a functional protein sequence can encode a kinase catalytic domain (Hanks and Hunter, 1995), producing a fusion protein capable of enzymatically adding a phosphate moiety to a specific amino acid, or can encode a Src homology 2 (SH2) domain (Sadowski et al., 1986; Mayer and Baltimore, 1993), producing a fusion protein that specifically binds phosphorylated tyrosine.

[0125] In some embodiments, the insert comprises an affinity domain, which comprises a peptide sequence capable of interacting with a binding partner, such as one immobilized on a solid support, useful for identification or purification. DNA sequences encoding multiple consecutive single amino acids, such as histidine, when fused to an expressed protein, can be used for one-step purification of recombinant proteins by binding with high affinity to a resin column, such as nickel sepharose. Exemplary affinity domains include HisV5 (HHHHH) (SEQ ID NO: 81), HisX6 (HHHHHH) (SEQ ID NO: 82), C-myc (EQKLISEEDL) (SEQ ID NO: 83), Flag (DYKDDDDK) (SEQ ID NO: 84), SteptTag (WSHPQFEK) (SEQ ID NO: 85), hemagglutinin, e.g., HA tag (YPYDVPDYA) (SEQ ID NO: 86), GST, thioredoxin, cellulose binding domain, RYIRS (SEQ ID NO: 87), Phe-His-His-Thr (SEQ ID NO: 88), chitin binding domain, S-peptide, T7 peptide, SH2 domain, C-end RNA tag, WEAAAREACCRECCARA (SEQ ID NO: 10), a metal binding domain, e.g., a zinc binding domain, or a calcium binding domain, such as from a calcium binding protein, e.g., calmodulin, troponin C, calcineurin B, myosin light chain, recoverin, S-modulin, visinin, VILIP, neurocalcin, hippocalcin, fryquenin, caltractin, calpain large subunit, S100 protein, parvalbumin, calbindin, D9K , Calbindin D 28K and calretinin, intein, biotin, streptavidin, MyoD, Id, leucine zipper sequences, maltose binding protein, and SPYTAG peptides or SPYCATCHER proteins (e.g., SYYHHHHHHDYDIPTTENLYFQGAMVTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEDGRELAGATMELRDSSGKTISTWISDGHVKDFYLYPGKYTFVETAAPDGYEVATPIEFTVNEDGQVTVDGEATEGDAHTGSSGS (SEQ ID NO: 89), SYYHHHHHHDYDIPTTENLYFQGAMVTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEDGRELAGATMELRDCSGKTISTWISDGHVKDFYLY PGKYTFVETAAPDGYEVATPIEFTVNEDGQVTVDGEATEGDAHTGSSGS (SEQ ID NO: 90), GSSHHHHHSSGLVPRGSRGVPHIVMVDAYKRYKGSGESGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSS (SEQ ID NO: 91).

[0126] In some embodiments, the insert is a fluorescent or luminescent protein. In some embodiments, the insert is a bioluminescent protein. In certain embodiments, the insert is luciferase. Suitable luciferase enzymes include Photinus pyralis or North American firefly luciferase, Luciola cruciata or Japanese firefly or Genji firefly luciferase, Luciola italic or Italian firefly luciferase, Luciola lateralis or Japanese firefly or Heike firefly luciferase, N. nambi luciferase, Luciola mingrelica or Eastern European firefly luciferase, Photuris pennsylvanica or Pennsylvania firefly luciferase, Pyrophorus plagiophthalamus or click beetle luciferase, Phrixothrix hirtus or railroad worm luciferase, Renilla reniformis or wild type Renilla luciferase, Renilla reniformis Rluc8 mutant Renilla luciferase, Renilla reniformis Green Renilla luciferase, Gaussia princeps wild type Gaussia luciferase, Gaussia princeps Gaussia-Dura luciferase, Cypridina noctiluca or Cypridina luciferase, Cypridina hilgendorfii or Cypridina or Vargula luciferase, Metridia longa or Metridia luciferase, TurboLuc (Auld et al. Biochemistry 2018, 57, 31, 4700-4706, incorporated by reference in its entirety), nanolanterns (Suzuki et al.Nature Communications volume 7, Article number: 13718 (2016), incorporated by reference in its entirety), and Oplophorus luciferase (e.g., Oplophorus gracilirostris (OgLuc luciferase), Oplophorus grimaldii, Oplophorus spinicauda, ​​Oplophorus foliaceus, Oplophorus noraezeelandiae, Oplophorus typus, Oplophorus noraezelandiae, or Oplophorus spinous). In some embodiments, the luciferase is selected from those found in Omphalotus olearius, fireflies (e.g., Photinini), Renilla reniformis, Aequoria, mutants thereof, portions thereof, variants thereof, and any other luciferase enzyme suitable for the systems and methods described herein.

[0127] In some embodiments, the bioluminescent reporter is a modified, enhanced luciferase enzyme from Oplophorus (e.g., the NANOLUC enzyme from Promega Corporation, SEQ ID NO:28, or a sequence having at least 70% identity thereto (e.g., greater than 70%, greater than 80%, greater than 90%, greater than 95%). Exemplary bioluminescent reporters are described, for example, in U.S. Patent Application No. 2010 / 0281552 and U.S. Patent Application No. 2012 / 0174242, both of which are incorporated by reference in their entireties.

[0128] In some embodiments, the modified dehalogenase comprises a peptide or polypeptide component of commercially available NanoLuc®-based technologies (e.g., NanoLuc® luciferase, NanoBiT, NanoTrip, NanoBRET, etc.), such as an insertion of loop 165, loop 180, or loop 194 / 195 of one of SEQ ID NOs: 29-31. PCT Application No. PCT / US2010 / 033449, U.S. Patent No. 8,557,970, PCT Application No. PCT / 2011 / 059018, and U.S. Patent No. 8,669,103 (each of which is incorporated herein by reference in their entirety for all purposes) describe compositions and methods comprising bioluminescent polypeptides used as heterologous sequences in the fusions herein. In some embodiments, the inserts are circularly permuted versions of NanoLuc®-based components (e.g., NanoLuc® luciferase, NanoBiT, NanoTrip, NanoBRET, etc.). Such polypeptides are used in embodiments herein and may be used in conjunction with the compositions and methods described herein. PCT Application No. PCT / US14 / 26354 and U.S. Patent No. 9,797,889, each of which is incorporated herein by reference in its entirety for all purposes, describe compositions and methods for the assembly of bioluminescent complexes, and such complexes, as well as their peptide and polypeptide components, may be used as heterologous sequences in embodiments herein and in conjunction with the compositions and methods described herein. In some embodiments, NanoBiT and other related technologies utilize peptide and polypeptide components in assembly into complexes that are significantly enhanced (e.g., 2-fold, 5-fold, 10 ... 2 Double, 10 3 Double, 10 4In some embodiments, the NanoBiT peptides and polypeptides are inserted into the modified dehalogenases herein. U.S. Patent Publication No. 2020 / 0270586 and International Application No. PCT / US19 / 36844, which are incorporated by reference in their entireties for all purposes, describe multipartite luciferase complexes (e.g., NanoTrip) that may be used as heterologous sequences in embodiments herein and in conjunction with the compositions and methods described herein.

[0129] In some embodiments, the insertions are circularly permuted versions of the protein or polypeptide insertions described herein. For example, the insertions (e.g., within loops 165, 180, or 194 / 195) are circularly permuted NanoLuc-, NanoBiT-, or NanoTrip-based peptides or polypeptides. SEQ ID NOs: 33-80 are exemplary constructs that include various cpNanoLucs inserted at various positions within loops 165, 180, or 194 / 195. Other combinations of cpNanoLuc and insertion sites herein are within the scope of the present invention. In some embodiments, NanoLuc-based polypeptides having cp sites between any of the following positions are inserted into the loop 165 / 180 insertion site: 6 / 7, 12 / 13, 24 / 25, 27 / 28, 49 / 50, 52 / 53, 55 / 56, 64 / 65, 667 / 68, 70 / 71, 79 / 80, 82 / 83, 84 / 85, 86 / 87, 103 / 104, 106 / 107, 120 / 121, 124 / 125, 130 / 131, 145 / 146, 148 / 149, or any other site within the NanoLuc or NanoLuc-based polypeptide. SEQ ID NOs: 91-120 are exemplary constructs containing various cpLgBiTs inserted into various positions within loop 165, 180, or 194 / 195. Other combinations of cpLgBiT and the insertion sites herein are within the scope of the present invention.

[0130] In some embodiments, provided herein are modified dehalogenases that include an insertion sequence(s) within loop 165 and / or 180. In some embodiments, the modified dehalogenases include an insertion sequence within both loop 165, loop 180, and loop 194 / 195. In some embodiments, the modified dehalogenases include an insertion sequence within one or both of loop 165 and loop 180, and further include a C-terminal and / or N-terminal fusion sequence. Any of the above insertions may also be used as terminal fusions to the extended loop modified dehalogenases described herein.

[0131] III. Substrate The modified dehalogenases herein utilize a haloalkane substrate. In some embodiments, the substrate is of formula (I): R-linker-AX, where R is a solid surface, one or more functional groups, or is absent, and the linker is a multi-atom straight or branched chain containing C, N, S, or O, or a group containing one or more rings, e.g., saturated or unsaturated rings, such as one or more aryl rings, heteroaryl rings, or any combination thereof, and where AX is a substrate for a dehalogenase, hydrolase, HALOTAG, or modified dehalogenase system herein (e.g., A is (CH2) 4~ 20, and X is a halide (e.g., Cl or Br). Suitable substrates are described, for example, in U.S. Pat. Nos. 11,072,812, 11,028,424, 10,618,907, and 10,101,332, which are incorporated by reference in their entireties. In certain embodiments, X in formula (I) is not a halide, but is a methylsulfonamide or trifluoromethylsulfonamide, and such embodiments result in exchangeable ligands that reversibly bind to modified dehalogenases (e.g., HALOTAG). Such ligands are described, for example, in Kompa et al. J. Am. Chem. Soc. 2023, 145, 5, 3075-3083, which is incorporated by reference in its entirety.

[0132] In some embodiments, R is one or more functional groups, such as a fluorophore, biotin, luminophore, or a fluorescent or luminescent molecule. Exemplary functional groups for use in the present invention include, but are not limited to, amino acids, proteins, such as enzymes, antibodies or other immunogenic proteins, radionuclides, nucleic acid molecules, drugs, lipids, biotin, avidin, streptavidin, magnetic beads, solid supports, electron opaque molecules, chromophores, MRI contrast agents, dyes, such as xanthene dyes, calcium sensitive dyes, such as 1-[2-amino-5-(2,7-dichloro-6-hydroxy-3-oxy-9-xanthenyl)-phenoxy]-2- (2'-amino-5'-methylphenoxy)ethane-N,N,N',N'-tetraacetic acid (Fluo-3), sodium sensitive dyes such as 1,3-benzenedicarboxylic acid, 4,4'-[1,4,10,13-tetraoxa-7,16-diazacyclooctadecane-7,16-diylbis(5-methoxy-6,2-benzofurandiyl)bis (PBFI), NO sensitive dyes such as 4-amino-5-methylamino-2',7'-difluorescein, or other fluorophores. In one embodiment, the functional group is an immunogenic molecule, i.e., one that is bound by an antibody specific for that molecule.

[0133] In some embodiments, a substrate of the invention is permeable to the plasma membrane of a cell (i.e., capable of passing from the outside of a cell (e.g., eukaryotic, prokaryotic) to the inside of the cell without chemical, enzymatic, or mechanical disruption of the cell membrane).

[0134] In some embodiments, the substrates herein include a cleavable linker, such as those described in US Pat. No. 10,618,907, which is incorporated by reference in its entirety.

[0135] In some embodiments, the substrate comprises a fluorescent functional group (R).Suitable fluorescent functional groups include, but are not limited to, stilbazolium derivatives (Marquesa et al. Mechanism-Based Strategy for Optimizing HaloTag Protein Labeling. ChemRxiv. Cambridge: Cambridge Open Engage; 2021; incorporated by reference in its entirety), xanthene derivatives (e.g., fluorescein, rhodamine, Oregon Green, eosin, Texas Red, etc.), cyanine derivatives (e.g., cyanine, indocarbocyanine, oxacarbocyanine, thiacarbocyanine, merocyanine, etc.), naphthalene derivatives (e.g., dansyl and prodan derivatives), oxadiazole derivatives (e.g., pyridyloxazole, nitrobenzoxadiazole, benzoxadiazole, etc.), pyrene derivatives (e.g., cascade blue), oxazine derivatives (e.g., Nile red, Nile blue, cresyl violet, oxazine 170, etc.), acridine derivatives (e.g., proflavine, acridine orange, acridine yellow, etc.), arylmethine derivatives (e.g., auramine, crystal violet, malachite green, etc.), tetrapyrrole derivatives (e.g., porphine, phthalocyanine, bilirubin, etc.), CF dyes (Biotium), BODIPY (Invitrogen), ALEXA FLOUR (Invitrogen), DYLIGHT FLUOR (Thermo Scientific, Pierce), ATTO and TRACY (Sigma Aldrich), FluoProbes (Interchim), DY and MEGASTOKES (Dyomics), SULFO CY dyes (CYANDYE, LLC), SETAU and SQUARE dyes (SETA BioMedicals), QUASAR and CAL FLUOR dyes (Biosearch Technologies), SURELIGHT dyes (APC, RPE, PerCP, phycobilisomes) (Columbia Biosciences), APC, APCXL, RPE, BPE (Phyco-Biotech), autofluorescent proteins (e.g., YFP, RFP, mCherry, mKate), quantum dot nanocrystals, and the like.

[0136] In some embodiments, the substrate comprises a fluorogenic functional group (R). The fluorescent functional group is a functional group that generates and enhances a fluorescent signal upon binding of the substrate to a target (e.g., binding of a haloalkane to a modified dehalogenase). By generating significantly increased fluorescence (e.g., 10-fold, 31-fold, 50-fold, 100-fold, 310-fold, 500-fold, 100-fold, or more) upon target engagement, problems with background signal are mitigated. Exemplary fluorogenic dyes for use in embodiments herein include the JANELIA FLUOR family of fluorophores, such as: [ka] [ka] (See, e.g., U.S. Pat. Nos. 9,933,417, 10,018,624, 10,161,932, and 10,495,632, each of which is incorporated by reference in its entirety.) In some embodiments, exemplary conjugates of JANELIA FLUOR 549 and JANELIA FLUOR 646 with haloalkane substrates of modified dehalogenases (e.g., HALOTAG) are commercially available (Promega Corp.). The use and design of fluorogenic functional groups, dyes, probes, and substrates are described, for example, in Grimm et al. Nat Methods. 3117 Oct;14(10):987-994., Wang et al. Nat Chem. 3120 Feb;12(2):165-172, each of which is incorporated by reference in its entirety.

[0137] In some embodiments, a "dual warhead" substrate is provided that includes a haloalkane moiety (e.g., a substrate for a modified dehalogenase (e.g., HALOTAG)) and a dimerization moiety that is a ligand (or capture element) for a second binding protein (capture element). For example, certain embodiments herein utilize haloalkanes linked to SNAP tag ligands (Cermakova & Hodges. Molecules 2018, 23(8), 1958 (incorporated by reference in its entirety)), haloalkanes linked to cTMP (Cermakova & Hodges. Molecules 2018, 23(8), 1958 (incorporated by reference in its entirety)), haloalkanes linked to rapamycin-like moieties capable of binding to FKBP or FRB (Chen et al. ACS Chem. Biol. 2021, 16, 12, 2808-2815; incorporated by reference in its entirety), or other haloalkane "dual warhead" ligands capable of binding to a modified dehalogenase (e.g., HALOTAG) and a second capture agent. In such embodiments, a system is provided that includes a modified dehalogenase as described herein, a dual warhead substrate, and a capture agent (e.g., FKBP, FRB, SNAP tag, eDHFR, etc.) capable of binding to the dimerization moiety. In some embodiments, the insert in the modified dehalogenase and the capture agent can interact (e.g., structurally or by energy transfer). In some embodiments, the dual warhead triggers the proximity of the inserted heterologous sequence and the capture agent by appending another protein-binding small molecule moiety to the haloalkane. Such embodiments provide for forced proximity of the insert and the capture agent. Any suitable linker can be used in the assembly of the dual warhead substrate. Linkers can be esters (-C(O)O-), amides (-C(O)NH-), carbamates (-NHC(O)O-), ureas (-NHC(O)NH-), phenylenes (e.g., 1,4-phenylene), linear or branched alkylenes, and / or oligo and polyethylene glycols (-(CH2CHO) xVarious combinations of such groups may be included to provide linkers having -) bonds. In some embodiments, the linker may include 2 or more atoms (e.g., 2-200 atoms, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 atoms, or any range therebetween (e.g., 2-20, 5-10, 15-35, 25-100, etc.). In some embodiments, the linker includes a combination of oligoethylene glycol and carbamate bonds. In some embodiments, the linker has the formula -O(CHCHO) z1 -C(O)NH-(CH2CH2O) z2 -C(O)NH-(CH2) z3 -(OCH2CH2) z4 O-, where z1, z2, z3, and z4 are each independently selected from the forms 0, 1, 2, 3, 4, 5, and 6. For example, in some embodiments, the linker has a formula selected from the following: [ka]

[0138] In some embodiments, the dual warheads used in the embodiments herein are haloalkanes linked to a ligand that can engage an E3 ubiquitin ligase (e.g., thalidomide, cereblon E3 ubiquitin ligase, von Hippel-Lindau (VHL) E3 ligase, or any other E3 ubiquitin ligase), otherwise known as protein degradation targeting chimeras (PROTACs). The haloalkane PROTACs can bind to a modified dehalogenase or modified dehalogenase complex and an E3 ubiquitin ligase, and recruitment of the E3 ligase results in the proteasome-mediated ubiquitination and subsequent degradation of the modified dehalogenase (complex) and any protein components fused thereto (e.g., target proteins). In some embodiments, the modified dehalogenase systems herein are used in applications such as assays / systems for measuring the kinetics of target protein ubiquitination, or in an end-point format, measuring compound dose-response curves. For example, in some embodiments, a sample is provided with a target protein expressed / provided as an insert within a modified dehalogenase, and the sample is contacted with a haloalkane PROTAC and a ligand that can engage an E3 ubiquitin ligase (e.g., thalidomide, cereblon E3 ubiquitin ligase, von Hippel-Lindau (VHL) E3 ligase, or any other E3 ubiquitin ligase), such that when the haloalkane is bound by the modified dehalogenase, the ligand is brought into proximity of the target protein, resulting in ubiquitination and targeting the fusion target to the proteasome for degradation.In some embodiments, the modified dehalogenase systems herein may be modified in a manner similar to a variety of other targeting chimera (TAC) systems, such as the phosphorylation-targeting chimera (PhosTAC; Chen et al. ACS Chem. Biol. 3121, 16, 12, 2808-2815; incorporated by reference in its entirety) system, the deubiquitinase-targeting chimera (DUBTAC; Henning et al. Deubiquitinase-Targeting Chimeras for Targeted Protein Stabilization. bioRxiv; 2021. DOI: 10.1101 / 2021.04.30.441959; incorporated by reference in its entirety) system, the lysosome-targeting chimera (LyTAC; Banik et al. Nature 584, 291-297 (2020); incorporated by reference in its entirety) system, the autophagy-targeting chimera (AUTAC; Takahashi et al. Mol. Cell. 2019 Dec 5;76(5):797-810.e10; incorporated by reference in its entirety) system, the Autophagy Linking Compound (ATTEC; Fu et al. Cell Research volume 31, pages 965-979(2021); incorporated by reference in its entirety) system, and oligo-based TACs. Dual warheads including haloalkanes and ligands for any of the above TAC systems can be used in the embodiments herein. For example, PhosTAC is similar to well-described PROTACs in its ability to induce ternary complexes, where PhosTAC focuses on recruiting Ser / Thr phosphatases to phosphosubstrates to mediate their dephosphorylation. PhosTAC extends the use of PROTAC technology beyond ubiquitination-mediated protein degradation to other post-translational modifications of proteins.For example, in some embodiments, a target protein is expressed / provided as an insert with a loop of a modified dehalogenase, and a sample is contacted with a ligand capable of engaging a haloalkane phosphorylation targeting chimera (PhosTAC) and a phosphatase enzyme, such that upon binding of the haloalkane by the modified dehalogenase, the ligand is brought into proximity of the target protein, resulting in phosphorylation of the target protein.

[0139] In some embodiments, the modified dehalogenase systems herein are used in combination with modified dehalogenases comprising a target protein into which a dual function ligand comprising a haloalkane and a ligand for a recruitable enzyme is inserted to direct the enzymatic activity of the recruitable enzyme to the target protein. Systems and methods comprising any combination of the above TAC systems / assays are within the scope of the present specification.

[0140] In some embodiments, the modified dehalogenase comprises a reporter protein inserted within loop 165, loop 180, or loop 194 / 195 that is capable of emitting energy (e.g., light) at a first wavelength, and the functional moiety (R) on the haloalkane substrate comprises a moiety that is capable of accepting energy at a first wavelength. In some embodiments, the acceptor moiety is a fluorophore. In other embodiments, the acceptor moiety is a photocatalyst that is activated by exposure to the emitted energy. In some embodiments, due to the location of the insertion site within the modified dehalogenase, the proximity / geometry between the inserted reporter and acceptor allows for optimized energy transfer.

[0141] In some embodiments, the functional moiety (R) on the haloalkane substrate comprises a fluorophore that can absorb light emitted from a luminophore (when interacting with a bioluminescent protein or complex (e.g., inserted into the loop of a modified dehalogenase)) and subsequently emit light. Suitable fluorophores include fluoresceins and fluorescein dyes (e.g., fluorescein isothiocyanate or FITC, naphthofluorescein, 4',5'-dichloro-2',7'-dimethoxyfluorescein, 6-carboxyfluorescein (e.g., FAM)), rhodamine dyes (e.g., carboxytetramethylrhodamine or TAMRA, carboxylrhodamine 6G, carboxy-X-rhodamine (ROX), Lissamine rhodamine B, rhodamine 6G, rhodamine green, rhodamine red, tetramethylrhodamine or TMR), coumarin and cumarin dyes (e.g., methoxycoumarin, dialkylaminocoumarin, hydroxycoumarin, and aminomethylcoumarin or AMCA), Oregon Green dyes (e.g., Oregon Green 488, Oregon Green 500, Oregon Green 514), Texas Red, Texas Red-X, SPECTRUM RED™, SPECTRUM GREEN™, cyanine dyes (e.g., CY-3™, CY-5™, CY-3.5™, CY-5.5™), Alexa Fluor dyes (e.g., Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 633, Alexa Fluor 660, and Alexa Fluor 680), BODIPY dyes (e.g., BODIPY FL, BODIPY R6G, BODIPY TMR, BODIPY TR, BODIPY 530 / 550, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY 581 / 591, BODIPY 630 / 650, BODIPY 650 / 665), IRDyes (e.g., IRD40, IRD 700, IRD 800), and the like.

[0142] In some embodiments, the functional moiety (R) on the haloalkane substrate comprises a photocatalyst that can absorb light emitted from a luminophore (when interacting with a bioluminescent protein or complex (e.g., inserted into the loop of a modified dehalogenase)) and subsequently activate an adjacent activatable label. Any compound or moiety that can receive light energy emitted from a bioluminescent protein or complex activated luminophore and function as a photocatalyst (e.g., transfer the energy to a target molecule (e.g., an activatable molecule)) can be used in the embodiments herein. In some embodiments, the excited photocatalyst transfers energy via Forster resonance energy transfer, Dexter energy transfer, single electron transfer, singlet oxygen, or any other suitable mechanism of energy or electron transfer. In some embodiments, the photocatalyst is an iridium- or ruthenium-based photocatalyst (Bevernaegie et al.'A Roadmap Towards Visible Light Mediated Electron Transfer Chemistry with Iridium(III) Complexes.'ChemPhotoChem 2021,5,217, incorporated by reference in its entirety). In some embodiments, the photocatalyst is an organic photoredox catalyst. In some embodiments, the organic photoredox catalyst is selected from quinones, pyryliums, acridiniums, xanthenes, and thiazines. In some embodiments, systems and methods are provided herein that include a modified dehalogenase that includes a bioluminescent protein or component of a bioluminescent complex inserted into a loop therein, a substrate for the modified dehalogenase that includes a photocatalyst as a functional group, and an activatable moiety capable of receiving energy transferred from the photocatalyst.

[0143] In addition to the haloalkane substrates described above and throughout this application (e.g., having R, linker, A, and X groups as described herein), exemplary substrates within the scope of the present specification include: [ka] In the formula, X=O, SiR 2 , or CR 2 , Wherein R=H, alkyl, or fluoroalkyl; In the formula, R 2 is H, alkyl, and can be cyclized with itself (R 2 ~R 2 ), or R 1 Cyclized onto R 1 is H, alkyl, R 2 or halogen; R 3 is H, F, or Cl, Ar is a halogen, OR, NR, as described in Wang et al. Nat. Chem. 12, 165-172 (2020), Nat. Chem. 12, 165-172 (2020), and Lardon et al. J. Am. Chem. Soc. 2021, 143, 14592-14600, which are incorporated by reference in their entireties. 2 , CO2R, CONR 2 , CN, alkyl, or haloalkyl optionally substituted aromatic ring (eg, phenyl).

[0144] IV. Nucleic acids, cells, etc. In some embodiments, provided herein are isolated nucleic acid molecules (polynucleotides) comprising a nucleic acid sequence encoding a modified dehalogenase described herein (e.g., having an internal insertion). In some embodiments, such polynucleotides comprise an open reading frame encoding a modified dehalogenase described herein. In some embodiments, such polynucleotides are within an expression vector or are integrated into the genomic material of a cell. In some embodiments, such polynucleotides further comprise a regulatory element, such as a promoter. Also provided are isolated nucleic acid molecules comprising a nucleic acid sequence encoding a fusion protein comprising a modified dehalogenase and one or more amino acid residues (e.g., peptide, polypeptide) inserted at a position within the 165 loop or 180 loop(s). In one embodiment, the modified dehalogenase comprises a sequence (e.g., at the N-terminus or C-terminus), e.g., for purification, e.g., a glutathione S-transferase (GST) or polyHis sequence, a sequence intended to alter the properties of the remainder of the fusion protein, e.g., a protein destabilizing sequence, or a sequence with distinguishing properties. In one embodiment, the isolated nucleic acid molecule comprises a nucleic acid sequence that is optimized for expression in at least one selected host. Optimized sequences include codon-optimized sequences, i.e., codons that are more frequently used in one organism relative to another organism, e.g., a distantly related organism, as well as modifications to add or modify Kozak sequences and / or introns, and / or modifications to remove undesirable sequences, e.g., potential transcription factor binding sites. In one embodiment, the polynucleotide comprises a nucleic acid sequence encoding a modified dehalogenase, which nucleic acid sequence is optimized for expression in a selected host cell. In one embodiment, an optimized polynucleotide no longer hybridizes to a corresponding non-optimized sequence, e.g., does not hybridize to a non-optimized sequence under medium or high stringency conditions.In another embodiment, the polynucleotides have less than 90%, e.g., less than 80%, nucleic acid sequence identity to a corresponding non-optimized sequence, and optionally encode a polypeptide having at least 80%, e.g., at least 85%, 90% or more amino acid sequence identity to a polypeptide encoded by the non-optimized sequence.

[0145] Constructs, e.g., vectors comprising the expression cassettes and isolated nucleic acid molecules, as well as host cells having one or more of the constructs, and kits comprising the isolated nucleic acid molecules, one or more of the constructs or vectors are also provided. Host cells include prokaryotic or eukaryotic cells, e.g., plant cells or vertebrate cells, e.g., mammalian cells, including, but not limited to, human, non-human primate, canine, feline, bovine, equine, ovine, sheep, or rodent (e.g., rabbit, rat, ferret, or mouse) cells. In some embodiments, the expression cassette comprises a promoter, e.g., a constitutive promoter or a regulatable promoter, operably linked to the nucleic acid molecule. In some embodiments, the expression cassette comprises an inducible promoter. In certain embodiments, the invention comprises a vector comprising a nucleic acid sequence encoding a fusion protein comprising a fragment of a dehalogenase. In some embodiments, an optimized nucleic acid sequence, e.g., a human codon-optimized sequence, encoding at least a fragment of a hydrolase, preferably a fusion protein comprising a fragment of a hydrolase, is used in the nucleic acid molecules of the invention. Optimization of nucleic acid sequences is known in the art, see, for example, WO02 / 16944, which is incorporated by reference in its entirety.

[0146] Also provided are cells comprising modified dehalogenases (e.g., having loop 165, loop 180, and / or loop 194 / 195 insertions), polynucleotides, expression vectors, etc. In some embodiments, the components described herein are expressed in cells. In some embodiments, the components described herein are introduced into cells via, for example, transfection, electroporation, infection, cell fusion, or any other means.

[0147] V. SYSTEMS AND METHODS In some embodiments, provided herein are systems and methods that include modified dehalogenases that contain or utilize an internal insertion within the 165 loop or the 180 loop, or a sequence corresponding thereto. In some embodiments, the systems and methods further include additional components, such as substrates, binding proteins (e.g., capable of binding to the insert), luminophores, complementary comparators (e.g., bioluminescent complexes having an insert of the modified dehalogenase), and other agents / reagents described herein. In some embodiments, the methods herein include contacting a modified dehalogenase described herein with a substrate and / or additional reagents (e.g., luminophores), detecting fluorescence / luminescence, isolating / purifying the components, etc.

[0148] Certain embodiments herein find use in energy transfer systems and applications. In some embodiments, modified dehalogenases herein containing internal insertions of components of bioluminescent proteins or bioluminescent complexes in the 165, 180, or 194 / 195 loops are useful for energy transfer to a suitable acceptor (e.g., an energy acceptor as a functional moiety (R) on a HALOTAG substrate). In some embodiments, the energy acceptor is a fluorophore or a photocatalyst. In some embodiments, the energy acceptor further transfers energy to a second acceptor. For example, in some embodiments, the first acceptor is a first fluorophore with an excitation spectrum that overlaps with the emission spectrum of the bioluminescent protein or bioluminescent complex, and the second acceptor is a second fluorophore with an excitation spectrum that overlaps with the emission spectrum of the first fluorophore. In some embodiments, upon contacting the bioluminescent protein or bioluminescent complex with a suitable luminophore, energy is transferred from the luminophore to the first fluorophore by BRET and from the first fluorophore to the second fluorophore by FRET. In other embodiments, the first acceptor is a photocatalyst with an excitation spectrum that overlaps with the emission spectrum of the bioluminescent protein or bioluminescent complex, and the second acceptor is an activatable target that is activated by the photocatalyst.

[0149] experiment Although loop regions in proteins are often more tolerant to sequence insertions, it was not immediately clear that the loop region in the commercially available modified dehalogenase HALOTAG would accommodate changes without disrupting protein folding or function. A known previous modification (Hiblot, J., et al. (2017) Angew Chem Int Ed Engl 56(46):14556-14560., incorporated by reference in its entirety) revealed that an insertion into loop 165 of the commercially available NANOLUC luciferase resulted in a functional HALOTAG, but not into loop 180. However, the activity of the resulting construct was highly dependent on the specific configuration in that example and was sensitive to the specific residues of the insertion, the linker, and whether NANOLUC was circularly permuted.

[0150] For example, the sequences of the peptides and polypeptides used in Examples 1-4 are provided in Table 1 (TABLE_1_Loop_HTs.txt, submitted herewith and incorporated by reference in its entirety) and Table 2. [Table 1-1] [Table 1-2]

[0151] Example 1 A circular permutation (CP) screen of HALOTAG was performed during the development of embodiments herein to systematically test the effect of circular permutation at all 297 individual positions. Data from the screen showed that HALOTAG could be circularly permuted and new N- and C-termini could be introduced into the 165 and 180 loops, retaining HALOTAG function and with minimal impact on protein stability. Screening data showed clear optimal positions for circular permutation within these loops, specifically after residues 165 and 180 within each loop, respectively. When the CP sites were moved, only two residues N- or C-terminal to these sites showed loss of activity or stability in HALOTAG, indicating the identification of the optimal positions.

[0152] Guided by the CP screening data, we tested the tolerance of sequence insertions at specific sites within the HALOTAG sequence by introducing 2, 5, or 10 residue stretches of glycine-serine separately into each loop. Both loop 165 and loop 180 tolerated these extensions and retained labeling activity with TMR and JF646 ligands (Figure 2, Figure 3). Longer Gly-Ser stretches of 5 to 10 residues within loop 165 resulted in decreased protein stability and reduced fluorescence activity with JF646 ligand, while retaining activity with the TMR ligand. This was the first evidence that insertions within these loops specifically modulate the activation of a fluorogenic dye.

[0153] Example 2 During development of the embodiments herein, experiments were performed to test the optimal positioning and composition of the loop insertion by sliding the insertion site of the 10x-Gly-Ser extension to loop 165 or loop 180 (Figure 4A). There was a clear preference for a specific extension site in terms of retaining expression and enzymatic activity in both E. coli lysates and purified protein. For loop 165, although the HaloTag version 7 ("v7") construct expressed best, purification of the constructs confirmed that the HaloTag version 6 ("v6") construct was optimal, inserting the loop immediately after D164 while simultaneously deleting residues Q165 and N166. For loop 180, the optimal site was the HaloTag version v2 ("v2") construct, where the loop is inserted immediately after V178 (Figure 4B). The results showed that insertion of sequences at different positions within loop 165 and loop 180 produced variants with different performance and characteristics, with some insertion points resulting in less expressed or active proteins.

[0154] Example 3 After establishing the optimal sites / configurations for extended loop insertion in loop 165 and loop 180, experiments were performed using libraries with randomized amino acids at the loop insertion sites to determine the tolerance of the extended loop-modifying dehalogenases to various amino acid loop compositions and their suitability for screening / selection to enable discovery of optimal sequences for specific applications (Table 1). Eight different library designs were tested. 1.165-7X / 180 = 7 randomized amino acids inserted into loop 165 2.165-11X / 180 = 11 randomized amino acids inserted into loop 165 3.165-15X / 180 = 15 randomized amino acids inserted into loop 165 4.165 / 180-7X = 7 randomized amino acids inserted into loop 180 5.165 / 180-11X = 11 randomized amino acids inserted into loop 180 6.165 / 180-15X = 15 randomized amino acids inserted into loop 180 7. 165-7X / 180-7X = 7 randomized amino acids inserted into both loop 165 and loop 180 8. 165-15X / 180-15X = 15 randomized amino acids inserted into both loop 165 and loop 180

[0155] The results are shown in Figure 5. The trends in library design indicate that in loop 165, longer loops as a group tend to show reduced total enzyme activity and JF646 activation, but there are clones similar to HALOTAG even with the insertion of 15-fold randomized loops indicating the influence of specific sequences at these sites (Figures 5A and 5B). For loop 180, more clones showed activity similar to HALOTAG, with roughly half showing reduced or no enzyme activity with the TMR and JF6464 ligands tested. Constructs with 7 or 15 randomized amino acids inserted in both loop 165 and loop 180 eliminated activity.

[0156] The library with randomized loops showed that diverse sequences can be inserted into loops 165 and 180 while retaining activity, demonstrating great flexibility in the possibility to engineer or screen for those that improve function for specific applications. Comparing the activity of individual loop variants between their activity with TMR and JF646 ligands, variants were found that showed tight binding to the TMR ligand and a range of activation levels of the fluorogenic JF646 ligand, from a complete loss of activation to a high level similar to unmodified HALOTAG (Figures 6A and 6B). This set of variants confirmed that sequence insertions at loop 165 or loop 180 can control the fluorescent activation of the dye without affecting the enzymatic function of HALOTAG, providing the ability to fine-tune the amount of fluorescent activation of the JF646 ligand using only changes to residues in the extended loop sequence. The experiments show that other activatable chemicals can also be tuned on the surface of HALOTAG, with changes to the proximal loop sequence modulating interactions that optimize activation.

[0157] More detailed characterization of several loop HALOTAG variants isolated through initial screening showed significant differences between the variants in their substrate specificity and kinetics. For example, a comparison of the various loop HALOTAG clone activities of JF646 versus Alexa488 ligands in Figures 7A and 7B shows that loop HALOTAG number 2 has low JF646 activity but high Alexa488 binding, whereas loop HALOTAG number 4 has high JF646 binding but low Alexa488 binding. This demonstrates that changes to the sequence in the loop alone are sufficient to alter the substrate specificity and binding kinetics of the loop HALOTAG variants.

[0158] Example 4 During development of embodiments herein, experiments were performed to engineer extended loop insertions in both loops 165 and 180 simultaneously, and it was observed in randomized libraries that dual insertions eliminated HALOTAG activity with a small sample size of randomized sequences tested. However, using specific sequences that retain full stability and function individually at each insertion site, combinations of sequences were tested to determine whether their stabilizing effects were synergistic (Figure 8A). It was observed that sequences providing highly active loop HALOTAG variants at either 165 or 180 could be combined together to provide active dual loop HALOTAG clones (Figure 8B).

[0159] Example 5 Experiments carried out during the development of embodiments herein indicate several possible mechanisms behind the activation effect observed for loop HALOTAG variants. In the direct interaction model, the extended loop sequence makes direct contact with the surface-exposed dye moiety of the ligand, and their interaction modulates the fluorescence activation. In the indirect interaction model, the extended loop insertion affects other protein:dye interactions or ligand binding, for example, changing the position of the adjacent Helix 8, which has close contact with the dye in the crystal structure, modulating its activation level during binding, or affecting contact with the chloroalkane moiety. In some embodiments, a combined direct / indirect model produces the effect.

[0160] Example 6 After establishing the tolerance of loop 165 and loop 180 to small 7-15 amino acid insertions, experiments were carried out to explore the feasibility of significantly larger insertions. To this end, different bioluminescent reporters were inserted into loop 165-V6 and loop 180-V2, including: NANOLUC (cpNLuc) circularly permuted at 1.67 / 68 positions 2. Thermostable NANOLUC (i.e., NanoLuc incorporating all LgBiT and HiBiT mutations) circularly permuted at positions 67 / 68 (cptsNLuc) 3. Thermostable NanoLuc (i.e., NanoLuc incorporating all LgBiT and HiBiT mutations) + the mutation F164C circularly permuted at positions 67 / 68 (cptsNLuc(F164C)

[0161] Comparing the resulting HALOTAG-NANOLUC chimera to the terminal HALOTAG-NANOLUC fusion (Figure 9), it was found that the chimera exhibited a slower binding kinetics to the HALOTAG TMR ligand (Figure 9B) and was dimmer in terms of NANOLUC emission (Figure 9C), but at the same time resulted in a greater BRET efficiency (Figure 9D), possibly through closer proximity and / or a conformation more favorable for energy transfer to the bound TMR ligand. Although both loops were tolerant to larger insertions, the resulting chimeras had different activity profiles. Consistent with previous results, the insertion of cpNLuc into loop 180-V2 had significantly less impact on the binding kinetics to the HALOTAG ligand compared to the same insertion into loop 165-V6 (Figure 9B). The insertion into loop 180-V2 also resulted in a greater increase in BRET efficiency indicating that the chimera was able to adopt a more favorable conformation for energy transfer to the bound TMR ligand (Figure 9D). The differential HALOTAG ligand binding rates and BRET efficiencies for insertion into the two loops can be further exploited towards orthogonality.

[0162] Comparing chimeras containing cpNLuc, cptsNLuc, and cptsNLuc(F164C) insertions into loop 180-V2, it was found that increased thermal stability of the inserted polypeptide (i.e., cptsNLuc) correlated with significantly slower binding rates to the HaloTag ligand (Figure 9B) and, to a lesser extent, lower BRET efficiency (Figure 9D), indicating that engineering greater flexibility / lower stability into the insertion may promote the adoption of a conformation that is favorable for both HALOTAG activity and energy transfer.

[0163] Example 7 During development of the embodiments herein, experiments were performed to further explore the ability of the chimeras to include inserting cpNLuc into loop 180 to provide increased intramolecular BRET efficiency not only with bound TMR ligands, but also with other bound fluorophores that exhibit extensive overlap between their excitation spectra and the bioluminescent reporter emission. Thus far, chimeras including insertion of cpNLuc into loop 180-V2 have provided increased BRET efficiency not only with bound TMR, but also with other fluorophores, including fluorogenic fluorophores (i.e., JF635 and JF646) and far-red fluorophores (i.e., Alexa 660), that exhibit minimal overlap between their excitation spectra and the bioluminescent reporter emission ( FIG. 10 ).

[0164] During development of embodiments herein, experiments were performed to further compare the purified NanoLuc-HaloTag fusion with a chimera containing cpNLuc inserted into loop 180 to provide intramolecular BRET efficiency to bound fluorophores that exhibited extensive overlap between their excitation spectra and the emission of the bioluminescent reporter. The emission of the bioluminescent energy donor and acceptor showed that the donor emission intensity (i.e., emission at 460 nm) of the chimera was significantly lower compared to the NLuc-HaloTag emission intensity for all six fluorophores, including the fluorogenic fluorophores (i.e., JF635 and JF646), but the far-red fluorophore (i.e., Alexa 660) was significantly higher, demonstrating the benefit provided by the chimera, likely due to close proximity and / or conformation more favorable for energy transfer to the bound fluorophore (Figure 11).

[0165] Example 8 During development of the embodiments herein, experiments were performed to further explore the insertion of a bioluminescent complementing reporter into loop 180, including: 1. Polypeptide components of the NANOLUC-based complementation system (LgBiT) 2. A polypeptide component of the NANOLUC-based complementation system (i.e., cpLgBiT) that was circularly permuted at positions 67 / 68. 3. The polypeptide component of the NANOLUC-based complementation system incorporating four LgTrip mutations (E4D, Q42M, M106K, T144D) circularly permuted at positions 67 / 68 (LgBiT+4) (i.e., cpLgBiT+4).

[0166] Although insertion of LgBiT or cpLgBiT into loop 180-V2 dramatically decreased the binding rate to HaloTag TMR ligand (Figure 12C), the resulting chimeras had very different binding rates when complemented with 10-fold excess of VS-HiBiT (the peptide component of the NANOLUC-based complementation system) (Figure 12). Overall, chimeras containing cpLgBiT or cpLgBiT+4 insertions showed significantly faster binding of TMR ligands upon complementation with 10-fold molar excess of VS-HiBiT (Figure 12D), suggesting that complementation promoted conformational adaptation favorable for HaloTag activity. Such a dependency of HaloTag binding on complementation could be further exploited as a HaloTag activity switch. Conversely, chimeras containing LgBiT insertions could be fully labeled after overnight incubation with 5-fold molar excess of TMR ligand (Figure 12B), but binding was not accelerated by pre-complementation with VS-HiBiT.

[0167] In addition, upon complementation with 10-fold excess VS-HiBiT, the three chimeras differed greatly in brightness and efficiency of intramolecular BRET towards bound TMR ligand (Figure 13). Although chimeras containing cpLgBiT or cpLgBIT+4 insertions were 2-log dimmers (Figure 13A), they yielded 20-fold higher BRET efficiency (Figure 13B), further suggesting that engineering greater flexibility / lower stability into the insertions may promote the adoption of conformations that favor not only complementation and HaloTag activity but also BRET.

[0168] Example 9 After determining that the insertion of NanoLuc into loop 180 of HaloTag resulted in both a functional enzyme and improved energy transfer through BRET, experiments were performed to test a panel of constructs containing different circularly permuted variants of NanoLuc inserted into HaloTag at loop 180 (Figure 14). All but two of the cpNanoLuc insertion designs tested significantly improved the BRET ratio over the unpermuted NanoLuc insertion control. Examples of the success of this strategy are highlighted by the insertion of cpNanoLuc49 and cpNanoLuc67, where both NanoLuc emission and energy transfer through BRET are significantly increased over the NanoLuc insertion control. More broadly, the overall number of successful circular permutations indicates that many configurations of protein insertion into the HaloTag loop are possible, creating efficient sites for positioning fusion partners in close proximity to the bound CA ligand of HaloTag.

[0169] Example 10 Given that positioning and geometric constraints are important for folding, activity, and potential efficiency of energy transfer between component enzymes in a fusion or chimera, during development of the embodiments herein, experiments were performed to test a panel of constructs based on the HaloTag-cpNanoLuc67 construct with insertions at loop 180 with flexible glycine-serine linkers of different sizes flanking the different components (Figure 15). Loop 180 continued to tolerate further modification beyond the insertion of cpNanoLuc67 by insertion of additional linkers ranging in length from 3 to 15 Gly-Ser residues. All of the linker variants were functional for HaloTag and NanoLuc activity and exhibited a range of BRET ratios. Notably, insertion of flexible linkers into the sequence immediately N-terminal to the cpNanoLuc67 insertion or within cpNanoLuc67 itself retained BRET ratios greater than 80% relative to constructs without linkers. This indicates that loop 180 of HaloTag can simultaneously accommodate the insertion of both a linker and a polypeptide, allowing flexibility in the nature and composition of the elements that can be positioned in proximity to the bound CA ligand.

[0170] Example 11 During development of the embodiments described herein, the loop 180 of the HaloTag (i.e., 178 -cpNLuc- 179 Experiments were performed to further characterize the lead HALOTAG-cpNANOLUC chimeras emerging from the screening of alternative circularly permuted sites of NanoLuc inserted at cpNLuc 49 / 50 and flexible linkers that could be incorporated between the components of the chimera (Figures 16-18). The structures of these chimeras incorporating NanoLuc circularly permuted between either amino acids 67 / 68 or 49 / 50, as well as flexible linkers containing three glycine-serine residues, are depicted in Figures 16A and 17A. Purified chimeras were compared for HaloTag-TMR ligand binding kinetics, brightness, and efficiency of intramolecular BRET to the bound TMR ligand (Figure 16). This evaluation revealed that chimeras incorporating cpNLuc 49 / 50 exhibited faster binding kinetics and were brighter, while chimeras incorporating cpNLuc 67 / 68 yielded better BRET efficiency. Furthermore, the binding rate of the chimera incorporating cpNLuc 67 / 68 could be slightly increased by the addition of the flexible linker (L1). Cell-based evaluation of the same chimera transiently expressed in HeLa cells (Figure 17) revealed lower expression of the chimera compared to the NanoLuc-HaloTag fusion. Among the chimeras, the chimera incorporating cpNLuc 49 / 50 had lower expression. The chimera incorporating cpNLuc 67 / 68 had higher expression, which was further increased by the addition of the flexible linker L1. Consistent with the biochemical evaluation, bioluminescence normalized to expression suggested that the chimera incorporating cpNLuc 49 / 50 was brighter but exhibited lower BRET efficiency for bound TMR ligand. This was further demonstrated in BRET imaging experiments (Figure 18) showing that the chimeras, especially the chimera incorporating cpNLuc 67 / 68, yielded significantly higher BRET efficiency for bound TMR ligand.

[0171] Example 12 During the development of the embodiments described herein, experiments were performed to evaluate the ability of loop 194 to tolerate large insertions. To this end, purified HALOTAG-cpNANOLUC chimeras containing insertions of cpNanoLuc 67 / 68 into surface loops 194 and 180 of HaloTag were compared for HaloTag-TMR ligand binding kinetics, brightness, and intramolecular BRET efficiency towards bound TMR ligand (Figure 19). This evaluation revealed that loop 194 can tolerate large insertions. Furthermore, the resulting chimeras showed faster binding kinetics and were brighter compared to chimeras generated by insertion into loop 180. At the same time, the BRET efficiency was significantly lower. These results further support the highly efficient BRET attributes provided by the insertion of circularly permuted NanoLuc into loop 180, likely through close proximity and / or a conformation more favorable for energy transfer to bound TMR ligand.

[0172] Example 13 During development of the embodiments described herein, experiments were performed to evaluate the tolerance of the HALOTAG-cpNANOLUC chimera to genetic fusions, as well as the incorporation of additional mutations in the HaloTag domain (Figure 20). Genetic fusions of the chimera to the N- or C-terminus of the model protein dCas12g1 were not only successfully expressed and purified from E. coli, but also showed brightness and BRET efficiency comparable to that of the unfused chimera, suggesting that the chimera can generally tolerate either N- or C-terminal fusions. dCas12g1-HaloTag 178 -cpNLuc- 179 Given the relatively slow binding kinetics of the fusion, it was selected as a template for incorporating additional mutations into the HaloTag domain. Evaluation of the purified fusion variants for brightness, BRET efficiency, and binding kinetics of the HaloTag-TMR ligand demonstrated that these mutations significantly improved the binding kinetics of the dCas12g1-HaloTag domain. 178 -cpNLuc- 179We found that either the L77I+V197A or P206A mutations had no or positive effect on fusion. Two variants incorporating either the L77I+V197A or P206A mutations showed a significant increase in brightness and binding rate of the HaloTag-TMR ligand, suggesting an increase in overall stability without compromising its ability to adopt a conformation favorable for efficient BRET. Together, these results further demonstrate the flexibility of the chimera to tolerate different compositions of elements.

[0173] Example 14 During development of the embodiments described herein, experiments were performed to evaluate the properties of different constructs incorporating circularly permuted NLuc, either as an insertion into loop 180 of HaloTag (i.e., HALOTAG-cpNANOLUC chimera) or as a fusion to HaloTag circularly permuted in the same loop (Figure 21). Biochemical evaluation of the two main circularly permuted sites, cpNLuc 67 / 68 and cpNLuc 49 / 50, as chimeras or fusions to cpHaloTag revealed that constructs incorporating the same cpNLuc had similar brightness. However, the chimeric constructs showed significantly better BRET efficiency, likely through proximity and / or conformation more favorable for energy transfer to the bound TMR ligand.

[0174] In addition, as expected, genetic fusion of the HALOTAG-cpNANOLUC chimera to NanoLuc resulted in increased brightness, however this construct showed a significantly smaller increase in BRET efficiency compared to the NanoLuc-HaloTag fusion.

[0175] Example 15 During development of the embodiments described herein, experiments were performed to optimize the properties of complementation-based chimeras by circular permutation of LgBiT+4 at the two major cp sites 67 / 68 and 49 / 50, as well as incorporating flexible glycine-serine linkers of different lengths between the components of the chimera (Figures 22 and 23). Biochemical evaluation of these chimeras (Figure 22) revealed that upon complementation with the VS-HiBiT peptide, chimeras incorporating cpLgBiT+4 49 / 50 exhibited faster binding kinetics of the HaloTag-TMR ligand as well as increased brightness, but all with low affinity for the VS-HiBiT peptide. HT- 178 cpLgBiT+4 49 / 50- 179 Incorporation of flexible linkers into the HT-HiBiT complex had no or only a small effect on binding kinetics, brightness, or BRET. However, these linkers, especially the long L1-15, did not significantly affect the binding rate, brightness, or BRET of HT-HiBiT complexes. 178 cpLgBiT+4 49 / 50- 179 The affinity of HT- 178 cpLgBiT+4 67 / 68- 179 Similar analyses of revealed that flexible linkers, especially short linkers, improved binding kinetics, brightness, and BRET, likely due to greater flexibility and the ability to adopt a more stable conformation upon complementation.

[0176] Cell-based evaluation of the same chimeras transfected into genome-edited HeLa cells expressing HiBiT-tagged GAPDH revealed significantly lower expression for the chimeras incorporating cpLgBiT+4 49 / 50 (Figure 23). Interestingly, evaluation in mammalian cells revealed a preference for the incorporation of either cpNLuc or cpLgBiT+4 circularly permuted between residues 67 / 68, as well as the additional inclusion of a flexible short L1-3 linker, for both HALOTAG-cpNANOLUC and HALOTAG-cpLGBIT chimeras.

[0177] Example 16 During development of the embodiments described herein, experiments were performed to optimize the properties of complementation-based chimeras by replacing the circularly permuted LgBiT+4 with the more stable circularly permuted LgTrip. Similar to Example 15, the inserted LgTrip was circularly permuted at the two major cp sites 67 / 68 and 49 / 50 to further explore the impact of a flexible glycine-serine linker between the components of the chimera (Figure 24). Biochemical evaluation of these chimeras revealed that upon complementation with the VS-HiBiT-Trip9 dipeptide, chimeras incorporating cpLgTrip 49 / 50 exhibited faster binding kinetics of the HaloTag-TMR ligand as well as increased brightness, all with low affinity for the dipeptide. HT- 178 cpLgTrip 49 / 50- 179 The incorporation of a flexible linker into HT- did not affect the binding kinetics, but generally reduced the brightness, BRET, and binding affinity, especially to the dipeptide. 178 cpLgTrip+4 67 / 68- 179 Similar analyses of revealed that flexible linkers, especially short linkers, improved binding kinetics, brightness, and BRET, likely due to greater flexibility and the ability to adopt a more stable conformation upon complementation.

[0178] Example 17 During the development of the embodiments described herein, experiments were carried out to evaluate the tolerance of complementation-based chimeras to additional mutations and the incorporation of linkers of different nature and length. Among the HALOTAG-cpLGBIT chimeras tested, HT- 178 (L1-3)cpLgBiT+4 67 / 68- 179The chimera was chosen as a template for incorporating additional mutations within the LgBiT domain as well as different configurations of the linker L-1 since it showed the highest expression, brightness, and BRET efficiency in mammalian cells (Figures 25-27). The ability to express and purify all these configurations as well as to obtain generally higher BRET efficiencies than those obtained with the LgBiT-HaloTag fusion demonstrated the flexibility of these complementation-based chimeras and their compatibility with different configurations. Notably, all chimeras incorporating additional LgTrip mutations showed a range of higher binding affinities for VS-HiBiT. Notably, the variant incorporating all three additional mutations (R112H+V127T+K123E) showed a significantly higher binding affinity to HT- 178 (L1-3)cpLgBiT+4 67 / 68- 179 The BRET efficiency of the VS-HIBiT antibody was similar to that of the VS-HIBiT antibody, but showed significantly higher affinity for VS-HIBiT.

[0179] array HT-SEQ ID NO:1 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARET FQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(loop165 insertion)-SEQ ID NO:2 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDX1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 X 15 X 16 X 17 X 18 X 19 X 20 X 21 X 22 X 23 X 24 X 25 VFIEGTLPMGVVRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT (Insertion at position 165) - SEQ ID NO: 3 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIX1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 X 15 X 16 X 17 X 18 X 19 X 20 X 21 X 22 X 23 X 24 X 25VFIEGTLMGVVRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT (loop 180 insertion) - SEQ ID NO: 4 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVX1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 X 15 X 16 X 17 X 18 X 19 X 20 X 21 X 22 X 23 X 24 X 25 RPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT (insertion at position 180)-SEQ ID NO:5 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEE VVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPX1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X13 X 14 X 15 X 16 X 17 X 18 X 19 X 20 X 21 X 22 X 23 X 24 X 25 LTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(1-163)-SEQ ID NO:6 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLII HT(1-164)-SEQ ID NO:7 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIID HT(1-165)-SEQ ID NO:8 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQ HT(1-166)-SEQ ID NO:9 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQN HT(164-297)-SEQ ID NO:10 DQNVFIEGTLMGVVRPLTEVMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(165-297)-SEQ ID NO:11 QNVFIEGTLMGVVRPLTEEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(166-297)-SEQ ID NO:12 NVFIEGTLMGVVRPLTEEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(167-297)-SEQ ID NO:13 VFIEGTLMGVVRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(1-176)-SEQ ID NO:14 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMG HT(1-177)-SEQ ID NO:15 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGV HT(1-178)-SEQ ID NO:16 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVV HT(1-179)-SEQ ID NO:17 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVR HT(1-180)-SEQ ID NO:18 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRP HT(1-181)-SEQ ID NO:19 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPL HT(1-182)-SEQ ID NO:20 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLT HT(177-297)-SEQ ID NO:21 VVRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(178-297)-SEQ ID NO:22 VRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(179-297)-SEQ ID NO:23 RPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(180-297)-SEQ ID NO:24 PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(181-297)-SEQ ID NO:25 LTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(182-297)-SEQ ID NO:26 TEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(183-297)-SEQ ID NO:27 EVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG NANOLUC-SEQ ID NO:28 MKHHHHHHAIAMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAV LgBiT-SEQ ID NO:29 MVFTLEDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIQRIVRSGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNMLNYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLITPDGSMLFRVTINSHHHHHH SmBiT - SEQ ID NO: 30 VTGYRLFEEIL HiBiT - SEQ ID NO:31 VSGWRLFKKIS Dual insert sequence - SEQ ID NO:32 NVFIEGTLPMG

[0180] Exemplary circularly permuted NanoLuc insertions into HaloTag loops Nomenclature: "HaloTag [HT residues preceding the insert]-cpNLuc [NLuc residues preceding the CP site / NLuc residues following the CP site]-[HT residues following the insert]" *All constructs contain a linker between the two NLuc domains, and constructs may also contain one or more linkers. HaloTag164-cpNLuc67 / 68-167-SEQ ID NO:33 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRN PERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDE RLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLVFIEGTLPMGV VRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag194-cpNLuc67 / 68-195-SEQ ID NO:34 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVRPLTEVEMDHYREPFLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFG RPYEGIAVFDGGKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIPYEGLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-NLuc-179- sequence number 35 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKI DIHVIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILARPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc67 / 68-179-SEQ ID NO:36 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cptsNLuc67 / 68-179-SEQ ID NO:37 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRN PERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAVFDGEKIT VTGTLWNGNKIIDERLITPDGSMLFRVTINGVSGWRLFKKISGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIQRIVRSGENALKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cptsNLuc67 / 68-179-SEQ ID NO:38 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRN PERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAVFDGEKIT VTGTLWNGNKIIDERLITPDGSMLFRVTINGVSGWRLCKKISGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIQRIVRSGENALKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc6 / 7-179-SEQ ID NO:39 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMPMGVVDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLS GDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTL ERPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc12 / 13-179-SEQ ID NO:40 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQ IEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGD WRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc24 / 25-179-SEQ ID NO:41 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVD DHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQV LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc27 / 28-179-SEQ ID NO:42 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHH FKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQ GRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc49 / 50-179-SEQ ID NO:43 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDY FGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVL SRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc52 / 53-179-SEQ ID NO:44 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGR PYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGE NRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc55 / 56-179-SEQ ID NO:45 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYE GIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGL KRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc64 / 65-179-SEQ ID NO: 46 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc70 / 71-179-sequence number 47 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGT LWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSG DRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc79 / 80-179-SEQ ID NO:48 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVKLIIDQNVFIEGTLMGVVKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIID ERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKI FRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc82 / 83-179-SEQ ID NO:49 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERL INPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKV VRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc84 / 85-179-SEQ ID NO:50 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLIN PDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVY PRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc86 / 87-179-SEQ ID NO:51 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPD GSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPV DRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc103 / 104-179-SEQ ID NO:52 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc106 / 107-179-sequence number 53 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCER ILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVT PRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc120 / 121-179-SEQ ID NO:54 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGG SMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGI ARPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc124 / 125-179-SEQ ID NO:55 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVF TLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFD GRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc130 / 131-179-SEQ ID NO:56 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFV GDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc145 / 146-179-SEQ ID NO:57 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVNPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLIRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-cpNLuc148 / 149-179-sequence number 58 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178-(L1-3)cpNLuc67 / 68-SEQ ID NO: 59 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L1-6) cpNLuc67 / 68-179 - SEQ ID NO: 60 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L1-9) cpNLuc67 / 68-179 - SEQ ID NO: 61 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L1-12) cpNLuc67 / 68-179 - SEQ ID NO: 62 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGGSSSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L1-15) cpNLuc67 / 68-179 - SEQ ID NO: 63 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGGSSSGGSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L2-3) cpNLuc67 / 68-SEQ ID NO: 64 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L2-6) cpNLuc67 / 68-179 - SEQ ID NO: 65 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGSGGGMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L2-9) cpNLuc67 / 68-179 - SEQ ID NO: 66 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGSGGGGSGMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L2-12) cpNLuc67 / 68-SEQ ID NO: 67 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGSGGGGSGGSSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L2-15) cpNLuc67 / 68 - SEQ ID NO: 68 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGSGGGGSGGSSSGGMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L3-3) cpNLuc67 / 68-179 - SEQ ID NO: 69 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLGGSRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L3-6) cpNLuc67 / 68-179 - SEQ ID NO: 70 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLGGSGGGRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L3-9) cpNLuc67 / 68-179 - SEQ ID NO: 71 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLGGSGGGGSGRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L3-12) cpNLuc67 / 68-179 - SEQ ID NO: 72 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLGGSGGGGSGGSSRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (L3-15) cpNLuc67 / 68-179 - SEQ ID NO: 73 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLGGSGGGGSGGSSSGGRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178 (Q165H + P174R) cpNLuc67 / 68179 - SEQ ID NO: 74 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDHNVFIEGTLRMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178(L771I)-cpNLuc67 / 68179-SEQ ID NO:75 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDIGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178(L77I)-cpNLuc67 / 68-179(V197A)-SEQ ID NO:76 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDIGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPADREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178(M22L)-cpNLuc67 / 68179-sequence number 77 GSEIGTGFPFDPHYVEVLGERLHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178(M69F) cpNLuc67 / 68 - SEQ ID NO: 78 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGFGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178(P206A)cpNLuc67 / 68-179-SEQ ID NO:79 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFANELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag178(W141E)cpNLuc67 / 68-179-SEQ ID NO:80 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEFPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITV TGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEG LRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(1-192)-SEQ ID NO:81 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEVEMDHYREP HT(1-193)-SEQ ID NO:82 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEVEMDHYREPF HT(1-194)-SEQ ID NO:83 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEVEMDHYREPFL HT(1-195)-SEQ ID NO:84 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEVEMDHYREPFLN HT(1-196)-SEQ ID NO:85 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPLTEVEMDHYREPFLNP HT(193-297)-SEQ ID NO:86 FLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(194-297)-SEQ ID NO:87 LNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(195-297)-SEQ ID NO:88 NPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(196-297)-SEQ ID NO:89 PVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HT(197-297)-SEQ ID NO:90 VDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG

[0181] Exemplary circularly permuted LgBiT insertion into the HaloTag loop Nomenclature: "HaloTag [HT residue preceding insert]-cpLgBiT [LgBiT residue preceding CP site / LgBiT residue following CP site]-[HT residue following insert]" *The construct contains a linker between the two LgBiT domains, and the construct may also contain one or more linkers. HaloTag-178-cpLgBiT 67 / 68-179-SEQ ID NO: 91 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNMLNYFGRPYEGIAVFD GKKITVTGTLWNGNKIIDERLITPDGSMLFRVTINSGGTGGSGGTGGSMVFTLEDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIQRIVRSGENALKIDIHVIIPYEGLRP LTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-cpLgBiT+4 67 / 68-(E4D,Q42M,M106K,T144D)179-SEQ ID NO:92 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAVFD GKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRP LTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179-SEQ ID NO:93 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3;L3-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179- seq. GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKR NPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAVFD GKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLGGS RPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-15)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179-SEQ ID NO:95 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGGSSSGGSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRP YEGIAVFDGKKITTVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYE GLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-cpLgBiT+4 49 / 50(E4D,Q42M,M106K,T144D)-179-SEQ ID NO:96 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVT PNKLNYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSRP LTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 49 / 50(E4D,Q42M,M106K,T144D)-179-SEQ ID NO:97 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGENALKIDIHVIIPYEGLSDQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDG VTPNKLNYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3;L3-3)cpLgBiT+4 49 / 50(E4D,Q42M,M106K,T144D)-179- seq. GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKR NPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVT PNKLNYFGRPYEGIAVFDGKKITTVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGGS RPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-15)cpLgBiT+4 49 / 50(E4D,Q42M,M106K,T144D)-179-SEQ ID NO:99 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGGSSSGGGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYG TLVIDGVTPNKLNYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIV RSRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-194-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-195-SEQ ID NO:100 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVRPLTEEVEMDHYREPFLGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVT PNKLNYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGEN ALKIDIHVIIPYEGLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,K123E)-179-SEQ ID NO:101 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGEKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,K123E,R112H)-179-SEQ ID NO:102 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAV FDGEKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,K123E,V127T)-179-SEQ ID NO:103 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGEKITTTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,R112H)-179-SEQ ID NO:104 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGPYEGIAV FDGKKITTVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,R112H,V127T)-179-SEQ ID NO:105 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAV FDGKKITTTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,R112H,V127T,K123E)-179-SEQ ID NO:106 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAV FDGEKITTTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,V127T)-179-SEQ ID NO:107 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITTTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,H86R)-179-SEQ ID NO:108 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDRHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITTVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,H86R,L142R)-179-SEQ ID NO:109 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDRHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERRIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,I73L,L142R)-179-SEQ ID NO:110 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQLEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERRIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,I73L,L142R,H86R)-179-SEQ ID NO:111 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQLEEVFKVVYPVDDRHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERRIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,I73L,H86R)-179-SEQ ID NO:112 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQLEEVFKVVYPVDDRHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITTVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,I73L)-179-SEQ ID NO:113 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQLEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D,L142R)-179-SEQ ID NO:114 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERRIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-GPR)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179- seq.number115 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGPRSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-GRP)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179- seq. GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGRPSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-RPG)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179- sequent number 117 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVRPGSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-VPR)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179- seq.number118 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVVPRSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-VRP)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179- seq 119 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAK RNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVVRPSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLR PLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-13*)cpLgBiT+4 67 / 68(E4D,Q42M,M106K,T144D)-179-SEQ ID NO:120 *Linker EPTTEDLYFQSDN GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNP ERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVEPTTEDLYFQSDNSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGRPY EGIAVFDGKKITVTGTLWNGNKIIDERLIDPDGSMLFRVTINSGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYE GLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-cpNLuc67 / 68-179-NLuc-SEQ ID NO: 121 MAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIID QNVFIEGTLMGVVSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQ NLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLRPLTEEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISGGSGMVFTLEDFV GDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGS cpHaloTag178 / 179-cpNLuc67 / 68-SEQ ID NO: 122 MLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISGGSSGGGSSGG EPTTENLYFQSDNGSSGGGSGGMAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHD WGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLMGVVRPSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAV FDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLSGENGLKIDIHVIIPYEGL cpHT178 / 179-cpNLuc49 / 50- SEQ ID NO: 123 MLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISGGSSGGGSSGGEPTTENLYFQSDNGSSGGGSSGGMAEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVRPGENGLKIDIHVIIPYEGLSGDQMGQIEKIFKVVYPVDDHHFKVILHYGTLVIDGVTPNMIDYFGRPYEGIAVFDGKKITVTGTLWNGNKIIDERLINPDGSLLFRVTINGVTGWRLCERILAGGTGGSGGTGGSMVFTLEDFVGDWRQTAGYNLDQVLEQGGVSSLFQNLGVSVTPIQRIVLS HaloTag-178-cpLgTrip67 / 68-179-SEQ ID NO:124 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgTrip67 / 68-SEQ ID NO: 125 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3;L3-3)cpLgTrip67 / 68-179-SEQ ID NO: 126 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHW AKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYE GIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLGGSRPL TEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-15)cpLgTrip67 / 68-179-SEQ ID NO: 127 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGGSSSGGSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGENALKIDIHVIIPYEGLRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-cpLgTrip49 / 50-179-SEQ ID NO:128 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVDDTHR.KCPEDRHPCHHPV.RSERRPNGPDRRGV.GGVPCG.SSL.GDPALWHTGNRRGYAEQAELFRTPV.RHRRVRRREDHYHRDPVERQQNYRRAPDRSRRRNRWQRWNREHGLHTRRFRWGLGTDSRLQPGPSP.TGRCVQFAAESRRVRNSDHEDCPEPLVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3)cpLgTrip49 / 50-179-SEQ ID NO: 129 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHW AKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYGT LVIDGVTPNKLNYFGHPYEGIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRRSRPLT EVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-3;L3-3)cpLgTrip49 / 50-179-SEQ ID NO: 130 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHW AKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYGTL VIDGVTPNKLNYFGHPYEGIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSGGSRPL TEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPPVKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG HaloTag-178-(L1-15)cpLgTrip49 / 50-179-SEQ ID NO: 131 GSEIGTGFPFDPHYVEVLGERMHYVDVGPRDGTPVLFLHGNPTSSYVWRNIIPHVAPTHRCIAPDLIGMGKSDKPDLGYFFDDHVRFMDAFIEALGLEEVVLVIHDWGSALGFHWAKRNPERVKGIAFMEFIRPIPTWDEWPEFARETFQAFRTTDVGRKLIIDQNVFIEGTLPMGVVGGSGGGGSGGSSSGGGENALKIDIHVIIPYEGLSADQMAQIEEVFKVVYPVDDHHFKVILPYGTLVIDGVTPNKLNYFGHPYEGIAVFDGEKITTTGTLWNGNKIIDERLIDPDGGTGGSGGTGGSMVFTLDDFVGDWEQTAAYNLDQVLEQGGVSSLLQNLAVSVTPIMRIVRSRPLTEVEMDHYREPFLNPVDREPLWRFPNELPIAGEPANIVALVEEYMDWLHQSPVPKLLFWGTPGVLIPPAEAARLAKSLPNCKAVDIGPGLNLLQEDNPDLIGSEIARWLSTLEISG

Table 2-1

Table 2-2

Table 2-3

Table 2-4

Table 2-5

Table 2-6

Table 2-7

Table 2-8

Table 2-9

Table 2-10

Table 2-11

Table 2-12

Table 2-13

Claims

【Request Item 1】 A composition comprising a polypeptide having at least 70% sequence identity with Sequence ID No. 2, wherein X 1 ~X 25 Each of them is independently selected from any amino acid, or is absent, X 1 ~X 25 The composition wherein at least five of the elements are not absent, and the polypeptide has less than 100% sequence identity with SEQ ID NO: 1.