A novel design for protein switches

By designing non-natural peptides containing helical bundles and amino acid linkers to form interactions between cage peptides and key peptides, the difficult problem of protein systems converting conformational states under external inputs was solved, and a modular and adjustable protein switch was realized to isolate or activate bioactive peptides.

CN112512544BActive Publication Date: 2025-09-09UNIV OF WASHINGTON
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980048359.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-04
Filing Date
2019-07-19
Publication Date
2025-09-09
Estimated Expiration
2039-07-19

AI Technical Summary

Technical Problem

The design of existing protein systems that switch conformational states under external input has not yet been achieved, especially the challenge of the free energy difference between multiple states is not sufficient to be switched by external input.

Method used

A non-naturally occurring polypeptide was designed, comprising a helix bundle and an amino acid linker connecting each α-helix to form a cage polypeptide. The conformational change is achieved through the interaction between the key polypeptide and the cage polypeptide, thereby isolating or activating the bioactive peptide.

Benefits of technology

A modular and adjustable protein switch was realized, which can undergo conformational transitions in response to protein binding, effectively isolate and activate bioactive peptides, and meet the switching needs of external input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0002905503330000191
    Figure BDA0002905503330000191
  • Figure BDA0002905503330000221
    Figure BDA0002905503330000221
  • Figure BDA0002905503330000222
    Figure BDA0002905503330000222
Patent Text Reader

Abstract

Disclosed are protein switches, components of such protein switches, and uses thereof, which can sequester biologically active peptides and / or binding domains, keeping them in an inactive ("off") state until combined with a second, designed polypeptide, termed a key, thereby inducing a conformational change that activates ("on") the biologically active peptide or binding domain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to earlier filed applications

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 700,681, filed on July 19, 2018, U.S. Provisional Patent Application Serial No. 62 / 785,537, filed on December 27, 2018, and U.S. Provisional Patent Application Serial No. 62 / 788,398, filed on January 4, 2019, each of which is incorporated herein by reference in its entirety.

[0003] Citations of Sequence Listings Submitted Electronically via EFS-Web

[0004] The contents of the electronically submitted sequence listing of the ASCII text file submitted with this application (Name: 18-1054-PCT_Sequence-Listing_ST25.txt; Size: 32,278 kb; and Creation Date: July 19, 2019) are incorporated herein by reference in their entirety. Background Art

[0005] Considerable progress has been made in the de novo design of stable protein structures, based on the principle that proteins fold into their lowest free energy state. These efforts focus on maximizing the free energy gap between the desired folded structure and all other structures. Designing proteins that can switch conformations is more challenging because the multiple states must have sufficiently low free energies to populate relative to the unfolded state, and the free energy differences between the states must be small enough that the state occupancy can be switched by external input. The de novo design of protein systems that switch conformational states in the presence of external input has not yet been achieved. Summary of the Invention

[0006] In a first aspect, a non-naturally occurring polypeptide is disclosed, comprising:

[0007] (a) a helical bundle comprising 2 to 7 α-helices; and

[0008] (b) The amino acid linker connecting each α-helix.

[0009] In one embodiment, each helix is ​​independently 18-60, 18-55, 18-50, 18-45, 22-60, 22-55, 22-50, 22-45, 25-60, 25-55, 25-50, 25-45, 28-60, 28-55, 28-50, 28-45, 32-60, 32-55, 32-50, 32-45, 35-60, 35-55, 35-50, 35-45, 38-60, 38-55, 38-50, 38-45, 40-60, 40-58, 40-55, 40-50, or 40-45 amino acids in length. In another embodiment, the length of each amino acid linker is independently 3-10, 4-10, 5-10, 6-10, 7-10, 8-10, 9-10, 2-9, 3-9, 4-9, 5-9, 6-9, 7-9, 8-9, 2-8, 3-8, 4-8, 5-8, 6-8, 7-8, 2-7, 3-7, 4-7, 5-7, 6-7, 2-6, 3-6, 4-6, 5-6, 2-5, 3-5, 4-5, 2-4, 3-4, 2-3, or 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids, which length does not include any other functional sequences that may be fused to the linker. In another embodiment, the polypeptide comprises one or more bioactive peptides in at least one alpha helix, wherein the one or more bioactive peptides are capable of selectively binding to a defined target, wherein the one or more bioactive peptides may comprise one or more bioactive peptides selected from the non-limiting group consisting of SEQ ID NOs: 60, 62-64, 66, 27052-27093, and 27118-27119.

[0010] In another aspect, the present disclosure provides a non-naturally occurring polypeptide comprising a polypeptide having at least 40% sequence identity along its length to an amino acid sequence of a cage polypeptide disclosed herein or selected from the group consisting of SEQ ID NOs: 1-49, 51-52, 54-59, 61, 65, 67-14317, 27094-27117, 27120-27125, 27278 to 27321, and a cage polypeptide listed in Table 2, Table 3, and / or Table 4, wherein the N-terminal and / or C-terminal 60 amino acids of the polypeptide are optional, wherein the sequence identity requirement excludes optional amino acid residues. In one aspect, the present disclosure provides a non-naturally occurring polypeptide comprising a polypeptide having at least 40% sequence identity along its length to an amino acid sequence of a cage polypeptide listed in Table 2, Table 3, and / or Table 4 (excluding optional amino acid residues). In one embodiment of each of these aspects, the polypeptide comprises an amino acid sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-91, SEQ ID NOs: 1-49, 51-52, 54-59, 61, 65, 67-14317, 27094-27117, 27120-27125, 27278 to 27321, and an amino acid sequence of a cage polypeptide listed in Table 2, Table 3 and / or Table 4 (excluding optional amino acid residues).

[0011] In one embodiment of any aspect of the present disclosure, the non-naturally occurring polypeptide further comprises one or more bioactive peptides within or in place of the latch region of the polypeptide, wherein the one or more bioactive peptides may include one or more bioactive peptides selected from the non-limiting group consisting of SEQ ID NOs: 60, 62-64, 66, 27052-27093 and 27118-27119.

[0012] In another aspect, non-naturally occurring polypeptides are provided, comprising polypeptides having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity (excluding optional amino acid residues) along their length to a key polypeptide disclosed herein or a key polypeptide selected from the group consisting of SEQ ID NOs: 14318-26601, 26602-27015 and 27016-27051, and the amino acid sequence of the key polypeptides listed in the Tables. In one embodiment, the polypeptide comprises an amino acid sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity (excluding optional amino acid residues) along its length to the amino acid sequence of a key protein selected from the group consisting of key polypeptides listed in Table 2, Table 3 and / or Table 4.

[0013] In another embodiment, the present disclosure provides a fusion protein comprising a cage polypeptide according to any embodiment or combination of embodiments of the present disclosure fused to a key polypeptide according to any embodiment or combination of embodiments of the present disclosure.

[0014] In other aspects, the present disclosure provides nucleic acids encoding the cage polypeptide and key polypeptide of the fusion protein according to any embodiment or combination of embodiments of the present disclosure; expression vectors comprising the nucleic acid operably linked to a promoter; and / or host cells comprising the nucleic acid and / or expression vector. In one embodiment, the nucleic acid or expression vector is integrated into the host cell chromosome. In another embodiment, the nucleic acid or expression vector is episomal. In another embodiment, the host cell comprises:

[0015] (a) a first nucleic acid encoding a cage polypeptide according to any embodiment or combination of embodiments of the present disclosure operably linked to a first promoter; and

[0016] (b) a second nucleic acid encoding a key polypeptide according to any embodiment or combination of embodiments of the present disclosure operably linked to a second promoter, wherein the second nucleic acid encodes a key polypeptide capable of binding to a structural region of a cage polypeptide encoded by the first nucleic acid, and wherein binding of the key polypeptide to the structural region of the cage polypeptide induces a conformational change in the cage polypeptide.

[0017] In another embodiment of the host cell of the present disclosure, the first nucleic acid comprises a plurality of first nucleic acids encoding a plurality of different cage polypeptides. In one embodiment, the second nucleic acid comprises a plurality of second nucleic acids encoding a plurality of different key polypeptides, wherein the plurality of different key polypeptides comprises one or more key polypeptides capable of binding to only a subset of the plurality of different cage polypeptides and inducing a conformational change in that subset. In another embodiment, the second nucleic acid encodes a single key polypeptide capable of binding to each different cage polypeptide and inducing a conformational change in each different cage polypeptide.

[0018] In another embodiment, the host cell may comprise:

[0019] (a) a first nucleic acid encoding a fusion protein according to any embodiment or combination of embodiments of the present disclosure operably linked to a first promoter; and

[0020] (b) a second nucleic acid encoding a fusion protein according to any embodiment or combination of embodiments of the present disclosure operably linked to a second promoter, wherein:

[0021] (i) a cage polypeptide encoded by a first nucleic acid is activated by a key polypeptide encoded by a second nucleic acid;

[0022] (ii) the cage polypeptide encoded by the first nucleic acid is not activated by the key polypeptide encoded by the first nucleic acid;

[0023] (iii) the cage polypeptide encoded by the second nucleic acid is activated by the key polypeptide encoded by the first nucleic acid; and

[0024] (iv) the cage polypeptide encoded by the second nucleic acid is not activated by the key polypeptide encoded by the second nucleic acid.

[0025] In another aspect, the present disclosure provides a kit comprising:

[0026] (a) one or more cage polypeptides according to any embodiment or combination of embodiments of the present disclosure;

[0027] (b) one or more key polypeptides according to any embodiment or combination of embodiments of the present disclosure; and

[0028] (c) Optionally, one or more fusion proteins according to any embodiment or combination of embodiments of the present disclosure.

[0029] In one aspect, the present disclosure provides a kit comprising:

[0030] (a) a first nucleic acid encoding one or more cage polypeptides according to any embodiment or combination of embodiments of the present disclosure;

[0031] (b) a second nucleic acid encoding one or more key polypeptides according to any embodiment or combination of embodiments of the present disclosure; and

[0032] (c) Optionally, a third nucleic acid encoding one or more fusion proteins according to any embodiment or combination of embodiments of the present disclosure.

[0033] In various embodiments, the first nucleic acid, the second nucleic acid, and / or the third nucleic acid comprises an expression vector.

[0034] In another aspect, the present disclosure provides a LOCKR switch comprising:

[0035] (a) a cage polypeptide comprising a structural region and a latch region further comprising one or more biologically active peptides, wherein the structural region interacts with the latch region to prevent the activity of the one or more biologically active peptides;

[0036] (b) optionally a key polypeptide that binds to the cage region, thereby displacing the latch region and activating the one or more biologically active peptides; and

[0037] (c) optionally, one or more effector polypeptides that bind to the one or more biologically active peptides when the one or more biologically active peptides are activated.

[0038] In one embodiment, an effector polypeptide is present, and wherein the effector polypeptide comprises an effector polypeptide that selectively binds to a biologically active peptide, including but not limited to Bcl2, GFP1-10, and a protease.

[0039] In various embodiments, a host cell, kit, or LOCKR according to any embodiment or combination of embodiments of the present disclosure comprises one or more cage polypeptides and one or more key polypeptides having amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity along their length to a cage polypeptide and a key polypeptide, respectively, in the same row of Table 1, Table 2, Table 3, and / or Table 4, or to a cage polypeptide and a key polypeptide, respectively, in the same row of Table 2, Table 3, and / or Table 4 (excluding optional residues). In other embodiments, the one or more bioactive peptides may include one or more bioactive peptides selected from the non-limiting group consisting of SEQ ID NOs: 60, 62-64, 66, 27052-27093, and 27118-27119.

[0040] In other embodiments, the host cell, kit, or LOCKR according to any embodiment or combination of embodiments of the present disclosure comprises:

[0041] (a) one or more cage polypeptides comprising one or more cage polypeptides having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along their length to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-49, 51-52, 54-59, 61, 65, 67-14317, 27094-27117, 27120-27125, 27278 to 27321, and a cage polypeptide listed in Table 2, Table 3 and / or Table 4, excluding optional amino acid residues, wherein the N-terminal and / or C-terminal 60 amino acids of the polypeptide are optional; and

[0042] (b) one or more key polypeptides selected from the group consisting of one or more polypeptides having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity (excluding optional amino acid residues) along its length to a key polypeptide selected from the group consisting of SEQ ID NOs: 14318-26601, 26602-27015 and 27016-27051, and a key polypeptide listed in Table 2, Table 3 and / or Table 4.

[0043] In another aspect, the present disclosure provides uses of the polypeptides, fusion proteins, nucleic acids, expression vectors, host cells, kits and / or LOCKR switches disclosed herein for sequestering biologically active peptides in caged polypeptides, maintaining them in an inactive ("off") state until combined with a key polypeptide to induce a conformational change that activates ("on") the biologically active peptide. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Design of the A-1E.LOCKR switch system. Figure 1 Figure A shows a thermodynamic model describing our design goal. The structural domains in the cage and the latch region form a switch that has a certain equilibrium between open and closed states. A key can bind to the cage to promote the open state, thereby allowing the target to bind to the latch. Figure 1 B shows two K from the model in (a) LTFigure 3 shows how fractional target binding is affected by the addition of a key (K CK =1 nM); curves of different colors show K 打开 =[On] / [Off] the effect of the logarithmic decreasing value. Figure 1 C shows the addition of a loop to the homotrimer 5L6HC3_1 5 to form monomeric five- and six-helical frameworks; the double mutant V217S / I232S weakens the latch, allowing it to be displaced by the key, resulting in a LOCKR system capable of binding exogenous keys. Figure 1 D shows chemical denaturation using guanidine chloride (Gdm) monitoring the mean residue ellipticity (MRE) at 222 nm. Figure 1 E shows small angle x-ray scattering (SAXS) which shows that the monomeric frameworks exhibit spectra that are very consistent with each other and with the original homotrimer. Figure 1 F, Pull-down assay showing key binding to the truncated five-helical framework and LOCKR (V217S / I232S), but not to the six-helical monomer; free GFP-key was added to the plate-immobilized monomeric framework via the hexahistidine tag; after a series of wash steps, binding was measured by GFP fluorescence (n = 2, error bars indicate standard deviation).

[0045] Figure 2 A-2D.BimLOCKR's design and activities. Figure 2 A. The free energy of the latch-cage interface is tuned by suboptimal Bim-cage interactions (left, shown as altered hydrophobic stacking and buried hydrogen bonds) and by exposing hydrophobic residues at the ends of the interface (right) as underpinnings. B) Introduction of the underpinning allows activation of 250 nM BimLOCKR via biolayer interferometry with the addition of 5 μM key ("on" lever). C) Biolayer interferometry shows key-dependent binding of Bcl2 with 250 nM BimLOCKR. Association takes from 0-500 s, followed by dissociation from 500-1700 s. D) Each point is the result of fitting the data in I and extracting the response at equilibrium. The curve shows similar data with a shorter key, indicating that the K of LOCKR can be tuned. CK and affects its activation range.

[0046] Figure 3 Design and validation of orthogonal BimLOCKR. A) Left: Cartoon representation of LOCKR. Cage with three different latches superimposed and hydrogen-bonding network labeled with markers. Right: Hydrogen-bonding network at the orthogonal LOCKR interface. b) BimLOCKR binding to Bcl2 in response to its cognate key on Octet. One replicate. c) Response of each switch (250 nM) and key (5 μM) pair to Octet after 500 seconds. Average of two replicates.

[0047] Figure 4 Experimentally determined X-ray crystal structures of asymmetric LOCKR switch designs. (A) Crystal structure of design 1-fix-short-BIM-t0, which contains the encoded BIM peptide. (B) Crystal structure of design 1-fix-short-noBim(AYYA)-t0 closely agrees with the design model in terms of (left) backbone, (center) hydrogen bond network, and (right) hydrophobic packing; the latch region that would encode Bim and Gfp11 is shown; electron density maps illustrate the network and hydrophobic cross-section (center and right).

[0048] Figure 5 In the absence of the key, the LOCKR switch prevents the split GFP11 from complementing GFP1-10. A. Crystal structure of GFP with chain 11 (pdb 2y0g) is shown. B. Crystal structure of a prototype switch in which GFP11 is stabilized as a helix (the mesh is the electron density). C. Computational design model to The root mean square deviation of the LOCKR switch (1fix-short-GFP-t0) was found to match the crystal structure. The experimentally determined X-ray crystal structure of the designed LOCKR switch, 1fix-short-GFP-t0, revealed that the 11th strand of the encoded GFP (GFP11) is an alpha helix and closely agrees with the designed model. GFP fluorescence was observed only in the presence of the key peptide, indicating that the switch is functional (off in the absence of the key and on in its presence).

[0049] Figure 6 .From Figure 4 Design of the GFP11-LOCKR switch, tuned for colocalization dependence. (A)

[0050] Schematic diagram of the test system in which colocalization dependence is determined by a connected Spycatcher TM / Spytag TM Fusion control. In this model, the key should only activate the LOCKR switch (produce fluorescence) when fused to Spytag, which would allow the key to colocalize with the cage (right). When added alone, the key should not be able to activate the LOCKR switch (middle). (B) Fluorescence data demonstrating the colocalization dependence of the LOCKR switch following the design schematic in (A). TM The design of 1fix-latch and 1fix-short is integrated with Spytag TM Homologous keys show more activation when mixed; Spytag is missing TM The keys showed significantly fewer activations.

[0051] Figure 7 Caged intein LOCKR switch. A designed LOCKR switch with a cage assembly encoding the VMAc intein exhibited successful activation when mixed with a designed key fused to sfGFP and the VMAn intein. SDS-PAGE demonstrated a successful VMAc-VMAn reaction, with bands corresponding to the correct molecular weight of the expected spliced ​​protein product.

[0052] Figure 8 Multiple sequence alignment (MSA) comparing the original LOCKR_a cage scaffold design to its asymmetric (1fix-short–noBim(AYYA)-t0) and orthogonal (LOCKRb-f) design counterparts. Across the MSA, only 150 (40.8%) sites are identical, with a pairwise % identity of 69.4%. The latch region (the C-terminal region starting at position labeled 311 in this MSA) has very little sequence identity / similarity. (From top to bottom, SEQ ID NOs: 17, 39, 7, 8, 9, 10, 11)

[0053] Figure 9 .1fix-short-noBim(AYYA)-t0( Figure 4 B) The crystal structure (white) used to prepare LOCKRa ( Figure 1 )Base bracket 5L6HC3_1 5 The superposition of the X-ray crystal structure (black) shows that the asymmetric mutation ( Figure 8 The variable positions shown in the MSA do not affect the three-dimensional structure of the protein. The backbone RMSD between the two proteins is 0.85 Å (from superposition of all backbone atoms between chains A).

[0054] Figure 10 Figure 3: GFP plate assay for mutation discovery in LOCKR. Different putative LOCKR constructs were attached to Ni-coated 96-well plates via a 6x-His tag, key-GFP was applied, and washed excessively. The resulting fluorescence represents key-GFP bound to the LOCKR construct. Truncations were used as positive controls because the key binds to the open interface. Monomers served as negative controls because they do not bind the key. Error bars represent the standard deviation of three replicates.

[0055] Figure 11: Orthogonal LOCKR GFP assay. A) In five redesigned LOCKR constructs (b to f), the latch was truncated from a 6x-His-tagged cage. The corresponding key was GFP-tagged. Key-cage binding was measured by Ni pull-down of the cage and measuring the resulting GFP fluorescence. Error bars are the standard deviation of three replicates. B) Each complete LOCKR construct that binds the key from (a) was given a nine-residue base and tested for binding to all four functional keys (a to d) in a GFP pull-down assay. Error bars are the standard deviation of five replicates. Key a was suspected to be promiscuous binding, but not activating, because LOCKR is generated from homotrimeric pseudosymmetry. Considering that the key from (a) and Figure 3 Results for b, LOCKRb shows no binding to its own key, which is attributed to the latch strength.

[0056] Figure 12 : Designed Mad1-SID LOCKR switch for key-dependent transcriptional repression. (a) Crystal structure of the interaction between the Mad1-SID domain (white) and the PAH2 domain (black) of the mSin3A transcriptional repressor (PDB ID: 1E91). Locking of the Mad1-SID domain should enable key-dependent recruitment of the transcriptional repressor mSin3A, thereby achieving precise epigenetic regulation. (b) Designed Mad1-LOCKR switch with cage components encoding Mad1-SID sequences at different positions (dark gray). (c) SDS-PAGE gel showing successfully purified 1fix_302_Mad1 (1), 1fix_309_Mad1 (2), and MBP_Mad1 (3). (d) Biolayer interferometry analysis of key-activated binding of the Mad1-LOCKR switch to the purified mSin3A-PAH2 domain. MBP-Mad1 is a positive control for mSin3a-PAH2 binding. 1fix_309_Mad1(309) in the key with design a 1fix_302_Mad1(302) showed very tight locking of the Mad1-SID domain, but in the key a Kinetic assays were performed by immobilizing 0.1 μg of biotin-mSin3A-PAH2 protein on a streptavidin biosensor tip (ForteBio). a In this case, the protein cages were tested at 50 nM.

[0057] Figure 13Caged STREPII-tagged LOCKR switches; Demonstration of the new 2plus1 and 3plus1 LOCKR switches. (A) Designs of 2+1 (left) and 3+1 (center) LOCKR switches were designed to encode the STREPII sequence WSHPQFEK (SEQ ID NO: 63). (B) Biolayer interferometry (Octet) data demonstrating functionality of the STREPII-LOCKR design: Anti-streptavidin antibodies were immobilized on anti-mouse Fc tips to assess binding of the STREPII tag. (B) The designed protein showed less binding than the positive control, indicating that STREPII was at least partially sequestered as expected. (C) Activation of the STREPII-3plus1_Lock_3 design by 3plus1_key_3: Curves are for 250 nM cage without key, compared to 250 nM cage in the presence of increasing concentrations of key, ranging from 121 nM to 6000 nM. (D) 250nM STREPII-3plus1_Lock_3 in the presence of 370nM, 1111nM, and 3333nM keys; 250nM cage without key is 250nM, and the other graphs are keys at the same concentrations (370nM, 1111nM, and 3333nM) but without cages. In all Octet graphs, the left half is the association (binding) step, and the right half is the dissociation step.

[0058] Figure 14. The 3plus1 LOCKR switch activates GFP fluorescence in response to expression of the key. A LOCKR switch was designed in which the 3plus1 cage was used to isolate chain 11 (GFP11) of GFP in an inactive conformation, thereby preventing the reconstruction of the split GFP (composed of GFP1-10 and GFP11), thereby generating fluorescence. Expression plasmids were prepared to inducibly express the cage (p15a replication origin, spectinomycin resistance, arabinose-inducible promoter controlling the expression of GFP1-10 and LOCKR caged GFP11) and the key (colE1 replication origin, kanamycin resistance, and IPTG-inducible promoter). Chemically competent Escherichia coli Stellar cells (Takarabio) were transformed with either a single cage plasmid or both the cage and key plasmids according to the manufacturer's protocol. These transformations were grown overnight at 37°C in liquid LB medium supplemented with spectinomycin (cage alone) or spectinomycin + kanamycin (cage and key). The resulting culture was diluted 1 / 100 into fresh LB medium supplemented with appropriate antibiotics and either arabinose alone (to induce expression of cage and GFP1-10) or both arabinose and IPTG (to induce expression of cage, GFP1-10, and key) and then grown at 37°C for 16 hours. 200uL of each expression culture was washed once in 200uL PBS, resuspended in 200uL PBS, and transferred to a black-walled 96-well plate. GFP fluorescence was assessed on a Biotek Synergy H1MF plate reader (excitation / emission 479 / 520nM). The fluorescence of the cage alone was minimal, confirming that the LOCKR protein prevented the activation of the split GFP in the absence of the key. Induction of key expression resulted in a significant increase in the fluorescence of SEQ ID NOs 27192, 27198, 27194, 27202, 27206, and 27210. These results demonstrate that the 3plus1 LOCKR architecture is capable of controlling the function of the bioactive peptide GFP11. DETAILED DESCRIPTION

[0059] All references cited are herein incorporated by reference in their entirety. In this application, unless otherwise indicated, the techniques used can be found in any of several well-known references, such as: Molecular Cloning: A Laboratory Manual (Sambrook et al., 1989, Cold Spring Harbor Laboratory Press), Gene Expression Technology (Methods in Enzymology, Vol. 185, edited by D. Goeddel, 1991. Academic Press, San Diego, CA), "Guide to Protein Purification" in Methods in Enzymology (MP Deutshcer, ed., (1990) Academic Press, Inc.); PCR Protocols: A Guide to Methods and Applications (Innis et al. 1990. Academic Press, San Diego, CA), Culture of Animal Cells: A Manual of Basic Technique, 2nd Edition (RI Freshney. 1987. Liss, Inc. New York, NY), Gene Transfer and Expression Protocols, pp. 109-128, EJ Murray, ed., The Humana Press Inc., Clifton, NJ) and the Ambion 1998 Catalog (Ambion, Austin, TX).

[0060] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. As used herein, "and" and "or" are used interchangeably unless expressly stated otherwise.

[0061] As used herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine ​​(Cys; C); glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G); histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).

[0062] Unless the context clearly dictates otherwise, all embodiments of any aspect of the disclosure may be used in combination.

[0063] Unless the context clearly requires otherwise, throughout the specification and claims, the words "comprise," "comprising," and the like are to be construed in an inclusive sense, and not in an exclusive or exhaustive sense; that is, in the sense of "including, but not limited to." Words using the singular or plural number also include the plural and singular, respectively. Additionally, the words "herein," "above," and "below," and words of similar import, when used in this application, shall refer to this entire application and not to any particular portions of this application.

[0064] The description of the embodiments of the present disclosure is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Although specific embodiments and examples of the present disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the present disclosure, as those skilled in the relevant art will recognize.

[0065] In a first aspect, the present disclosure provides a non-naturally occurring polypeptide comprising:

[0066] (a) a helical bundle comprising 2 to 7 α-helices; and

[0067] (b) The amino acid linker connecting each α-helix.

[0068] The non-naturally occurring polypeptides of this first aspect of the present disclosure can be used, for example, as the "cage" polypeptide component (which may also be referred to herein as a "lock") of the novel protein switches disclosed in detail herein. Protein switches can be used, for example, to sequester biologically active peptides within the cage polypeptide, maintaining them in an inactive ("off") state until combined with a second component (a "key" polypeptide) of the novel protein switches disclosed herein; the key polypeptide induces a conformational change that activates ("turns on") the biologically active peptide (see Figure 1A) The polypeptides described herein include the first completely new designed polypeptides that can undergo conformational transitions in response to protein binding. Furthermore, no known natural proteins can undergo conformational transitions in the modular, tunable manner of the polypeptides disclosed herein.

[0069] The polypeptide is "non-naturally occurring" because the complete polypeptide is not found in any naturally occurring polypeptide. It should be understood that components of the polypeptide can be naturally occurring, including but not limited to biologically active peptides that can be included in some embodiments.

[0070] The cage polypeptide comprises a helical bundle comprising 2 to 7 alpha helices. In various embodiments, the helical bundle comprises 3-7, 4-7, 5-7, 6-7, 2-6, 3-6, 4-6, 5-6, 2-5, 3-5, 4-5, 2-4, 3-4, 2-3, 2, 3, 4, 5, 6, or 7 alpha helices.

[0071] The design of the helical cage polypeptides disclosed herein can be performed in any suitable manner. In one non-limiting embodiment, the design can be based on the Crick expression of the coiled coil using Rosetta Stone. TM BundleGridSampler in the program TM to generate backbone geometry and allow efficient parallel sampling of a regular grid of coiled coil expression parameter values ​​corresponding to a continuum of peptide backbone conformations. This can be supplemented by designing hydrogen bond networks using any suitable means including, but not limited to, those described by Boyken et al. (Science 352, 680-687 (2016)), followed by Rosetta TM Side Chain Design. In another non-limiting embodiment, the best scoring design based on the overall score, the number of unsatisfied hydrogen bonds, and the lack of voids in the protein core can be selected for helix bundle cage polypeptide design.

[0072] Each alpha helix can have any suitable length and amino acid composition suitable for the intended use. In one embodiment, the length of each helix is ​​independently 38 to 58 amino acids. In various embodiments, the length of each helix is ​​independently 18-60, 18-55, 18-50, 18-45, 22-60, 22-55, 22-50, 22-45, 25-60, 25-55, 25-50, 25-45, 28-60, 28-55, 28-50, 28-45, 32-60, 32-55, 32-50, 32-45, 35-60, 35-55, 35-50, 35-45, 38-60, 38-55, 38-50, 38-45, 40-60, 40-58, 40-55, 40-50 or 40-45 amino acids.

[0073] The amino acid linker connecting each alpha helix can have any suitable length or amino acid composition suitable for the intended use. In a non-limiting embodiment, the length of each amino acid linker is independently 2 to 10 amino acids, excluding any other functional sequence that can be fused to the linker. In various non-limiting embodiments, the length of each amino acid linker is independently 3-10, 4-10, 5-10, 6-10, 7-10, 8-10, 9-10, 2-9, 3-9, 4-9, 5-9, 6-9, 7-9, 8-9, 2-8, 3-8, 4-8, 5-8, 6-8, 7-8, 2-7, 3-7, 4-7, 5-7, 6-7, 2-6, 3-6, 4-6, 5-6, 2-5, 3-5, 4-5, 2-4, 3-4, 2-3, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. In all embodiments, the linker can be structured or flexible (eg, poly-GS). These linkers can encode other functional sequences, including but not limited to protease cleavage sites or split halves of the intein system (see sequences below).

[0074] The polypeptide of this first aspect includes a region known as a "latch region" for inserting a biologically active peptide. Thus, a cage polypeptide comprises a latch region and a structural region (i.e., the remaining portion of the cage polypeptide that is not the latch region). When the latch region is modified to include one or more biologically active peptides, the structural region of the cage polypeptide interacts with the latch region to prevent the activity of the biologically active peptide. Upon activation by the key polypeptide, the latch region dissociates from its interaction with the structural region, exposing the biologically active peptide, thereby allowing the peptide to function.

[0075] As used herein, a "bioactive peptide" is any peptide of any length or amino acid composition that is capable of selectively binding to a defined target (i.e., capable of binding to an "effector" polypeptide). Such bioactive peptides may include

[0076] Peptides of all three types of secondary structures: alpha helices, beta strands, and loops in an inactive conformation. Polypeptides of this aspect can be used to control the activity of a wide range of functional peptides. The ability to exploit these biological functions using strict inducible control can be used, for example, to engineer cells (inducible activation of function, engineering complex logic behaviors and circuits, etc.), develop sensors, develop therapeutics based on inducible proteins, and create new biomaterials.

[0077] The latch region may be present near either end of the cage polypeptide. In one embodiment, the latch region is placed at the C-terminal helix to position the biologically active peptide for maximal burial of functional residues that need to be isolated to maintain the biologically active peptide in an inactive state, while burying hydrophobic residues and promoting solvent exposure / compensatory hydrogen bonding of polar residues. In various embodiments, the latch region may comprise a portion or all of a single α-helix in the cage polypeptide at the N-terminal or C-terminal portion. In various other embodiments, the latch region may comprise a portion or all of the first, second, third, fourth, fifth, sixth, or seventh α-helix in the cage polypeptide. In other embodiments, the latch region may comprise all or part of two or more different α-helices in the cage polypeptide; for example, the C-terminal portion of one α-helix and the N-terminal portion of the next α-helix, all two consecutive α-helices, etc.

[0078] In another embodiment, the polypeptide of the first aspect comprises one or more biologically active peptides within the latch region. In this embodiment, the biologically active peptide can replace one or more amino acids in the latch region, or can be added to the latch region without removing any amino acid residues from the latch region. In various non-limiting embodiments, the biologically active peptide can comprise one or more of the following of SEQ ID NOs: 60, 62-64, 66, 27052-27093 and 27118-27119 (all sequences in brackets are optional) or variants thereof:

[0079] GFP11 fluorescent peptide and GFP1-10 binding peptide: RDHMVLHEYVNAAGIT (SEQ ID NO: 27052)

[0080] BIM binding peptide and BCL-2 apoptotic peptide: IxxxLRxIGDxFxxxY (SEQ ID NO: 50), where x is any amino acid; in one embodiment, the peptide is EIWIAQELRRIGDEFNAYYA (SEQ ID NO: 60)

[0081] Designed peptide that binds to BCL-2: KMAQELIDKVRAASLQINGDAFYAILRAL (SEQ ID NO: 62)

[0082] StreptagII binding peptide or antibody to streptactin: (N) WSHPQFEK (SEQ ID NO: 63)

[0083] TEV protease cleavage site: ENLYFQ(G)-X (SEQ ID NO: 64), where (G) can also be S, and the last position, -X, can be anything except proline

[0084] Thrombin protease cleavage site: LVPRGS (SEQ ID NO: 66)

[0085] Cathepsin cleavage site: RLVGFE (SEQ ID NO: 27053)

[0086] Spycatcher's Spytag covalently cross-linked peptide: AHIVMVDAYK (PTK) (SEQ ID NO: 27054)

[0087] NLS peptide that targets proteins to the nucleus: AAAKRARTS (SEQ ID NO: 27055)

[0088] NES1 peptide that excludes proteins from the nucleus: LALKLAGLDIN (SEQ ID NO: 27056)

[0089] NES2 peptide that excludes proteins from the nucleus: ELAEKLAGLDIN (SEQ ID NO: 27057)

[0090] NES3 peptide that excludes proteins from the nucleus: ELAEKLRAGLDLN (SEQ ID NO: 27058)

[0091] EZH2 binding peptide that recruits DNA methylases: TMFSSNRQKILERTETLNQEWKQRRIQ (SEQ ID NO: 27059)

[0092] MDM2 binding peptide that recruits p53: ETFSDLWKLL (SEQ ID NO: 27060)

[0093] ·CP5 binding peptide: GELDELVYLLDGPGYDPIHSDVVTRGGSHLFNF (SEQ ID NO:27061)

[0094] 9aaTAD1 for transcriptional activation: TMDDVYNYLFDD (SEQ ID NO: 27062)

[0095] 9aaTAD2 for transcriptional activation: LLTGLFVQYLFDD (SEQ ID NO: 27063)

[0096] 9aaTAD3 for transcriptional activation: DDAVVESFFSS (SEQ ID NO: 27064)

[0097] 9aaTAD4 for transcriptional activation: GDFLSDLFD (SEQ ID NO: 27065)

[0098] 9aaTAD5 for transcriptional activation: GDVLSDLVD (SEQ ID NO: 27066)

[0099] Mad1-SID-epigenetic modification: NIQMLLEAADYLE (SEQ ID NO: 27067)

[0100] Mad1-SID (3A mutant) - epigenetic modification: NIAMLLAAAAYLE (SEQ ID NO: 27068)

[0101] RHIM domain 1 from ZBP1: IQIG (SEQ ID NO: 27069)

[0102] RHIM domain 2 from ZBP1: VQLG (SEQ ID NO: 27070)

[0103] nanoBit split luciferase: VSGWRLFKKIS (SEQ ID NO: 27071)

[0104] ·CC-A:GLEQEIAALEKENAALEWEIAALEQGG(SEQ ID NO:27072)

[0105] ·CC-B:GLKQKIAALKYKNAALKKKIAALKQGG(SEQ ID NO:27073)

[0106] ·GCN4:RMKQLEDKVEELLSKNYHLENEVARLKKLVGER(SEQ ID NO:27074)

[0107] ·CC-Di:GEIAALKQEIAALKKENAALKWEIAALKQG(SEQ ID NO:27075)

[0108] Membrane-disrupting / cell-penetrating peptides:

[0109] GALA for membrane disruption: WEAALAEALAEALAEHLAEALAEALEALAA (SEQ ID NO: 27076)

[0110] ·Aurein 1.2:GLFDIIKKIAESF(SEQ ID NO:27077)

[0111] ·Magainin-1:GIGKFLHSAGKFGKAFVGEIMKS(SEQ ID NO:27078)

[0112] ·Magainin-2:GIGKFLHSAKKFGKAFVGEIMNS(SEQ ID NO:27079)

[0113] Melittin: GIGAVLKVLTTGLPALISWIKRKRQQ (SEQ ID NO: 27080)

[0114] ·Mastoparan X:INWKGIAAMAKKLL(SEQ ID NO:27081)

[0115] Cecropin A: KWKLFKKIEKVGQNIRDGIIKAGPAVAVVGQATQIAK (SEQ ID NO: 27082)

[0116] Cecropin P1: SWLSKTAKKLENSAKKRISEGIAIAIQGGPR (SEQ ID NO: 27083)

[0117] ·Citropin 1.1: GLFDVIKKVASVIGGL (SEQ ID NO: 27084)

[0118] ·Temporin-1Lb: NFLGTLINLAKKIL (SEQ ID NO:27085)

[0119] ·HPV33 L2 peptide: SYFILRRRRKRFPYFFTDVRVAA (SEQ ID NO:27086)

[0120] Adenovirus pVI membrane fusion domain: AFSWGSLWSGIKNFGSTVKNY (SEQ ID NO: 27087)

[0121] Gamma-1 peptide from animal shed virus: ASMWERVKSIIKSSLAAASNI (SEQ ID NO: 27088)

[0122] Poliovirus 2B pore-forming peptide: VTSTITEKLLKNLIKIISSLVIITRNYEDTTTVLATLALLGCDASPWQWL (SEQ ID NO: 27089)

[0123] Rhinovirus pore-forming peptide: IAQNPVENYIDEVLNEVLVVPNIN (SEQ ID NO: 27090)

[0124] Influenza HA2 pore-forming peptide: FLGIAEAIDIGNGWEGMEFG (SEQ ID NO: 27091)

[0125] Influenza HA2 derivative: GLFGAIAGFIENGWEGMIDG (SEQ ID NO: 27092)

[0126] HA-derived INF6: GLFGAIAGFIENGWEGMIDGWYG (SEQ ID NO: 27093)

[0127] KRAB domain - epigenetic modification: (SEQ ID NO: 27118) MDAKSLTAWSRTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLTKPDVILRLEKGEEPWLV

[0128] Minimal Krab domain (KOX1 11-55) - epigenetic modification: (SEQ ID NO: 27119) RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGY

[0129] In one embodiment, the dynamic range of activation by the key polypeptide can be adjusted by shortening the latch region to be shorter than the α-helix in the structural region, thereby weakening the cage polypeptide-latch region interaction and opening an exposed region on the cage polypeptide to which the key polypeptide can bind to act as a "bottom" ( Figure 2 Similarly, the dynamic range of activation by the key polypeptide can be adjusted in a similar manner by designing mutations in the latch that weaken the cage polypeptide-latch region interaction ( Figure 1-2 and 10). In other embodiments, the latch region can be one or more helices with a total length of between 18-150 amino acids, between 18-100 amino acids, between 18-58 amino acids, or any range encompassed by these ranges. In other embodiments, the latch region can be composed of a helical secondary structure, a beta strand secondary structure, a loop secondary structure, or a combination thereof.

[0130] In a second aspect, the present disclosure provides non-naturally occurring polypeptides comprising sequences along their length that are identical to the cage polypeptides disclosed herein (such as SEQ ID NOs: 1-49, 51-52, 54-59, 61, 65, 67-91, 92-2033 (submitted as Appendix 1 in U.S. Provisional Application Serial No. 62 / 700681, filed on July 19, 2018, and / or U.S. Provisional Application Serial No. 62 / 785537, filed on December 27, 2018), SEQ ID NOs: Nos. 2034-14317 (submitted as Appendix 2 to U.S. Provisional Application Serial No. 62 / 700681, filed July 19, 2018, and / or U.S. Provisional Application Serial No. 62 / 785537, filed December 27, 2018), 27094-27117, 27120-27125, 27278 to 27321, and Table 2 (with even-numbered SEQ ID NOs. 27126 and 27276).

[0015] In some embodiments, the present invention provides a polypeptide having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an amino acid sequence of a polypeptide (e.g., a polypeptide of the invention) or a cage polypeptide listed in Table 3 and / or Table 4, excluding optional amino acid residues, and optionally excluding amino acid residues in the latch region. In each embodiment, the N-terminal and / or C-terminal 60 amino acids of each cage polypeptide can be optional, as the terminal 60 amino acid residues can comprise a latch region that can be modified, such as by replacing all or part of the latch with a biologically active peptide. In one embodiment, the N-terminal 60 amino acid residues are optional; in another embodiment, the C-terminal 60 amino acid residues are optional; and in another embodiment, each of the N-terminal 60 amino acid residues and the C-terminal 60 amino acid residues is optional. In one embodiment, these optional N-terminal and / or C-terminal 60 residues are not included in determining the percent sequence identity.In another embodiment, optional residues may be included in determining the percent sequence identity.

[0131] As disclosed herein, the bioactive peptide to be isolated by the polypeptide of the present disclosure is located within the latch region. The latch region is represented by brackets in the sequence of each cage polypeptide. The bioactive peptide can be added to the latch region without removing any residues in the latch region, or one or more (1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) amino acid residues in the cage scaffold latch region can be replaced to produce the final polypeptide. Therefore, when comprising a bioactive peptide, the latch region can be significantly modified. In one embodiment, the optional residues are not included in the determination of the percentage of sequence identity. In another embodiment, the latch region residues can be included in the determination of the percentage of sequence identity. In another embodiment, each of the optional residues and the latch residues can be excluded in the determination of the percentage of sequence identity.

[0132] In one embodiment of this second aspect, the polypeptide is a polypeptide according to any embodiment or combination of embodiments of the first aspect, and further has the desired 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length to the amino acid sequence of a reference cage polypeptide disclosed herein. In another embodiment, the polypeptide further comprises (or replaces) a biologically active peptide within the latch region of the cage polypeptide.

[0133] The cage polypeptide can be a cage scaffold polypeptide (i.e., without a biologically active peptide), e.g., see SEQ ID NOs: 1-17, 2034-14317, and certain cage polypeptides listed in Table 2, Table 3, and / or Table 4, or can further include a sequestering biologically active peptide in the latch region of the cage scaffold polypeptide (present as a fusion to the cage scaffold polypeptide), as described in more detail herein (e.g., see SEQ ID NOs: 18-49, 51-52, 54-59, 61, 65, 67-2033, 27094-27117, 27120-27125, and certain cage polypeptides listed in Tables 2, 3, and / or 4). In a specific embodiment, the cage polypeptide shares 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length with the amino acid sequence of a cage polypeptide in Table 2, Table 3 and / or Table 4.

[0134] In another specific embodiment, the cage polypeptide shares 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length with the amino acid sequence of a cage polypeptide in Table 3. In another specific embodiment, the cage polypeptide shares 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length with the amino acid sequence of a cage polypeptide in Table 4. In one embodiment of each of these embodiments, the optional N-terminal and / or C-terminal 60 residues are not included in determining the percent sequence identity. In another embodiment, optional residues can be included in determining the percent sequence identity.

[0135] As disclosed in the Examples below, exemplary cage and key polypeptides of the present disclosure have been identified and subjected to mutational analysis. Furthermore, different designs starting from the same exemplary cage and key polypeptides produce different amino acid sequences while maintaining the same intended function. In various embodiments, a given amino acid can be replaced by a residue with similar physiochemical characteristics, for example, by substituting one aliphatic residue for another (such as substitution of Ile, Val, Leu, or Ala with one another), or by substituting one polar residue for another (such as between Lys and Arg; between Glu and Asp; or between Gln and Asn). Other such conservative substitutions (e.g., substitutions of entire regions with similar hydrophobicity characteristics) are known. Polypeptides comprising conservative amino acid substitutions can be tested in any of the assays described herein to confirm retention of the desired activity. Amino acids can be grouped according to the similarity of their side chain properties (ALLehninger, Biochemistry, 2nd ed., pp. 73-75, Worth Publishers, New York (1975)): (1) Nonpolar: Ala (A), Val (V), Leu (L), Ile (I), Pro (P), Phe (F), Trp (W), Met (M); (2) Uncharged polar: Gly (G), Ser (S), Thr (T), Cys (C), Tyr (Y), Asn (N), Gln (Q); (3) Acidic: Asp (D), Glu (E); (4) Basic: Lys (K), Arg (R), His (H). Alternatively, naturally occurring residues can be grouped based on common side chain properties: (1) hydrophobic: norleucine, Met, Ala, Val, Leu, Ile; (2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues affecting chain orientation: Gly, Pro; (6) aromatic: Trp, Tyr, Phe. Non-conservative substitutions would require exchanging a member of one of these classes for another. Specific conservative substitutions include, for example: Ala to Gly or to Ser; Arg to Lys; Asn to Gln or to His; Asp to Glu; Cys to Ser; Gln to Asn; Glu to Asp; Gly to Ala or to Pro; His to Asn or to Gln; Ile to Leu or to Val; Leu to Ile or to Val; Lys to Arg, to Gln, or to Glu; Met to Leu, to Tyr, or to Ile; Phe to Met, to Leu, or to Tyr; Ser to Thr; Thr to Ser; Trp to Tyr; Tyr to Trp; and / or Phe to Val, to Ile, or to Leu.

[0136] Exemplary cage polypeptides (see also SEQ ID NOs: 92-14317, 27094-27117, 27120-27125, 27728-27321, and the cage polypeptides listed in Table 2, Table 3, and / or Table 4):

[0137] 1) Exemplary reference cage polypeptide; latch region indicated by brackets [ ]

[0138] 6His-MBP-TEV, 6His-TEV and flexible linker sequences are underlined text

[0139] The functional domains of the fusion (DARPin, split intein components and fluorescent protein) are in bold text

[0140] Functional peptides are italicized and underlined text

[0141] Exemplary positions that have been mutated to any amino acid to adjust responsiveness are underlined in bold text. These positions are exemplary and not an exhaustive list of residues that can adjust responsiveness.

[0142] The brackets contain C-terminal sequences that can be removed to adjust responsiveness. Starting from the C-terminus and removing consecutive residues therein, you can remove a range of one (1) to all residues in the brackets.

[0143] All sequences in brackets are optional

[0144] >SB76L (SEQ ID NO: 1)

[0145]

[0146] >SB76L_17 (SEQ ID NO: 2)

[0147] ( MGSSHHHHHHSSGLVPRGSHM )GSKEAVTKLMALNLKLAEKLLEAIARLQELNIALVYLATELTDPERIREEIRKVKEESARIVEEAEEEIRRAAARSEDILREGSGSGSDAVAELQRLNLELAELLLRAAAKLQELNIDLVRLLTEL TDPKTIRDAIERVKAESERIVREAERLIREAKADSERILREGSGSGDPDVARLQELFIELARELLEALARLQELNIDLVRLASELTDP[DTIRDAIRRVKEESARIVEDARRLIKEAAEEAEKISRE]

[0148] >SB76L_18(SEQ ID NO:3)

[0149] ( MGSSHHHHHHSSGLVPRGSHM )GSKRAVTELQKLNIELARKLLRALAELMELNIALVYLAVELTDPRRIREEIRKVKEKSDEIVKRAEDEIRKAAASEKILREGSGSGSDAVALQRLNLELAKLLLEAIAKLQALNIDLVRLLTELDTPETIRRAIKRVKDESARIVEEAEKLIRAAKDKAREIIDKGSGSGDPDVARLQELNIELARELLEAAARLQELFIDLVRLASELTDP[DEARKAIERVKREAERIVREAERLIREAKRASKEISDE]

[0150] >LOCKR_extend5(SEQ ID NO:4)

[0151] ( MGSSHHHHHHSSGLVPRGSHM ()

[0152] >LOCKR_extend9(SEQ ID NO:5)

[0153] ( MGSSHHHHHHSSGLVPRGSHM() ATIREAIRKVKEDSERIVAEAERLIAAAKAESIIREAERLIAAAAGSGSGSIELARELLRDVARLQELNIELARELLRAAAELQELNIKLVELASELTDP[DEARKAIARVKRESKRIVEDAERLIREAAAASEKISREAERLIREAA]

[0154] >LOCKR_extend18(SEQ ID NO:6)

[0155] ( MGSSHHHHHHSSGLVPRGSHM )SKEAVTKLQALNIKLAEKLLEAVTKLQALNIKLAEKLLEALARLQELNIALVYLAVELTDPKRIADEIKKVKDKSKEIVERAEEIARAAAESKKILDEAEEEIARAAAESKKILDEGSGSGSGSDAVAELQALNLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTKLTDPATIREAIRKVKEDSERIVAEAERLIAAKAESERIIREEARLIAAKAESERIIREGSGSGDPDVARLQELNIELARELLRDVARLQELNIELARELLRAAAELQELNIKLVELASELTDP[DEARKAIARVKRESKRIVEDAERLIREAAAASEKISREAERLIREAAAASEKISRE]

[0156] >LOCKRb(SEQ ID NO:7)

[0157] ( MGSSHHHHHHSSGLVPRGSHM)SHAAVIKLSDLNIRLLDKLLQAVIKLTELNAELNRKLIEALQRLFDLNVALVHLAAELTDPKRIADEIKKVKDKSKEIVERAEEEEIARAAAESKKILDEAEEEIARAAAESKKILDEGSGSGSDAVAELQALNLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTTKLTD PATIREAIRKVKEDSERIVAEAERLIAAAKAESERIIREAERLIAAAAKAESERIIREGSGSNDPQVAQNQETFIELARDALRLVAENQEAFIEVARLTLRAAALAQEVAIKAVEAASEGGSGSG[NKEEIEKLAKEAREKLKKAEKEHKEIHDKLRKKKKAREDLKKADELRETNKRVN]

[0158] >LOCKRc(SEQ ID NO:8)

[0159] ( MGSSHHHHHHSSGLVPRGSHM )SLEAVLKLAELNLKLSDKLAEAVQKLAALLNKLLEKLSEALQRLFELNVALVTLAIELTDPKRIADEIKKVKDKSKEIVERAEEIARAAAESKKILDEAEEEIARAAAESKKILDEGSGSGSDAVALQALNLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTKLTDPATIREAIRKVKEDSERIVAEAERLIAAKAESERIIREAERLIAAKAESERIIREGSGSNDPLVARLQELLIEHARELLRLVATSQEIFIELARAFLANAAAQLQEAAIKAVEAASENGSGSG[SSEKVRRELKESLKENHKQNQKLLKDHKRAQEKLNRELEELKKKHKKTLDDIRRES]

[0160] >LOCKRd(SEQ ID NO:9)

[0161] ( MGSSHHHHHHSSGLVPRGSHM)SLEAVLKLFELNHKLSEKLLEAVLKLHALNQKLSQKLLEALARLLELNVALVELAIELTDPKRIADEIKKVKDKSKEIVERAEEIARAAAESKKILDEAEEEIARAAAESKKILDEGSGSGSDAVAELQALNLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTKLTDPATIREAIRKVKEDSERIVAEAERLIAAKAESERIIREREAERLIAAKAESERIIREGSGSGDPEVARLQEAFIEQAREILRNVAAAQEALIEQARRLALAALAQEAAIKAVELASEHGSGSG[DTVKRILEELRRRFEKLAKDLDDIARKLLEDHKKHNKELKDKQRKIKKEADDAARS]

[0162] >LOCKRe(SEQ ID NO:10)

[0163] ( MGSSHHHHHHSSGLVPRGSHM )SLEAVLKLQDLNSKLSEKLSEAQLKLQALNNKLLRKLLEALLRLQDLNQALVNLALQLTDPKRIADEIKKVKDKSKEIVERAEEEEIARAAAESKKILDEAEEIARAAAESKKILDEGSGSGSDAVAELQALNLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTKLTD PATIREAIRKVKEDSERIVAEAERLIAAAKASERIIREAERLIAAKASERIIREGSGSGDPDVAKSQEHLIEHARELLRQVAKSQELFIELARQLLLRLAAKSQELAIKAVELASEAGSGSG[DDVERRLRKANKESKKEAEELTEEAKKANEKTKEDSKELTKENRKTNKTIKDEARS]

[0164] >LOCKRf(SEQ ID NO:11)

[0165] ( MGSSHHHHHHSSGLVPRGSHM)SREAVEKLAELNHKLSHKLQQAQQKLQALNLKLLQKLLEALDRLQDLNNALVKLAQRLTDPKRIADEIKKVKDKSKEIVERAEEEIARAAAESKKILDEAEEEIARAAAESKKILDEGSGSGSDAVAELQALNLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTKLTD PATIREAIRKVKEDSERIVAEAERLIAAAKAESERIIREAERLIAAAAKAESERIIREGSGSGDPDVARQQETLIEQARRLLRNVAESQELFIEAARTVLRLAAKLQEINIKQVELASEAGSGSG[DDEERRSEKTVQDAKREIKKVEDDLQRLNEEQKKKVKKQEDENQKTLKKHKDDARS]

[0166] >miniLOCKRa_1(SEQ ID NO:12)

[0167] ( MGSSHHHHHHSSGLVPRGSHM )NKEDATEAQKKAIRAAEELLKDVTRIQERAIREAEKALERLARVQEEAIRRVYEAVESKNKEELKKVKEEIEELLRRLKRELDELEREREIRELLKEIKEKADRLEKEIRDLIERIRRDRNASDEVVTRLARLNEELIRELREDVRRLAELNKELLRELERAARELARLNEKLLELADRVETE[EEARKAIARVKRESKRIVEDAERLIREAAAASEKISRE]

[0168] >miniLOCKRa_2(SEQ ID NO:13)

[0169] ( MGSSHHHHHHSSGLVPRGSHM)DERLKRLNERLADELDKDLERLLRLNEELARELTRAAEELRELNEKLVELAKKLQGGRSREVAERAEKEREKIRRKLEEIKKEIKEDADRIKKRADELRRRLEKTLEDAARALEKLKREPRTEELKRKATELQKEAIRRAEELLKEVTDVQRRAIERAEELLEKLARLQEEAIRTVYLLVELNKV[DRAKAIARVKRESKRIVEDAERLIREAAAASEKISRE]

[0170] >miniLOCKRc_1(SEQ ID NO:14)

[0171] ( MGSSHHHHHHSSGLVPRGSHM )LIERLTRLEKEHVRELKRLLDTSLEILRRLVEAFETNLRQLKEALKRALEAANLHNEEVEEVLRKLEEDLRRLEEELRKTLDDVRKEVKRLKEELDKRIKEVEDELRKIKEKLKKGDKNEKRVLEEILRLAEDVLKKSDKLAKDVQERARELNEILEELSRKLQELFERVVEEVTRNVPT[TERIEKVRRELKESLKENHKQNQKLLKDHKRAQEKNLRELEELKKKHTLDDIRRES]

[0172] >miniLOCKRc_2(SEQ ID NO:15)

[0173] ( MGSSHHHHHHSSGLVPRGSHM )SEERVLELAEEALRLSDEAAKEIQELARRLNEELEKLSKELQDLFERIVETVTRLIDADEETLKRAAEIKKRLEDARKKAKEAADKAREELDRARKKLKELVDEIRKKAKDALEKAGADEELVARLLRLLEEHARELERLLRTSARIIERLLDAFRRNLEQLKEAADKAVEAAEEAVRRVED[VRVWSEKVRRELKESLKENHKQNQKLLKDHKRAQEKLNRELEKKKHKKTLDIRRES]

[0174] >1fix-short-noBim-t0 (SEQ ID NO:16)

[0175] ( MGSHHHHHHGSGSENLYFQGSGG )SELARKLLEASTKLQRLNIRLAEALLEAIARLQELNLELVYLAVELTDPKRIRDEIKEVKDKSKEIIRRAEKEIDDAAKESEKILEEAREAISGSGSELAKLLLKAIAETQDLNLRAAKAFLEAAAKLQELNIRAVELLVKLTDPATIREALEHAKRRSKEIIDEAERAIRAAKRESERIIEEARRLIEKGSGSGS[ELARELLRAHAQLQRLNLELLRELLRALAQLQELNLDLLRLASELTDPDEARKAIARVKRESKRIVEDAERLIREAAAASEKISREAERLIR]

[0176] >1fix - short - noBim(AYYA)-t0(SEQ ID NO:17)

[0177] ( MGSHHHHHHGSGSENLYFQGSGG )SELARKLLEASTKLQRLNIRLAEALLEAIARLQELNLELVYLAVELTDPKRIRDEIKEVKDKSKEIIRRAEKEIDDAAKESEKILEEAREAISGSGSELAKLLLKAIAETQDLNLRAAKAFLEAAAKLQELNIRAVELLVKLTDPATIREALEHAKRRSKEIIDEAERAIRAAKRESERIIEEARRLIEKGSGSGS[ELARELLRAHAQLQRLNLELLRELLRALAQLQELNLDLLRLASELTDPDEARKAIARVKRESNAYYADAERLIREAAAASEKISREAERLIR]

[0178] “(3) Functional LOCKR cage design with bioactive peptides encoded into the latch”

[0179] >aBcl2LOCKR(SEQ ID NO:18)

[0180]

[0181] >pBimLOCKR(SEQ ID NO:19)

[0182]

[0183] >BimLOCKR_extend5(SEQ ID NO:20)

[0184]

[0185]

[0186] >BimLOCKR_extend9(SEQ ID NO:21)

[0187] ( MGSSHHHHHHSSGLVPRGSHM )KLAEKLLEAVTKLQALNIKLAEKLLEALARLQELNIALVYLAVELTDPKRIADEIKKVKDKSKEIVERAEEEIARAAAESKKILDEAEEEIARAGSGSGSLKLAELLLEAVAELQALNLKLAELLLEAIAKLQELNIKLVELLTKLTDPATIREAIRKVKEDSERIVAEAERLIAAAKAESERIIREAERLIAAAAGSGSGSIELARELLRDVARLQELNIELARELLRAAAELQELNIKLVELASELTD[ EIWIAQELRRIGDEFNAYYA DAERLIREAAAASEKISREAERLIREAA]

[0188] >BimLOCKR_extend18(SEQ ID NO:22)

[0189]

[0190] >BimLOCKRb(SEQ ID NO:23)

[0191]

[0192] >BimLOCKRc(SEQ ID NO:24)

[0193]

[0194] >BimLOCKRd(SEQ ID NO:25)

[0195]

[0196]

[0197] >StrepLOCKRa_300(SEQ ID NO:26)

[0198]

[0199] >strepLOCKRa_306(SEQ ID NO:27)

[0200]

[0201] >strepLOCKRa_309(SEQ ID NO:28)

[0202]

[0203] >strepLOCKRa_312(SEQ ID NO:29)

[0204]

[0205]

[0206] >strepLOCKRa_313(SEQ ID NO:30)

[0207]

[0208] >strepLOCKRa_317(SEQ ID NO:31)

[0209]

[0210] >strepLOCKRa_320(SEQ ID NO:32)

[0211]

[0212] >strepLOCKRa_323(SEQ ID NO:33)

[0213]

[0214] >strepLOCKRa_329(SEQ ID NO:34)

[0215]

[0216] >SB13_LOCKR(SEQ ID NO:35)

[0217]

[0218] >ZCX12_LOCKR(SEQ ID NO:36)

[0219]

[0220] >SB13_LOCKR_extend18(SEQ ID NO:37)

[0221]

[0222] >ZCX12_LOCKR_extend18(SEQ ID NO:38)

[0223]

[0224] >fretLOCKRa(SEQ ID NO:39)

[0225]

[0226] >fretLOCKRb(SEQ ID NO:40)

[0227]

[0228] >fretLOCKRc(SEQ ID NO:41)

[0229]

[0230]

[0231] >fretLOCKRd(SEQ ID NO:42)

[0232]

[0233] >tevLOCKR(SEQ ID NO:43)

[0234]

[0235] >spyLOCKR(SEQ ID NO:44)

[0236]

[0237]

[0238] >1_nesLOCKR(SEQ ID NO:45)

[0239]

[0240] >2_nesLOCKR(SEQ ID NO:46)

[0241]

[0242] >3_nesLOCKR (SEQ ID NO:47)

[0243]

[0244] >nlsLOCKR (SEQ ID NO:48)

[0245]

[0246] >ezh2LOCKR (SEQ ID NO:49)

[0247]

[0248] >1fix_VMAc_C_BIMlatcht9(SEQ ID NO:51)

[0249]

[0250] >sfGFP_VMAn_1fix_BIM_t0_latch(SEQ ID NO:52)

[0251]

[0252] Asymmetric functional cage encoding Bim and GFP11 (i.e., bioactive peptide)

[0253] (6His-MBP-TEV, 6His-TEV and flexible linker sequences are underlined text)

[0254] (Colocalization domains are in bold text)

[0255] (Functional peptides are italicized and underlined text)

[0256] (Positions that can be mutated to any amino acid to adjust responsiveness are in underlined bold text)

[0257] (C-terminal sequences that can be removed to adjust responsiveness are in italics)

[0258] (All sequences in brackets are optional)

[0259] >1fix-long-BIM-t0(SEQ ID NO:54)

[0260]

[0261] >1fix-long-GFP-t0(SEQ ID NO:55)

[0262]

[0263] >1fix-short-BIM-t0(SEQ ID NO:56)

[0264]

[0265] >1fix-short-GFP-t0(SEQ ID NO:57)

[0266]

[0267] >Spycatcher-1fix-long-GFP-t0(SEQ ID NO:58)

[0268]

[0269] >Spycatcher-1fix-short-GFP-t0(SEQ ID NO:59)

[0270]

[0271] >1fix-latch_Mad1SID_t0_1(SEQ ID NO:61)

[0272]

[0273] >1fix-latch_Mad1SID_T0_2(SEQ ID NO:65)

[0274]

[0275] >1fix-short-Bim-t0-relooped(SEQ ID NO:67)

[0276]

[0277]

[0278] >1fix-short-spytag-t0_2(SEQ ID NO:68)

[0279]

[0280] >1fix-short-spytag-t0_8(SEQ ID NO:69)

[0281]

[0282] >1fix-short-TEV-t0_1(SEQ ID NO:70)

[0283]

[0284] >1fix-short-TEV-t0_6(SEQ ID NO:71)

[0285]

[0286] >1fix-short-nanoBit-t0_1(SEQ ID NO:72)

[0287]

[0288]

[0289] >1fix-short-nanoBit-t0_3(SEQ ID NO:73)

[0290]

[0291] >1fix-short-RHIM-t0_8(SEQ ID NO:74)

[0292]

[0293] >1fix-short-RHIM-t0_19(SEQ ID NO:75)

[0294]

[0295] >1fix-short-RHIM-t0_22(SEQ ID NO:76)

[0296]

[0297] >1fix-short-gcn4-t0_4(SEQ ID NO:77)

[0298] ( MGSSHHHHHHSSGLVPRGSHM)SELARKLLEASTKLQRLNIRLAEALLEAIARLQELNLELVYLAVELTDPKRIRDEIKEVKDKSKEIIRRAEKEIDDAAKESEKILEEAREAISGSGSELAKLLLKAIAETQDLNLRAAKAFLEAA AKLQELNIRAVELLVKLTDPATIREALEHAKRRSKEIIDEAERAIRAAKRESERIIEARRLIEKGSGSGSELARELLRAHAQLQRLNLELLRELLRALAQLQELNLDLLRLASELTDP[DESVKE( LEDKVEELLSKNYHLENEVARLKKLVGER) SREAERLIR]

[0299] >1fix-short-ccDi-t0_6(SEQ ID NO:78)

[0300]

[0301] >1fix-short-cc-a-t0_6(SEQ ID NO:79);

[0302]

[0303] >1fix-short-cc-b-t0_6(SEQ ID NO:80)

[0304]

[0305] STREPII-LOCKR FIGURE:

[0306] >STEPII-2plus1_LOCK_1(SEQ ID NO:81)

[0307]

[0308] >STEPII-2plus1_LOCK_2(SEQ ID NO:82)

[0309]

[0310] >STEPII-2plus1_LOCK_3(SEQ ID NO:83)

[0311]

[0312] >STEPII-2plus1_LOCK_4C(SEQ ID NO:84)

[0313]

[0314] >STREPII-2plus1_LOCK_4N(SEQ ID NO:85)

[0315]

[0316] >STREPII-3plus1_LOCK_1(SEQ ID NO:86)

[0317]

[0318] >STREPII-3plus1_LOCK_2(SEQ ID NO:87)

[0319]

[0320] >STREPII-3plus1_LOCK_3(SEQ ID NO:88)

[0321]

[0322] >STREPII-3plus1_LOCK_4(SEQ ID NO:89)

[0323]

[0324] >STREPII-3plus1_LOCK_3-relooped(SEQ ID NO:90)

[0325]

[0326] >STREPII-2plus1_LOCK_3-relooped(SEQ ID NO:91)

[0327]

[0328] >BimLOCKR_a_short_Nterm(SEQ ID NO:27094)

[0329]

[0330] >BimLOCKR_g(SEQ ID NO:27095)

[0331]

[0332] >reloop_strepLOCKRh(SEQ ID NO:27096)

[0333]

[0334] >reloop_strepLOCKRi(SEQ ID NO:27097)

[0335]

[0336] >spyLOCKRa_2(SEQ ID NO:27098)

[0337]

[0338] >spyLOCKRa_8(SEQ ID NO:27099)

[0339]

[0340] >tevLOCKRa_1(SEQ ID NO:27100)

[0341]

[0342] >tevLOCKRa_6(SEQ ID NO:27101)

[0343]

[0344] >lucLOCKRa_1(SEQ ID NO:27102)

[0345]

[0346] >lucLOCKRa_3(SEQ ID NO:27103)

[0347]

[0348] >rhimLOCKRa_8(SEQ ID NO:27104)

[0349]

[0350] >rhimLOCKRa_19(SEQ ID NO:27105)

[0351]

[0352]

[0353] >rhimLOCKRa_22(SEQ ID NO:27106)

[0354]

[0355] >gcn4LOCKRa_4(SEQ ID NO:27107)

[0356]

[0357] >cc-DiLOCKRa_6(SEQ ID NO:27108)

[0358]

[0359] >cc-aLOCKRa_6(SEQ ID NO:27109)

[0360]

[0361] >cc-bLOCKRa_6(SEQ ID NO:27110)

[0362]

[0363] >tev-spyLOCKRa_short_40(SEQ ID NO:27111)

[0364]

[0365]

[0366] >tev-spyLOCKRa_short_57(SEQ ID NO:27112)

[0367]

[0368] >tev-spyLOCKRa_short_63(SEQ ID NO:27113)

[0369]

[0370] >tev-spyLOCKRa_29(SEQ ID NO:27114)

[0371]

[0372] >tev-spyLOCKRa_32(SEQ ID NO:27115)

[0373]

[0374] >Bim-fretLOCKRa_short(SEQ ID NO:27116)

[0375]

[0376]

[0377] >fretLOCKRa_short(SEQ ID NO:27117)

[0378]

[0379] E18_KRAB_full(SEQ ID NO:27120)

[0380]

[0381] E18_KRAB_N13t(SEQ ID NO:27121)

[0382]

[0383] E18_KRAB_C9t(SEQ ID NO:27122)

[0384]

[0385] E18_KRAB_Cterm1(SEQ ID NO:27123)

[0386]

[0387] E18_KRAB_Cterm2(SEQ ID NO:27124)

[0388]

[0389] E18_KRAB_Cterm3(SEQ ID NO:27125)

[0390]

[0391] >3plus1_cage_Nterm_GFP11_668(SEQ ID NO:27278)

[0392] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRDHMVLHEYVNAAGITLEELRRGSLDAKELLKTLEDLLREVLEVARRVVETLKELNRRVLEVVREDIEANERLLRRVLDTLRRGGVDERRIKDLERLIRESLKKAEEVLREAAEKSREIVDEIREVLKRADEALKRIIKKIRETRGADALSRLLEELLRVVDDLIRKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0393] >3plus1_cage_Cterm_GFP11_668(SEQ ID NO:27279)

[0394] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRLEKELKELRDLKRRIEEKLEELRRGSLDAKELLLKTLEDLREVLEVARRVVETLKELNRRVLEVVREDIERNERLLRRVLDTLRRGGVDERRIKDLERLIRESLKKAEEVLREAAEKSREIVDEIREVLKRADEALKRIIKKIRETRGADADHMVLHEYVNAAGITIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0395] >3plus1_cage_Cterm_GFP11_668(SEQ ID NO:27280)

[0396] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRLEKELKELRDLKRRIEEKLEELRRGSLDAKELLLKTLEDLREVLEVARRVVETLKELNRRVLEVVREDIERNERLLRRVLDTLRRGGVDERRIKDLERLIRESLKKAEEVLREAAEKSREIVDEIREVLKRADEALKRIIKKIRETRGADARDHMVLHEYVNAAGITRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0397] >3plus1_cage_Cterm_GFP11_668(SEQ ID NO:27281)

[0398] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRLEKELKELRDLKRRIEEKLEELRRGSLDAKELLLKTLEDLREVLEVARRVVETLKELNRRVLEVVREDIERNERLLRRVLDTLRRGGVDERRIKDLERLIRESLKKAEEVLREAAEKSREIVDEIREVLKRADEALKRIIKKIRETRGADALSRDHMVLHEYVNAAGITLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0399] >3plus1_cage_Cterm_GFP11_668(SEQ ID NO:27282)

[0400] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRLEKELKELRDLKRRIEEKLEELRRGSLDAKELLLKTLEDLREVLEVARRVVETLKELNRRVLEVVREDIERNERLLRRVLDTLRRGGVDERRIKDLERLIRESLKKAEEVLREAAEKSREIVDEIREVLKRADEALKRIIKKIRDHMVLHEYVNAAGITLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0401] >3plus1_cage_Nterm_GFP11_669(SEQ ID NO:27283)

[0402] SEKKLLKESEEEVRRLRRTLEELLRKYREVLERLRDHMVLHEYVNAAGITRLKEVLDRSGLDIDTIIKEVEDLLKTVLDRLRELLDKIARLTKEAIEVVREIIERIVRHAERVKDELRKGGADKRKLDRVDRLIKENTRHLKEILDRIEDLVRRRSEKKLDIIREVRRLIEELRKKAEEIKKDPDERLVKTLIEDVERVIKRILELITRVAEDNERVLERIIRELTDNLERHLKIVREIVK

[0403] >3plus1_cage_Nterm_GFP11_670(SEQ ID NO:27284)

[0404] SEKEDLARKLRKLVEELTREYEELVKKLERLIEEIERDHMVLHEYVNAAGITISEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0405] >3plus1_cage_Cterm_GFP11_670(SEQ ID NO:27285)

[0406] SEKEDLARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLREISEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEARDHMVLHEYVNAAGITRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0407] >3plus1_cage_Cterm_GFP11_670(SEQ ID NO:27286)

[0408] SEKEDLARKLRKLVEELTREEEELVKKLERLIEEIEKVSEESVRKLEKLLREISEVRKLGTDERVLKRLERLRRIIEEDHELNTELLKRLLDLLLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDERDHMVLHEYVNAAGITIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0409] >3plus1_cage_Cterm_GFP11_670(SEQ ID NO:27287)

[0410] SEKEDLARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLREISEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGRDHMVLHEYVNAAGITRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0411] >3plus1_cage_Nterm_GFP11_670(SEQ ID NO:27288)

[0412] SEKEDLARKLRKLVEELTREYEELVKKLERLIEIRDHMVLHEYVNAAGITEISEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0413] >3plus1_cage_Cterm_GFP11_670(SEQ ID NO:27289)

[0414] SEKEDLARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLREISEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTRDHMVLHEYVNAAGITRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0415] >3plus1_cage_Nterm_GFP11_670(SEQ ID NO:27290)

[0416] SEKEDLARKLRKLVEELTREYEELVKKLERLIERDHMVLHEYVNAAGITLREISEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0417] >3plus1_cage_Nterm_GFP11_670(SEQ ID NO:27291)

[0418] SEKEDLARKLRKLVEELTREYEELVKKLERLIEIEKRDHMVLHEYVNAAGITSEEVRKLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0419] >3plus1_cage_Nterm_GFP11_670(SEQ ID NO:27292)

[0420] SEKEDLARKLRKLVEELTREEEELVKKLERLIEEIEKVSEESRDHMVLHEYVNAAGITKLGTDERVLKRLERRLRRIIEEDHELNTELLKRLLDLLLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEALRKLVELLVEVLRLRIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0421] >3plus1_cage_Nterm_GFP11_670(SEQ ID NO:27293)

[0422] SEKEDLARKLRKLVEELTREEEELVKKLERLIEEIEKVSEESVRDHMVLHEYVNAAGITLGTDERVLKRLLERLRRIIEEDHELNTELLKRLLDLLLKEILDTSRELLKRLLDILRKGVRDEEVLRDLERTLREVLEENERAIEEAERVLRKVLEDSERAVRDARRVLAEVDKSPTGDEALRKLVELLVEVLRLRIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0423] >3plus1_cage_Cterm_GFP11_671(SEQ ID NO:27294)

[0424] SEEEDLLERVKRVLDELIEIVDRNHELNRVVETSAALVERLLEEVERALETLEREIPGSSLLDKAIKDLRDVLRRVKEKVKRSIEELKEVLEESRRVLEEVVRKLREVIDRVRRLVEKGVDLRDLIRELKRVLEEAVKLIERLVRLNTRAAEKDNESLRELVRAIKEALKRAVDAVRKGGLDSRAVKKLDRDHMVLHEYVNAAGITNEELWRALVELNKESVRRLREIVERVARDLEETAR

[0425] >3plus1_cage_Cterm_GFP11_671(SEQ ID NO:27295)

[0426] SEEDLLERKRVLDELIEIVDRNHELNRRVVETSAALVERLLEEVERALETLEREIPGSSLLDKAIKDLRDVLRRVKEKVKRSIEELKEVLEESRRVLEEVVRKLREVIDRVRRLVEKGVDLRDLIRELKRVLEEAVKLIERLVRLNTRAAEKDNESLRELVRAIKEALKRAVDAVRKGGDRARDHMVLHEYVNAAGITDVVRRNEELWRALVELNKESVRRLREIVERVARDLEETAR

[0427] >3plus1_cage_Cterm_GFP11_671(SEQ ID NO:27296)

[0428] SEEDLLERKRVLDELIEIVDRNHELNRRVVETSAALVERLLEEVERALETLEREIPGSSLLDKAIKDLRDVLRRVKEKVKRSIEELKEVLEESRRVLEEVVRKLREVIDRVRRLVEKGVDLRDLIRELKRVLEEAVKLIERLVRLNTRAAEKDNESLRELVRAIKEALKRAVDAVRKGDGLDSRDHMVLHEYVNAAGITEDVVRRNEELWRALVELNKESVRRLREIVERVARDLEETAR

[0429] >3plus1_cage_Cterm_GFP11_671(SEQ ID NO:27297)

[0430] SEEDLLERKRVLDELIEIVDRNHELNRRVVETSAALVERLLEEVERALETLEREIPGSSLLDKAIKDLRDVLRRVKEKVKRSIEELKEVLEESRRVLEEVVRKLREVIDRVRRLVEKGVDLRDLIRELKRVLEEAVKLIERLVRLNTRAAEKDNESLRELVRAIKEALKRAVDAVRKGLRDHMVLHEYVNAAGITKKLEDVVRRNEELWRALVELNKESVRRLREIVERVARDLEETAR

[0431] >3plus1_cage_Cterm_GFP11_672(SEQ ID NO:27298)

[0432] DETEEVIKKIAKLIDELIETLKKIVEDSAKLIKELIDLLEEIIKRSADLIRELEKVSKRGDRDRLKKLIDRMRRETERALEELRRVIEELKEATERALKTIEDLIREIEEAIRRELGDEDLAREILEVIRRILELLREDMELNLRALEEEIRLLEEVIKLNKDAIDEILELLKRDHMVLHEYVNAAGITESKKLVRDVVEEAKRLLKENRDSTRKIIEDIRRLLRKIEDSTR

[0433] >3plus1_cage_Cterm_GFP11_672(SEQ ID NO:27299)

[0434] DETECTIVEKIAKLIETLKKIVEDSAKKELIDLLEEIIKRSADLIRELEKVSKRGDRDRLKKLIDRMRRETERALEELRRVIEELKEATERALKTIEDLIREIEEAIRRELGDEDLAREILEVIRRILELLREDMELNLRALEEEILRLLEEVIKLNKDAIDEILELLKGRDHMVLHEYVNAAGITSSKKLVRDVVEEAKRLLKENRDSTRKIIEDIRRLLRKIEDSTR

[0435] >3plus1_cage_Nterm_GFP11_67>3(SEQ ID NO:27300)

[0436] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLARDHMVLHEYVNAAGITEELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVIKEVIDEDEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDEAAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKKILTEILDALRLVEKIEK

[0437] >3plus1_cage_Nterm_GFP11_67>3(SEQ ID NO:27301)

[0438] SKEETLRKEAEDLLRRLEELTRRLEKARELLEKKLRSRRDHMVLHEYVNAAGITGVPEELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVKEVIDESEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDEAAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKILTEILDALRRLVEKIEK

[0439] >3plus1_cage_Cterm_GFP11_67>3(SEQ ID NO:27302)

[0440] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGVPEELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVIKEVIDEDEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDDEARDHMVLHEYVNAAGITRKNKRILEELLKAIETLSDENKKILTEILDALRRLVEKIEK

[0441] >3plus1_cage_Nterm_GFP11_67>3(SEQ ID NO:27303)

[0442] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERARDHMVLHEYVNAAGITELREKVPEELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVIKEVIDEDEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDEAAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKKILTEILDALRLVEKIEK

[0443] >3plus1_cage_Nterm_GFP11_67>3(SEQ ID NO:27304)

[0444] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKKLSRDHMVLHEYVNAAGITKGVPEELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVIKEVIDEDEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDEAAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKKITEILDALRRLVEKIEK

[0445] >3plus1_cage_Nterm_GFP11_67>3(SEQ ID NO:27305)

[0446] SKEETLRKEAEDLLRRLEELTRRLEKARELLERELLAKLRDHMVLHEYVNAAGITEGVPEELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVKEVIDESEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDEAAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKKILTEILDARRLVEKIEK

[0447] >3plus1_cage_Nterm_GFP11_67>3(SEQ ID NO:27306)

[0448] SKEETLRKEAEDLLRRLEELTRRLEKKARERELELLEKLSRDLRDHMVLHEYVNAAGIPELIKVLEEVAKTSIDLNKEVIDVIRRVIEESLELNKRVLDVLRRVLDIAARKGRDKKAIRDVVREAARVKEVIDESEKVNRDLKRVIERVIDDDRKALERLRKAVEEVASRGDDEAAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKILTEILDALRRLVEKIEK

[0449] >3plus1_cage_Nterm_GFP11_674(SEQ ID NO:27307)

[0450] SEREEVKELDRLLEEVEKTVRELKREHDELLKEKLVRDLKRDHMVLHEYVNAAGITKEILDVIREHVRTNKEILDRVLEVVEEHLRRNKEILDKLLDDIRKVVEDAKRILLGIGDDETLRRAVRRILEERKLVEDIRKKVKDSLETLERALEEAEERIRRSLEDLKRVLKEAKDKTKDKDRLDKVELVKKLLEDTKRTVDRVRELVRKILKKSRETLEERLIEKILRELEKDAR

[0451] >3plus1_cage_Cterm_GFP11_674(SEQ ID NO:27308)

[0452] SEREEVKELKLDRLLEEVEKTVRELKREHDELLKEVEKLVRDLKKEHDELLKKVKDDGVPKEILDVIREHVRTNKEILDRVLEVVEEHLRRNKEILDKLLDDIRKVVEDAKRILGIGDDETLRRAVRRILEELRKLVEDIRKKKVKDSLETLERALEEAEERIRRRSLEDLKRVLKEAKDKTKDKDRDHMVLHEYVNAAGITKRTVDRVRELVRKILKKSRETLEELERLIEKILRELEKDAR

[0453] >3plus1_cage_Nterm_GFP11_675(SEQ ID NO:27309)

[0454] SERETVKRRLEELLKEVKRTLDKLKEEHDRLLEDVRRVVEELRDHMVLHEYVNAAGITPEELLRVIAKVLETNKRILDDLLRVVKKHVDLNKEIRLEMIKEIEVERKRVLGDGDEKTLRDKIRDIIRRLEDAAREAEEERVRRSLEELKKAVEKIRKKIEDSLRELEEALKRVRDKEEDDKRLEDISRLVKRLLDESRRVLRELEETIRKRAEESKRVLEEVKRLVEKLIRELRKEAE

[0455] >3plus1_cage_Nterm_GFP11_676(SEQ ID NO:27310)

[0456] SEDEIIKKIIEDLRRVLKEVEIHKEVEERLDKRDHMVLHEYVNAAGITDRVLDEVKRIGDVETVLRLAIEAVRRALEIVRKVTEDLIRVVRKLLEIVDAVNKRVKKAVEDVLKLLRDTLKGSKDVERVLDEARRLLEKLHRDSRRILKRIAELLDELKELSERILERIKRLLEEVERRGSDDERLLEELIELKDLIRRHIELLRRLIERLKELLKKHIEDLRRLLEEIEKLLEEARK

[0457] >3plus1_cage_Nterm_GFP11_677(SEQ ID NO:27311)

[0458] SREEVLRELEEVIEDNRRLLEELIEKSKKVLDESLKLIDELLRRDHMVLHEYVNAAGITSKGDDEETARRIVTKIRKELEDVLDEVRRIVKEILDRDRRVAEDLKRVVDENLEKLKKVLDELDKKLGDRRLVEILRDIVRLLIEINERALRELLKLHEEHTRIVAELLRELERVVEELRRLVELIAGEISEDDLKRVVDEVEKKLRELKRRYKEALERIKEKIKELKDRYERAVREVVAELRKTTK

[0459] >3plus1_cage_Cterm_GFP11_677(SEQ ID NO:27312)

[0460] SREEVLRELEEVIEDNRRLLEELIEKSKKVLDESLKLIDELLRRLEEVLERVLRLLEELSKGDDEETARRIVTKIRKELEDVLDEVRRIVKEILDRDRRVAEDLKRVVDENLEKLKKVLDELDKKLGDRRLVEILRDIVRLLIEINERALRELLKLHEEHTRIVAELLRELERVVEELRRLVELIAGEISEDDRDHMVLHEYVNAAGITRRYKEALERIKEKIKELKKDRYERAVREVVAELRKTTK

[0461] >3plus1_cage_Cterm_GFP11_677(SEQ ID NO:27313)

[0462] SREEVLRELEEVIEDNRRLLEELIEKSKKVLDESLKLIDELLRRLEEVLERVLRLLEELSKGDDEETARRIVTKIRKELEDVLDEVRRIVKEILDRDRRVAEDLKRVVDENLEKLKKVLDELDKKLGDRRLVEILRDIVRLLIEINERALRELLKLHEEHTRIVAELLRELERVVEELRRLVELIARGERDHMVLHEYVNAAGITLRELKRRYKEALERIKEKIKELKKDRYERAVREVVAELRKTTK

[0463] >3plus1_cage_Nterm_GFP11_678(SEQ ID NO:27314)

[0464] SKAEEIAEKDLRLLEENRRALEEITTRLDDLLRRDHMVLHEYVNAAGITLDDLRRGGIREELLKRIADVIRDIMRLLKELHDHTAEVIKTIKKKLLKELHDINKEIIERLKRLKDGNVPKEELLKRVEELVRTSARLTTEVLKTVEKLIRDDKRSEILKRVKELIEELKRGVDSERVKEILERILRVVEEAVRLNEESLRILDVVRKAVKLDRESLKKILDVVEEAVR

[0465] >3plus1_cage_Cterm_GFP11_678(SEQ ID NO:27315)

[0466] SKAEEIAEKDLRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGIREELLKRIADVIRDIMRLLKELHDHTAEVIKTIKKKLLKELHDINKEIIERLKRLKDGNVPKEELLKRVEELVRTSARLTTEVLKTVEKLIRDDKRSEILKRVKELIEELRDHMVLHEYVNAAGITLRVVEEAVRLNEESLRILDVVRKAVKLDRESLKKILDVVEEAVR

[0467] >3plus1_cage_Nterm_GFP11_678(SEQ ID NO:27316)

[0468] SKAEEIAEKDLRLLEENRRALEEITTRLDDLLRRNKDRDHMVLHEYVNAAGITRRGGIREELLKRIADVIRDIMRLLKELHDHTAEVIKTIKKKLLKELHDINKEIIERLKRLKDGNVPKEELLKRVEELVRTSARLTTEVLKTVEKLIRDDKRSEILKRVKELIELKRGVDSERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[0469] >3plus1_cage_Cterm_GFP11_678(SEQ ID NO:27317)

[0470] SKAEEIAEKDLRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGIREELLKRIADVIRDIMRLLKELHDHTAEVIKTIKKKLLKELHDINKEIIERLKRLKDGNVPKEELLKRVEELVRTSARLTTEVLKTVEKLIRDDKRSEILKRVKELIEELKRGRDHMVLHEYVNAAGITVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[0471] >3plus1_cage_Nterm_GFP11_678(SEQ ID NO:27318)

[0472] SKAEEIAEKDLRLLEENRRALEEITTRLDDLLRRNKRDHMVLHEYVNAAGITLRRGGIREELLKRIADVIRDIMRLLKELHDHTAEVIKTIKKKLLKELHDINKEIIERLKRLKDGNVPKEELLKRVEELVRTSARLTTEVLKTVEKLIRDDKRSEILKRVKELIEELKRGVDSERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[0473] >3plus1_cage_Nterm_GFP11_678(SEQ ID NO:27319)

[0474] SKAEEIAEKDLRLLEENRRALEEITTRLDDLLRRNKDALRDHMVLHEYVNAAGITGGIREELLKRIADVIRDIMRLLKELHDHTAEVIKTIKKKLLKELHDINKEIIERLKRLKDGNVPKEELLKRVEELVRTSARLTTEVLKTVEKLIRDDKRSEILKRVKELIELKRGVDSERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[0475] >3plus1_cage_Cterm_GFP11_679(SEQ ID NO:27320)

[0476] SRVEELKKLIEDILRISREVVERIKRVAEDIHRINRRVLDDLRKLIEDILRTVEEILARKVGDTEIAERLRDTIARVVDEIAKLLEEHEKRSRELLEEIRKLLEDILRRSERAVEEIRELLKKGVSTKDVLRIIEEILREHLELLERVVRRIEIILRELLKTIEEIVKRIKILEELKEVLKRGRVKDDEVERDHMVLHEYVNAAGITYRRLLEEIKRKLEIILRRVEELHRRLRRKLEEIDR

[0477] >3plus1_cage_Nterm_GFP11_679(SEQ ID NO:27321)

[0478] SRVEELKKLIEDILRISREVVERIKRVAEDIHRINRRVRDHMVLHEYVNAAGITEILARKVGDTEIAERLRDTIARVVDEIAKLLEEHEKRSRELLEEIRKLLEDILRRSERAVEEIRELLKKGVSTKDVLRIIEEILREHLELLERVVRRIEIILRELLKTIEIVKRIKILEELKEVLKRGRVKDDEVEREIRRVKEDLDRILEEYRRLLEEIKRKLEIILRRVEELHRRLRRKLEEIDR

[0479] In a fourth aspect, the present disclosure provides a non-naturally occurring polypeptide comprising a key polypeptide disclosed herein along its length, or selected from SEQ ID NOs: 14318-26601 (submitted as Appendix 3 in U.S. Provisional Application Serial No. 62 / 700681 filed on July 19, 2018 and / or U.S. Provisional Application Serial No. 62 / 785537 filed on December 27, 2018), 26602-27015 (submitted as Appendix 3 in U.S. Provisional Application Serial No. 62 / 700681 filed on July 19, 2018 and / or U.S. Provisional Application Serial No. 62 / 785537 filed on December 27, 2018), plus1_GFP11_key_Cterm_1″ No. 1-67 and 97-117, 3plus1_GFP11_key_Nterm_1″ No. 68-96 and 118-140, 2plus1_GFP11_key_Cterm_″ No. 1-173 and GFP11_key_Nterm_″ No. 174-274 submitted), 27016-27050, 27322 to 27358, and Table 2 (with SEQ 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence of the key polypeptide in Table 3 and / or Table 4 (excluding optional amino acid residues).

[0480] As disclosed herein, polypeptides of this aspect can be used, for example, as key polypeptides that bind to cage polypeptides to displace the latch through competitive intermolecular binding that induces a conformational change, thereby exposing the encoded biologically active peptide or domain and thereby activating the system (see, e.g., Figure 1 ).

[0481] As noted in the present disclosure, key polypeptides may include optional residues; these residues are provided in brackets and, in one embodiment, are not included in determining the percent sequence identity. In another embodiment, optional residues may be included in determining the percent sequence identity.

[0482] In another embodiment, the non-naturally occurring polypeptide includes a polypeptide having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length to the amino acid sequence of a key polypeptide selected from the group consisting of SEQ ID NOs: 26602-27050 and 27322 to 27358, as described in detail below.

[0483] Key sequence is plain text

[0484] 6His-MBP-TEV, 6His-TEV and flexible linker sequences are underlined text

[0485] The sequences in bold and italics are optional residues required for MBP_key biotinylation

[0486] All sequences in brackets are optional

[0487] Any number of consecutive amino acids at the N or C terminus of the non-optional key sequence can be removed to adjust the responsiveness

[0488] >SB76_C-helix(SEQ ID NO:27016)

[0489] DEARKAIARVKRESKRIVEDAERLIREAAAASEKIS

[0490] >SB76_C-helix-biotin(SEQ ID NO:27017)

[0491] DEARKAIARVKRESKRIVEDAERLIREAAAASEKISGSGK-Biotin

[0492] >p5_MBP (SEQ ID NO: 27018)

[0493]

[0494] >p9_MBP (SEQ ID NO: 27019)

[0495]

[0496]

[0497] >p18_MBP (SEQ ID NO: 27020)

[0498]

[0499] >MBP_p18(aka.p76)(SEQ ID NO:27021)

[0500]

[0501] >key_b(SEQ ID NO:27022)

[0502] (M)NKEEIEKLAKEAREKLKKAEKEHKEIHDKLRKKNKKAREDLKKKADELRETNKRVN( GSENLYFQ GSGSGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYA QSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSAL MFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETA MTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPL GAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNLEHHHHHH)

[0503] >key_c(SEQ ID NO:27023)

[0504] (M)SSEKVRRELKESLKENHKQNQKLLKDHKRAQEKLNRELEELKKKHKKTLDDIRRES( GSENLYFQ GSGSGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYA QSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSAL MFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETA MTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPL GAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNLEHHHHHH)

[0505] >key_d(SEQ ID NO:27024)

[0506] (M)DTVKRILEELRRRFEKLAKDLDDIARKLLEDHKKHNKELKDKQRKIKKEADDAARS( GSENLYFQ GSGSGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYA QSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSAL MFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETA MTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPL GAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNLEHHHHHH)

[0507] >key_e(SEQ ID NO:27025)

[0508] (M)DDVERRLRKANKESKKEAEELTEEAKKANEKTKEDSKELTKENRKTNKTIKDEARS( GSENLYFQ GSGSGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYA QSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSAL MFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETA MTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPL GAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNLEHHHHHH)

[0509] >key_f(SEQ ID NO:27026)

[0510] (M)DDEERRSEKTVQDAKREIKKVEDDLQRLNEEQKKKVKKQEDENQKTLKKHKDDARS( GSENLYFQ GSGSGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYA QSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSAL MFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETA MTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPL GAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNLEHHHHHH)

[0511] Additional keys:

[0512] The key sequence is plain text

[0513] (6His-MBP-TEV, 6His-TEV and flexible linker sequences are underlined text)

[0514] (Colocalization domains are in bold text)

[0515] (Positions that can be mutated to any amino acid to adjust responsiveness are in underlined bold text. These are exemplary and not exhaustive.)

[0516] (Any number of consecutive amino acids at the N or C terminus of the non-optional key sequence can be removed to adjust the responsiveness)

[0517] (All sequences in brackets are optional)

[0518] >p76-long (SEQ ID NO: 27027)

[0519]

[0520] >p76-short (SEQ ID NO: 27028)

[0521]

[0522] >k76-long (SEQ ID NO: 27029)

[0523]

[0524] >k76-short (SEQ ID NO: 27030)

[0525]

[0526]

[0527] >p76_GLISE (SEQ ID NO: 27031)

[0528] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVKRESKRIVEDAEGLISEAAAASEKISREAERLIREAAAASEKISRE

[0529] >p76_GSSEKIS(SEQ ID NO:27032)

[0530] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVKRESKRIVEDAERLIREAAGSSEKISREAERLIREAAAASEKISRE

[0531] >p76_R26G(SEQ ID NO:27033)

[0532] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVKRESKRIVEDAERLIGEAAAASEKISREAERLIREAAAASEKISRE

[0533] >p76-short_E19G(SEQ ID NO:27034)

[0534] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVKRESKRIVGDAERLIREAAAASEKISREAERLIR

[0535] >p76-short_GLISE_E01_EGFR(SEQ ID NO:27035)

[0536] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVKRESKRIVEDAEGLISEAAAASEKISREAERLIR

[0537] >p76-short_AE_EGFR(SEQ ID NO:27036)

[0538] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVAEESKRIVEDAERLIREAAAASEKISREAERLIR

[0539] >p76-short_AAE_EGFR(SEQ ID NO:27037)

[0540] ( MGSHHHHHHGSGSENLYFQGSGGS )DEAAKAIARVAEESKRIVEDAERLIREAAAASEKISREAERLIR

[0541] >p76-short_EE_EGFR(SEQ ID NO:27038)

[0542] ( MGSHHHHHHGSGSENLYFQGSGGS )DEARKAIARVKRESKRIVEDAERLIREAAEASEEISREAERLIR

[0543] >p76-spytag (SEQ ID NO: 27039)

[0544]

[0545]

[0546] >p76-short-spytag(SEQ ID NO:27040)

[0547]

[0548] >sfGFP_VMAn_p18(SEQ ID NO:27041)

[0549]

[0550] >p18_VMAc_mCherry(SEQ ID NO:27042)

[0551]

[0552] (Homologous key for 2plus1 and 3plus1 STREPII-LOCKR functional cage designs):

[0553] >2plus1_key_100000.fasta alt_STREP_2plus1_1(SEQ ID NO:27043)

[0554] DKVRKVAEVAEKVLRDIDKLDRESKEAFRRTNGEISKLDEDTRRVAERVKKAIEDLAK

[0555] >2plus1_key_2(SEQ ID NO:27044)

[0556] SEVDEIIADNERALDEVRREVEEIDKENAERLKEWVEEAREILDRLAKALEEIR

[0557] >2plus1_key_3(SEQ ID NO:27045)

[0558] PEEALSKAIKDVRDIVKKVKDELKEWRDRNKELVDRLSEELKEWLKDVERVLKELTDKDR

[0559] >2plus1_key_4(SEQ ID NO:27046)

[0560] DERVREELKKLLTRVEEEHRKVLETDKKILKEAHKESKEVNDRDRELLERLEESVR

[0561] >3plus1_key_1(SEQ ID NO:27047)

[0562] SRLVKKLDEIVKEVAKKLEDVVRANEELWRKLVELNKESVARLREAVERVARDLEETAR

[0563] >3plus1_key_2(SEQ ID NO:27048)

[0564] SDEERLEKVVKDVIEKVRRILEKWKKDIDKVVKELRRILEEWEKIIREVLDKVR

[0565] >3plus1_key_3(SEQ ID NO:27049)

[0566] DKDAVIKVIEKLIRANAAVWDALLKINEDLVRVNKTVWKELLRVNEKLARDLERVVK

[0567] >3plus1_key_4(SEQ ID NO:27050)

[0568] SLVDELRKSLERNVRVSEEVARRLKEALKRWVDVVRKVVEDLIRLNEDVVRVVEK

[0569] SEQ ID NO:26602-27015:

[0570] >3plus1_GFP11_key_Cterm_1(SEQ ID NO:26602)

[0571] SGSKEVLDILERAVEVVRRVIKALKEVLERHVDATREVIERVKRVNKRLLEAVREVVT

[0572] >3plus1_GFP11_key_Cterm_2(SEQ ID NO:26603)

[0573] GVPEEIDRELKRVVEELRRLHEEIKERLDDVARRSEEELRRIIKKLKEVVKEIRKKLK

[0574] >3plus1_GFP11_key_Cterm_3(SEQ ID NO:26604)

[0575] DLLRKLEEELRRIKEKLRKALEELEREHRELEKELDKLHDESRKEHERIEEELRR

[0576] >3plus1_GFP11_key_Cterm_4(SEQ ID NO:26605)

[0577] DEDLLEKIKRVIREHIKALEKLARDLKEILRRHIEALKELARDLAEVIRKLLEDVKR

[0578] >3plus1_GFP11_key_Cterm_5(SEQ ID NO:26606)

[0579] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0580] >3plus1_GFP11_key_Cterm_6(SEQ ID NO:26607)

[0581] DLEDILRKNLDRLRKLLERLREILRENLEALKKTLKRLEDVVREILEDLKRERK

[0582] >3plus1_GFP11_key_Cterm_7(SEQ ID NO:26608)

[0583] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0584] >3plus1_GFP11_key_Cterm_8(SEQ ID NO:26609)

[0585] SGSKEVLDILERAVEVVRRVIKALKEVLERHVDATREVIERVKRVNKRLLEAVREVVT

[0586] >3plus1_GFP11_key_Cterm_9(SEQ ID NO:26610)

[0587] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0588] >3plus1_GFP11_key_Cterm_10(SEQ ID NO:26611)

[0589] RLIEEVVRLLRENLDVVRRILEALAKLIKELLEALEEVLRRNKELIRELLRVLDEALK

[0590] >3plus1_GFP11_key_Cterm_11(SEQ ID NO:26612)

[0591] DIVRAMEEVIRRLIEILRRDVELNLDVAKKLLELLKEDSKLNLDVARELLELLDR

[0592] >3plus1_GFP11_key_Cterm_12(SEQ ID NO:26613)

[0593] DIVRAMEEVIRRLIEILRRDVELNLDVAKKLLELLKEDSKLNLDVARELLELLDR

[0594] >3plus1_GFP11_key_Cterm_13(SEQ ID NO:26614)

[0595] RLIEEVVRLLRENLDVVRRILEALAKLIKELLEALEEVLRRNKELIRELLRVLDEALK

[0596] >3plus1_GFP11_key_Cterm_14(SEQ ID NO:26615)

[0597] RLIEEVVRLLRENLDVVRRILEALAKLIKELLEALEEVLRRNKELIRELLRVLDEALK

[0598] >3plus1_GFP11_key_Cterm_15(SEQ ID NO:26616)

[0599] DLLRKLEEELRRIKEKLRKALEELEREHRELEKELDKLHDESRKEHERIEEELRR

[0600] >3plus1_GFP11_key_Cterm_16(SEQ ID NO:26617)

[0601] DLLRKLEEELRRIKEKLRKALEELEREHRELEKELDKLHDESRKEHERIEEELRR

[0602] >3plus1_GFP11_key_Cterm_17(SEQ ID NO:26618)

[0603] ELAREVERVIKELLDKSKEILERIERAIDELLKVSEEILKLSEDASEELLKILREFAK

[0604] >3plus1_GFP11_key_Cterm_18(SEQ ID NO:26619)

[0605] DVKDIIRTILEVARDLLRLLEEDSRTNSEVVKRLLDLLREDSKANSEVVKRLLDVLRE

[0606] >3plus1_GFP11_key_Cterm_19(SEQ ID NO:26620)

[0607] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0608] >3plus1_GFP11_key_Cterm_20(SEQ ID NO:26621)

[0609] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0610] >3plus1_GFP11_key_Cterm_21(SEQ ID NO:26622)

[0611] RLIEEVVRLLRENLDVVRRILEALAKLIKELLEALEEVLRRNKELIRELLRVLDEALK

[0612] >3plus1_GFP11_key_Cterm_22(SEQ ID NO:26623)

[0613] DLEDILRKNLDRLRKLLERLREILRENLEALKKTLKRLEDVVREILEDLKRERK

[0614] >3plus1_GFP11_key_Cterm_23(SEQ ID NO:26624)

[0615] DLLRKLEEELRRIKEKLRKALEELEREHRELEKELDKLHDESRKEHERIEEELRR

[0616] >3plus1_GFP11_key_Cterm_24(SEQ ID NO:26625)

[0617] DEDLLEKIKRVIREHIKALEKLARDLKEILRRHIEALKELARDLAEVIRKLLEDVKR

[0618] >3plus1_GFP11_key_Cterm_25(SEQ ID NO:26626)

[0619] ELVRIAIEVLKRLLEIIEELVRLNNEILERLLKIVRELHKDNIKILEDLLRIIEEVLR

[0620] >3plus1_GFP11_key_Cterm_26(SEQ ID NO:26627)

[0621] ELVRIAIEVLKRLLEIIEELVRLNNEILERLLKIVRELHKDNIKILEDLLRIIEEVLR

[0622] >3plus1_GFP11_key_Cterm_27(SEQ ID NO:26628)

[0623] RLARLLKALADKLIRVLEEILKINEELNRKIIKFARENLERNRRVNKKVIEVLREAAR

[0624] >3plus1_GFP11_key_Cterm_28(SEQ ID NO:26628)

[0625] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0626] >3plus1_GFP11_key_Cterm_29(SEQ ID NO:26630)

[0627] ELVRIAIEVLKRLLEIIEELVRLNNEILERLLKIVRELHKDNIKILEDLLRIIEEVLR

[0628] >3plus1_GFP11_key_Cterm_30(SEQ ID NO:26631)

[0629] ELVRIAIEVLKRLLEIIEELVRLNNEILERLLKIVRELHKDNIKILEDLLRIIEEVLR

[0630] >3plus1_GFP11_key_Cterm_31(SEQ ID NO:26632)

[0631] DIVRAMEEVIRRLIEILRRDVELNLDVAKKLLELLKEDSKLNLDVARELLELLDR

[0632] >3plus1_GFP11_key_Cterm_32(SEQ ID NO:26633)

[0633] RKIAKIIEELKRLLEDLARDTRRVIEEAKRLLKEWRDRNKEVADTLKKLLEDLIRKIR

[0634] >3plus1_GFP11_key_Cterm_33(SEQ ID NO:26634)

[0635] DLLRKLEEELRRIKEKLRKALEELEREHRELEKELDKLHDESRKEHERIEEELRR

[0636] >3plus1_GFP11_key_Cterm_34(SEQ ID NO:26635)

[0637] DLLRKLEEELRRIKEKLRKALEELEREHRELEKELDKLHDESRKEHERIEEELRR

[0638] >3plus1_GFP11_key_Cterm_35(SEQ ID NO:26636)

[0639] RKIAKIIEELKRLLEDLARDTRRVIEEAKRLLKEWRDRNKEVADTLKKLLEDLIRKIR

[0640] >3plus1_GFP11_key_Cterm_36(SEQ ID NO:26637)

[0641] ELVRIAIEVLKRLLEIIEELVRLNNEILERLLKIVRELHKDNIKILEDLLRIIEEVLR

[0642] >3plus1_GFP11_key_Cterm_37(SEQ ID NO:26638)

[0643] TVRRLREALKKLEDDLRKIERDAEREYKKLKDELEELTERYRREIRKLKEELKADRK

[0644] >3plus1_GFP11_key_Cterm_38(SEQ ID NO:26639)

[0645] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0646] >3plus1_GFP11_key_Cterm_39(SEQ ID NO:26640)

[0647] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0648] >3plus1_GFP11_key_Cterm_40(SEQ ID NO:26641)

[0649] DLEDILRKNLDRLRKLLERLREILRENLEALKKTLKRLEDVVREILEDLKRERK

[0650] >3plus1_GFP11_key_Cterm_41(SEQ ID NO:26642)

[0651] DLERLRRKVEELEDRLRRLLEKLARDSAELMRELERILDRYARESEELDRRLAE

[0652] >3plus1_GFP11_key_Cterm_42(SEQ ID NO:26643)

[0653] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0654] >3plus1_GFP11_key_Cterm_43(SEQ ID NO:26644)

[0655] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0656] >3plus1_GFP11_key_Cterm_44(SEQ ID NO:26645)

[0657] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0658] >3plus1_GFP11_key_Cterm_45(SEQ ID NO:26646)

[0659] DKAVEELEKALEEIKRRLKEVIDRYEDELRKLRKEYKEKIDKYERKLEEIERRERT

[0660] >3plus1_GFP11_key_Cterm_46(SEQ ID NO:26647)

[0661] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0662] >3plus1_GFP11_key_Cterm_47(SEQ ID NO:26648)

[0663] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0664] >3plus1_GFP11_key_Cterm_48(SEQ ID NO:26649)

[0665] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0666] >3plus1_GFP11_key_Cterm_49(SEQ ID NO:26650)

[0667] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0668] >3plus1_GFP11_key_Cterm_50(SEQ ID NO:26651)

[0669] EVKRRLEEKERRIRTRYEELRRRLRKRVKDYEDKLREIEKKVRRDAERIEEELERAKK

[0670] >3plus1_GFP11_key_Cterm_51(SEQ ID NO:26652)

[0671] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0672] >3plus1_GFP11_key_Cterm_52(SEQ ID NO:26653)

[0673] KIAEEIERELEELRRMIKRLHEDLERKLKESEDELREIEARLEEKIRRLEEKLERKRR

[0674] >3plus1_GFP11_key_Cterm_53(SEQ ID NO:26654)

[0675] KIAEEIERELEELRRMIKRLHEDLERKLKESEDELREIEARLEEKIRRLEEKLERKRR

[0676] >3plus1_GFP11_key_Cterm_54(SEQ ID NO:26655)

[0677] DKAVEELEKALEEIKRRLKEVIDRYEDELRKLRKEYKEKIDKYERKLEEIERRERT

[0678] >3plus1_GFP11_key_Cterm_55(SEQ ID NO:26656)

[0679] KIAEEIERELEELRRMIKRLHEDLERKLKESEDELREIEARLEEKIRRLEEKLERKRR

[0680] >3plus1_GFP11_key_Cterm_56(SEQ ID NO:26657)

[0681] ELVRIAIEVLKRLLEIIEELVRLNNEILERLLKIVRELHKDNIKILEDLLRIIEEVLR

[0682] >3plus1_GFP11_key_Cterm_57(SEQ ID NO:26658)

[0683] DEVEREIRRVKEDLDRILEEYRRLLEEIKRKLEEILRRVEELHRRLRRKLEEIDR

[0684] >3plus1_GFP11_key_Cterm_58(SEQ ID NO:26659)

[0685] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0686] >3plus1_GFP11_key_Cterm_59(SEQ ID NO:26660)

[0687] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0688] >3plus1_GFP11_key_Cterm_60(SEQ ID NO:26661)

[0689] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0690] >3plus1_GFP11_key_Cterm_61(SEQ ID NO:26662)

[0691] TLREVVRKVLEEAKRLLDELEEVHKRVKKELEDIIEENRRVVKRVRDELREIKRELDE

[0692] >3plus1_GFP11_key_Cterm_62(SEQ ID NO:26663)

[0693] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0694] >3plus1_GFP11_key_Cterm_63(SEQ ID NO:26664)

[0695] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0696] >3plus1_GFP11_key_Cterm_64(SEQ ID NO:26665)

[0697] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0698] >3plus1_GFP11_key_Cterm_65(SEQ ID NO:26666)

[0699] DEAERRRRELKDKLDRLREEHEEVKRRLEEELTRLRETHKKIEKELREALKRVRDRST

[0700] >3plus1_GFP11_key_Cterm_66(SEQ ID NO:26667)

[0701] KIAEEIERELEELRRMIKRLHEDLERKLKESEDELREIEARLEEKIRRLEEKLERKRR

[0702] >3plus1_GFP11_key_Cterm_67(SEQ ID NO:26668)

[0703] DVKRALEELVSRLRKLLEDVKKASEDIVREVERIVRELAKRSDEILKKLEDIVEKLRE

[0704] >3plus1_GFP11_key_Nterm_68(SEQ ID NO:26669)

[0705] SEAERLADEVRKAVKKSEEDNETLVREVEKAVRELKKNNKTWVDEVRKLMKRLVDLLR

[0706] >3plus1_GFP11_key_Nterm_69(SEQ ID NO:26670)

[0707] SEAERLADEVRKAVKKSEEDNETLVREVEKAVRELKKNNKTWVDEVRKLMKRLVDLLR

[0708] >3plus1_GFP11_key_Nterm_70(SEQ ID NO:26671)

[0709] DKDKRLEELLKRLKELNDKTFEELERILEELKRANEASLREAERILEELRARIEGGNL

[0710] >3plus1_GFP11_key_Nterm_71(SEQ ID NO:26672)

[0711] SEAEDLEELIKELAELLKDVIRKLEKINRRLVKILEDIIRRLKEISKEAEEELRKGTV

[0712] >3plus1_GFP11_key_Nterm_72(SEQ ID NO:26673)

[0713] SDKEEIKRRVEKTARDLETEHDKIKKRLEDTVRDIKRELDELLEKYERVLRKIEKTLR

[0714] >3plus1_GFP11_key_Nterm_73(SEQ ID NO:26674)

[0715] SEAEKIREALETNLRLLEELIKRLKEILDTHNELLRRVIETLERLLKELLELLEEGGL

[0716] >3plus1_GFP11_key_Nterm_74(SEQ ID NO:26675)

[0717] SEAEKIREALETNLRLLEELIKRLKEILDTHNELLRRVIETLERLLKELLELLEEGGL

[0718] >3plus1_GFP11_key_Nterm_75(SEQ ID NO:26676)

[0719] SKEERLREVAEKHKKDLEDIVKRVDEAAKETARRLEEILKRLEEVLKKILDDLEKGPD

[0720] >3plus1_GFP11_key_Nterm_76(SEQ ID NO:26677)

[0721] SLEEITKRLLELVEENLARHEEILRELLELAKRLAKEDRDILEEVLKLIEELLKLLED

[0722] >3plus1_GFP11_key_Nterm_77(SEQ ID NO:26678)

[0723] SKEETLKRLLDELEKRNRETVERLERLLKELEDRNRASLEELEAVLEELERKIEESGL

[0724] >3plus1_GFP11_key_Nterm_78(SEQ ID NO:26679)

[0725] SKEETLKRLLDELEKRNRETVERLERLLKELEDRNRASLEELEAVLEELERKIEESGL

[0726] >3plus1_GFP11_key_Nterm_79(SEQ ID NO:26680)

[0727] SKEETLKRLLDELEKRNRETVERLERLLKELEDRNRASLEELEAVLEELERKIEESGL

[0728] >3plus1_GFP11_key_Nterm_80(SEQ ID NO:26681)

[0729] STREKAKKVLDTLRADNEDMKRVVEKILRALKRTNERAEKLAREITEEIKRILKEVGV

[0730] >3plus1_GFP11_key_Nterm_81(SEQ ID NO:26682)

[0731] DAEEVVKRLADVLRENDETIRKVVEDLVRIAEENDRLWKKLVEDIAEILRRIVELLRR

[0732] >3plus1_GFP11_key_Nterm_82(SEQ ID NO:26683)

[0733] SKEETLKRLLDELEKRNRETVERLERLLKELEDRNRASLEELEAVLEELERKIESGL

[0734] >3plus1_GFP11_key_Nterm_83(SEQ ID NO:26684)

[0735] STREKAKKVLDTLRADNEDMKRVVEKILRALKRTNERAEKLAREITEEIKRILKEVGV

[0736] >3plus1_GFP11_key_Nterm_84(SEQ ID NO:26685)

[0737] SKEEVEKVLRKWEILRRLIENKRANDKIRREYEEELVKEIRRFLEIKEVAERLGV

[0738] >3plus1_GFP11_key_Nterm_85(SEQ ID NO:26686)

[0739] DREKSVRDIEEDLKRVLDKLRRRVETSKEELKKVLKADKENADELEKTLRDVVRELDR

[0740] >3plus1_GFP11_key_Nterm_86(SEQ ID NO:26687)

[0741] SDKEEIKRRVEKTARDLETEHDKIKKRLEDTVRDIKRELDELLEKYERVLRKIEKTLR

[0742] >3plus1_GFP11_key_Nterm_87(SEQ ID NO:26688)

[0743] STREKAKKVLDTLRADNEDMKRVVEKILRALKRTNERAEKLAREITEEIKRILKEVGV

[0744] >3plus1_GFP11_key_Nterm_88(SEQ ID NO:26689)

[0745] SKDEELARLLEELVERWRKIVEDLERDHRRLVKEIRELVERIRKKLEELVDRIRKNGI

[0746] >3plus1_GFP11_key_Nterm_89(SEQ ID NO:26690)

[0747] SEAERLADEVRKAVKKSEEDNETLVREVEKAVRELKKNNKTWVDEVRKLMKRLVDLLR

[0748] >3plus1_GFP11_key_Nterm_90(SEQ ID NO:26691)

[0749] SKDEELARLLEELVERWRKIVEDLERDHRRLVKEIRELVERIRKKLEELVDRIRKNGI

[0750] >3plus1_GFP11_key_Nterm_91(SEQ ID NO:26692)

[0751] KEIEETLKELEDLNREMVETNRRVLEETRRLNKETVDRVKATLDELAKMLKKLVDDVR

[0752] >3plus1_GFP11_key_Nterm_92(SEQ ID NO:26693)

[0753] SEAERLADEVRKAVKKSEEDNETLVREVEKAVRELKKNNKTWVDEVRKLMKRLVDLLR

[0754] >3plus1_GFP11_key_Nterm_93(SEQ ID NO:26694)

[0755] SKEETLKRLLDELEKRNRETVERLERLLKELEDRNRASLEELEAVLEELERKIEESGL

[0756] >3plus1_GFP11_key_Nterm_94(SEQ ID NO:26695)

[0757] DKAEVLREALKLLKDLLEELIKIHEESLKRILDLIDTLVKVHEDALRALKELLERSGL

[0758] >3plus1_GFP11_key_Nterm_95(SEQ ID NO:26696)

[0759] SKEEEVEKVLRKWEEILRRLIEENKRANDKIRREYEELVKEIRRVLEEIKEVAERLGV

[0760] >3plus1_GFP11_key_Nterm_96(SEQ ID NO:26697)

[0761] SKEETLKRLLDELEKRNRETVERLERLLKELEDRNRASLEELEAVLEELERKIEESGL

[0762] >3plus1_GFP11_key_Cterm_97(SEQ ID NO:26698)

[0763] SERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[0764] >3plus1_GFP11_key_Cterm_98(SEQ ID NO:26699)

[0765] SERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[0766] >3plus1_GFP11_key_Cterm_99(SEQ ID NO:26700)

[0767] DERRIAERIRELLRESKKLVRDVVEEAKRLLKENRDSTRKIIEDIRRLLRKIEDSTR

[0768] >3plus1_GFP11_key_Cterm_100(SEQ ID NO:26701)

[0769] DALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0770] >3plus1_GFP11_key_Cterm_101(SEQ ID NO:26702)

[0771] DERRIAERIRELLRESKKLVRDVVEEAKRLLKENRDSTRKIIEDIRRLLRKIEDSTR

[0772] >3plus1_GFP11_key_Cterm_102(SEQ ID NO:26703)

[0773] EALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0774] >3plus1_GFP11_key_Cterm_103(SEQ ID NO:26704)

[0775] EALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0776] >3plus1_GFP11_key_Cterm_104(SEQ ID NO:26705)

[0777] AAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKKILTEILDALRRLVEKIEK

[0778] >3plus1_GFP11_key_Cterm_105(SEQ ID NO:26706)

[0779] EALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0780] >3plus1_GFP11_key_Cterm_106(SEQ ID NO:26707)

[0781] EALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[0782] >3plus1_GFP11_key_Cterm_107(SEQ ID NO:26708)

[0783] DALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0784] >3plus1_GFP11_key_Cterm_108(SEQ ID NO:26709)

[0785] DALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0786] >3plus1_GFP11_key_Cterm_109(SEQ ID NO:26710)

[0787] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[0788] 3plus1_GFP11_key_Cterm_110(SEQ ID NO:26711)

[0789] DRLDKVEELVKKLLEDTKRTVDRVRELVRKILKKSRETLEELERLIEKILRELEKDAR

[0790] >3plus1_GFP11_key_Cterm_111(SEQ ID NO:26712)

[0791] DALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[0792] >3plus1_GFP11_key_Cterm_112(SEQ ID NO:26713)

[0793] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[0794] >3plus1_GFP11_key_Cterm_113(SEQ ID NO:26714)

[0795] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[0796] >3plus1_GFP11_key_Cterm_114(SEQ ID NO:26715)

[0797] SEDDLKRVVDEVEKKLRELKRRYAEALERIKEKIKELKDRYERAVREVVAELRKTTK

[0798] >3plus1_GFP11_key_Cterm_115(SEQ ID NO:26716)

[0799] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[0800] >3plus1_GFP11_key_Cterm_116(SEQ ID NO:26717)

[0801] DEVEREIRRVKEDLDRILEEYRRLLEEIKRKLEEILRRVEELHRRLRRKLEEIDR

[0802] >3plus1_GFP11_key_Cterm_117(SEQ ID NO:26718)

[0803] SEDDLKRVVDEVEKKLRELKRRYAEALERIKEKIKELKDRYERAVREVVAELRKTTK

[0804] >3plus1_GFP11_key_Nterm_118(SEQ ID NO:26719)

[0805] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRLEKELRDLKRRIEEKLEELRRGSL

[0806] >3plus1_GFP11_key_Nterm_119(SEQ ID NO:26720)

[0807] SECEDAARKLRKLVEELTREEELVKKLERLIEIEEKVSEESVRKLEKLLAEISEEEVR

[0808] >3plus1_GFP11_key_Nterm_120(SEQ ID NO:26721)

[0809] SEDEIIKKIIEDLRRVLKEVEEIHKEVEERLDKVLKEAEEMHKEVLKELDRVLDEVKR

[0810] >3plus1_GFP11_key_Nterm_121(SEQ ID NO:26722)

[0811] SKAEEIAEKLDRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKKLKRLLDDLRRGGI

[0812] >3plus1_GFP11_key_Nterm_122(SEQ ID NO:26723)

[0813] SEKEKLLKESEEEVRRLRRTLEELLRKYREVLERLRKELREIEERVRDVVRRLKEVLD

[0814] >3plus1_GFP11_key_Nterm_123(SEQ ID NO:26724)

[0815] SECEDAARKLRKLVEELTREEELVKKLERLIEIEEKVSEESVRKLEKLLAEISEEEVR

[0816] >3plus1_GFP11_key_Nterm_124(SEQ ID NO:26725)

[0817] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[0818] >3plus1_GFP11_key_Nterm_125(SEQ ID NO:26726)

[0819] SECONDARY CLERKLVELTREEELVKKLERLIEEIEKVSEESVRKKLLAEISEVR

[0820] >3plus1_GFP11_key_Nterm_126(SEQ ID NO:26727)

[0821] SKAEEIAECLDRLLEENRRLEITTRLDLLRRNKDALRKVMEKKRLLDLLRRGGI

[0822] >3plus1_GFP11_key_Nterm_127(SEQ ID NO:26728)

[0823] SECONDARY CLERKLVELTREEELVKKLERLIEEIEKVSEESVRKKLLAEISEVR

[0824] >3plus1_GFP11_key_Nterm_128(SEQ ID NO:26729)

[0825] SKAEEIAECLDRLLEENRRLEITTRLDLLRRNKDALRKVMEKKRLLDLLRRGGI

[0826] >3plus1_GFP11_key_Nterm_129(SEQ ID NO:26730)

[0827] SKEETLRKEAEDLLRRLEELTRRLKARELERALKLSRDLAEELKRLLKELREKGV

[0828] >3plus1_GFP11_key_Nterm_130(SEQ ID NO:26731)

[0829] SRVEELKKLIEDILRICE REVERIKKRVAEDIHRINRRVLDLRKLIEDILRTVEILA

[0830] >3plus1_GFP11_key_Nterm_131(SEQ ID NO:26732)

[0831] SKEETLRKEAEDLLRRLEELTRRLKARELERALKLSRDLAEELKRLLKELREKGV

[0832] >3plus1_GFP11_key_Nterm_132(SEQ ID NO:26733)

[0833] SKAEEIAEKLDRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGI

[0834] >3plus1_GFP11_key_Nterm_133(SEQ ID NO:26734)

[0835] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[0836] >3plus1_GFP11_key_Nterm_134(SEQ ID NO:26735)

[0837] SEKEDAARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLAEISEEVR

[0838] >3plus1_GFP11_key_Nterm_135(SEQ ID NO:26736)

[0839] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[0840] >3plus1_GFP11_key_Nterm_136(SEQ ID NO:26737)

[0841] SERETVKRRLEELLKEVKRTLDKLKEEHDRLLEDVRRVVEELKREHDKLLKEVKDSGV

[0842] >3plus1_GFP11_key_Nterm_137(SEQ ID NO:26738)

[0843] SEKEDAARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLAEISEEVR

[0844] >3plus1_GFP11_key_Nterm_138(SEQ ID NO:26739)

[0845] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[0846] >3plus1_GFP11_key_Nterm_139(SEQ ID NO:26740)

[0847] KEREEVKEKLDRLLEEVEKTVRELKREHDELLKEVEKLVRDLKKEHDELLKKVKDDGV

[0848] >3plus1_GFP11_key_Nterm_140(SEQ ID NO:26741)

[0849] SREEVLRELEEVIEDNRRLLEELIEKSKKVLDESLKLIDELLRRLEEVLERVLRLLEE

[0850] >2plus1_GFP11_key_Cterm_1(SEQ ID NO:26742)

[0851] DEVVKRVRDLLDTVRRRNEKVNEDVKRMNDKLRRDNEDVIRRVEKLLRELEEKRRT

[0852] >2plus1_GFP11_key_Cterm_2(SEQ ID NO:26743)

[0853] SEDSVERIARELERNLDDLARVLKESEDDLAEILRRLKEVLEESERDLERVEREVRK

[0854] >2plus1_GFP11_key_Cterm_3(SEQ ID NO:26744)

[0855] SKELLEKAKAVVDEIKRLAEESLKRLEDLSRDHKRRAKELNDEIAKVVDELAKRAT

[0856] >2plus1_GFP11_key_Cterm_4(SEQ ID NO:26745)

[0857] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0858] >2plus1_GFP11_key_Cterm_5(SEQ ID NO:26746)

[0859] DEVVKRVRDLLDTVRRRNEKVNEDVKRMNDKLRRDNEDVIRRVEKLLRELEEKRRT

[0860] >2plus1_GFP11_key_Cterm_6(SEQ ID NO:26747)

[0861] DIKTLLDRVRKLAEEDAERLDRLRRESEELNERVRRVDKKLLEEIRRKAKKVEDDTR

[0862] >2plus1_GFP11_key_Cterm_7(SEQ ID NO:26748)

[0863] DAETLLRELEKLSRDNKELLKKIEKEIRDLIKEDKERNIELSERLRKLVEELKKKAT

[0864] >2plus1_GFP11_key_Cterm_8(SEQ ID NO:26749)

[0865] DEVVKRVRDLLDTVRRRNEKVNEDVKRMNDKLRRDNEDVIRRVEKLLRELEEKRRT

[0866] >2plus1_GFP11_key_Cterm_9(SEQ ID NO:26750)

[0867] DIKTLLDRVRKLAEEDAERLDRLRRESEELNERVRRVDKKLLEEIRRKAKKVEDDTR

[0868] >2plus1_GFP11_key_Cterm_10(SEQ ID NO:26751)

[0869] SEELSAEVKKLLDEVRKALARHKDENDKLLKEIEDSLRRHKEENDRLLEKLKESTR

[0870] >2plus1_GFP11_key_Cterm_11(SEQ ID NO:26752)

[0871] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0872] >2plus1_GFP11_key_Cterm_12(SEQ ID NO:26753)

[0873] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0874] >2plus1_GFP11_key_Cterm_13(SEQ ID NO:26754)

[0875] DEVVKRVRDLLDTVRRRNEKVNEDVKRMNDKLRRDNEDVIRRVEKLLRELEEKRRT

[0876] >2plus1_GFP11_key_Cterm_14(SEQ ID NO:26755)

[0877] SEELSAEVKKLLDEVRKALARHKDENDKLLKEIEDSLRRHKEENDRLLEKLKESTR

[0878] >2plus1_GFP11_key_Cterm_15(SEQ ID NO:26756)

[0879] DAETVLRSAEDIVAKNRKLAEEVLRRVKKIVEENRKIASEVLDDVRKLVEDVLARAS

[0880] >2plus1_GFP11_key_Cterm_16(SEQ ID NO:26757)

[0881] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0882] >2plus1_GFP11_key_Cterm_17(SEQ ID NO:26758)

[0883] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0884] >2plus1_GFP11_key_Cterm_18(SEQ ID NO:26759)

[0885] DEEKLKDLIRKLRDILRRAAEAHKKLIDDARESLERAKREHEKLIDRLKKILEELER

[0886] >2plus1_GFP11_key_Cterm_19(SEQ ID NO:26760)

[0887] DIKTLLDRVRKLAEEDAERLDRLRRESEELNERVRRVDKKLLEEIRRKAKKVEDDTR

[0888] >2plus1_GFP11_key_Cterm_20(SEQ ID NO:26761)

[0889] DATRVIEEAKRILDEARKLNEETIRRSEELVRRIERVIEEIIKRSEKLLEDVARESK

[0890] >2plus1_GFP11_key_Cterm_21(SEQ ID NO:26762)

[0891] SEELSAEVKKLLDEVRKALARHKDENDKLLKEIEDSLRRHKEENDRLLEKLKESTR

[0892] >2plus1_GFP11_key_Cterm_22(SEQ ID NO:26763)

[0893] DAETVLRSAEDIVAKNRKLAEEVLRRVKKIVEENRKIASEVLDDVRKLVEDVLARAS

[0894] >2plus1_GFP11_key_Cterm_23(SEQ ID NO:26764)

[0895] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0896] >2plus1_GFP11_key_Cterm_24(SEQ ID NO:26765)

[0897] SKELLEKAKAVVDEIKRLAESLKRLEDLSRDHKRAKENLNDEIAKVVDELAKRAT

[0898] >2plus1_GFP11_key_Cterm_25(SEQ ID NO:26766)

[0899] DEEVLKKLAEIVRRVKEENRKKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0900] >2plus1_GFP11_key_Cterm_26(SEQ ID NO:26767)

[0901] DEEVLKKLAEIVRRVKEENRKKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0902] >2plus1_GFP11_key_Cterm_27(SEQ ID NO:26768)

[0903] DKLLKEARDLIREEIKRLEELKRVEKLTEDAKRDLERSNREHKELADRIKETAR

[0904] >2plus1_GFP11_key_Cterm_28(SEQ ID NO:26769)

[0905] DKDSARELERIVKENAELARVFREVEKIRENTKLAEDSPRELKRLVELEKKRAK

[0906] >2plus1_GFP11_key_Cterm_29(SEQ ID NO:26770)

[0907] SKEKIDRIIRELERILEEAKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0908] >2plus1_GFP11_key_Cterm_30(SEQ ID NO:26771)

[0909] DEEKLKDLIRKLRDILRRAAEAHKKLIDDARESLERAKREHEKLIDRLKKILEER

[0910] >2plus1_GFP11_key_Cterm_31(SEQ ID NO:26772)

[0911] DEVVKRVRDLLDTVRRRNEKVNEDVKRMNDKLRRDNEDVIRRVEKLLRELEEKRRT

[0912] >2plus1_GFP11_key_Cterm_32(SEQ ID NO:26773)

[0913] DEEVLRTLEEIIRRLTKELEDVLREYERELRRLEEENKRVIDKTEEEIRRLADRLRR

[0914] >2plus1_GFP11_key_Cterm_33(SEQ ID NO:26774)

[0915] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[0916] >2plus1_GFP11_key_Cterm_34(SEQ ID NO:26775)

[0917] SEELSAEVKKLLDEVRKALARHKDENDKLLKEIEDSLRRHKEENDRLLEKLKESTR

[0918] >2plus1_GFP11_key_Cterm_35(SEQ ID NO:26776)

[0919] LPEEVLRELEELLKESEERIKRIEEEIKKIIDKSREDIKRVLEEIERLNAKAADDLRK

[0920] >2plus1_GFP11_key_Cterm_36(SEQ ID NO:26777)

[0921] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0922] >2plus1_GFP11_key_Cterm_37(SEQ ID NO:26778)

[0923] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0924] >2plus1_GFP11_key_Cterm_38(SEQ ID NO:26779)

[0925] DKLLKEARDLIREIEKRLEELLKRVEKLTEDAKRDLERSNREHKELADRIKETAR

[0926] >2plus1_GFP11_key_Cterm_39(SEQ ID NO:26780)

[0927] DEEVLRTLEEIIRRLTKELEDVLREYERELRRLEEENKRVIDKTEEEIRRLADRLRR

[0928] >2plus1_GFP11_key_Cterm_40(SEQ ID NO:26781)

[0929] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[0930] >2plus1_GFP11_key_Cterm_41(SEQ ID NO:26782)

[0931] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[0932] >2plus1_GFP11_key_Cterm_42(SEQ ID NO:26783)

[0933] SEELSAEVKKLLDEVRKALARHKDENDKLLKEIEDSLRRHKEENDRLLEKLKESTR

[0934] >2plus1_GFP11_key_Cterm_43(SEQ ID NO:26784)

[0935] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0936] >2plus1_GFP11_key_Cterm_44(SEQ ID NO:26785)

[0937] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0938] >2plus1_GFP11_key_Cterm_45(SEQ ID NO:26786)

[0939] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[0940] >2plus1_GFP11_key_Cterm_46(SEQ ID NO:26787)

[0941] TLRELARSIRKLSAENKERLKELLRELKKLSDENKERIKKLLSDAEKIIEDVARRAK

[0942] >2plus1_GFP11_key_Cterm_47(SEQ ID NO:26788)

[0943] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[0944] >2plus1_GFP11_key_Cterm_48(SEQ ID NO:26789)

[0945] EKLKELRDVIAEVAKRIDELDEYTRESIRRAKKEIERLNRETKKVIEEVVKRIEEERK

[0946] >2plus1_GFP11_key_Cterm_49(SEQ ID NO:26790)

[0947] DERVREELKKLLTRVEEEHRKVLETDKKILKEAHKESKEVNDRDRELLERLEESVR

[0948] >2plus1_GFP11_key_Cterm_50(SEQ ID NO:26791)

[0949] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0950] >2plus1_GFP11_key_Cterm_51(SEQ ID NO:26792)

[0951] TVKRLLDELRELLERLKRTIEELLKRNRDLLADAEEKARRLLEENRKLLKAARDTAT

[0952] >2plus1_GFP11_key_Cterm_52(SEQ ID NO:26793)

[0953] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0954] >2plus1_GFP11_key_Cterm_53(SEQ ID NO:26794)

[0955] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0956] >2plus1_GFP11_key_Cterm_54(SEQ ID NO:26795)

[0957] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[0958] >2plus1_GFP11_key_Cterm_55(SEQ ID NO:26796)

[0959] DATRVIEEAKRILDEARKLNEETIRRSEELVRRIERVIEEIIKRSEKLLEDVARESK

[0960] >2plus1_GFP11_key_Cterm_56(SEQ ID NO:26797)

[0961] EAAREIIKRLREVNKRTKEKLDELIKHSEEVLERVKRLIDELRKHSEEVLEDLRRRAK

[0962] >2plus1_GFP11_key_Cterm_57(SEQ ID NO:26798)

[0963] EKLKELRDVIAEVAKRIDELDEYTRESIRRAKKEIERLNRETKKVIEEVVKRIEEERK

[0964] >2plus1_GFP11_key_Cterm_58(SEQ ID NO:26799)

[0965] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[0966] >2plus1_GFP11_key_Cterm_59(SEQ ID NO:26800)

[0967] SKAIKDVRDIVKKVKDELKEWRDRNKELVDRLSEELKEWLKDVERVLKELTDKDR

[0968] >2plus1_GFP11_key_Cterm_60(SEQ ID NO:26801)

[0969] DERVREELKKLLTRVEEEHRKVLETDKKILKEAHKESKEVNDRDRELLERLEESVR

[0970] >2plus1_GFP11_key_Cterm_61(SEQ ID NO:26802)

[0971] DIDKLLKELRDLVEKIKKDLKELLERYEEIVRRIKELLKDLNREAEEVVRRLKEELR

[0972] >2plus1_GFP11_key_Cterm_62(SEQ ID NO:26803)

[0973] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0974] >2plus1_GFP11_key_Cterm_63(SEQ ID NO:26804)

[0975] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0976] >2plus1_GFP11_key_Cterm_64(SEQ ID NO:26805)

[0977] EREEELKEVADRVKEKLDRLNRENEKSSEELKRELDKINDENRETSERLKREIDETTR

[0978] >2plus1_GFP11_key_Cterm_65(SEQ ID NO:26806)

[0979] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0980] >2plus1_GFP11_key_Cterm_66(SEQ ID NO:26807)

[0981] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[0982] >2plus1_GFP11_key_Cterm_67(SEQ ID NO:26808)

[0983] TKDLLDENSKRSNEISREVKKDLERTVRENKKIVDEVAKALEDTVDKNRRIVEEVTT

[0984] >2plus1_GFP11_key_Cterm_68(SEQ ID NO:26809)

[0985] DEVVKRVRDLLDTVRRRNEKVNEDVKRMNDKLRRDNEDVIRRVEKLLRELEEKRRT

[0986] >2plus1_GFP11_key_Cterm_69(SEQ ID NO:26810)

[0987] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[0988] >2plus1_GFP11_key_Cterm_70(SEQ ID NO:26811)

[0989] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[0990] >2plus1_GFP11_key_Cterm_71(SEQ ID NO:26812)

[0991] SEELSAEVKKLLDEVRKALARHKDENDKLLKEIEDSLRRHKEENDRLLEKLKESTR

[0992] >2plus1_GFP11_key_Cterm_72(SEQ ID NO:26813)

[0993] DIDKLLKELRDLVEKIKKDLKELLERYEEIVRRIKELLKDLNREAEEVVRRLKEELR

[0994] >2plus1_GFP11_key_Cterm_73(SEQ ID NO:26814)

[0995] DADDVLARVEELAKRAHDENERLIREVEELVRAHNKRNKELVDEVKRLVEKVIEEER

[0996] >2plus1_GFP11_key_Cterm_74(SEQ ID NO:26815)

[0997] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[0998] >2plus1_GFP11_key_Cterm_75(SEQ ID NO:26816)

[0999] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[1000] >2plus1_GFP11_key_Cterm_76(SEQ ID NO:26817)

[1001] SKEKIDRIIRELERILEEAKKKHEDVLRRLEDSLRRVAELLKAALDRLREIVDRLRR

[1002] >2plus1_GFP11_key_Cterm_77(SEQ ID NO:26818)

[1003] SEELREELKKLERKIEKVAKEIHDHDKEVTERLEDLLRRITEHARKSDREIEETAR

[1004] >2plus1_GFP11_key_Cterm_78(SEQ ID NO:26819)

[1005] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[1006] >2plus1_GFP11_key_Cterm_79(SEQ ID NO:26820)

[1007] DATRVIEEAKRILDEARKLNEETIRRSEELVRRIERVIEEIIKRSEKLLEDVARESK

[1008] >2plus1_GFP11_key_Cterm_80(SEQ ID NO:26821)

[1009] EKLKELRDVIAEVAKRIDELDEYTRESIRRAKKEIERLNRETKKVIEEVVKRIEEERK

[1010] >2plus1_GFP11_key_Cterm_81(SEQ ID NO:26822)

[1011] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1012] >2plus1_GFP11_key_Cterm_82(SEQ ID NO:26823)

[1013] DIDKLLKELRDLVEKIKKDLKELLERYEEIVRRIKELLKDLNREAEEVVRRLKEELR

[1014] >2plus1_GFP11_key_Cterm_83(SEQ ID NO:26824)

[1015] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1016] >2plus1_GFP11_key_Cterm_84(SEQ ID NO:26825)

[1017] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[1018] >2plus1_GFP11_key_Cterm_85(SEQ ID NO:26826)

[1019] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[1020] >2plus1_GFP11_key_Cterm_86(SEQ ID NO:26827)

[1021] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[1022] >2plus1_GFP11_key_Cterm_87(SEQ ID NO:26828)

[1023] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[1024] >2plus1_GFP11_key_Cterm_88(SEQ ID NO:26829)

[1025] DLKRVEERAREVSRRNEESMRRVKEDADRVSEANKEVLDRVREEVKRLIEEVRETLR

[1026] >2plus1_GFP11_key_Cterm_89(SEQ ID NO:26830)

[1027] EKLKELRDVIAEVAKRIDELDEYTRESIRRAKKEIERLNRETKKVIEEVVKRIEEERK

[1028] >2plus1_GFP11_key_Cterm_90(SEQ ID NO:26831)

[1029] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1030] >2plus1_GFP11_key_Cterm_91(SEQ ID NO:26832)

[1031] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1032] >2plus1_GFP11_key_Cterm_92(SEQ ID NO:26833)

[1033] LPEEVLRELEELLKESEERIKRIEEEIKKIIDKSREDIKRVLEEIERLNAKAADDLRK

[1034] >2plus1_GFP11_key_Cterm_93(SEQ ID NO:26834)

[1035] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1036] >2plus1_GFP11_key_Cterm_94(SEQ ID NO:26835)

[1037] DEEVLKKLAEIVRRVKEENRKVNEEVEKRLRELEEENKKVIEDLKSTVEELVERLR

[1038] >2plus1_GFP11_key_Cterm_95(SEQ ID NO:26836)

[1039] DKLLKEARDLIREIEKRLEELLKRVEKLTEDAKRDLERSNREHKELADRIKETAR

[1040] >2plus1_GFP11_key_Cterm_96(SEQ ID NO:26837)

[1041] DKLLKEARDLIREIEKRLEELLKRVEKLTEDAKRDLERSNREHKELADRIKETAR

[1042] >2plus1_GFP11_key_Cterm_97(SEQ ID NO:26838)

[1043] DIVRKIERIVETIEREVRESVKKVEEIARDIRRKVDESVKNVEKLLRDVDKKARDRKK

[1044] >2plus1_GFP11_key_Cterm_98(SEQ ID NO:26839)

[1045] DEIKRIVDEVRERLKRIVDENAKIVEDARRALEKIVKENEEILRRLKKELRELRK

[1046] >2plus1_GFP11_key_Cterm_99(SEQ ID NO:26840)

[1047] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[1048] >2plus1_GFP11_key_Cterm_100(SEQ ID NO:26841)

[1049] DLKRVEERAREVSRRNEESMRRVKEDADRVSEANKEVLDRVREEVKRLIEEVRETLR

[1050] >2plus1_GFP11_key_Cterm_101(SEQ ID NO:26842)

[1051] DATRVIEEAKRILDEARKLNEETIRRSEELVRRIERVIEEIIKRSEKLLEDVARESK

[1052] >2plus1_GFP11_key_Cterm_102(SEQ ID NO:26843)

[1053] DAETIERVVRELLEENKEVLRKTEEAVKRSTETNKRLLEASKEVADRLRERIKEAAK

[1054] >2plus1_GFP11_key_Cterm_103(SEQ ID NO:26844)

[1055] EKLKELRDVIAEVAKRIDELDEYTRESIRRAKKEIERLNRETKKVIEEVVKRIEEERK

[1056] >2plus1_GFP11_key_Cterm_104(SEQ ID NO:26845)

[1057] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1058] >2plus1_GFP11_key_Cterm_105(SEQ ID NO:26846)

[1059] DEVVERAERISEENKRRVEDVARKSKELVEDVRRHSEEVVRRVEELVKEVEERVR

[1060] >2plus1_GFP11_key_Cterm_106(SEQ ID NO:26847)

[1061] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1062] >2plus1_GFP11_key_Cterm_107(SEQ ID NO:26848)

[1063] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1064] >2plus1_GFP11_key_Cterm_108(SEQ ID NO:26849)

[1065] EAVRRLKEILERLKEEVRRSLEELRKEVERLKKEVEDSLRELKKSLEEWVKSLEEATR

[1066] >2plus1_GFP11_key_Cterm_109(SEQ ID NO:26850)

[1067] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[1068] >2plus1_GFP11_key_Cterm_110(SEQ ID NO:26851)

[1069] DATRVIEEAKRILDEARKLNEETIRRSEELVRRIERVIEEIIKRSEKLLEDVARESK

[1070] >2plus1_GFP11_key_Cterm_111(SEQ ID NO:26852)

[1071] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1072] >2plus1_GFP11_key_Cterm_112(SEQ ID NO:26853)

[1073] EKLKELRDVIAEVAKRIDELDEYTRESIRRAKKEIERLNRETKKVIEEVVKRIEEERK

[1074] >2plus1_GFP11_key_Cterm_113(SEQ ID NO:26854)

[1075] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1076] >2plus1_GFP11_key_Cterm_114(SEQ ID NO:26855)

[1077] DKVERVVREVEKLHEEDRKRLEESTRSVRKLLEELKRELEKSTRSVKALVDELRERVR

[1078] >2plus1_GFP11_key_Cterm_115(SEQ ID NO:26856)

[1079] DEVVERAERISEENKRRVEDVARKSKELVEDVRRHSEEVVRRVEELVKEVEERVR

[1080] >2plus1_GFP11_key_Cterm_116(SEQ ID NO:26857)

[1081] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1082] >2plus1_GFP11_key_Cterm_117(SEQ ID NO:26858)

[1083] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1084] >2plus1_GFP11_key_Cterm_118(SEQ ID NO:26859)

[1085] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[1086] >2plus1_GFP11_key_Cterm_119(SEQ ID NO:26860)

[1087] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[1088] >2plus1_GFP11_key_Cterm_120(SEQ ID NO:26861)

[1089] DLKRVEERAREVSRRNEESMRRVKEDADRVSEANKEVLDRVREEVKRLIEEVRETLR

[1090] >2plus1_GFP11_key_Cterm_121(SEQ ID NO:26862)

[1091] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1092] >2plus1_GFP11_key_Cterm_122(SEQ ID NO:26863)

[1093] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1094] >2plus1_GFP11_key_Cterm_123(SEQ ID NO:26864)

[1095] SKAIKDVRDIVKKVKDELKEWRDRNKELVDRLSEELKEWLKDVERVLKELTDKDR

[1096] >2plus1_GFP11_key_Cterm_124(SEQ ID NO:26865)

[1097] DKVERVVREVEKLHEEDRKRLEESTRSVRKLLEELKRELEKSTRSVKALVDELRERVR

[1098] >2plus1_GFP11_key_Cterm_125(SEQ ID NO:26866)

[1099] DIDKLLKELRDLVEKIKKDLKELLERYEEIVRRIKELLKDLNREAEEVVRRLKEELR

[1100] >2plus1_GFP11_key_Cterm_126(SEQ ID NO:26867)

[1101] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1102] >2plus1_GFP11_key_Cterm_127(SEQ ID NO:26868)

[1103] DKLLKEARDLIREIEKRLEELLKRVEKLTEDAKRDLERSNREHKELADRIKETAR

[1104] >2plus1_GFP11_key_Cterm_128(SEQ ID NO:26869)

[1105] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1106] >2plus1_GFP11_key_Cterm_129(SEQ ID NO:26870)

[1107] DEIKRIVDEVRERLKRIVDENAKIVEDARRALEKIVKENEEILRRLKKELRELRK

[1108] >2plus1_GFP11_key_Cterm_130(SEQ ID NO:26871)

[1109] DRIEEELKRLIDTLREKNREVEKRARDSNRDLKRTNDEIAKEVRELIKKLREDLK

[1110] >2plus1_GFP11_key_Cterm_131(SEQ ID NO:26872)

[1111] DERILRELEERVKELEKEAREILKRSEDETDKLREKAERILEDLERANRRTMDEARR

[1112] >2plus1_GFP11_key_Cterm_132(SEQ ID NO:26873)

[1113] DAETIERVVRELLEENKEVLRKTEEAVKRSTETNKRLLEASKEVADRLRERIKEAAK

[1114] >2plus1_GFP11_key_Cterm_133(SEQ ID NO:26874)

[1115] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1116] >2plus1_GFP11_key_Cterm_134(SEQ ID NO:26875)

[1117] DKVERVVREVEKLHEEDRKRLEESTRSVRKLLEELKRELEKSTRSVKALVDELRERVR

[1118] >2plus1_GFP11_key_Cterm_135(SEQ ID NO:26876)

[1119] DIERILRELEAVLKKLTDESERLNREVERVSRDTKKKSKELNEELKAVLDEVKRKAD

[1120] >2plus1_GFP11_key_Cterm_136(SEQ ID NO:26877)

[1121] DEVVERAERISEENKRRVEDVARKSKELVEDVRRHSEEVVRRVEELVKEVEERVR

[1122] >2plus1_GFP11_key_Cterm_137(SEQ ID NO:26878)

[1123] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1124] >2plus1_GFP11_key_Cterm_138(SEQ ID NO:26879)

[1125] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1126] >2plus1_GFP11_key_Cterm_139(SEQ ID NO:26880)

[1127] DEIKRIVDEVRERLKRIVDENAKIVEDARRALEKIVKENEEILRRLKKELRELRK

[1128] >2plus1_GFP11_key_Cterm_140(SEQ ID NO:26881)

[1129] DRRIEKVLKEIEEKIREVIKEWERVHREVEELLKRLIDENRKVLDEIRKLLEEKSK

[1130] >2plus1_GFP11_key_Cterm_141(SEQ ID NO:26882)

[1131] DLKRVEERAREVSRRNEESMRRVKEDADRVSEANKEVLDRVREEVKRLIEEVRETLR

[1132] >2plus1_GFP11_key_Cterm_142(SEQ ID NO:26883)

[1133] DAETIERVVRELLEENKEVLRKTEEAVKRSTETNKRLLEASKEVADRLRERIKEAAK

[1134] >2plus1_GFP11_key_Cterm_143(SEQ ID NO:26884)

[1135] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1136] >2plus1_GFP11_key_Cterm_144(SEQ ID NO:26885)

[1137] ELLRRIKKLLDEIKKAIEDSSREIKRLLEESERVMKRSSEDIKRTLDDTRRVVEEVRR

[1138] >2plus1_GFP11_key_Cterm_145(SEQ ID NO:26886)

[1139] DKVERVVREVEKLHEEDRKRLEESTRSVRKLLEELKRELEKSTRSVKALVDELRERVR

[1140] >2plus1_GFP11_key_Cterm_146(SEQ ID NO:26887)

[1141] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1142] >2plus1_GFP11_key_Cterm_147(SEQ ID NO:26888)

[1143] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1144] >2plus1_GFP11_key_Cterm_148(SEQ ID NO:26889)

[1145] DEIKRIVDEVRERLKRIVDENAKIVEDARRALEKIVKENEEILRRLKKELRELRK

[1146] >2plus1_GFP11_key_Cterm_149(SEQ ID NO:26890)

[1147] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1148] >2plus1_GFP11_key_Cterm_150(SEQ ID NO:26891)

[1149] DEVTKVKKVADDVLAEIKKLDDETRRVIEDTNKKIADLDKATRDVVRKVLEEVKKLEK

[1150] >2plus1_GFP11_key_Cterm_151(SEQ ID NO:26892)

[1151] DIDKLLKELRDLVEKIKKDLKELLERYEEIVRRIKELLKDLNREAEEVVRRLKEELR

[1152] >2plus1_GFP11_key_Cterm_152(SEQ ID NO:26893)

[1153] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1154] >2plus1_GFP11_key_Cterm_153(SEQ ID NO:26894)

[1155] DEIKRIVDEVRERLKRIVDENAKIVEDARRALEKIVKENEEILRRLKKELRELRK

[1156] >2plus1_GFP11_key_Cterm_154(SEQ ID NO:26895)

[1157] RLVREVEDLVRRLVRRSEKSNEEVKRTVEELVRRMEESNDRVRDLVRRLVEELKRAVD

[1158] >2plus1_GFP11_key_Cterm_155(SEQ ID NO:26896)

[1159] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1160] >2plus1_GFP11_key_Cterm_156(SEQ ID NO:26897)

[1161] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1162] >2plus1_GFP11_key_Cterm_157(SEQ ID NO:26898)

[1163] RLVREVEDLVRRLVRRSEKSNEEVKRTVEELVRRMEESNDRVRDLVRRLVEELKRAVD

[1164] >2plus1_GFP11_key_Cterm_158(SEQ ID NO:26899)

[1165] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1166] >2plus1_GFP11_key_Cterm_159(SEQ ID NO:26900)

[1167] DEVTKVKKVADDVLAEIKKLDDETRRVIEDTNKKIADLDKATRDVVRKVLEEVKKLEK

[1168] >2plus1_GFP11_key_Cterm_160(SEQ ID NO:26901)

[1169] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1170] >2plus1_GFP11_key_Cterm_161(SEQ ID NO:26902)

[1171] DLKRVEERAREVSRRNEESMRRVKEDADRVSEANKEVLDRVREEVKRLIEEVRETLR

[1172] >2plus1_GFP11_key_Cterm_162(SEQ ID NO:26903)

[1173] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1174] >2plus1_GFP11_key_Cterm_163(SEQ ID NO:26904)

[1175] DEVTKVKKVADDVLAEIKKLDDETRRVIEDTNKKIADLDKATRDVVRKVLEEVKKLEK

[1176] >2plus1_GFP11_key_Cterm_164(SEQ ID NO:26905)

[1177] DKVERVVREVEKLHEEDRKRLEESTRSVRKLLEELKRELEKSTRSVKALVDELRERVR

[1178] >2plus1_GFP11_key_Cterm_165(SEQ ID NO:26906)

[1179] EAKKKLDEVLERAKRTIDRLLETSDRSLEKVEADLRRLNEELDRSLERAERTIRELAK

[1180] >2plus1_GFP11_key_Cterm_166(SEQ ID NO:26907)

[1181] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1182] >2plus1_GFP11_key_Cterm_167(SEQ ID NO:26908)

[1183] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1184] >2plus1_GFP11_key_Cterm_168(SEQ ID NO:26909)

[1185] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1186] >2plus1_GFP11_key_Cterm_169(SEQ ID NO:26910)

[1187] DEVTKVKKVADDVLAEIKKLDDETRRVIEDTNKKIADLDKATRDVVRKVLEEVKKLEK

[1188] >2plus1_GFP11_key_Cterm_170(SEQ ID NO:26911)

[1189] DKVERVVREVEKLHEEDRKRLEESTRSVRKLLEELKRELEKSTRSVKALVDELRERVR

[1190] >2plus1_GFP11_key_Cterm_171(SEQ ID NO:26912)

[1191] TAERARETLKRLLDENRDRSKKVKEEIRRILEDLTRTTERVKREIAKLLKELEDTAR

[1192] >2plus1_GFP11_key_Cterm_172(SEQ ID NO:26913)

[1193] DKARKVAEVAEKVLRDIDKLDRESKEAFRATNEEIAKLDEDTARVAERVKKAIEDLAK

[1194] >2plus1_GFP11_key_Cterm_173(SEQ ID NO:26914)

[1195] RLVREVEDLVRRLVRRSEKSNEEVKRTVEELVRRMEESNDRVRDLVRRLVEELKRAVD

[1196] >2plus1_GFP11_key_Nterm_174(SEQ ID NO:26915)

[1197] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1198] >2plus1_GFP11_key_Nterm_175(SEQ ID NO:26916)

[1199] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1200] >2plus1_GFP11_key_Nterm_176(SEQ ID NO:26917)

[1201] SVDEVLKEIEDALRRLKEEVERVLKENEDELRRLEEEVRRVLKEDEELLESLKRGVGE

[1202] >2plus1_GFP11_key_Nterm_177(SEQ ID NO:26918)

[1203] SEVDEIIKELERLLAEIARENERIIRESRKLADEVRKRNEDAIRKLEELVARLADAVR

[1204] >2plus1_GFP11_key_Nterm_178(SEQ ID NO:26919)

[1205] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1206] >2plus1_GFP11_key_Nterm_179(SEQ ID NO:26920)

[1207] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1208] >2plus1_GFP11_key_Nterm_180(SEQ ID NO:26921)

[1209] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1210] >2plus1_GFP11_key_Nterm_181(SEQ ID NO:26922)

[1211] KEVEDAVKELEDLLRANEDKTRSIVEDMRASNKDLEDHSRASEEEVRKLLDDLRRAGV

[1212] >2plus1_GFP11_key_Nterm_182(SEQ ID NO:26923)

[1213] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1214] >2plus1_GFP11_key_Nterm_183(SEQ ID NO:26924)

[1215] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1216] >2plus1_GFP11_key_Nterm_184(SEQ ID NO:26925)

[1217] SESDDVIRKLRELLEELRTHVEKSIRDLRKILEDSTRHAKRSIEELERLLEEVRKKPG

[1218] >2plus1_GFP11_key_Nterm_185(SEQ ID NO:26926)

[1219] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1220] >2plus1_GFP11_key_Nterm_186(SEQ ID NO:26927)

[1221] SEAEKAKETIDRLADRVRKLLEEIKRSLDDSRRKSKETVEENEKTLDRMRKEVDAAKR

[1222] >2plus1_GFP11_key_Nterm_187(SEQ ID NO:26928)

[1223] SEAEKAKETIDRLADRVRKLLEEIKRSLDDSRRKSKETVEENEKTLDRMRKEVDAAKR

[1224] >2plus1_GFP11_key_Nterm_188(SEQ ID NO:26929)

[1225] SEVEELIKRLAKVLKELVDKVRKVIEDTKELLERLKRRSEDHIRKLREVLKEAKDQPI

[1226] >2plus1_GFP11_key_Nterm_189(SEQ ID NO:26930)

[1227] SELEEIEKKVRELTKRHRELVERVRKTVKELIETNRRLLETLTERIKRVLEEVRDLER

[1228] >2plus1_GFP11_key_Nterm_190(SEQ ID NO:26931)

[1229] SSEERLRAVIEDLKRLAEESRKRHKELIDELAKAVERIERRHKKLLDEIKAVVDDIRR

[1230] >2plus1_GFP11_key_Nterm_191(SEQ ID NO:26932)

[1231] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1232] >2plus1_GFP11_key_Nterm_192(SEQ ID NO:26933)

[1233] STAETVAEEVERVLKHSDDLIKEVEDVNRRVEEEIKRVIRELEEENERLVAEVRKGVK

[1234] >2plus1_GFP11_key_Nterm_193(SEQ ID NO:26934)

[1235] SEVDEIIKELERLLAEIARENERIIRESRKLADEVRKRNEDAIRKLEELVARLADAVR

[1236] >2plus1_GFP11_key_Nterm_194(SEQ ID NO:26935)

[1237] SEIDEVLTRLRKISKDLNETSDRVNERARKIIDDIKKESKRVNDEAREIVERLKREID

[1238] >2plus1_GFP11_key_Nterm_195(SEQ ID NO:26936)

[1239] SEDEDLDRVAEKLAREHKKSVEEIKRVLKSADEESKKLVRDTERVIEEIKREVEEARR

[1240] >2plus1_GFP11_key_Nterm_196(SEQ ID NO:26937)

[1241] SSVEELLERLRRISEENKRRIEKLLREVEKVLRELKDRHRKLLKRVEEIIRKVKEEIK

[1242] >2plus1_GFP11_key_Nterm_197(SEQ ID NO:26938)

[1243] SAADEVVERMKELVATVKRENDEVVKELKKLVKELEDDNRRVVEESKKSVEDLARRVG

[1244] >2plus1_GFP11_key_Nterm_198(SEQ ID NO:26939)

[1245] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1246] >2plus1_GFP11_key_Nterm_199(SEQ ID NO:26940)

[1247] KEVEDAVKELEDLLRANEDKTRSIVEDMRASNKDLEDHSRASEEEVRKLLDDLRRAGV

[1248] >2plus1_GFP11_key_Nterm_200(SEQ ID NO:26941)

[1249] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1250] >2plus1_GFP11_key_Nterm_201(SEQ ID NO:26942)

[1251] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1252] >2plus1_GFP11_key_Nterm_202(SEQ ID NO:26943)

[1253] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1254] >2plus1_GFP11_key_Nterm_203(SEQ ID NO:26944)

[1255] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1256] >2plus1_GFP11_key_Nterm_204(SEQ ID NO:26945)

[1257] SKLEEVEKAVRKVIEDSRRVNEEVNRRSEEVVRELEKVHREVNDASRRVVEKARRVLK

[1258] >2plus1_GFP11_key_Nterm_205(SEQ ID NO:26946)

[1259] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1260] >2plus1_GFP11_key_Nterm_206(SEQ ID NO:26947)

[1261] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1262] >2plus1_GFP11_key_Nterm_207(SEQ ID NO:26948)

[1263] SESDDVIRKLRELLEELRTHVEKSIRDLRKILEDSTRHAKRSIEELERLLEEVRKKPG

[1264] >2plus1_GFP11_key_Nterm_208(SEQ ID NO:26949)

[1265] DEVRELLERNRRLLEEIKKTVKDLIRANEELLKRIEDDAKRLIDRNEELLDELEKGLS

[1266] >2plus1_GFP11_key_Nterm_209(SEQ ID NO:26950)

[1267] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1268] >2plus1_GFP11_key_Nterm_210(SEQ ID NO:26951)

[1269] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1270] >2plus1_GFP11_key_Nterm_211(SEQ ID NO:26952)

[1271] DEEEDLERAIKKLLDENRELLKRIAEELRRLLEELRRLTEESADRLRRLLKELKDRGV

[1272] >2plus1_GFP11_key_Nterm_212(SEQ ID NO:26953)

[1273] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1274] >2plus1_GFP11_key_Nterm_213(SEQ ID NO:26954)

[1275] SKEDRLREELKKLLARLAEEIERLKRALEESNKDLKRTIDASEKHLRDVNEDVKRGGV

[1276] >2plus1_GFP11_key_Nterm_214(SEQ ID NO:26955)

[1277] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1278] >2plus1_GFP11_key_Nterm_215(SEQ ID NO:26956)

[1279] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1280] >2plus1_GFP11_key_Nterm_216(SEQ ID NO:26957)

[1281] SESDDVIRKLRELLEELRTHVEKSIRDLRKILEDSTRHAKRSIEELERLLEEVRKKPG

[1282] >2plus1_GFP11_key_Nterm_217(SEQ ID NO:26958)

[1283] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1284] >2plus1_GFP11_key_Nterm_218(SEQ ID NO:26959)

[1285] SEVDEIIKELERLLAEIARENERIIRESRKLADEVRKRNEDAIRKLEELVARLADAVR

[1286] >2plus1_GFP11_key_Nterm_219(SEQ ID NO:26960)

[1287] SSVEELLERLRRISEENKRRIEKLLREVEKVLRELKDRHRKLLKRVEEIIRKVKEEIK

[1288] >2plus1_GFP11_key_Nterm_220(SEQ ID NO:26961)

[1289] SELEEVLRRIEALVRKAHKENEDVLREIERLVRTAHRLNKKVDDDSAKIAEDLKRGGR

[1290] >2plus1_GFP11_key_Nterm_221(SEQ ID NO:26962)

[1291] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1292] >2plus1_GFP11_key_Nterm_222(SEQ ID NO:26963)

[1293] SESDDVIRKLRELLEELRTHVEKSIRDLRKILEDSTRHAKRSIEELERLLEEVRKKPG

[1294] >2plus1_GFP11_key_Nterm_223(SEQ ID NO:26964)

[1295] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1296] >2plus1_GFP11_key_Nterm_224(SEQ ID NO:26965)

[1297] DEVRELLERNRRLLEEIKKTVKDLIRANEELLKRIEDDAKRLIDRNEELLDELEKGLS

[1298] >2plus1_GFP11_key_Nterm_225(SEQ ID NO:26966)

[1299] STEEVLDEIRKLHKTLTEDIKRVLREIEELHRRTIEENKEVLDKIAEDYKRVIDDVRT

[1300] >2plus1_GFP11_key_Nterm_226(SEQ ID NO:26967)

[1301] SEIEKILKEIEDLARRDEEVSKKIVEDIRRLAKEVEDTSRDIVRKIEELAKRVLDRLR

[1302] >2plus1_GFP11_key_Nterm_227(SEQ ID NO:26968)

[1303] SEAERLEARARELLRANEELMDDLRAKAEELLKRNDRLVKEIEKKVREVLAAIEELKR

[1304] >2plus1_GFP11_key_Nterm_228(SEQ ID NO:26969)

[1305] DDLERAREEVADLIRKHEEKTRRILEESRRLNERHRELSARILDEIRKLAERIEELIK

[1306] >2plus1_GFP11_key_Nterm_229(SEQ ID NO:26970)

[1307] DEEEDLERAIKKLLDENRELLKRIAEELRRLLEELRRLTEESADRLRRLLKELKDRGV

[1308] >2plus1_GFP11_key_Nterm_230(SEQ ID NO:26971)

[1309] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1310] >2plus1_GFP11_key_Nterm_231(SEQ ID NO:26972)

[1311] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1312] >2plus1_GFP11_key_Nterm_232(SEQ ID NO:26973)

[1313] STAETVEKKVEEVIRENEKSMRESEEKVDRSTKRIEDVLRRLEETIRKTSDDIAKGVK

[1314] >2plus1_GFP11_key_Nterm_233(SEQ ID NO:26974)

[1315] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1316] >2plus1_GFP11_key_Nterm_234(SEQ ID NO:26975)

[1317] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1318] >2plus1_GFP11_key_Nterm_235(SEQ ID NO:26976)

[1319] REVEEMIKELEELLKDLKEKNERASKRNRELVRRLEEENKRVIEELKKLVKELEDLVR

[1320] >2plus1_GFP11_key_Nterm_236(SEQ ID NO:26977)

[1321] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1322] >2plus1_GFP11_key_Nterm_237(SEQ ID NO:26978)

[1323] SKLEEVEKAVRKVIEDSRRVNEEVNRRSEEVVRELEKVHREVNDASRRVVEKARRVLK

[1324] >2plus1_GFP11_key_Nterm_238(SEQ ID NO:26979)

[1325] DEVEDVLRKIEKILDDHRKRIEKNSRDMARIIDEHRRKVEENSREMKKLVDDLKKAVD

[1326] >2plus1_GFP11_key_Nterm_239(SEQ ID NO:26980)

[1327] SSVEELLERLRRISEENKRRIEKLLREVEKVLRELKDRHRKLLKRVEEIIRKVKEEIK

[1328] >2plus1_GFP11_key_Nterm_240(SEQ ID NO:26981)

[1329] DEVEKVLEEIKRALDDLRKKVEESKREIKEALKAVEKHTRDSDTANKRTLAEIERGVK

[1330] >2plus1_GFP11_key_Nterm_241(SEQ ID NO:26982)

[1331] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1332] >2plus1_GFP11_key_Nterm_242(SEQ ID NO:26983)

[1333] KEVEDAVKELEDLLRANEDKTRSIVEDMRASNKDLEDHSRASEEEVRKLLDDLRRAGV

[1334] >2plus1_GFP11_key_Nterm_243(SEQ ID NO:26984)

[1335] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1336] >2plus1_GFP11_key_Nterm_244(SEQ ID NO:26985)

[1337] STAETVAEEVERVLKHSDDLIKEVEDVNRRVEEEIKRVIRELEEENERLVAEVRKGVK

[1338] >2plus1_GFP11_key_Nterm_245(SEQ ID NO:26986)

[1339] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1340] >2plus1_GFP11_key_Nterm_246(SEQ ID NO:26987)

[1341] SEVDEIIKELERLLAEIARENERIIRESRKLADEVRKRNEDAIRKLEELVARLADAVR

[1342] >2plus1_GFP11_key_Nterm_247(SEQ ID NO:26988)

[1343] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1344] >2plus1_GFP11_key_Nterm_248(SEQ ID NO:26989)

[1345] SKLEEVEKAVRKVIEDSRRVNEEVNRRSEEVVRELEKVHREVNDASRRVVEKARRVLK

[1346] >2plus1_GFP11_key_Nterm_249(SEQ ID NO:26990)

[1347] SAEEVKEELKRIATKLKEEIKENIRRLEESVEKIAKELAENIKRLEDILRDVKRGLRD

[1348] >2plus1_GFP11_key_Nterm_250(SEQ ID NO:26991)

[1349] SDVDRVLEEIRKLLEDLKRHSEKVSEENEDLLRANTELNKRVSEDNERLLEELKRLRE

[1350] >2plus1_GFP11_key_Nterm_251(SEQ ID NO:26992)

[1351] DEVEDVLRKIEKILDDHRKRIEKNSRDMARIIDEHRRKVEENSREMKKLVDDLKKAVD

[1352] >2plus1_GFP11_key_Nterm_252(SEQ ID NO:26993)

[1353] SSVEELLERLRRISEENKRRIEKLLREVEKVLRELKDRHRKLLKRVEEIIRKVKEEIK

[1354] >2plus1_GFP11_key_Nterm_253(SEQ ID NO:26994)

[1355] DREREVKKRLDEVRERIERLLRRVEEESRRVAEEIRRLIEEVRRRNKKVTEEIRELLK

[1356] >2plus1_GFP11_key_Nterm_254(SEQ ID NO:26995)

[1357] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1358] >2plus1_GFP11_key_Nterm_255(SEQ ID NO:26996)

[1359] SEIEKILKEIEDLARRDEEVSKKIVEDIRRLAKEVEDTSRDIVRKIEELAKRVLDRLR

[1360] >2plus1_GFP11_key_Nterm_256(SEQ ID NO:26997)

[1361] SEAEKAKETIDRLADRVRKLLEEIKRSLDDSRRKSKETVEENEKTLDRMRKEVDAAKR

[1362] >2plus1_GFP11_key_Nterm_257(SEQ ID NO:26998)

[1363] SEVEELIKRLAKVLKELVDKVRKVIEDTKELLERLKRRSEDHIRKLREVLKEAKDQPI

[1364] >2plus1_GFP11_key_Nterm_258(SEQ ID NO:26999)

[1365] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1366] >2plus1_GFP11_key_Nterm_259(SEQ ID NO:27000)

[1367] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1368] >2plus1_GFP11_key_Nterm_260(SEQ ID NO:27001)

[1369] STAETVEKKVEEVIRENEKSMRESEEKVDRSTKRIEDVLRRLEETIRKTSDDIAKGVK

[1370] >2plus1_GFP11_key_Nterm_261(SEQ ID NO:27002)

[1371] SRVEEIIEDLRRLLEEIRKENEDSIRRSKELLDRVKEINDTIIAELERLLKDIEKEVR

[1372] >2plus1_GFP11_key_Nterm_262(SEQ ID NO:27003)

[1373] REVEEMIKELEELLKDLKEKNERASKRNRELVRRLEEENKRVIEELKKLVKELEDLVR

[1374] >2plus1_GFP11_key_Nterm_263(SEQ ID NO:27004)

[1375] DAVEEAEKLIRKVIADSEKLLRDLADLNAKSIRRSEKLVEDDRRANEDVIRKLEELRR

[1376] >2plus1_GFP11_key_Nterm_264(SEQ ID NO:27004)

[1377] SEVDDVLRRLEELIKTLEDINAKSLEDIKKLIDDLAKILEDALRKHEKLIRELREAKK

[1378] >2plus1_GFP11_key_Nterm_265(SEQ ID NO:27006)

[1379] SEIERVKKRLEELLAEVEESTRRLEERLKRLLEEAKRSSEEVEKELRRLLEAVRRGLS

[1380] >2plus1_GFP11_key_Nterm_266(SEQ ID NO:27007)

[1381] SDVDRVLEEIRKLLEDLKRHSEKVSEENEDLLRANTELNKRVSEDNERLLEELKRLRE

[1382] >2plus1_GFP11_key_Nterm_267(SEQ ID NO:27008)

[1383] DEVEDVLRKIEKILDDHRKRIEKNSRDMARIIDEHRRKVEENSREMKKLVDDLKKAVD

[1384] >2plus1_GFP11_key_Nterm_268(SEQ ID NO:27009)

[1385] SESDEVIRDLARLLDELARHVDDSVRRMDEVVKRSTREADELAKRLDELVKEVEKKPG

[1386] >2plus1_GFP11_key_Nterm_269(SEQ ID NO:27010)

[1387] SRAETVLKEVTDKIKKLADSSDELLRRNKENIDELKKSSEELLRRLTKAIEEIEKGSV

[1388] >2plus1_GFP11_key_Nterm_270(SEQ ID NO:27011)

[1389] DEEEDLERAIKKLLDENRELLKRIAEELRRLLEELRRLTEESADRLRRLLKELKDRGV

[1390] >2plus1_GFP11_key_Nterm_271(SEQ ID NO:27012)

[1391] DEIRKVVKEITDLLKASNDKNRKVVEEIRDLLRKSKKLADELVERLRALVEDLRRRID

[1392] >2plus1_GFP11_key_Nterm_272(SEQ ID NO:27013)

[1393] DAVEEAEKLIRKVIADSEKLLRDLADLNAKSIRRSEKLVEDDRRANEDVIRKLEELRR

[1394] >2plus1_GFP11_key_Nterm_273(SEQ ID NO:27014)

[1395] SEDEDLDRVAEKLAREHKKSVEEIKRVLKSADEESKKLVRDTERVIEEIKREVEEARR

[1396] >2plus1_GFP11_key_Nterm_274(SEQ ID NO:27015)

[1397] SKEDRLREELKKLLARLAEEIERLKRALEESNKDLKRTIDASEKHLRDVNEDVKRGGV

[1398] >3plus1_key_668_Nterm(SEQ ID NO:27322)

[1399] DEAKELLDEIRKAVKESEDRLEKLLRDYEKELRRLEKELRDLKRRIEEKLEELRRGSL

[1400] >3plus1_key_668_Cterm(SEQ ID NO:27323)

[1401] RGADALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[1402] >3plus1_key_668_Cterm(SEQ ID NO:27324)

[1403] SEKEKLLKESEEEVRRLRRTLEELLRKYREVLERLRKELREIEERVRDVVRRLKEVLD

[1404] >3plus1_key_668_Cterm(SEQ ID NO:27325)

[1405] SEKEDAARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLAEISEEVR

[1406] >3plus1_key_668_Cterm(SEQ ID NO:27326)

[1407] EALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[1408] >3plus1_key_669_Nterm(SEQ ID NO:27327)

[1409] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[1410] >3plus1_key_670_Nterm(SEQ ID NO:27328)

[1411] SDERRIAERIRELLRESKKLVRDVVEEAKRLLKENRDSTRKIIEDIRRLLRKIEDSTR

[1412] >3plus1_key_670_Cterm(SEQ ID NO:27329)

[1413] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[1414] >3plus1_key_670_Cterm(SEQ ID NO:27330)

[1415] AAKRLVEELLKAVTDLSRKNKRILEELLKAIETLSDENKKILTEILDALRRLVEKIEK

[1416] >3plus1_key_670_Cterm(SEQ ID NO:27331)

[1417] KEREEVKEKLDRLLEEVEKTVRELKREHDELLKEVEKLVRDLKKEHDELLKKVKDDGV

[1418] >3plus1_key_670_Nterm(SEQ ID NO:27332)

[1419] DRLDKVEELVKKLLEDTKRTVDRVRELVRKILKKSRETLEELERLIEKILRELEKDAR

[1420] >3plus1_key_670_Cterm(SEQ ID NO:27333)

[1421] SERETVKRRLEELLKEVKRTLDKLKEEHDRLLEDVRRVVEELKREHDKLLKEVKDSGV

[1422] >3plus1_key_670_Nterm(SEQ ID NO:27334)

[1423] SEDEIIKKIIEDLRRVLKEVEEIHKEVEERLDKVLKEAEEMHKEVLKELDRVLDEVKR

[1424] >3plus1_key_670_Nterm(SEQ ID NO:27335)

[1425] SREEVLRELEEVIEDNRRLLEELIEKSKKVLDESLKLIDELLRRLEEVLERVLRLLEE

[1426] >3plus1_key_670_Nterm(SEQ ID NO:27336)

[1427] ISEDDLKRVVDEVEKKLRELKRRYAEALERIKEKIKELKDRYERAVREVVAELRKTTK

[1428] >3plus1_key_670_Nterm(SEQ ID NO:27337)

[1429] SKAEEIAEKLDRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGI

[1430] >3plus1_key_671_Cterm(SEQ ID NO:27338)

[1431] VDSERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[1432] >3plus1_key_671_Cterm(SEQ ID NO:27339)

[1433] VKDDEVEREIRRVKEDLDRILEEYRRLLEEIKRKLEEILRRVEELHRRLRRKLEEIDR

[1434] >3plus1_key_671_Cterm(SEQ ID NO:27340)

[1435] SRVEELKKLIEDILRISREVVERIKRVAEDIHRINRRVLDDLRKLIEDILRTVEEILA

[1436] >3plus1_key_671_Cterm(SEQ ID NO:27341)

[1437] RGADALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[1438] >3plus1_key_672_Cterm(SEQ ID NO:27342)

[1439] RGADALSRLLEELLRVVDDLIRVLKELIDKSRKVIEELLELLKRINEENLKVLAEIIK

[1440] >3plus1_key_67>3_Nterm(SEQ ID NO:27343)

[1441] EALRKLVELLVEVLRRLIRVNRELVKLLREVLERLLRILRESVKKLKRLIEKVIKDAT

[1442] >3plus1_key_67>3_Cterm(SEQ ID NO:27344)

[1443] SEKEDAARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLAEISEEVR

[1444] >3plus1_key_67>3_Nterm(SEQ ID NO:27345)

[1445] SEKEDAARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLAEISEEVR

[1446] >3plus1_key_674_Nterm(SEQ ID NO:27346)

[1447] SEKEDAARKLRKLVEELTREYEELVKKLERLIEEIEKVSEESVRKLEKLLAEISEEVR

[1448] >3plus1_key_674_Cterm(SEQ ID NO:27347)

[1449] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[1450] >3plus1_key_675_Nterm(SEQ ID NO:27348)

[1451] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[1452] >3plus1_key_676_Nterm(SEQ ID NO:27349)

[1453] RAVKKLDEIVKEVAKKLEDVVRANEELWRALVELNKESVRRLREIVERVARDLEETAR

[1454] >3plus1_key_677_Nterm(SEQ ID NO:27350)

[1455] SDERRIAERIRELLRESKKLVRDVVEEAKRLLKENRDSTRKIIEDIRRLLRKIEDSTR

[1456] >3plus1_key_677_Cterm(SEQ ID NO:27351)

[1457] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[1458] >3plus1_key_678_Nterm(SEQ ID NO:27352)

[1459] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[1460] >3plus1_key_678_Cterm(SEQ ID NO:27353)

[1461] SKEETLRKEAEDLLRRLEELTRRLEKKARELLERAKKLSRDLAEELKRLLKELREKGV

[1462] >3plus1_key_678_Cterm(SEQ ID NO:27354)

[1463] ISEDDLKRVVDEVEKKLRELKRRYAEALERIKEKIKELKDRYERAVREVVAELRKTTK

[1464] >3plus1_key_678_Nterm(SEQ ID NO:27355)

[1465] SKAEEIAEKLDRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGI

[1466] >3plus1_key_678_Nterm(SEQ ID NO:27356)

[1467] VDSERVKEILERILRVVEEAVRLNEESLRRILDVVRKAVKLDRESLKKILDVVEEAVR

[1468] >3plus1_key_679_Cterm(SEQ ID NO:27357)

[1469] SKAEEIAEKLDRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGI

[1470] >3plus1_key_679_Nterm(SEQ ID NO:27358)

[1471] SKAEEIAEKLDRLLEENRRALEEITTRLDDLLRRNKDALRKVMEKLKRLLDDLRRGGI

[1472] In a specific embodiment, the key polypeptide shares 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length with the amino acid sequence of a key polypeptide in Table 2 (polypeptides having an odd number of SEQ ID NOs between SEQ ID NOs: 27127 and 27277), Table 3 and / or Table 4. In another specific embodiment, the key polypeptide shares 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length with the amino acid sequence of a key polypeptide in Table 3. In another specific embodiment, the key polypeptide shares 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along its length with the amino acid sequence of a key polypeptide in Table 4. In one embodiment of each of the above, the percent identification can be determined without the optional N-terminal and C-terminal 60 amino acids; in another embodiment, the percent identification can be determined with the optional N- and C-terminal 60 amino acids.

[1473] The polypeptides of the present disclosure (i.e., cage polypeptides and key polypeptides) may include additional residues at the N-terminus, C-terminus, internally, or a combination thereof; these additional residues are not included in determining the percent identity of a polypeptide of the present invention relative to a reference polypeptide. Such residues may be any residue suitable for the intended use, including but not limited to tags. As used herein, "tag" includes general detectable moieties (i.e., fluorescent proteins, antibody epitope tags, etc.), therapeutic agents, purification tags (His tags, etc.), linkers, ligands suitable for purification purposes, ligands that drive polypeptide localization, peptide domains that add functionality to a polypeptide, and the like. Examples are provided herein.

[1474] In one embodiment, the polypeptide is a fusion protein comprising a cage polypeptide disclosed herein fused to a key polypeptide disclosed herein. In one embodiment, the fusion protein comprises a cage polypeptide fused to a key polypeptide, wherein the cage polypeptide is not activated by the key polypeptide. As noted herein, the orthogonal LOCKR design (see Figure 3 ) is represented by a lowercase subscript: LOCKR a By Cage a and keys a Composition, and LOCKR b By Cage b and keys b Composition, etc., so that the cage a Only by key a Activate and cage b Only by key b activation, etc. Thus, for example, a fusion protein may comprise b Peptide-fused cage a For example, such embodiments can be used in combination to improve control of orthogonal LOCKR designs (e.g., LOCKR 1 contains a cage a -key b Fusion polypeptide, and LOCKR 2 contains a cage b -key a fusion polypeptides, which can then be expressed in the same cell)

[1475] In one embodiment of the fusion proteins disclosed herein, the cage polypeptide and key polypeptide components of the fusion protein comprise at least one cage polypeptide and at least one key polypeptide whose amino acid sequences have at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along their length to the cage polypeptide and key polypeptide, respectively, in different rows of Table 1, Table 2, Table 3 and / or Table 4 (i.e., each cage polypeptide in row 1, column 1 of the table can be fused to any key polypeptide in row 1, column 2, and so on).

[1476] Table 1

[1477]

[1478]

[1479]

[1480]

[1481]

[1482]

[1483]

[1484]

[1485]

[1486]

[1487]

[1488]

[1489]

[1490]

[1491]

[1492]

[1493]

[1494]

[1495]

[1496]

[1497]

[1498]

[1499]

[1500]

[1501]

[1502] As used throughout this application, the term "polypeptide" is used in its broadest sense to refer to a subunit amino acid sequence. The polypeptides of the present invention may comprise L-amino acids + glycine, D-amino acids + glycine (which are resistant to L-amino acid specific proteases in vivo), or a combination of D- and L-amino acids + glycine. The polypeptides described herein may be chemically synthesized or recombinantly expressed. The polypeptides may be linked to other compounds to promote an increase in half-life in vivo, such as by pegylation, glycosylation, phosphorylation, glycosylation, or may be produced as Fc fusions or deimmunized variants. As will be appreciated by those skilled in the art, such linkages may be covalent or non-covalent.

[1503] In a fifth aspect, the present disclosure provides nucleic acids encoding polypeptides of any embodiment or combination of embodiments of the various aspects disclosed herein. The nucleic acid sequence may comprise single-stranded or double-stranded RNA or DNA, or DNA-RNA hybrids in the form of genomic or cDNA, each of which may include chemically or biochemically modified, non-natural or derived nucleotide bases. Such nucleic acid sequences may comprise additional sequences for promoting expression and / or purification of the encoded polypeptide, including but not limited to polyA sequences, modified Kozak sequences, and sequences encoding epitope tags, export and secretion signals, nuclear localization signals, and plasma membrane localization signals. Based on the teachings herein, it will be clear to those skilled in the art which nucleic acid sequence will encode the polypeptide of the present disclosure.

[1504] In a sixth aspect, the present disclosure provides an expression vector comprising a nucleic acid of any aspect of the present disclosure operably linked to a suitable control sequence. An "expression vector" includes a vector in which a nucleic acid coding region or gene is operably linked to any control sequence capable of achieving expression of a gene product. A "control sequence" operably linked to a nucleic acid sequence of the present disclosure is a nucleic acid sequence that can affect the expression of a nucleic acid molecule. A control sequence need not be adjacent to a nucleic acid sequence, as long as the control sequence functions to direct the expression of the nucleic acid sequence. Thus, for example, an intermediate untranslated but transcribed sequence may exist between the promoter sequence and the nucleic acid sequence, and the promoter sequence may still be considered to be "operably linked" to the coding sequence. Other such control sequences include, but are not limited to, polyadenylation signals, termination signals, and ribosome binding sites. Such expression vectors may be of any type, including but not limited to plasmids and viral-based expression vectors. The control sequences used to drive expression of the disclosed nucleic acid sequences in mammalian systems can be constitutive (driven by any of a variety of promoters, including but not limited to CMV, SV40, RSV, actin, EF) or inducible (driven by any of a number of inducible promoters, including but not limited to tetracycline, ecdysone, steroid responsiveness). The expression vector must be replicable in the host organism either as an episome or by integration into the host chromosomal DNA. In various embodiments, the expression vector can comprise a plasmid, a viral-based vector, or any other suitable expression vector.

[1505] In a seventh aspect, the present disclosure provides a host cell comprising a nucleic acid or expression vector disclosed herein (i.e., episomal or chromosomally integrated), wherein the host cell can be prokaryotic or eukaryotic. Cells can be transiently or stably engineered to incorporate the expression vector of the present disclosure using techniques including, but not limited to, bacterial transformation, calcium phosphate coprecipitation, electroporation, or liposome-mediated, DEAE dextran-mediated, polycation-mediated, or viral-mediated transfection. In one embodiment, the recombinant host cell comprises:

[1506] (a) a first nucleic acid encoding a polypeptide of any embodiment or combination of embodiments of the cage polypeptide of aspects 1 to 3 of the present disclosure operably linked to a first promoter; and

[1507] (b) a second nucleic acid encoding a polypeptide of any embodiment or combination of embodiments of the key polypeptide of the fourth aspect of the present disclosure, wherein the key polypeptide is capable of binding to a structural region of the cage polypeptide to induce a conformational change in the cage polypeptide, wherein the second nucleic acid is operably linked to the second promoter.

[1508] The recombinant host cell can contain a single cage polypeptide-encoding nucleic acid and a single key polypeptide-encoding nucleic acid, or can contain multiple (i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) first and second nucleic acids. In one such embodiment, each second nucleic acid can encode a key polypeptide capable of binding to a structural region and inducing a conformational change in a different cage polypeptide encoded by the plurality of first nucleic acids. In another embodiment, each second nucleic acid can encode a key polypeptide capable of binding to a structural region and inducing a conformational change in more than one cage polypeptide encoded by the plurality of first nucleic acids.

[1509] Thus, in one embodiment, the first nucleic acid comprises a plurality of first nucleic acids encoding a plurality of different cage polypeptides. In one such embodiment, the second nucleic acid comprises a plurality of second nucleic acids encoding a plurality of different key polypeptides, wherein the plurality of different key polypeptides comprises one or more key polypeptides capable of binding to only a subset of the plurality of different cage polypeptides and inducing a conformational change in that subset. In another such embodiment, the second nucleic acid encodes a single key polypeptide capable of binding to each of the different cage polypeptides and inducing a conformational change therein.

[1510] In another embodiment, the host cell comprises a nucleic acid encoding an expression vector capable of expressing a fusion protein disclosed herein and / or the expression vector, wherein the host cell comprises:

[1511] (a) a first nucleic acid encoding a first fusion protein (i.e., a cage polypeptide fused to a key polypeptide) linked to a first promoter; and

[1512] (b) a second nucleic acid encoding a second fusion protein operably linked to a second promoter, wherein:

[1513] (i) a cage polypeptide encoded by a first nucleic acid is activated by a key polypeptide encoded by a second nucleic acid;

[1514] (ii) the cage polypeptide encoded by the first nucleic acid is not activated by the key polypeptide encoded by the first nucleic acid;

[1515] (iii) the cage polypeptide encoded by the second nucleic acid is activated by the key polypeptide encoded by the first nucleic acid; and

[1516] (iv) the cage polypeptide encoded by the second nucleic acid is not activated by the key polypeptide encoded by the second nucleic acid.

[1517] In all of these embodiments, the first and / or second nucleic acid may, for example, be in the form of an expression vector.In other embodiments, the first and / or second nucleic acid may be in the form of a nucleic acid that is integrated into the genome of the host cell.

[1518] Methods for producing polypeptides according to the present disclosure are another aspect of the present disclosure. In one embodiment, the method comprises the steps of: (a) culturing a host according to this aspect of the present disclosure under conditions conducive to polypeptide expression, and (b) optionally, recovering the expressed polypeptide. The expressed polypeptide can be recovered from a cell-free extract or from the culture medium. In another embodiment, the method comprises chemically synthesizing the polypeptide.

[1519] In an eighth aspect, the present disclosure provides a kit. In one embodiment, the kit comprises:

[1520] (a) one or more polypeptides (i.e., cage polypeptides) according to any embodiment or combination of embodiments of the first to third aspects of the present disclosure;

[1521] (b) one or more polypeptides according to any embodiment or combination of embodiments of the fourth aspect of the present disclosure (i.e., key polypeptides); and

[1522] (c) Optionally, one or more fusion proteins of any embodiment disclosed herein.

[1523] In another embodiment, the kit comprises:

[1524] (a) a first nucleic acid encoding a cage polypeptide according to any embodiment or combination of embodiments of the first to third aspects of the present disclosure;

[1525] (b) a second nucleic acid encoding a key polypeptide according to any embodiment or combination of embodiments of the fourth aspect of the present disclosure; and

[1526] (c) Optionally, a third nucleic acid encoding the fusion protein of any embodiment disclosed herein.

[1527] In another embodiment, the kit comprises:

[1528] (a) a first expression vector comprising a first nucleic acid encoding a cage polypeptide according to any embodiment or combination of embodiments of aspects 1 to 3 of the present disclosure, wherein the first nucleic acid is operably linked to a first promoter; and

[1529] (b) a second expression vector comprising a second nucleic acid encoding the key polypeptide according to any embodiment or combination of embodiments of the fourth aspect of the present disclosure, wherein the second nucleic acid is operably linked to a second promoter.

[1530] In each kit embodiment, the first nucleic acid, the second nucleic acid, the first expression vector, and / or the second expression vector can comprise a single nucleic acid encoding an expression vector capable of expressing a cage or key polypeptide, or can comprise a plurality (i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) of the first nucleic acid, the second nucleic acid, the first expression vector, and / or the second expression vector. In various such embodiments, each second nucleic acid can encode, or each second expression vector can express, a key polypeptide capable of binding to a structural region and inducing a conformational change in a different cage polypeptide encoded by the plurality of first nucleic acids, or capable of being expressed by the plurality of first expression vectors. In other embodiments, each second nucleic acid can encode, or each second expression vector can express, a key polypeptide capable of binding to a structural region and inducing a conformational change in more than one cage polypeptide encoded by the plurality of first nucleic acids, or capable of being expressed by the plurality of first expression vectors.

[1531] In one embodiment, the promoter operably linked to the nucleic acid encoding the cage polypeptide (first promoter) is different from the promoter operably linked to the nucleic acid encoding the key polypeptide (second promoter), thereby allowing for regulatable control of the cage polypeptide and any functional polypeptide domains by controlling the expression of the key polypeptide. In other embodiments, the promoter operably linked to the nucleic acid encoding the cage polypeptide (first promoter) is the same as the promoter operably linked to the nucleic acid encoding the key polypeptide (second promoter). In other embodiments, the first promoter and / or the second promoter can be an inducible promoter.

[1532] In a ninth aspect, the present disclosure provides a LOCKR switch comprising:

[1533] (a) a cage polypeptide comprising a conformational domain and a latch domain further comprising one or more biologically active peptides, wherein the conformational domain interacts with the latch domain to prevent the activity of the one or more biologically active peptides;

[1534] (b) optionally a key polypeptide that binds to the cage domain, thereby displacing the latch domain and activating the one or more biologically active peptides; and

[1535] (c) optionally, one or more effector polypeptides that bind to the one or more biologically active peptides when the one or more biologically active peptides are activated.

[1536] Cage polypeptides and their structural regions and latch regions have been discussed above, as have bioactive peptides, effector peptides, and key peptides. Any of the embodiments of cage polypeptides, bioactive peptides, and key peptides disclosed herein can be used in the LOCKR switches and kits disclosed herein. For example, in one embodiment, the cage polypeptide comprises:

[1537] (a) a helical bundle comprising 2 to 7 α-helices; and

[1538] (b) an amino acid linker connecting each α-helix;

[1539] In one embodiment, a key polypeptide is present, and can comprise a key polypeptide of any embodiment or combination of embodiments disclosed herein.

[1540] In another embodiment, an effector polypeptide is present and comprises a polypeptide that selectively binds to a bioactive peptide. Depending on the bioactive peptide of interest, any suitable effector polypeptide may be used. In various non-limiting embodiments, the effector peptide may include Bcl2, GFP1-10, proteases, and the like.

[1541] The present disclosure also provides a LOCKR switch comprising a cage polypeptide and a key polypeptide as described herein. In some aspects, the LOCKR switch comprises: (a) a cage polypeptide comprising a structural region and a latch region further comprising one or more bioactive peptides, and (b) a key polypeptide bound to the cage structural region. In some aspects, the LOCKR switch comprises: (a) a cage polypeptide comprising a structural region and a latch region further comprising one or more bioactive peptides, and (b) a key polypeptide bound to the cage structural region, wherein the one or more bioactive peptides in the latch region bind to or interact with one or more effector polypeptides. In other aspects, the LOCKR switch comprises: (a) a cage polypeptide comprising a structural region and a latch region further comprising one or more bioactive peptides, wherein the structural region interacts with the latch region to prevent the activity of the one or more bioactive peptides; (b) a key polypeptide that binds to the cage structural region, thereby displacing the latch region and activating the one or more bioactive peptides. In some other aspects, LOCKR further comprises one or more effector polypeptides that bind to the one or more bioactive peptides when the one or more bioactive peptides are activated.

[1542] In some aspects, both the latch region and the key polypeptide can bind to or interact with a structural region in a corresponding cage polypeptide. The interaction between the latch region and the structural region in the cage polypeptide can be an intramolecular interaction, while the interaction between the key polypeptide and the structural region of the corresponding cage polypeptide can be an intermolecular interaction. However, in some aspects, in the absence of an effector polypeptide, the affinity of the latch region for the structural region of the cage polypeptide is higher than the affinity of the key polypeptide for the structural region of the cage polypeptide.

[1543] In some aspects, the affinity of the latch region for the structural region of the cage polypeptide is at least about 1.5-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 11-fold, at least about 12-fold, at least about 13-fold, at least about 14-fold, at least about 15-fold, at least about 16-fold, at least about 17-fold, at least about 18-fold, at least about 19-fold, at least about 20-fold, at least about 21-fold, at least about 22-fold, at least about 23-fold, at least about 24-fold, at least about 25-fold, at least about 26-fold, at least about 27-fold, at least about 28-fold, at least about 29-fold, or at least about 30-fold greater than the affinity of the key polypeptide for the structural region of the cage polypeptide in the absence of the effector polypeptide. In some aspects, the affinity of the latch region for the structural region of the cage polypeptide is at least about 1.1 fold, at least about 1.2 fold, at least about 1.3 fold, at least about 1.4 fold, at least about 1.5 fold, at least about 1.6 fold, at least about 1.7 fold, at least about 1.8 fold, at least about 1.9 fold, at least about 2.0 fold, at least about 2.1 fold, at least about 2.2 fold, at least about 2.3 fold, at least about 2.4 fold, at least about 2.5 fold, at least about 2.6 fold, at least about 2.7 fold, at least about 2.8 fold, at least about 2.9 fold, or at least about 3.0 fold greater than the affinity of the key polypeptide for the structural region of the cage polypeptide in the absence of the effector polypeptide. In some aspects, the affinity of the latch region for the structural region of the cage polypeptide is at least about 30-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, at least about 100-fold, at least about 110-fold, at least about 120-fold, at least about 130-fold, at least about 140-fold, at least about 150-fold, at least about 160-fold, at least about 170-fold, at least about 180-fold, at least about 190-fold, at least about 200-fold, at least about 210-fold, at least about 220-fold, at least about 230-fold, at least about 240-fold, at least about 250-fold, at least about 260-fold, at least about 270-fold, at least about 280-fold, at least about 290-fold, at least about 300-fold, at least about 310-fold, at least about 320-fold, at least about 330-fold, at least about 340-fold, at least about 350-fold, at least about 360-fold, at least about 370-fold, at least about 380-fold, at least about 390-fold, at least about 400-fold, at least about 410-fold, at least about 420-fold, at least about 430-fold, at least about 440-fold At least about 160 times, at least about 170 times, at least about 180 times, at least about 190 times, at least about 200 times, at least about 210 times, at least about 220 times, at least about 230 times, at least about 240 times, at least about 250 times, at least about 260 times, at least about 270 times, at least about 280 times, at least about 290 times, at least about 300 times, for example, from about 30 times to about 300 times, for example, from about 100 times to about 300 times, from about 50 times to about 100 times.

[1544] In other embodiments, the intramolecular latch-cage affinity is higher than the intermolecular key-cage affinity, and in the presence of the effector protein, the intermolecular key-cage affinity is higher than the intramolecular latch-cage affinity. Thus, the function of the bioactive peptide depends on the presence of the cage, key, and effector protein.

[1545] In certain embodiments, in the absence of an effector protein, the intermolecular key-cage interaction can outperform the latch-cage interaction. In the absence of a key, the latch-cage affinity is higher than the latch-effector affinity (via binding of the bioactive peptide to the effector protein), and in the presence of a key, the latch-effector affinity (via binding of the bioactive peptide to the effector protein) is higher than the latch-cage affinity. Thus, the function of the bioactive peptide depends on the presence of the cage, key, and effector protein.

[1546] As disclosed herein, by using a cage polypeptide to sequester a bioactive peptide in a latch region, the cage polypeptide can be used with a key polypeptide to hold the bioactive peptide in an inactive state in that region until the key polypeptide displaces the latch through competitive intermolecular binding that induces a conformational change, thereby exposing the encoded bioactive peptide or domain and activating the system (see Figure 1 The combined use of cage and key polypeptides is described in more detail herein in the Examples below and is referred to as a LOCKR switch. LOCKR stands for Latch Orthogonal Cage-Key Protein; each LOCKR design consists of a cage polypeptide and a key polypeptide, which are two separate polypeptide chains. Orthogonal LOCKR designs (see Figure 3 ) is represented by a lowercase subscript: LOCKR a By Cage a and keys a Composition, and LOCKR b By Cage b and keys b Composition, etc., so that the cage a Only by key a Activate and cage b Only by key b Activation, etc. The prefix in the peptide and LOCKR names indicates the functional group encoded and controlled by the LOCKR switch. For example, BimLOCKR refers to a designed switch encoding the Bim peptide, while GFP11-LOCKR refers to a designed switch encoding GFP11 (the 11th chain of GFP). Figure 8 Sequence alignment of the original LOCKR_a cage scaffold design was compared with its asymmetric (1fix-short–noBim(AYYA)-t0) and orthogonal (LOCKRb-f) design counterparts.

[1547] In another embodiment, the nomenclature of the cage is identified by 1fix-short and 1fix-latch, indicating that the cage as defined above a Similar but different implementations of aActivation. The functional group encoded in the latch is identified by the third part of the name, while the suffix indicates the presence of a base. For example, 1fix-short-Bim-t0 encodes Bim on a 1fix-short scaffold without a base. In another example, 1fix-latch_Mad1SID_T0_2 indicates that the 1fix-latch scaffold is used to encode Mad1SID without a residue. The suffix 2 indicates that there are two versions, in which the functional sequence is encoded at different positions in the latch region.

[1548] In one embodiment of the eighth and ninth aspects of the present disclosure, the one or more cage polypeptides and the one or more key polypeptides comprise at least one cage polypeptide and at least one key polypeptide from the same row of Tables 1, 2, 3, and / or 4. As will be understood by one of skill in the art based on the teachings herein, such kits can include multiple (i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, or more) cage and key polypeptide pairs that can be used together as a LOCKR switch.

[1549] In one embodiment of the kit or switch disclosed herein, the one or more cage polypeptides and the one or more key polypeptides comprise at least one cage polypeptide and at least one key polypeptide having an amino acid sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along their length to the cage polypeptide and key polypeptide in the same row of Table 1, Table 2, Table 3 and / or Table 4, respectively.

[1550] In a tenth aspect, the present disclosure provides uses of the polypeptides, kits, and / or LOCKR switches disclosed herein for sequestering bioactive peptides in caged polypeptides, maintaining them in an inactive ("off") state until combined with a key polypeptide to induce a conformational change that activates ("on") the bioactive peptide. Details of exemplary such uses and methods are disclosed throughout.

[1551] In one embodiment of the kit or switch disclosed herein, the one or more cage polypeptides and the one or more key polypeptides comprise at least one cage polypeptide and at least one key polypeptide having an amino acid sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity along their length to the cage polypeptide and key polypeptide in the same row of Table 1, Table 2, Table 3 and / or Table 4, respectively.

[1552] In another embodiment of the eighth and ninth aspects of the present disclosure, the one or more cage polypeptides and the one or more key polypeptides comprise at least one cage polypeptide and at least one key polypeptide matched by identification numbers in the naming convention used herein. As noted above, the orthogonal LOCKR design (see Figure 3 ) is represented by a lowercase subscript: LOCKR a By Cage a and keys a Composition, and LOCKR b By Cage b and keys b Composition, etc., so that the cage a Only by key a Activate and cage b Only by key b Activation, etc. The prefixes in the polypeptide and LOCKR names indicate the functional groups encoded and controlled by the LOCKR switch. In one embodiment, all 3 plus 1 (3 + 1) and 2 plus 1 (2 + 1) cage and key polypeptides disclosed herein are matched by identification number.

[1553] In some embodiments, the cage and key polypeptide names include the prefix 2plus1 or 3plus1, which defines the helical architecture, where the first number defines the number of helices in the structural region and the second number defines the number of helices in the latch region. The Nterm or Cterm suffix defines whether the latch on the cage component of the kit comprises the N- or C-terminus, respectively, as indicated by brackets []. The Nterm and Cterm, along with the numerical suffix, correspond to the same suffix on the key that activates it. For example, cage 2plus1_cage_Cterm_2406 (SEQ ID NO: 27126) is activated by 2plus2_key_Cterm_2406 (SEQ ID NO: 27127).

[1554] Example

[1555] Overview:

[1556] We have developed a general approach to design novel protein switches that can sequester bioactive peptides and / or binding domains, keeping them in an inactive ("off") state until combined with a second, designed polypeptide, termed a key, which induces a conformational change that activates ("on") the bioactive peptide or binding domain.

[1557] Define the naming and structural characteristics of the LOCKR switch:

[1558] LOCKR stands for Latch-Orthogonal Cage-Key Protein; each LOCKR design consists of a cage protein and a key protein, which are two separate polypeptide chains.

[1559] The cage encodes an isolated bioactive peptide or binding domain in a region of the cage scaffold represented as a latch. The general strategy is to optimize the position of the encoded peptide or binding domain for maximum burial of the functional residues to be isolated, while optimizing burial of hydrophobic residues and solvent exposure / compensatory hydrogen bonding for polar residues.

[1560] The key displaces the latch through competitive intermolecular binding that induces conformational changes, thereby exposing the encoded bioactive peptide or domain and activating the system ( Figure 1 ).

[1561] Orthogonal LOCKR design ( Figure 3 ) is represented by a lowercase subscript: LOCKR a By Cage a and keys a Composition, and LOCKR b By Cage b and keys b Composition, etc., so that the cage a Only by key a Activate and cage b Only by key b Activation, etc.

[1562] The prefix indicates the functional group encoded and controlled by the LOCKR switch. For example, BimLOCKR refers to a designed switch encoding the Bim peptide, while GFP11-LOCKR refers to a designed switch encoding GFP11 (the 11th strand of GFP).

[1563] Bottom: The dynamic range of LOCKR activation by the key can be adjusted by shortening the latch length, weakening the cage-latch interaction, and opening exposed areas on the cage to which the key can bind to act as a "bottom" ( Figure 2 LOCKR can also be tuned in a similar manner by engineering mutations in the latch that weaken the cage-latch interaction. Figure 1-2 、 Figure 10 The length of the underpinning is included as a suffix to the design name: for example, "-t0" means no underpinning, while "-t9" means a 9-residue underpinning (i.e., the latch is truncated by 9 residues).

[1564] • If the term "lock" refers to a single polypeptide chain (rather than to the LOCKR acronym), it is assumed to be synonymous with "cage".

[1565] These designs include the first completely new designed proteins that can undergo conformational transitions in response to protein binding. They are modular because they can encode bioactive peptides of all three types of secondary structures in inactive conformations: alpha helix, beta strand, loop, and are adjustable because their responsiveness can be adjusted over a larger dynamic range by changing the length (length of the cage scaffold and / or latch base) and / or mutating the residues in the cage-latch interface. The designed LOCKR switch can be used to control the activity of a wide range of functional peptides. The ability to utilize these biological functions using strict inducible control can be used, for example, for engineering cells (inducible activation of function, engineering complex logic behaviors and circuits), developing sensors, developing therapeutic agents based on inducible proteins, and creating new biomaterials.

[1566] Design of the LOCKR switch

[1567] We set out to design novel switchable protein systems guided by the following general considerations. First, the free energy adjustments required to achieve maximal dynamic range upon addition of a switch-triggering input are more straightforward in systems governed by competition between inter- and intramolecular interactions at the same site rather than at distant sites (as is common in allosteric biological systems). Second, stable protein frameworks with extended binding surfaces available for competing interactions have advantages over frameworks that become ordered only upon binding, as the former are more programmable and less likely to engage in off-target interactions. These features are driven by Figure 1 Abstract depiction of the system depicted in a, which undergoes thermodynamically driven switching between binding-incompetent and binding-competent states. A latch (blue) contains a peptide sequence (orange) that can bind a target (yellow) unless blocked by intramolecular interactions with the cage (cyan); a tighter-binding key (magenta) outcompetes the latch, allowing the peptide to bind the target. The behavior of this system is governed by the binding equilibrium constants of the individual subreactions ( Figure 1 a):K 打开 , dissociation of latch and cage; K LT , the binding of the latch to the target; and K CK , the binding of the key to the cage. The solutions to this set of equations show that when the latch-cage interaction is too weak (red and orange curves), the system is leaky and the folding induced by the key is low; whereas when the latch-cage interaction is too strong (purple curve), the system is only partially activated even at high key concentrations. The latch-cage interaction affinity for optimal switching ( Figure 1 b, left blue curve, right green curve) is a function of latch-target binding affinity. We use this model to guide the design of optimal switchable protein systems, as described in the following sections.

[1568] LOCKR Design Strategy

[1569] To design such a switchable system, we selected structural features that allow us to tune the affinities of the cage-latch and cage-key interactions over a wide dynamic range. Alpha helices offer an advantage over beta strands because the interhelical interface is governed by side-chain-side-chain interactions, which are much easier to tune than the cooperative backbone hydrogen bonds required for beta sheets. To better control the relative affinities of the cage-latch and cage-key interactions, we chose to design interfaces containing buried hydrogen bond networks: as demonstrated by Watson-Crick base pairing, considerable specificity can be achieved with relatively small changes in the positions of hydrogen bond donors and acceptors. 4,5 As a starting point, we chose a designed α-helical hairpin homotrimer with a specific subunit-subunit interaction mediated by a hydrogen bond network (5L6HC3_1). 5 By designing short, unstructured loops connecting the subunits, we generated monomeric protein frameworks with five or six helices and 40 residues per helix ( Figure 1 c) In the five-helical framework, there is an open binding site for the trans-added sixth helix, whereas in the six-helical framework this site is filled by the cis-acting sixth helix.

[1570] When recombinantly expressed in E. coli, the five-helix (cage) and six-helix (cage plus latch) designs were soluble, and the proteins purified were largely monomeric by size exclusion chromatography with multi-angle light scattering and were very thermostable, remaining folded at 95°C and 5 M guanidine hydrochloride ( Figure 1 d). Small-angle X-ray scattering (SAXS) spectra are in good agreement with the designed model and previous original trimer ( Figure 1 e), indicating that the structure is not altered by the loop. In the pull-down assay, the five-helical framework, but not the six-helical framework, bound the sixth helix fused to GFP ( Figure 1 f); the latter result is expected because if the interfaces are identical, the intramolecular interaction K 打开 should outperform its intermolecular counterpart K CK , because the entropic cost of forming intramolecular interactions is reduced. In order to adjust K 打开 We screened for destabilizing mutations in the latch (more hydrophobic to amino acids or serines, and alanine residues to more hydrophobic or serines) and, using a GFP pull-down assay, identified mutants with a limited affinity for the key. The double mutant V223S / I238S bound the key as strongly as the five-helix cage without the latch ( Figure 1e, 10); the two serines may weaken the cage-latch interaction because of the desolvation penalty associated with the buried side chain hydroxyl group and because they reduce the helical tendency of the latch. SAXS and CD spectra show that in the absence of the key, V223S, I238S is a well-folded six-helix bundle with a structure similar to the original monomer ( Figure 1 d) We call this cage-latch-key system LOCKR, for latching orthogonal cage-key proteins.

[1571] Control Bim-Bcl2 binding and adjust the dynamic range of activation:

[1572] To install functionality into LOCKR, we chose the Bim-Bcl2 interaction at the center of apoptosis as a model system, attempting to lock Bim so that binding to Bcl2 occurs only in the presence of a key. We designed constructs with two possible Bim-related sequences on the latch: a designed Bcl2-binding peptide (aBcl2LOCKR) or only Bim residues that are critical for Bcl2 binding (pBimLOCKR). Each has a different affinity for Bcl2, which allowed us to sample a range of K values ​​in an initial series of designs. LT By sampling different helical registers, Bim-related sequences were grafted onto the latch, sequestering residues involved in Bcl2 binding within the cage-latch interface (data not shown), thereby optimizing the burial of hydrophobic residues and the surface exposure of polar residues. K can be tuned by nonoptimal interactions between the cage and Bim residues or by varying the length of the latch. 打开 ( Figure 2 a) The initial design was tested for binding to Bcl2 by biolayer interferometry and showed minimal Bcl2 binding even in the presence of the key, or even minimal Bcl2 binding in the absence of the key. In this case, the K obtained for this system is 打开 and K CK The range of values ​​is obviously the same as K LT Mismatch: Key-induced response distance Figure 1 The ideal curve in b is far away.

[1573] We hypothesized that the system could be improved by expanding the interface area presented on the cage: expanding the latch could increase affinity to the cage (reducing K 打开 ), making the system more "closed" in the absence of a key, and extending the key to a longer length allows it to outperform a latch (relative to K 打开 Lower K CK), making the system more inductive. Exploiting the modular nature of the new parametric helix bundle, the cage, latch, and key were each extended by 5, 9, or 18 residues. To enable the key to outcompete the latch, the latter was truncated by four to nine residues to achieve a range of K 打开 This forms a "base" on the cage for the key to bind. An 18-residue stretch with a 9-residue base results in strongly induced binding ( Figure 2 b, c; The signal in biolayer interferometry is not due to the key binding to Bcl2, nor is it due to the key increasing the inactive LOCKR volume. The activation of the key binding is about 40-fold ( Figure 2 c), comparable to or superior to many naturally occurring processes regulated by protein interactions.

[1574] Since the interaction energy is roughly proportional to the total surface area of ​​the interacting residues, K can be tuned by varying the length of the key CK The key concentration range for BimLOCKR activation can be controlled. The EC50 of the 58-length key is 55.6+ / -34nM ( Figure 2 c, d), while the EC50 of the 45-residue key is 230 + / - 58 nM. Truncation of an additional five residues completely abolishes key activation, indicating that the equilibrium is very sensitive to small changes in free energy, as expected from our model ( Figure 2 d) In order to LT To examine the function of BimLOCKR within the context of BimLOCKR, we investigated key-induced binding to the Bcl2 homologs BclB and Bak (Kd for Bim binding were 0.17 nM (Bcl2), 20 nM (BclB), and 500 nM (Bak), respectively). 6 Biolayer interferometry experiments were performed with the target immobilized and assayed against the switch with or without the key in solution, and with the key immobilized and assayed against the switch alone or with the target in solution. From these results, we can obtain the fraction of target or key bound as a function of the switch, key, and target concentrations. The model has an influence on K 打开 , K CK and K LT A global fit of these data yields K 打开 =.01+ / -0.0033, K CK =2.1+ / -1.1nM, K LT (Bcl2) = 28 + / - 7.8 nM and K LT (BclB) = estimated value of 32 + / - 22 nM, where K LTNo estimate was found for (Bak) because little switch activation was observed. The RMSE (root mean square error) of this fit to the observed BLI data was 0.072 nM. These estimates are broadly consistent with the Kd for Bim binding (not used in the fit), indicating that the thermodynamic model ( Figure 1 a) The system is well represented, but small features of the system that may affect target binding may be missing.

[1575] Next, we aimed to design a series of orthogonal LOCKR systems with the goal of engineering a variety of switches that can be selectively activated in heterogeneous mixtures. Specificity is designed to utilize different hydrogen bonding networks at the cage-key interface. a The model deleted the latch helix and generated the backbone of a new sixth helix by parametrically sampling the radius, helical phase, and z-offset of the new latch / key helix. The resulting structure was scanned for a new hydrogen bond network spanning the interface between the new sixth helix and the cage, in which all buried polar atoms were involved in hydrogen bonds; the model was constructed using Rosetta Stone. TM The remaining interfaces around the network were designed for full sequence and side chain rotamer optimization. Five designs were selected based on packing quality, sequence diversity, and the absence of buried polar atoms not involved in hydrogen bonding. Figure 1 The truncated and underpinned variants were analyzed for cognate and non-target key binding by GFP pull-down assays. The three new designs were found to bind their cognate keys ( Figure 11 ), and are orthogonal to each other. All combined keys a To some extent still unknown. With the original design BimLOCKR a Similarly, screw the Bim sequence on the latches of these three designs ( Figure 2 ). BimLOCKR b and BimLOCKR c In the case where its cognate key gives a nine-residue base on the latch, it shows 22-fold and 8-fold activation, respectively ( Figure 3 a, b). BimLOCKR a 、BimLOCKR b and BimLOCKR c are also orthogonal; each is activated only by its cognate key at concentrations up to 5uM ( Figure 3 c). The fact that three of the six designed BimLOCKR proteins were successfully switched and could be orthogonally activated, with a success rate of 50% starting from a single scaffold, illustrates the power of the buried hydrogen bond network approach to achieve specificity.

[1576] Asymmetric LOCKR switch

[1577] Original LOCKR switch design ( Figure 1-2 ) was constructed from a newly designed symmetrical homotrimer 5L6HC3_1 (which contains 6 helices) 5 The symmetrical repeat motif creates opportunities for misfolding and aggregation. To mitigate these effects, we redesigned the original LOCKR switch to be asymmetric (sequences are listed at the end of this document). The asymmetric design performed better and was more monomeric, and we experimentally resolved the binding of the encoded BIM peptide ( Figure 4 A) and without the BIM peptide ( Figure 4 B) X-ray crystal structure ( Figure 4 ). The experimental structure without BIM is almost the same as the computational design model ( Figure 4 B), demonstrating the atomic-level accuracy of our design strategy. Details of the computational design and experimental testing are provided in Methods. Figure 4 B) Superimposed on the basic scaffold 5L6HC3_1 used to prepare LOCKRa 5 (dark) X-ray crystal structure ( Figure 1 ), see Figure 9 .

[1578] gfpLOCKR (GFP11-LOCKR)

[1579] Using the asymmetric design as a starting point, we successfully encoded the 11th strand of GFP into the designed LOCKR switch ( Figure 5 The common split-type GFP consists of two parts: chains 1-10 and 11; when mixed, 1-10 combines with 11 to produce fluorescence. Here, we demonstrate that chain 11 is isolated in the absence of the key and cannot combine with GFP-1-10, but readily produces fluorescence when mixed with the key in the presence of GFP-1-10 ( Figure 5 We experimentally determined the X-ray crystal structure of the designed protein, which showed that GFP-11 is structurally encoded as an α-helix with a conformation that is almost identical to the computationally designed model ( Figure 5 ); these results highlight the functionality and modularity of the LOCKR system and suggest that we can encode bioactive peptides with a propensity for non-helical secondary structures.

[1580] Adjusting for colocalization dependencies

[1581] Figure 1-2 We show that the dynamic range of LOCKR activation can be tuned predictably, suggesting that the system can be modulated to respond only when the cage and key are co-located, which would be advantageous for a variety of functions. Figure 4We demonstrated that this is indeed the case using GFP11-LOCKR and that Spycatcher TM / Spytag TM Fusion tunes the designed LOCKR switch to colocalization dependence ( Figure 6 ). Spycatcher TM With Spytag TM Covalent fusion; when Spycatcher TM Fusion cage and its Spytag TM When the fused key is mixed, it is displayed with its unfused Spytag TM The keys are noticeably more fluorescent when mixed ( Figure 6 ).

[1582] Caged intein LOCKR switch

[1583] The designed LOCKR switch with a cage assembly encoding the VMAc intein showed successful activation when mixed with the designed key fused to sfGFP and the VMAn intein ( Figure 7 ). SDS-PAGE showed a successful VMAc-VMAn reaction with bands corresponding to the correct molecular weight of the expected spliced ​​protein product ( Figure 7 ).

[1584] Large-scale high-throughput design of LOCKR switches

[1585] Original LOCKR switch design ( Figure 1-2 ) was constructed from a newly designed symmetrical homotrimer 5L6HC3_1 (which contains 6 helices) 5 We thought we should make even smaller LOCKR switches, made of three or four helices. Using everything we learned from testing and experimental validation of the original LOCKR switch, we created a computational pipeline that automatically designs thousands of new LOCKR switch scaffolds from scratch by thoroughly sampling the Crick helix parameters. 4,9 .

[1586] These 2plus1 and 3plus1 LOCKR switches have smaller payloads than the original design (facilitating battery engineering efforts) and, due to their lack of symmetry, are likely to be well-behaved and less prone to aggregation. (See the Methods section for details of the computational design and experimental testing.)

[1587] strepLOCKR (STREPII-LOCKR)

[1588] We designed and tested new LOCKR scaffolds encoding and controlling the STREPII sequence (N)WSHPQFEK (SEQ ID NO: 63) using new 2+1 and 3+1 LOCKR scaffolds from large-scale high-throughput design (see Methods for details). Figure 13 B) compared to the design ( Figure 13 A) Isolates the STREPII tag and can be used in the presence of a key ( Figure 13 CD), as determined by biolayer interferometry (Octet) data.

[1589] Figure 12 The data in Figure 2 demonstrate the locking of the PAH2 domain of the mSin3A transcriptional repressor. See the figure legend for details.

[1590] Figure 14 The data in indicate that the 3plus1 LOCKR switch activates GFP fluorescence in response to expression of the key.

[1591] See legend for details

[1592] discuss

[1593] Here, we demonstrate the power of the LOCKR platform by locking protein-protein interactions that can be activated by key-induced activation. We show in vitro data demonstrating that the locked Bim peptide binds to its family members, a 1-10 construct with a complete truncation of GFP strand 11, and an anti-StrepTag TM II antibody binds to the locked StrepTag TM II. The modularity and ultrastability of the newly designed protein enable tuning of switch activation over a wide dynamic range by adjusting the strength of competing cage-key and cage-latch interfaces. Using this approach, we can now design switches beyond these proof-of-concept designs for more complex applications. LOCKR can be used to control natural signaling networks and, in general, to control biological functions through completely synthetic networks of novel signaling molecules.

[1594] LOCKR brings the modularity of DNA-switching technology to proteins, but with the advantage of being able to control and couple a wide range of biochemical functions that can be performed by proteins and bioactive peptides (which are more diverse and widespread than nucleic acids).

[1595] method

[1596] PCR mutagenesis and isothermal assembly

[1597] All mutagenic primers were purchased from Integrated DNA Technologies (IDT). Designed mutagenic primers annealed >18 bp on either side of the site to induce the desired mutation encoded in the mutagenic primer. PCR was used to generate fragments overlapping >20 bp with the desired pET vector upstream and downstream of the mutation site. The resulting amplicons were isothermally assembled into pET21b, pET28b, or pET29b restricted by XhoI and NdeI and transformed into chemically competent E. coli XL1-Blue cells. Monoclonal colonies were sequenced using Sanger sequencing. Sequence-verified plasmids were purified using a Qiagen miniprep kit and transformed into chemically competent E. coli BL21 (DE3) Star, BL21 (DE3) Star-pLysS cells (Invitrogen), or Lemo21 (DE3) cells (NEB) for protein expression.

[1598] Synthetic gene construction

[1599] Synthetic genes were purchased from Genscript Inc. (Piscataway, NJ, USA) and delivered in pET 28b+, pET21b+, or pET29b+ E. coli expression vectors, inserted into the NdeI and XhoI sites of each vector. For the pET28b+ construct, the synthetic DNA was cloned in a frame with an N-terminal hexa-histidine tag and a thrombin cleavage site, and a stop codon was added at the C-terminus. For the pET21b+ construct, a stop codon was added at the C-terminus so that the protein was expressed without the hexa-histidine tag. For the pET29b+ construct, the synthetic DNA was cloned in a frame with a C-terminal hexa-histidine tag. The plasmids were transformed into chemically competent E. coli BL21(DE3)Star, BL21(DE3)Star-pLysS cells (Invitrogen), or Lemo21(DE3) cells (NEB) for protein expression.

[1600] Bacterial protein expression and purification

[1601] Starter cultures were grown in lysogeny broth (LB) or Terrific in the presence of 50 μg / mL carbenicillin (pET21b+) or 30 μg / mL (for LB) to 60 μg / mL (for TBII) kanamycin (pET28b+ and pET29b+). TMThe cell culture medium of 500mL is grown overnight in broth II (TBII).Starter culture is used to inoculate 500mL and contains the antibiotic Studier TBM-5052 autoinduction culture medium, and is grown 24 hours down at 37 ℃.By under 4 ℃, come harvested cell with 4000rcf centrifugal 20 minutes, and be resuspended in lysis buffer (20mM Tris, 300mM NaCl, 20mM imidazoles, pH 8.0 room temperature), then under 1mM PMSF exists, pass through microfluidization cracking.By under 4 ℃, come clearing lysate with 24000rcf centrifugal at least 30 minutes.Supernatant is applied to Ni-NTA (Qiagen) post of pre-equilibration in lysis buffer. The column was washed twice with 15 column volumes (CV) of wash buffer (20 mM Tris, 300 mM NaCl, 40 mM imidazole, pH 8.0 room temperature), followed by 15 CV of high salt wash buffer (20 mM Tris, 1 M NaCl, 40 mM imidazole, pH 8.0 room temperature), and then 15 CV of wash buffer. The protein was eluted with 20 mM Tris, 300 mM NaCl, 250 mM imidazole at pH 8.0 at room temperature. FPLC and Superdex chromatography were used. TM The protein was further purified by gel filtration on a 75% increase 10 / 300 GL (GE) size exclusion column and the fractions containing monomeric protein were pooled.

[1602] Size exclusion chromatography, multi-angle light scattering (SEC-MALS)

[1603] SEC-MALS experiments were performed using the miniDAWN TM Superdex connected to TREOS multi-angle static light scattering TM 75 Add 10 / 300GL (GE) size exclusion column and Optilab T-rEX TM The protein samples were injected at a concentration of 3-5 mg / mL in TBS (pH 8.0). TM The data were analyzed by COMBIFLASH® (Wyatt Technologies) software to estimate the weight average molar mass (Mw) of the eluted material, and the number average molar mass (Mn) to estimate the monodispersity via the polydispersity index (PDI) = Mw / Mn.

[1604] Circular dichroism (CD) measurement

[1605] CD wavelength scans (260 to 195 nM) and temperature melting (25 to 95°C) were measured using an AVIV model 420 CD spectrometer. Temperature melting monitored the absorbance signal at 222 nM and was performed at a heating rate of 4°C / min. The protein sample was 0.3 mg / mL in PBS, pH 7.4, in a 0.1 cm cuvette. Guanidine chloride (GdmCl) titrations were performed on the same spectrometer using an automated titrator in PBS, pH 7.4, at 25°C, monitoring the concentration of 0.03 mg / mL protein in a 1 cm cuvette with a stir bar at 222 nM. Each titration consisted of at least 40 evenly spaced concentration points, with a mixing time of one minute per step. The titrant solution consisted of equal concentrations of protein in PBS + GdmCl. The GdmCl concentration was determined by the refractive index.

[1606] Small-angle X-ray scattering (SAXS)

[1607] The samples were exchanged into SAXS buffer (20 mM Tris, 150 mM NaCl, 2% glycerol, pH 8.0 at room temperature) via gel filtration. Scattering measurements were performed at the SIBYLS TM 12.3.1 is carried out at the beamline. The X-ray wavelength (λ) is And the sample to detector distance of Mar165 detector is 1.5m, corresponding to 0.01 to The scattering vector q (q = 4π*sin(θ / λ), where 2θ is the scattering angle) range. The data set was collected using 34 exposures of 0.2 seconds over 7 seconds at 11 keV, with a protein concentration of 6 mg / mL. Data were also collected at a concentration of 3 mg / mL to determine concentration dependence; all presented data were collected at the higher concentration as no concentration-dependent aggregation was observed. The 32 exposures were averaged for the Gunier, Parod, and Wide-q regions, respectively, based on the signal quality on each region and frame. Using The software package Analyze Means was used to analyze the data and report statistics. FoXS was used to compare the designed model with the experimental scattering curve and calculate the quality of fit (X) value. TM The hexahistidine tag and thrombin cleavage site at the N-terminus of the LOCKR protein were modeled so that the designed sequence matched the experimentally tested protein. To capture the conformational flexibility of these residues, 100 independent models were generated, clustered, and the cluster center of the largest cluster was selected as the representative model for unbiased FoXS fitting.

[1608] GFP pull-down assay

[1609] Following the above protocol, His-tagged LOCKR was expressed from pET28b+, while Key was expressed fused to superfolded GFP (sfGFP) without a His tag in pET21b+. His-tagged LOCKR was purified to completion and dialyzed into TBS (20 mM Tris, 150 mM NaCl, pH 8.0 room temperature); Key-GFP remained as the lysate for this assay. 100 μL of LOCKR at >1 uM was applied to a 96-well black Nickel-coated plates (ThermoFisher) were plated and incubated at room temperature for 1 hour. The samples were discarded from the plates and washed 3 times with 200 μL TBST (TBS + 0.05% Tween-20). 100 μL of lysate containing key-GFP was added to each well and incubated at room temperature for 1 hour. The samples were discarded from the plates and washed 3 times with 200 μL TBST (TBS + 0.05% Tween-20). The plates were washed once with TBS and 100 μL of TBS was added to each well. sfGFP fluorescence was measured on a Molecular Devices SpectraMax TM Fluorescence was measured on a BioTek Synergy Neo2 plate reader; fluorescence was measured at 485 nm excitation and 530 nm emission with a 20 nm bandwidth for excitation and emission.

[1610] Biolayer Interferometry (BLI)

[1611] BLI measurements were performed on a streptavidin (SA)-coated biosensor. The assay was performed on a RED96 system (ForteBio) and all analyses were performed in ForteBio Data Analysis Software version 9.0.0.10. Proteins diluted into HBS-EP+ buffer (10 mM HEPES, 150 mM NaCl, 3 mM EDTA, 0.05% v / v surfactant P20, 0.5% skim milk powder, pH 7.4 room temperature) from GE were used for the assay. Biotinylated Bcl2 was loaded onto the SA tip with a threshold of 0.5 nm, which was programmed into the machine protocol. A baseline was obtained by immersing the loaded biosensor in HBS-EP+ buffer; association kinetics were observed by immersing in a hole containing a defined concentration of LOCKR and a key, and dissociation kinetics were then observed by immersing in the buffer used to obtain the baseline. Kinetic constants and equilibrium responses were calculated by fitting a 1:1 binding model.

[1612] Thermodynamic LOCKR model

[1613] Figure 1The thermodynamic model in a illustrates three free parameters for the five equilibrium regions. This defines three equations that relate the concentrations of all species at equilibrium (open or closed switch, key, target, switch-key, switch-target, and switch-key-target).

[1614] K 打开 =[Switch 打开 ] / [switch 关闭 ]

[1615] K CK =[Switch 打开 [Key] / [Switch-Key]=[Switch-Target] [Key] / [Switch-Key-Target]

[1616] K LT =[switch-key][target] / [switch-key-target]=[switch 打开 ][Target] / [Switch-Target]

[1617] The total amount of each component (switch, key, and target) is also constant, and the value of each species at equilibrium is constrained. This introduces the following equation into the model.

[1618] [switch] 总计 =[Switch 打开 ]+[Switch 关闭 ]+[switch-key]+[switch-target]+[switch-key-target]

[1619] [key] 总计 =[key]+[switch-key]+[switch-key-target]

[1620] [Target] 总计 =[target]+[switch-target]+[switch-key-target]

[1621] These six equations are fed into sympy.nsolve() to find a solution for the given six constants (three equilibrium constants, three total concentrations). The fractions are extracted from this solution and the corresponding graphs are plotted.

[1622] Porting functional sequences to LOCKR using Rosetta

[1623] A functional LOCKR model was prepared by grafting a biologically active sequence onto the latch, which was then expressed using Rosetta TM XML design to sample the graft starting from each helical register on the latch. This scheme uses two Rosetta movers, SimpleThreadingMover to change the amino acid sequence on the latch, and FastRelax with default settingsTM Used to find the lowest energy structure under a given functional mutation. TM Designs were selected by eye in 2.0, and high-quality grafts had important binding residues that interacted with the cage and minimized the number of buried unsatisfied hydrogen-bonding residues.

[1624] Rosetta design of orthogonal LOCKR

[1625] Using Rosetta with scoring function β_nov16, LOCKR a Redesigned as an orthogonal cage-key pair. We extracted the five-helical cage model from the extended LOCKR model and used Rosetta TM The BundleGridSampler module generates backbone ensembles for new latch geometries. The BundleGridSampler generates backbone geometries based on the Crick mathematical representation of the coiled coil and allows efficient parallel sampling of a regular grid of coiled coil representation parameter values ​​that corresponds to a continuum of peptide backbone conformations. For each parameterized latch conformation sampled, Rosetta TM Residue selectors were used to specify cage and latch interfaces for designing hydrogen bond networks (HBNet) and subsequently Rosetta TM Side chain design. The interface of the latch and cage was selected via the InterfaceByVector residue selector, and residues were selected for design via the Rosetta residue selector. This residue selection was passed to both HBNet and the side chain design to rigorously design the conversion interface while keeping the cage to its original LOCKR sequence. Hydrogen bond networks were designed on the residues selected at the interface using HBNetStapleInterface. The output contained designs with two or three hydrogen bond networks spanning the three helices that make up the interface. All outputs from HBNet were then designed using PackRotamersMover to place residues at the interface while maintaining hydrogen bond networks. Two rounds of design were performed. The first round used β_soft to actively populate the interface with potentially colliding rotamers while optimizing the interaction energy at the interface, and then minimized the structure using β to resolve potentially colliding atoms according to the full Rosetta scoring function. The final round of design combined the rotors with the full βRosetta scoring function to finally optimize the interactions at the cage-latch interface.

[1626] Candidate orthogonal LOCKR designs were selected based on the lack of unsatisfied buried hydrogen-bonding residues, the count of alanine residues as a proxy for packing quality, and sequence dissimilarity as a metric for finding polar / hydrophobic patterns dissimilar enough to be orthogonal. Unsatisfied hydrogen-bonding atoms were filtered out using the BuriedUnsatHbonds filter, and unsatisfied polar atoms were disallowed based on the filter's metric. Packing quality was determined by counting alanine residues at the interface, as high alanine counts indicate poor cross-bonding of residues. A maximum of 15 alanine residues were allowed throughout the three-helical interface. Pairwise sequence differences for each designed latch were scored using BLOSUM62 by aligning the sequences using the Bio.pairwise2 package from BioPython, as shown in seq_alignment.py. Alignment was performed with large opening and extension penalties, without allowing gaps within the sequences, similar to the structural alignment of two helices, to find the most similar stacking based on hydrophobic polarity patterns. Each score was subtracted from the maximum score to convert the score to a distance metric; the most diverse sequences had the lowest BLOSUM62 scores, which were converted to the maximum distance. The sequences were then clustered using HeirClust_fromRMSD.py with a cutoff of 170, resulting in 13 clusters. The centers of each cluster were selected by maximizing the distance between the 13 selected centers. The clusters were then analyzed in PyMol TM These 13 candidates were filtered by eye in 2.0 for unsatisfactory hydrogen-bonded atoms and qualitative packing quality. The five best designs for these three metrics were selected as LOCKR b-f .

[1627] Asymmetric LOCKR switch

[1628] Using Rosetta with HBNet TM Redesign of the original LOCKR a Switch; Residues known to be important for LOCKR function were kept fixed, and the remaining residues were optimized to preserve hydrophobic packing while introducing sequence diversity that minimized the number of repetitive amino acid sequences and motifs. Synthetic DNA encoding the design was obtained as described above, and the design was expressed, purified, and biophysically characterized as described above. Crystallization experiments were performed as described in the next section.

[1629] X-ray crystallography

[1630] Crystallization of protein samples

[1631] The purified protein samples were concentrated to 12-50 mg / ml in 20 mM Tris pH 8.0 and 100 mM NaCl. The samples were screened using a 5-deck mosquito crystallizer (ttplabtech) with an active humidity chamber using the following crystallization screens: JCSG+ (Qiagen), JCSG Core I-IV (Qiagen), PEG / Ion (Hampton Research), and Morpheus (Molecular Dimensions). The optimal conditions for crystallization of the different designs were found to be as follows:

[1632] · 1-fix-short-BIM-t0: 0.1M Tris pH 8.5, 5% (w / v) PEG 8000, 20% (v / v) PEG 300, 10% (v / v) glycerol (no need to freeze)

[1633] · 1fix-short-GFP-t0: 0.2 M sodium chloride, 0.1 M sodium cacodylate pH 6.5, 2.0 M ammonium sulfate (plus 20% glycerol for freezing)

[1634] · 1fix-short-noBim(AYYA)-t0: 0.2M disodium tartrate, 20% (w / v) PEG 3350 (no cryogen added)

[1635] X-ray data collection and structure determination

[1636] The crystals of the designed protein were cyclized and placed in the corresponding reservoir solution, which contained 20% (v / v) glycerol if it did not contain a cryoprotectant, and snap-frozen in liquid nitrogen. The X-ray data sets were collected at the Advanced Light Source at Lawrence Berkeley National Laboratory using beamlines 8.2.1 and 8.2.2. XDS was used. 35 or HKL2000 36 Index and scale the dataset. Use the designed model as the initial search model and use Phenix TM Software Suite 38 PHASER TM 37 , generate the initial model by molecular replacement. Refine it by simulated annealing using Phenix.refine or, if the resolution is sufficient, by using Phenix.autobuild 40 In-place reconstruction was set to false, simulated annealing, and perfusion and switching phasing were used in an effort to reduce model bias. Manual construction in COOT and Phenix were used. TMThe final model was produced through iterative rounds of refinement in

[15] . Due to the high self-similarity of coiled-coil proteins, the dataset of reported structures has a high degree of pseudo-translational non-crystallographic symmetry, as reported by Phenix. TM , complicating structure refinement and likely explaining the higher-than-expected R values ​​reported. TM The RMSD of bond lengths, angles, and dihedral angles of the ideal geometry were calculated using the program MOLPROBITY. TM The overall quality of all final models was assessed.

[1637] gfpLOCKR:(GFP11-LOCKR) switch design and characterization

[1638] Using Asymmetric LOCKR a Design the bracket as described in the previous section "Using Rosetta TM As described in "Grafting Functional Sequences into LOCKR", the 11th strand of GFP was encoded into the latch sequence of the cage, and a synthetic gene encoding the designed protein was obtained as described above. The protein was purified and biophysical characterized as described above. To test for fluorescence induction after addition of the key, the protein was mixed by pipetting and immediately assayed on a black 96-well plate using a Biotek Synergy Neo2 plate reader to monitor relative GFP fluorescence (Ex: 488, Em: 508, with a reading interval of 10 minutes). Cage leakage was assessed by measuring GFP fluorescence over time in the absence of the key.

[1639] In vitro colocalization-dependent switching with gfpLOCKR (GFP11-LOCKR)

[1640] SpyCatcher fused to its N-terminus via a floppy linker TM Clone the fpLOCKR cage with SpyTag fused to its C-terminus via a floppy linker TM The gfpLOCKR key and GFP1-10 were cloned into their own pET21 vectors. These proteins were expressed overnight in E. coli Lemo21 cells using Studier's auto-induction medium at 18°C. After expression, the producer cells were harvested by centrifugation and lysed by microfluidization. The desired proteins were purified from the clarified lysate by Ni-NTA affinity chromatography and purified by A-PCR on a nanodrop. 280Proteins were diluted in PBS to final concentrations (GFP1-10: 1.9uM in all samples; cage: 1.5uM, 0.8uM, 0.4uM, 0.2uM, 0.094uM; key: 1.5uM, 0.8uM, 0.4uM, 0.2uM, 0.094uM) and pooled as follows: SpyCatcher alone TM - Cage (keyless), SpyCatcher using bare key TM -Cage (without SpyTag TM ) and SpyCatcher-cage using SpyTag-key. The proteins were mixed by pipetting and immediately assayed on a black 96-well plate using a Biotek Synergy Neo2 plate reader to monitor relative GFP fluorescence (Ex: 488, Em: 508, reading interval 10 minutes). Cage leakage was assessed by measuring GFP fluorescence over time in the absence of the key. TM -keys activated GFP fluorescence faster than naked keys, confirming the colocalization dependence.

[1641] Caged intein LOCKR switch

[1642] The VMA intein sequence was designed to encode LOCKR a The VMAn intein sequence and the key a The fusion constructs were cloned and purified according to the previous LOCKR design described above. Intein activity (splicing) was assessed by SDS-PAGE.

[1643] Large-scale high-throughput design of LOCKR switches

[1644] The computational process for designing thousands of new LOCKR switch scaffolds from scratch was as follows: the backbones of 3-helix bundles (expressed as 2plus1 or 2+1 due to a 2-helix scaffold plus a 1-helix latch) and 4-helix bundles (expressed as 3plus1 or 3+1 due to a 3-helix plus a 1-helix latch) were exhaustively sampled using Crick helix parameters; the sampled parameters included z-offsets (-1.51, 0, and 1.51), helical phases every 10 degrees between 0 and 100, and superhelical radii of each helix from the central superhelical axis (z-axis) ranging from 5 to 10 Å; based on the success of the original LOCKR design, we focused on designs with straight helices and no superhelices (superhelical twist fixed to 0.0). The length of each generated helix was 58 residues; the Rosetta loop closure method was used to add loops that connect all helices into a single polypeptide chain (cage scaffold). HBNet, MC-HBNet, and RosettaDesign were used to design the LOCKR switch scaffolds. TMSequence and side chain designs were performed. Additional designs were generated by truncating the helical bundle to shorter scaffolds, preparing versions with latches as N-terminal or C-terminal helices, and by trying different buttress lengths (truncating latch helices ending in polar residues and removing at least one or two hydrophobic stacking residues from the original design). Designs were selected based on computational methods learned from iterative testing and design of previous LOCKR scaffolds and HBNet helical bundles: important metrics included secondary structure shape complementarity (ss_sc) > 0.65 (ss_sc > 0.7 for the best design); Rosetta Holes TM Filtering was performed in the region surrounding the hydrogen-bonding network to eliminate designs with large cavities near the hydrogen-bonding network in the scaffold core; designs were required to have at least 2 distinct hydrogen-bonding networks spanning all helices of the design model (i.e., each helix had to contribute at least one amino acid side chain to the network); the number of Ile, Leu, and Val residues and the number of contacts formed by these amino acid types compared to Ala (a smaller amino acid) were also included as proxies that correlate well with designs having tight, interdigitated hydrophobic packing, which is important for generating stable protein scaffolds.

[1645] strepLOCKR (STREPII-LOCKR) computational design:

[1646] A LOCKR switch encoding the STREPII tag (N)WSHPQFEK (SEQ ID NO: 63) was designed using the 2plus1 and 3plus1 switches from the large, high-throughput LOCKR design set. This sequence is difficult to encode because Pro (which kinks the α-helix) and Trp and His (if buried, must likely participate in hydrogen bonding). To address these issues, rather than sampling all helical residues, a large design set was mined to find LOCKR scaffolds containing Trp (W), His (H) that were pre-organized into the designed hydrogen bond network. Designs with pre-organized Phe (F) were also considered.

[1647] strepLOCKR (STREPII-LOCKR) experimental testing:

[1648] Using biolayer interferometry ( The ability of the purified protein to sequester the STREPII sequence in the absence of the key and to activate it in the presence of the key was tested using the RED96 System (PALL ForteBio). TMNWSHPQFEK (SEQ ID NO: 63) tag antibody (mAb mouse, Genscript A01732-200) was loaded onto anti-mouse IgG Fc capture (AMC) biosensor (PALL ForteBio); Octet Assay Buffer Tips were preconditioned by cycling between: HBS-EP+ buffer from GE (10 mM HEPES, 150 mM NaCl, 3 mM EDTA, 0.05% v / v surfactant P20, 0.5% nonfat dry milk, pH 7.4 room temperature). Protein samples were diluted into Octet assay buffer, keeping the dilution factor consistent to minimize noise. Antibody-loaded tips were reused up to 8 times using the recommended regeneration protocol of cycling between glycine pH 1.65 and Octet assay buffer (minimal loading loss was observed when the tip was preconditioned, and the signal threshold was set to ensure consistent tip loading each time).

[1649] The 5 μg / mL concentration of the 5 μg / mL ...2 μg / mL 1 μg / mL 2 μg / mL TM NWSHPQFEK (SEQ ID NO: 63) tag antibody (mAb mouse, Genscript A01732-200); antibody stock solution was made up to 0.5 mg / mL with 400 ul mqH2O, aliquoted and stored at -80°C, and thawed immediately before use.

[1650] Purification of proteins from bacterial preparations not described above:

[1651] In the presence of 50 μg / ml carbenicillin (pET21-NESG) or 50 μg / ml kanamycin (pET-28b+), the starter culture was grown overnight in Luria-Bertani (LB) medium at 37°C or grown for 8 hours in Terrific broth. The starter culture was used to inoculate 500 mL LB (induced at an OD600 of approximately 0.6-0.9 with 0.2 mM IPTG) or Studier autoinduction medium containing antibiotics. The culture was expressed overnight at 18°C ​​(many designs were also expressed for 4 hours at 37°C, but there was no significant difference in yield). By centrifuging for 15 minutes with 5000rcf at 4 ℃, harvesting cells, and being resuspended in lysis buffer (20mM Tris, 300mM NaCl, 20mM imidazoles, pH 8.0 room temperature), then in the presence of lysozyme, DNAse and the cocktail protease inhibitor (Roche) or 1mM PMSF that do not contain EDTA, by microfluidization cracking.By centrifuging for at least 30 minutes with 18000rpm at 4 ℃, clear lysate, and be applied to Ni-NTA (Qiagen) post pre-equilibrated in lysis buffer.With the lavage buffer (20mM Tris, 300mM NaCl, 40mM imidazoles, pH 8.0 room temperature) of 5 column volumes (CV), be 3-5CV high salt lavage buffer (20mM Tris, 1M NaCl, 40mM imidazoles, pH 8.0 room temperature) subsequently, and be that 5CV lavage buffer washes described post three times then. The protein was eluted with 20 mM Tris, 300 mM NaCl, and 250 mM imidazole at room temperature. No reducing agent was added because the designed proteins did not contain cysteine.

[1652] References

[1653] 1.Huang,P.-S.,Boyken,SE&Baker,D.The coming of age of de novoprotein design.Nature 537,320–327(2016).

[1654] 2.Joh, NH et al., De novo design of a transmembrane Zn 2+ -transportingfour-helix bundle.Science 346,1520–1524(2014).

[1655] 3. Davey, J. A., Damry, A. M., Goto, N. K. & Chica, R. A. Rational design of proteins that exchange on functional timescales. Nat. Chem. Biol. 13, 1280–1285 (2017).

[1656] 4. Huang, P.-S. et al., High thermodynamic stability of parametrically designed helical bundles. Science 346, 481–485 (2014).

[1657] 5. Boyken, S. E. et al., De novo design of protein homo-oligomers with modular hydrogen-bond network-mediated specificity. Science 352, 680–687 (2016).

[1658] 6. Berger, S. et al., Computationally designed high specificity inhibitors delineate the roles of BCL2 family proteins in cancer. Elife 5, 1422 (2016).

[1659] 7. Leaver-Fay, A. et al., ROSETTA3: an object-oriented software suite for the simulation and design of macromolecules. Meth Enzymol 487, 545–574 (2011).

[1660] 8. Kuhlman, B. & Baker, D. Native protein sequences are close to optimal for their structures. Proc Natl Acad Sci USA 97, 10383–10388 (2000).

[1661] 9. Crick, F. H. C. The packing of [alpha]-helices: simple coiled-coils. Acta Cryst (1953). Q6, 689 - 697 [doi:10.1107 / S0365110X53001964] 6, 1–9 (1953).

[1662] 10. Glantz, S. T. et al., Functional and topological diversity of LOV domain photoreceptors. Proc Natl Acad Sci USA 113, E1442–51 (2016).

[1663] 11. Dueber, J. E., Mirsky, E. A. & Lim, W. A. Engineering synthetic signaling proteins with ultrasensitive input / output control. 25, 660–662 (2007).

[1664] 12. Dueber, J. E., Yeh, B. J., Chak, K. & Lim, W. A. Reprogramming control of an allosteric signaling switch through modular recombination. Science 301, 1904–1908 (2003).

[1665] 13. Huang, P.-S. et al., RosettaRemodel: A Generalized Framework for Flexible Backbone Protein Design. PLoS ONE 6, e24109 (2011).

[1666] 14. Schneidman-Duhovny, D., Hammel, M. & Sali, A. FoXS: a web server for rapid computation and fitting of SAXS profiles. Nucleic Acids Res 38, W540–4 (2o10).

[1667] It should be noted that in the original text, "2o10" in item 15 seems to be a typo and should probably be "2010". The translation has been done as accurately as possible based on the provided text.15.Maguire,J.B.,Boyken,S.E.,Baker,D.&Kuhlman,B.Rapid Sampling ofHydrogen Bond Networks for Computational Protein Design.J.Chem.Theory Comput14,2751–2760(2018).

Claims

1. A LOCKR switch comprising (a) a cage polypeptide comprising a structural region, a latch region, and one or more biologically active peptides, wherein the cage polypeptide comprises a helical bundle, the structural region comprises 2 to 7 α-helices, and wherein the latch region comprises at least one α-helix; (b) a key polypeptide that binds to the structural region of the cage polypeptide in the presence of the cage polypeptide, in, The cage polypeptide and the key polypeptide are selected from the following amino acid combinations: (i) SEQ ID NO: 27192 and SEQ ID NO: 27193; (ii) SEQ ID NO: 27198 and SEQ ID NO: 27199; (iii) SEQ ID NO: 27194 and SEQ ID NO: 27195; (iv) SEQ ID NO: 27202 and SEQ ID NO: 27203; (v) SEQ ID NO: 27206 and SEQ ID NO: 27207; or (vi) SEQ ID NO: 27210 and SEQ ID NO: 27211, wherein, in the absence of the key polypeptide, the binding affinity between the latch region and the structural region is higher than the binding affinity between the latch region and the effector polypeptide that binds to the one or more biologically active peptides, wherein, in the absence of the key polypeptide, the structural region of the cage polypeptide interacts with the latch region to prevent the activity of the one or more biologically active peptides, and Wherein, in the presence of the key polypeptide, the structural region of the cage polypeptide preferentially binds to the key polypeptide to displace the latch region and activate the one or more biologically active peptides.

2. The LOCKR switch according to claim 1, wherein: The helical bundle of the cage polypeptide comprises an amino acid linker connecting each alpha helix.

3. The LOCKR switch according to claim 1, wherein: The latch region of the cage polypeptide comprises the one or more biologically active polypeptides.

4. The LOCKR switch of claim 1 , further comprising one or more effector polypeptides that bind to the one or more bioactive peptides when the one or more bioactive peptides are activated.

5. The LOCKR switch according to claim 4, wherein: The effector polypeptide selectively binds to the biologically active peptide, and wherein the effector polypeptide comprises Bcl2, GFP1-10 or a protease.

6. The LOCKR switch according to claim 1, wherein: The one or more bioactive peptides include an amino acid sequence selected from the group consisting of SEQ ID NOs: 60, 62-64, 66, 27052-27093, and 27118-27119.

7. A nucleic acid comprising a first nucleic acid encoding a cage polypeptide as defined in any one of claims 1 to 6 and a second nucleic acid encoding a key polypeptide as defined in any one of claims 1 to 6.

8. An expression vector comprising the nucleic acid according to claim 7 operably linked to a promoter.

9. The expression vector according to claim 8, wherein The first nucleic acid encoding the cage polypeptide is operably linked to a first promoter, and the second nucleic acid encoding the key polypeptide is operably linked to a second promoter.

10. The expression vector according to claim 8, wherein The expression vector comprises a plurality of nucleic acids encoding cage polypeptides.

11. The expression vector according to claim 8, wherein The expression vector comprises a plurality of nucleic acids encoding key polypeptides.

12. The expression vector according to claim 10, wherein The expression vector comprises a single nucleic acid encoding a key polypeptide that binds to each of the plurality of cage polypeptides, wherein each of the plurality of cage polypeptides is different.

13. The expression vector according to claim 12, wherein The expression vector comprises a first nucleic acid encoding the plurality of cage polypeptides and a second nucleic acid encoding the single key polypeptide.

14. The expression vector according to claim 13, wherein The first nucleic acid is operably linked to a first promoter, and the second nucleic acid is operably linked to a second promoter.

15. A host cell comprising the LOCKR switch according to any one of claims 1-6, the nucleic acid according to claim 7, or the expression vector according to any one of claims 8-12.

16. The host cell according to claim 15, wherein The nucleic acid or the expression vector is integrated into the host cell chromosome.

17. The host cell according to claim 15, wherein The nucleic acid or the expression vector is episomal.

18. Use of the LOCKR switch according to any one of claims 1 to 6, the nucleic acid according to claim 7, the expression vector according to any one of claims 8 to 12, or the host cell according to any one of claims 15 to 17, for sequestering one or more biologically active peptides in the cage polypeptide, maintaining the one or more biological activities in an inactive state, wherein when combined with the key polypeptide, the cage polypeptide displaces the latch region to activate the one or more biologically active peptides.

Citation Information

Patent Citations

  • Movable contact for electric switches.

    US922033A

  • Dimeric vstm3 fusion proteins and related compositions and methods

    CN103168048A

  • Generation of libraries of antibodies in yeast and uses thereof

    CN1444651A