Split luciferase-coupled system for detecting protein-protein interactions

The split luciferase-coupled system using engineered intein fragments addresses limitations of existing PPI detection methods by enhancing sensitivity and specificity, facilitating high-throughput screening and drug discovery.

WO2026090735A1PCT designated stage Publication Date: 2026-05-07THE GOVERNING COUNCIL OF THE UNIV OF TORONTO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE GOVERNING COUNCIL OF THE UNIV OF TORONTO
Filing Date
2025-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current methods for studying protein-protein interactions (PPIs) are limited by low sensitivity, specificity, and high operational costs, making them unsuitable for high-throughput screening and drug discovery applications.

Method used

A novel split luciferase-coupled system using engineered GP41-1 intein fragments (β9 and β10 peptides) fused to proteins of interest, allowing reconstitution into functional NanoLuc for luminescence detection in a homogeneous phase, enabling sensitive and specific PPI analysis.

Benefits of technology

The system provides improved sensitivity and specificity for PPI detection, compatible with high-throughput screening and drug discovery, reducing operational costs and simplifying sample manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025051428_07052026_PF_FP_ABST
    Figure CA2025051428_07052026_PF_FP_ABST
Patent Text Reader

Abstract

System for detecting interactions between a first a second protein comprising: (a) a first construct comprising the first protein fused to a first handle comprising a β10 peptide or a β9 peptide of a nanoluciferase fused to the N-terminal of a modified N-terminus intein (IN) of GP41-1; and (b) a second construct comprising the second protein fused to a second handle comprising a modified C-terminus intein (IC) of GP41-1 fused at its C-terminal to a β9 peptide when the β10 is fused to the modified IN, or to the β10 peptide when the β9 peptide is fused to the modified IN. Interaction between the first and second proteins reconstitutes the IN and IC into the GP41-1 that induces splicing of the reconstituted GP41-1 from the first handle and the second handle, ligating the first and the second proteins and arranging the β10 peptide and β9 peptide in a fused tandem.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SPLIT LUCIFERASE-COUPLED SYSTEM FOR DETECTING PROTEIN-PROTEIN INTERACTIONS CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to United States Provisional Patent Application No. 63 / 713,170, filed October 29, 2024, the contents of which are incorporated herein by reference in their entirety.

[0003] REFERENCE TO ELECTRONIC SEQUENCE LISTING

[0004] The application contains a Sequence Listing which has been submitted electronically in .XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on October 23, 2025, is named “180354.0098. xml” and is 27,448 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety. FIELD OF THE DISCLOSURE

[0005] The present disclosure relates to novel split luciferase-coupled systems and methods for detecting protein-protein interactions with an in vivo genetic system.

[0006] BACKGROUND OF THE DISCLOSURE

[0007] Protein-protein interactions (PPIs) are fundamental biochemical steps in cellular processes1-6. Their alterations are involved in various diseases such as cancer, making PPIs attractive therapeutic targets78. However, it is notoriously difficult to target PPIs directly due to the common characteristics of interacting surfaces, which are usually large and flat. Accordingly, PPIs were long considered as “undruggable”. This scenario has been subject to change, however, due to an evolving understanding of protein interaction interfaces and the recent development of new targeting strategies9-11. Indeed, the milestone PPI inhibitor, venetoclax12, which targets the BCL2 / BAX interaction, has been approved by the FDA for treating chronic lymphocytic leukemia, small lymphocytic leukemia, and acute myeloid leukemia. Numerous PPI inhibitors are also currently under development, with many undergoing clinical trials13. In line with inhibitors, therapeutic potential was also demonstrated for other types of small molecule PPI modulators such as molecular glues that rewire cellular functions through artificially enhancing or enforcing target PPIs14, Proteolysis Targeting Chimeras (PROTACs) that direct targets to the ubiquitin proteasome system for degradation15’16, or other chemicals operating through the modality of enhanced proximity17Comprehensive examination of PPIs, both in-depth mechanistic investigation and proteome-scale exploration, is essential for driving biological understanding and drug discovery. However, PPI research heavily depends on technology development. Although numerous techniques for the study of PPIs are available, each one is accompanied by certain limitations2’18. Moreover, many methods, whether biophysical or biochemical, are not ideal for applications in PPI-targeted drug discovery. Thus, continued technology advancement is needed to maximize our capabilities in exploring PPIs. To meet this demand, we developed a method called Split-Intein Medicated Protein Ligation (SIMPL)19For this, we first engineered the GP41-1 split intein and created a version with markedly reduced intrinsic affinity between its two fragments - N-terminal intein (IN) and C-terminal intein (IC) - but without deterioration of its intein activity. Next, we fused the engineered IN and IC, along with two different peptide tags, respectively to two proteins of interest (which we refer to generally as ‘B’ and ‘P’ for convenience). If a PPI occurs between them, spatial proximity allows the IN and IC to reconstitute into a functional enzyme, which splices the B and P proteins into an intact polypeptide or transfers an associated tag from one protein to its partner (depending on the configuration of the constructs). SIMPL demonstrates marked sensitivity and specificity and can be applied to different cellular compartments and various organisms. As part of our initial development of the technology, we demonstrated that SIMPL is compatible with high-throughput screening (HTS) via coupling to the enzyme-linked immunosorbent assay (ELISA). This SIMPL-ELISA platform can also be used for characterizing PPI inhibitors. However, the cumbersome procedures involved in ELISA and the associated expense do not make SIMPL-ELISA an ideal choice for use as an HTS assay. Alternatively, SIMPL can be coupled with homogeneous time-resolved fluorescence (HTRF)20, a platform with significantly improved operability. Unfortunately, the cost of reagents and equipment for HTRF limits its widespread application.

[0008] A recently developed tri-part strategy splits NanoLuc® (NLuc), the brightest luciferase identified so far, into three fragments: two short peptides (β9 and β10 each containing 11 amino acids) and one 16 kDa fragment (Δ11S)21’22. When β9 and β10 are close, these peptides rapidly bind to Δ11S, and NanoLuc assembles and achieves the catalytic activity of luciferase. To obtain a better luciferase signal, it would be advantageous to present β9 and β10 as one peptide fragment instead of two separate peptide fragments.

[0009] SUMMARY OF THE DISCLOSURE

[0010] In one embodiment, the present disclosure provides for a system for detecting or testing interactions between a first protein or fragment thereof (bait protein) and a second protein or fragment thereof (prey protein), the system comprising two separate constructs: (a) a first construct comprising the bait protein and a first handle connected to the bait protein, the first handle comprising (i) a β10 peptide of a nanoluciferase (NanoLuc) fused to the N-terminal side of a modified N-terminus intein of GP41-1 (IN), or (ii) a β9 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and (b) a second construct comprising the prey protein and a second handle fused to the prey protein, the second handle comprising a modified C-terminus intein of GP41-1 (IC) fused at the C-terminal side (i) to the β9 peptide of the NanoLuc when the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified IN, or (ii) to the β10 peptide of the NanoLuc when the β9 peptide of the NanoLuc is fused to the N-terminal side of the modified IN.

[0011] In one embodiment of the system of the present disclosure, the first handle comprises the β10 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and the second handle comprises the modified IC fused at the C-terminal side to the β9 peptide of the NanoLuc.

[0012] In one embodiment of the system of the present disclosure, the first handle comprises the β10 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and the second handle comprises the modified IC fused at the C-terminal side to the β9 peptide of the NanoLuc.

[0013] In another embodiment of the system of the present disclosure, the first handle comprises the β9 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and the second handle comprises the modified IC fused at the C-terminal side to the β10 peptide of the NanoLuc.

[0014] In another embodiment of the system of the present disclosure, (i) the modified IC includes amino acid residues at positions 13 to 37 of wild type IC of GP41-1, and the modified IN includes amino acid residues at positions 1 to 88 of the wild type IN of GP41 -1 fused to amino acid residues 1 to 12 of wild type IC of GP41 -1 (C25 GP41-1 split intein), or (ii) the modified IC includes amino acids at positions 14 to 37 of wild type IC of GP41-1 and the modified IN includes amino acid residues at positions 1 to 88 of wild type IN of GP41-1 fused to amino acid residues at positions 1 to 13 of wild type IC of GP41 -1 (C24 GP41 -1 split intein), or (iii) the modified IC includes amino acids at positions 15 to 37 of wild type IC of GP41-1 and the modified IN includes amino acid residues at positions 1 to 88 of wild type IN of GP41-1 fused to amino acid residues at positions 1 to 14 of wild type IC of GP41-1 (C23 GP41-1 split intein).

[0015] In another embodiment of the system of the present disclosure, the modified IN comprises SEQ ID NO: 3 and the modified IC comprises SEQ ID NO: 4, or the modified IN comprises SEQ ID NO: 5 and the modified IC comprises SEQ ID NO: 6, or the modified IN comprises SEQ ID NO: 7 and the modified IC comprises SEQ ID NO: 8.

[0016] In another embodiment of the system of the present disclosure, the β9 peptide comprises SEQ ID NO: 12 and the β10 peptide comprises SEQ ID NO: 14.

[0017] In another embodiment of the system of the present disclosure, the first handle comprises SEQ ID NO: 16 and the second handle comprises SEQ ID NO: 18. In another embodiment of the system of the present disclosure, the first handle is connected to the N-terminus of the bait protein.

[0018] In another embodiment of the system of the present disclosure, the first handle is connected to the C-terminus of the bait protein.

[0019] In another embodiment of the system of the present disclosure, the second handle is connected to the N-terminus of the prey protein.

[0020] In another embodiment of the system of the present disclosure, the second handle is connected to the C-terminus of the prey protein.

[0021] In another embodiment of the system of the present disclosure, the first construct and the second construct further comprise a tag, wherein the tag is one or more of a V5 tag, an HA tag, and a 3xFLAG tag.

[0022] In another embodiment, the present disclosure relates to a method for detecting the interaction between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein) comprising: (a) providing the system according to any embodiment of the present disclosure; (b) incubating the first construct and the second construct in a homogenous liquid phase under conditions that allow the formation of the GP41 -1; and (c) adding to the incubate of (b) a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc, whereby detection of the luminescence signal being indicative that the bait protein or part thereof and the prey protein or part thereof interact.

[0023] In one embodiment of the method for detecting the interaction between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein) of the present disclosure, the method further comprises comparing the luminescence signal with the luminescence signal of a control to determine a binding strength between the bait protein and the prey protein. In another embodiment, the present disclosure relates to a method for screening a binding inhibitor between two proteins, the method comprising: (a) providing the system according to any embodiment of the present disclosure, the bait protein and the prey protein being proteins known to bind; (b) (i) incubating the bait construct, the prey construct and a compound to be tested in a homogenous liquid phase under conditions that allow the formation of the GP41-1 (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound to be tested under conditions that allow the formation of the GP41-1 (second incubate); (c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and (d) measuring a difference in luminescence signal between the first incubate and the second incubate, wherein greater luminescence signal in the second incubate is indicative that the compound is a binding inhibitor between the bait protein and the prey protein.

[0024] In another embodiment, the present disclosure relates to a method for determining a binding strength between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein), the method comprising: (a) providing the system according to any embodiment of the present disclosure, the bait protein and the prey protein being proteins known to bind; (b) (i) incubating the bait construct, the prey construct and a compound known to inhibit the binding between the bait protein and the prey protein in a homogenous liquid phase under conditions that allow the formation of the GP41 -1 (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound under conditions that allow the formation of the GP41-1 (second incubate); (c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and (d) determining the binding strength between the bait protein and prey protein in the presence of the compound and the binding strength between the bait protein and the prey protein in the absence of the compound based on the luminescence signal of the first incubate and the second incubate. In one embodiment, the A11S peptide of the methods of the present disclosure comprises SEQ ID NO: 10.

[0025] In another embodiment, the present disclosure provides for a recombinant polynucleotide, wherein the recombinant polynucleotide comprises SEQ ID NO: 15 or SEQ ID NO: 17, or a recombinant polynucleotide having at least 90% sequence identity to SEQ ID NO: 15 or to SEQ ID NO: 17.

[0026] In another embodiment, the present disclosure provides for a recombinant peptide, wherein the recombinant peptide comprises an amino acid sequence selected from SEQ ID NO: 16 or SEQ ID NO: 18, or a recombinant peptide having at least 90% sequence identity to SEQ ID NO: 16 or to SEQ ID NO: 18.

[0027] In another embodiment, the present disclosure provides for an isolated polynucleotide, wherein the isolated polynucleotide comprises SEQ ID NO: 15 or SEQ ID NO: 17, or an isolated polynucleotide having at least 90% sequence identity to SEQ ID NO: 15 or to SEQ ID NO: 17.

[0028] In another embodiment, the present disclosure provides for an isolated peptide, wherein the isolated peptide comprises an amino acid sequence selected from SEQ ID NO: 16 or SEQ ID NO: 18, or an isolated peptide having at least 90% sequence identity to SEQ ID NO: 16 or to SEQ ID NO: 18.

[0029] In another embodiment, the present disclosure provides for a kit comprising (a) a cell line expressing a first construct and a second construct, (b) a Δ11S peptide of NanoLuc, and (d) a substrate that generates a luminescence signal in the presence of the NanoLuc, wherein the first construct comprises a first protein and a first handle connected to the first protein, the first handle comprising (i) a β10 peptide of NanoLuc fused to the N-terminal side of a modified N-terminus intein of GP41-1 (IN), or (ii) a β9 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and (b) a second construct comprising a second protein and a second handle fused to the second protein, and the second handle comprising a modified C-terminus intein of GP41-1 (IC) fused at the C-terminal side (i) to the β9 peptide of the NanoLuc when the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified IN, or (ii) to the β10 peptide of the NanoLuc when the β9 peptide of the NanoLuc is fused to the N-terminal side of the modified IN.

[0030] In one embodiment of the kit of the present disclosure, the modified IN comprises SEQ ID NO: 3 and the modified IC comprises SEQ ID NO: 4, or the modified IN comprises SEQ ID NO: 5 and the modified IC comprises SEQ ID NO: 6, or the modified IN comprises SEQ ID NO: 7 and the modified IC comprises SEQ ID NO: 8.

[0031] BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Embodiments will now be described, by way of example only, with reference to the drawings, in which:

[0033] Figs. 1 A to 1 D: Overview and validation of SIMPL2 system. 1 A: In the SIMPL2 system, β10 and β9 tags derived from tNLuc are respectively included in the SIMPL constructs: immediately N-terminal to IN and immediately C-terminal to IC. Interaction-induced splicing conjugates the two tags and allows them to bind the third component of tNLuc, Δ11S. In presence of the substrate furimazine, the reconstituted tNLuc produces luminescence signal. 1B: Specificity of the tNLuc.

[0034] β9, β10 or β10-β9 peptide fused to recombinant protein (protein G) was serially diluted and incubated with Δ11S and furimazine, followed by luminescence measurements. The experiment was performed in quadruplicates and each replicate is presented as a single dot. The data were fitted to a model of specific binding with Hill slope, which is shown as curves. 1C: Schematic diagram of the cell-based SIMPL2 assay protocol. A liquid mixture of detergent, furimazine and Δ11S is added directly to the cultured cells expressing the investigated SIMPL2 constructs, followed by incubation and luminescence detection. 1D: Comparison of the SIMPL2 and SIMPL-ELISA assays performed using rapamycin-mediated FRB / FKBP1A interaction. Cells expressing FRB-β10-IN and IC-β9-FKBP1A were subject to rapamycin treatment with varied concentrations for two hours followed by measurement of reporter signal. The experiment was performed in quadruplicates and each replicate is presented as a single dot. The data were fitted to a three-parameter agonist-response model, which is shown as curves. SIMPL- ELISA data were extracted from our previous study (Nature Communications (2020) 11:2440). Signal to background (S / B) ratios were calculated based on the derived models.

[0035] Figs. 2A to 2I: Reference set benchmarking of SIMPL2 as a PPI discovery tool. SIMPL2 DNA constructs in B-β10-IN and IC-β9-P format encoding each protein pair from the RRS (n = 80) and PRS (n = 60) were co-transfected into HEK 293 cells. Interaction between B and P proteins was measured as luminescence signal. In parallel, each B-β10-IN construct was co-expressed with IC handle and the obtained luminescence signal (SB) was used as an estimate of protein B expression. Similarly, protein P expression (SP) was estimated by IC-β9-P and IN handle co-expression. Experiments were performed in triplicates and their average is presented as a single point. (2A, 2B) Comparison of RRS and PRS results using raw data (in the logarithmic scale) was performed via box-whisker plot (2A) and ROC analysis (2B). AUC: area under curve. Cutoff value for PRS to keep all RRS pairs as negative (false positive rate of zero) is highlighted. (2C) Comparison of interaction signals (logarithmic scale) and expression signals (SBP, product of B and P handle signals in logarithmic scale) is presented in the scatter plot. Simple linear regression analysis of RRS data was performed and the resulting fitted function is presented. The interaction signals were then normalized with B and P expressions using the formula Snormalized — Sraw ~ SBP in which S rai is the raw SIMPL2 signal and all terms are in logarithmic scale, or equivalently calculated as Snormalized = Sraw I (sb x sp) in which all terms are in linear scale. Analyses of normalized data are presented in box-whisker plot (2D), ROC plot (2E) and scatter plot (2F). Alternatively, interaction signals were also transformed according to the empirical function derived from the linear distribution of RRS signals in (2C), using the formula Stransformed = Sraw + 0.9426 – 0.6678SBP. Analyses of transformed data are presented in a box-whisker plot (2G), ROC plot (2H) scatter plot (2I).

[0036] Figs. 3A to 3D: Evaluation of SIMPL2 as a drug discovery tool for PPI inhibitors. 3A: Characterization using sotorasib, the KRAS G12C-specific inhibitor. HEK 293 cells stably expressing IC-β9-KRAS (HEK 293 WT, G12C, G12D, G12R, G12V or Q61H mutant) or IC-β9-HRAS and RBD-β10-IN, or cells expressing EGFR-IN-β10 and IC-β9-SHC1 were treated with doxycycline and sotorasib at the indicated concentrations for 16 hours, followed by SIMPL2 measurement of luciferase signal. SIMPL2 readouts were then normalized as described in Methods. The experiment was performed in quadruplicates and each replicate is presented as a single dot. The normalized data were fitted into a four-parameter inhibitor-response model. 3B-3D: Characterization using EGFR inhibitors. HEK 293 cells stably expressing EGFR(WT)-β10-IN / IC-β9-SHC1 or EGFR(L858R)-β10-IN / IC-β9-SHC1, or control cells expressing IC-β9-KRAS(G12C) / RBD-β10-IN were treated with EGFR inhibitor AG1478 (3B) or afatinib (3C), or MEK inhibitor selumetinib (treatment control) (3D) at the indicated concentrations together with doxycycline (1 pg / ml) for 16 hours, followed by SIMPL2 measurement of luciferase signal. SIMPL2 readouts were then normalized as described in Methods. The normalized data were then fitted into a four-parameter inhibitor-response model, with calculated IC50 values for each inhibitor / PPI pair presented in Table 2. The 95% Cl of each IC50 is also included in parenthesis. The experiment was performed in quadruplicates and each replicate is presented as a single dot.

[0037] Figs. 4A to 4B: Evaluation of SIMPL2 as a drug discovery tool using molecular glue and PROTACs. 4A: HEK 293 cells stably expressing MDM2-β10-IN / IC-β9-MDM4 or control cells expressing IC-β9-KRAS(G12C) / RBD-β10-IN were treated with RO-5963 or selumetinib (compound control) at the indicated concentrations together with doxycycline (1 pg / ml) for 16 hours, followed by SIMPL2 measurement of luciferase signal. Samples were measured and normalized as described in Methods. The experiment was performed in quadruplicates and each replicate is presented as a single dot. 4B: β10-IN-CRBN / IC-β9-BRD4 expression in the stable HEK 293 cells was induced with doxycycline (1 pg / ml) for 16 hours, followed by 3 hours treatment of ARV -825 at different concentrations, in the presence and absence of MG132 proteasome inhibitor (1 pM). The cells were then subjected to SIMPL2 luciferase measurement, and the raw values are plotted. The experiment was performed in triplicates and each replicate is presented as a single dot. Figs. 5A-5C: SIMPL2 constructs. 5A. A protein of interest (B) can be tagged with β10-IN (IN handle) on either its N- or C-terminus. Similarly, a second protein of interest (P) can be tagged with IC-β9 (IC handle) on either terminus. IN / IC handles are used for estimating protein expression. 5B-5C. Comparison of the SIMPL2 and SIMPL-ELISA assays performed using rapamycin-mediated FRB / FKBP1A interaction. Cells expressing FRB-β10-IN and IC-β9-FKBP1A (B) or FRB-β10-IN and FKBP1A-IC-β9 (C) were subject to rapamycin treatment with varied concentrations for two hours followed by measurement of SIMPL2 signal. Data of SIMPL-ELISA with corresponding configuration combinations were extracted from previous study (Nature Communications (2020) 11:2440). The data were fitted to a four-parameter agonist-response model, with corresponding curves shown. The raw data are normalized with corresponding maximal signals and presented as mean ± s.d. with n = 4 for both SIMPL2 and SIMPL-ELISA in (B), n=3 for SIMPL2 and n = 4 for SIMPL-ELISA in (C). Signal to background (S / B) ratios were calculated based on the derived models.

[0038] Figs. 6A-6G: Behavior of reference set in the SIMPL2 system. 6A: β10-IN-tagged Protein ‘B’ constructs were co-expressed with IC handle in HEK 293 cells and the cell lysates were subjected to Western blot analysis. The total expression level of B was detected using a-V5 antibody, while levels of spliced product caused by non-specific interaction with IC handle were detected using a-FLAG antibody.

[0039] 6B: IC-β9-tagged Protein ‘P’ constructs were co-expressed with IN handle, followed by Western blot analysis using a-V5 antibody and a-FLAG antibody, to detect protein expression levels as described in 6A. 6C-6D, Interaction signals of RRS and PRS in Fig. 2 were plotted against ‘handle’ signal correlated with B protein expression (SB in logarithmic scale) (6C) or P protein expression (SP in logarithmic scale) (6D). 6E-6F, Evaluation of different data processing methods. Based on the ROC curves obtained in Fig. 2B, 2E, 2H, sensitivities of different data processing methods at selected specificities (1.00, 0.98, 0.95 or 0.90) are shown in (6E). Similarly, specificities of the methods at selected sensitivities (0.2, 0.4, 0.4 or 0.8) are shown in (6F). For each selected specificity (6E) and sensitivity (6F) each bar, from left to right, represents raw data, normalized data and transformed data. 6G, Comparison of SIMPL2 and SIMPL-ELISA. ROC curve of a SIMPL-ELISA assay with a PRS set (n = 88) and an RRS set (n = 88) (published in Nat Commun 11: 2440) is compared with that from a SIMPL2 assay using the same reference sets.

[0040] Figs. 7A-7B: Validation of SIMPL2 as method for detecting PPI inhibition. 7A. Comparison of SIMPL2 and NanoBiT in evaluating the KRAS inhibitor BI2865. SIMPL2 stable cells expressing IC-β9-KRAS(WT) / β10-IN-RBD were simultaneously treated with both doxycycline and BI2865 for 5 hours at the indicated concentrations followed by luminescence measurement. In contrast, the expression of SmBiT-KRAS(WT) / LgBiT-RBD was induced for 16 hours in NanoBiT stable cells, followed by BI2865 treatment. Data are presented as mean ± s.d. with n = 4 and were fitted into a four-parameter inhibitor-response model. 7B. The expression of EGFR (WT or L858R)-β10-IN and IC-β9-SHC1 in stable cells was induced with doxycycline plus treatment with AG1478 (AG, 100 nM) or afatinib (Af, 100 nM). After 16 hours of treatment, the cells were lysed and subjected to Western blot analysis. Both blots with short exposure (S. E.) and long exposure (L. E.) are presented. Interaction-induced splicing produced a fusion protein of EGFR-SHC1 (*).

[0041] Figs. 8A-8C: Validation of SIMPL2 as method for detecting PROTACs using a Western blot readout. 8A. Expression of β10-IN-CRBN / IC-β9-BRD4 in stable cells was induced with doxycycline for 16 hours followed by treatment with ARV-825 at the indicated concentrations for 3 hours. The cells were lysed and subjected to Western blot analysis. Stable cells expressing only β10-IN-CRBN were used as a control. ARV-825 induced splicing allowed V5 tag transfer to BRD4, which presents a BRD4 band (*) in the V5 blot. 8B. Comparison of SIMPL2 and NanoBiT in evaluating PROTAC. Stable cells expressing β10-IN-CRBN / IC-β9-BRD4 or SmBiT-CRBN / LgBiT-BRD4 were treated with ARV-825 at the indicated concentration for three hours followed by signal measurement. 8C. Performance of SIMPL2 and NanoBiT in measuring PROTAC activity was further compared by plotting the luminescence readings of control and ARV-825 (1000 nM) treated samples obtained from both assays. S / B: signal / background ratio. Data are presented as mean ± s.d. with n = 4. S / B: signal / background ratio.

[0042] Fig. 9: DNA constructs used in the Examples. Plasmid vectors created for SIMPLE2 and NanoBiT assays.

[0043] DETAILED DESCRIPTION OF THE DISCLOSURE

[0044] Definitions

[0045] In this specification and in the claims that follow, reference will be made to a number of terms that shall be defined to have the meanings below. All numerical designations, e.g., dimensions and weight, including ranges, are approximations that typically may be varied ( + ) or ( - ) by increments of 0.1, 1.0, or 10.0, as appropriate. All numerical designations may be understood as preceded by the term “about”.

[0046] Throughout this specification and the claims, the terms “comprise,” “comprises,” and “comprising” are used in a non-exclusive sense, except where the context requires otherwise. Likewise, the terms “include”, “has” and their grammatical variants are intended to be non-limiting, such that recitation of items in a list is not to the exclusion of other like items that can be substituted or added to the listed items. “Consisting essentially of’ when used to define systems, compositions and methods, shall mean excluding other elements of any essential significance to the combination for the intended use. Thus, a system or composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers. “Consisting of” shall mean excluding more than trace elements of other ingredients and substantial method steps for using the split inteins of this disclosure. Embodiments defined by each of these transition terms are within the scope of this disclosure.

[0047] For the purposes of this specification and appended claims, unless otherwise indicated, all numbers expressing amounts, sizes, dimensions, proportions, shapes, formulations, parameters, percentages, parameters, quantities, characteristics, and other numerical values used in the specification and claims, are to be understood as being modified in all instances by the term “about” even though the term “about” may not expressly appear with the value, amount or range. The term “about,” particularly in reference to a given quantity, is meant to encompass deviations of plus or minus five percent.

[0048] The term “IN” as used herein refers to the N-terminal portion of an intein protein of the present disclosure or of GP41 -1.

[0049] The term “IC” is used to refer to the C-terminal portion of an intein protein of the present disclosure or of GP41 -1.

[0050] “Bait” as used in this document is a test peptide or polypeptide or protein whose interaction or binding to another peptide, polypeptide or protein (prey as defined below) is being studied.

[0051] In one embodiment, the term "bait construct" or “bait fusion protein” as used in this document defines a fusion protein between a first test protein or bait peptide (bait), and one or more other polypeptides, one of which is IN, including modified IN. In one embodiment, the bait construct includes a tag (the tag in the bait fusion protein may be referred to as the “first tag”). The first tag may be located between the first test protein and the IN (bait-tag-IN) or at an end of the IN opposite to the end linked to the bait (i.e. bait-IN-tag).

[0052] The term "bait vector" as used in this document refers to a nucleic acid construct which contains sequences encoding the bait construct and regulatory sequences that are necessary for the transcription and translation of the encoded sequences by the host cell, and preferably regulatory sequences that are needed for the propagation of the nucleic acid construct in mammalian cells.

[0053] In one embodiment, the term “prey construct” or "prey fusion protein" as used in this document defines a fusion between a second test peptide or prey peptide (prey), and one or more other polypeptides, one of which is IC, including modified IC. In one embodiment, the prey construct includes a tag (the tag in the prey construct may be referred to as the “second tag”). The second tag may be located between the second test peptide and the IC (prey-tag-IC) or at an end of the IC opposite to the end linked to the prey (i.e. prey-IC-tag).

[0054] “Prey” as used in this document is a test peptide or polypeptide or protein whose interaction or binding to another peptide, polypeptide or protein (bait, as defined above) is being studied.

[0055] The terms "prey vector" and "library vector" as used herein refer to a nucleic acid construct which contains sequences encoding the prey construct and regulatory sequences that are necessary for the transcription and translation of the encoded sequences encoding by the host cell.

[0056] The term "tag" as used in this document refers to a nucleic acid sequence or its translation product, which allows the immunological isolation, detection and / or purification of a polypeptide bound to the tag by means of an antibody directed specifically against the tag. Examples of tags that can be used in the present application include V5 tag, HA tag, 3xFLAG tag.

[0057] “Test polypeptide” is a polypeptide whose interaction with another polypeptide is being studied with the system of the present disclosure.

[0058] By “isolated” is meant, when referring to a polypeptide, that the indicated molecule is separate and discrete from the whole organism with which the molecule is found in nature or is present in the substantial absence of other biological macro molecules of the same type. The term “isolated” with respect to a polynucleotide is a nucleic acid molecule devoid, in whole or part, of sequences normally associated with it in nature; or a sequence, as it exists in nature, but having heterologous sequences in association therewith; or a molecule disassociated from the chromosome.

[0059] The term "peptide" as used herein is defined as a chain of amino acid residues, usually having a defined sequence. As used herein the term "peptide" is mutually inclusive of the terms "peptides”, “polypeptides” and "proteins".

[0060] The terms “polynucleotide,” “oligonucleotide,” “nucleic acid” (NA) and “NA molecule” are used herein to include a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. An “oligonucleotide” refers to a single stranded polynucleotide having less than about 100 nucleotides, less than about, e.g., 75, 50, 25, or 10 nucleotides.

[0061] The term “recombinant” when used in association with a polynucleotide means a polynucleotide of genomic, cDNA, semi-synthetic, or synthetic origin, which either does not occur in nature or is linked to another polynucleotide in a non-natural arrangement.

[0062] The term “recombinant” when used in association with a peptide or protein means a peptide or protein of semi-synthetic or synthetic origin, which either does not occur in nature or is linked to another peptide or protein in a non-natural arrangement.

[0063] The "percentage of identity" between two amino acid sequences or two nucleotide sequences is to be understood as the percentage of the sequence positions with identical amino acids or nucleotides. The percentage of identity between sequences may be calculated by means of "sequence alignment". The sequence alignment may be local or global. In the sense of the present disclosure the percentage of identity will be calculated, preferably, over a global alignment, among the entire sequence or an entire active fragment of the sequence. Global alignments are more useful when the sequences are similar and have approximately the same size (length). There are several algorithms available in the state of the art for performing these global alignments. There are also bioinformatics tools using such algorithms to obtain the percentage of identity and similarity between sequences. As an example, global alignment between sequences may be performed by means of the well-known GGSEARCH or GLSEARCH software.

[0064] “SIMPL” means Split Intein-Mediated Protein Ligation.

[0065] Overview

[0066] The present disclosure relates to a novel approach for protein-protein interaction (PPI) using SIMPL in solution. In this approach, a split intein is used as a sensor for protein interactions. The present disclosure enables analysis in solution of interactions occurring in various cellular compartments as well as their responses to pharmacological challenges such as enzymatic and PPI inhibitors.

[0067] The present disclosure relates to the incorporation of tri-part Nanoluciferase (tNLuc)21’22into the SIMPL system. This new SIMPL-tNLuc (SIMPL2) platform of the present disclosure allows detection of splicing in a liquid phase, including a homogeneous liquid phase, thereby simplifying sample manipulation and markedly improving assay quantifiability. SIMPL2 readout improves sensitivity and specificity. The modularity of the cell-based SIMPL2 assay platform makes it an excellent tool for characterizing PPI modulators, exemplified with several known small molecule inhibitors, molecular glues and PROTACs.

[0068] In one embodiment, the present disclosure relates to a system for detecting or testing whether two proteins interact with one another. In one embodiment, the present disclosure provides for a system for detecting interactions between a first protein or fragment thereof (bait protein) and a second protein or fragment thereof (prey protein) comprising two separate constructs:

[0069] (a) a first construct comprising the bait protein and a first handle fused to the bait protein, the first handle comprising one of a β10 peptide of a nanoluciferase (NanoLuc) or a β9 peptide of the NanoLuc fused to the N-terminal side of a modified N-terminus intein of GP41-1; and

[0070] (b) a second construct comprising the prey protein and a second handle fused to the prey protein, the second handle comprising a modified C-terminus intein of GP41-1 fused at its C-terminal side either (i) to the β9 peptide of the NanoLuc when the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified N-terminus intein, or (ii) to the β10 of the NanoLuc when the β9 peptide of the NanoLuc is fused to the N-terminal side of the modified N-terminus intein. If the bait protein and the prey protein interact, this interaction reconstitutes the modified IN and modified IC into the GP41 -1 that induces splicing of the reconstituted GP41 -1 from the first handle and the second handle, ligating the bait protein and the prey protein together and arranging the β10 peptide and β9 peptide in a fused tandem position. The NanoLuc is reconstituted in presence of Δ11S peptide, allowing for production of luminescence in the presence of a suitable substrate.

[0071] Although the figures illustrate an embodiment in which the β10 peptide is fused to the modified IN and the modified IC is fused to the β9 peptide, in another embodiment of the present disclosure the β9 peptide is fused to the modified IN and the modified IC is fused to the β10 peptide.

[0072] In one embodiment, the first handle (also referred to as IN handle) 50a comprises, or alternatively consists essentially of, or alternatively consists of, the β10 peptide 54 of NanoLuc fused to the N-terminus of a modified intein N-terminal fragment 51 derived from a split intein GP41-1, and the second handle, also referred to as IC handle, 50b comprises or alternatively consists essentially of, or alternatively consists of, a modified C-terminal fragment 59 derived from the split intein GP41-1 fused at its C-terminal side (i.e., the C-terminal side of the modified C-terminal fragment) to the β9 peptide 55 of the NanoLuc (see Fig. 5A).

[0073] In another embodiment, the first handle (also referred to as IN handle) comprises, or alternatively consists essentially of, or alternatively consists of, the [39 peptide of the NanoLuc fused to the N-terminal side of a modified intein N-terminal fragment derived from a split intein GP41-1, and the second handle, also referred to as IC handle, comprises or alternatively consists essentially of, or alternatively consists of, a modified C-terminal fragment derived from the split intein GP41-1 fused at its C-terminal side (i.e., the C-terminal side of the modified C-terminal fragment) to the β10 peptide of NanoLuc.

[0074] With reference to Fig. 5A, the first and second handles 50a, 50b are separately fused to a bait protein 57 and a prey protein 58 (in one embodiment the first handle is fused to the prey and the second handle is fused to the bait, and in another embodiment the first handle is fused to the bait and the second handle is fused to the prey protein). In one embodiment, the split intein GP41-1 is re-engineered with markedly reduced intrinsic affinity between its two IN and IC fragments, but without deterioration of its intein activity. Although Fig. 5A illustrates the embodiment in which the β10 peptide is fused to the modified IN and the modified IC is fused to the β9 peptide, it is understood that in another embodiment the first and second handles are separately fused to the bait and prey proteins when the β9 peptide is fused to the modified IN and the modified IC is fused to the β10 peptide.

[0075] The bait / prey binding reconstitutes the intein, which splices the bait and prey peptides into a single intact fused protein that can be detected by the production of luminescence in the presence of Δ11S and a substrate (for example, coelenterazine, furimazine or its derivatives such as but not limited to diphenylterazine, selenoterazine, Ad-FMZ-OH, hydrofurimazine, fluorofurimazine, Methylfluorofurimazine, hikarazine-001, hikarazine-003, hikarazine-097, cephalofurimazine, vivazine, and endurazine). Only when the bait and prey protein interact, the luciferase is reconstituted in presence of Δ11S, allowing for production of luminescence in the presence of the substrate (see Fig. 1 A).

[0076] In one embodiment, with reference to Fig. 1A, the β10 peptide 24 and the β9 25 peptide are oriented in the first construct 20a and the second construct 20b such that when the bait 27 and prey 28 ligate, β10 and β9 naturally rejoin into one peptide 22. Rejoined peptide 22 has increased affinity for Δ11S 23 than when β10 and β9 are not rejoined but are in close proximity. Having rejoined peptide 22 results in a stronger luminescence signal 26 with less background.

[0077] In one embodiment, the re-engineered or artificial split intein of the present disclosure comprises a modified IN and a modified IC. In one embodiment the modified IC includes amino acid residues at positions 13 to 37 of wild type IC of GP41-1, and the modified N-terminus fragment includes amino acid residues at positions 1 to 88 of the wild type IN of GP41-1 fused to amino acid residues 1 to 12 of wild type IC of GP41 -1 (C25 GP41 -1 split intein). In another embodiment, the modified IC includes amino acids at positions 14 to 37 of wild type IC of GP41-1 and the modified IN includes amino acid residues at positions 1 to 88 of wild type IN of GP41-1 fused to amino acid residues at positions 1 to 13 of wild type IC of GP41 -1 (C24 GP41 -1 split intein). In another embodiment, the modified IC includes amino acids at positions 15 to 37 of wild type IC of GP41-1 and the modified IN includes amino acid residues at positions 1 to 88 of wild type IN of GP41-1 fused to amino acid residues at positions 1 to 14 of wild type IC of GP41-1 (C23 GP41-1 split intein).

[0078] It should be understood that the modified IC and the modified IN of the artificial split intein of the present disclosure are separate peptides with markedly reduced intrinsic affinity between each other relative to the affinity of wild type IN to wild type IC, but without deterioration of its intein activity.

[0079] SIMPL2

[0080] The SIMPL2 construct or system serves to detect protein-to-protein interactions. In one embodiment, the SIMPL2 construct or system for detecting interactions between a first protein or fragment thereof (bait protein) and a second protein or fragment thereof (prey protein) comprises: (a) a bait construct or first construct comprising the bait protein and an IN handle fused to the C-terminus or to the N-terminus of the bait protein, the IN handle comprising a β10 peptide of a nanoluciferase (NanoLuc) fused to the N-terminal side of a modified IN; and (b) a prey construct or second construct comprising the prey protein and an IC handle fused to the C-terminus or to the N-terminus of the prey protein, the IC handle comprising a modified IC fused at its C-terminal side to a β9 peptide of the NanoLuc. The modified IN comprises at least a fragment of a N-terminal intein of GP41 -1 and the modified IC comprises at least a fragment of a C-terminal intein of GP41 -1. If the bait protein and prey protein interact, then the interaction between the bait protein and the prey protein reconstitutes the modified IN and modified IC into GP41-1 that induces splicing of the reconstituted GP41-1, ligating the bait protein and the bait protein together and arranging the β10 peptide and β9 peptide in a fused tandem position.

[0081] In another embodiment, the SIMPL2 construct or system for detecting or testing interactions between a first protein or fragment thereof (bait protein) and a second protein or fragment thereof (prey protein) comprises: (a) a bait construct or first construct comprising the bait protein and an IN handle fused to the C-terminus or to the N-terminus of the bait protein, the IN handle comprising a β9 peptide of a nanoluciferase (NanoLuc) fused to the N-terminal side of a modified IN; and (b) a prey construct or second construct comprising the prey protein and an IC handle fused to the C-terminus or to the N-terminus of the prey protein, the IC handle comprising a modified IC fused at its C-terminal side to a β10 peptide of the NanoLuc. The modified IN comprises at least a fragment of a N-terminal intein of GP41 -1 and the modified IC comprises at least a fragment of a C-terminal intein of GP41 -1. If the bait protein and prey protein interact, then interaction between the bait protein and the prey protein reconstitutes the IN and IC into GP41-1 that induces splicing of the reconstituted GP41-1, ligating the bait protein and the bait protein together and arranging the β10 peptide and β9 peptide in a fused tandem position.

[0082] The SIMPL2 design includes (a) a bait construct or bait fusion protein carrying a bait protein, and (b) prey construct or a prey fusion protein carrying a prey protein. To investigate the interaction of the test proteins in vivo both the bait construct and the prey construct are expressed in a cell line of interest, including mammalian and non-mammalian cells, preferably mammalian cells. If the bait protein and the prey protein interact, then the association of the bait protein in the bait construct and the prey protein in the prey construct brings the modified IN and the modified IC into close proximity, allowing them to reconstitute into a fully functional intein, which then catalyzes its excision and the concurrent ligation of the bait and the prey (as well as their respective β10 peptide and the β9 peptide) into an intact protein. The resulting spliced protein can be resolved by regular analytical procedures such as Western blot analysis due to its altered mobility, while the presence of the β10 peptide and the β9 peptide allows, in the presence of Δ11S and a substrate, for visualization of the interaction through luminescence.

[0083] In one embodiment, the C-terminus of the bait protein is connected to the N-terminus of the β10 peptide, which in turn is connected to the N-terminus of the modified IN (see Fig. 1A and Fig. 5A (B- β10-IN)).

[0084] In another embodiment, the N-terminus of the bait protein is connected to the C-terminus of the modified IN, which in turn is connected to the C-terminus of the β10 peptide (see Fig. 5A (β10- IN-B)). In one embodiment, the N-terminus of the prey protein is connected to the C-terminus of the β9 peptide, which in turn is connected to the C terminus of the modified IC (see Fig. 1A and Fig. 5 A (IC- β9-P)).

[0085] In one embodiment, the C-terminus of the prey protein is connected to the N-terminus of the modified IC, which in turn is connected to the N-terminus of the β9 peptide (Fig. 5 A (P-IC- β9)).

[0086] Although Fig. 5A illustrates an embodiment in which the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified IN and the modified IC is fused at its C-terminal end to the β9 peptide, it should be understood that in another embodiment of the present disclosure the β9 peptide of the NanoLuc is fused to the N-terminal end of the modified IN and the modified IC is fused at its C-terminal end to the β10 peptide.

[0087] In one embodiment, the C-terminus of the bait protein is connected to the N-terminus of the β9 peptide, which in turn is connected to the N-terminus of the modified IN (B- β9-IN).

[0088] In another embodiment, the N-terminus of the bait protein is connected to the C-terminus of the modified IN, which in turn is connected to the C-terminus of the β9 peptide (β9-IN-B).

[0089] In another embodiment, the N-terminus of the prey protein is connected to the C-terminus of the β10 peptide, which in turn is connected to the C terminus of the modified IC (IC-β10-P).

[0090] In one embodiment, the C-terminus of the prey protein is connected to the N-terminus of the modified IC, which in turn is connected to the N-terminus of the β10 peptide (P-IC-β10).

[0091] In one embodiment of the present disclosure, the modified IN comprises SEQ ID NO: 3 and the modified IC comprises SEQ ID NO: 4, or the modified IN comprises SEQ ID NO: 5 and the modified IC comprises SEQ ID NO: 6, or the modified IN comprises SEQ ID NO: 7 and the modified IC comprises SEQ ID NO: 8. In another embodiment of the present disclosure, the β9 peptide comprises SEQ ID NO: 12 and the β10 peptide comprises SEQ ID NO: 14.

[0092] In another embodiment of the present disclosure, the first handle comprises SEQ ID NO: 16 and the second handle comprises SEQ ID NO: 18.

[0093] In another embodiment of the present disclosure, the first handle is fused to the N-terminus of the bait protein.

[0094] In another embodiment of the present disclosure, the first handle is fused to the C-terminus of the bait protein.

[0095] In another embodiment of the present disclosure, the second handle is fused to the N-terminus of the prey protein.

[0096] In another embodiment of the present disclosure, the second handle is fused to the C-terminus of the prey protein.

[0097] In another embodiment of the present disclosure, the first construct and the second construct further comprise a tag, wherein the tag is one or more of a V5 tag, an HA tag, and a 3xFLAG tag.

[0098] As such, in another embodiment, the present disclosure provides for a recombinant polynucleotide, wherein the recombinant polynucleotide comprises, or consists essentially of, or consists of SEQ ID NO: 15 or SEQ ID NO: 17, or a recombinant polynucleotide having at least 90% sequence identity to SEQ ID NO: 15 or to SEQ ID NO: 17.

[0099] In another embodiment, the present disclosure provides for a recombinant peptide, wherein the recombinant peptide comprises, or consists essentially of, or consists of an amino acid sequence selected from SEQ ID NO: 16 or SEQ ID NO: 18, or a recombinant peptide having at least 90% sequence identity to SEQ ID NO: 16 or to SEQ ID NO: 18. In another embodiment, the present disclosure provides for an isolated polynucleotide, wherein the isolated polynucleotide comprises, or consists essentially of, or consists of SEQ ID NO: 15 or SEQ ID NO: 17, or an isolated polynucleotide having at least 90% sequence identity to SEQ ID NO: 15 or to SEQ ID NO: 17.

[0100] In another embodiment, the present disclosure provides for an isolated peptide, wherein the isolated peptide comprises, or consists essentially of, or consists of an amino acid sequence selected from SEQ ID NO: 16 or SEQ ID NO: 18, or an isolated peptide having at least 90% sequence identity to SEQ ID NO: 16 or to SEQ ID NO: 18.

[0101] High Throughput Constructs

[0102] Use of the SIMPL construct of the present disclosure allows for high-throughput, quantifiable measurements of PPI.

[0103] The constructs / systems and methods of the present disclosure provide for detection of physiological PPIs (see Fig. 3A) and detection and follow up of weak or transient PPIs.

[0104] Applications

[0105] (1) Protein-to-Protein Interactions (PPI)

[0106] The SIMPL2 constructs / systems of the present disclosure may be used as a high-throughput screening technology for the identification of PPI of any proteins.

[0107] In addition SIMPL2 is sensitive enough to detect subtle changes in protein interactions, which can differ slightly depending on the presence or absence of various stimuli, like hormones or agonists, or inhibitory drugs. Specifically, SIMPL2 follows the kinetic process of kinase / substrate interactions.

[0108] In one embodiment, the present disclosure provides for a method for detecting the interaction between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein) comprising:

[0109] (a) providing the system according to an embodiment of the present disclosure; (b) incubating the first construct and the second construct in a homogenous liquid phase under conditions that allow the formation of the GP41 -1; and

[0110] (c) adding to the incubate of (b) a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc, whereby detection of the luminescence signal being indicative that the bait protein or part thereof and the prey protein or part thereof interact.

[0111] In one embodiment, the method further comprises comparing the luminescence signal with the luminescence signal of a control to determine a binding strength between the bait protein and the prey protein.

[0112] (2) Drug Screening Platform

[0113] SIMPL2 can be used as a drug screening platform suitable for the identification of small molecule inhibitors or enhancers that alter a defined set of membrane protein interactions in their natural environment.

[0114] In another embodiment, the present disclosure provides for a kit of reagents for detecting binding between a first protein (membrane or soluble) or part thereof and a second protein or part thereof (membrane bound or soluble). The kit, in one embodiment, may include: (a) a host cell; (b) a first bait vector (bait), which may be maintained episomally or integrated into the genome of the host cell, comprising a first nucleic acid coding for a bait protein or part thereof, β10 peptide and a modified IN or a modified IC, the first bait vector may further include a promoter; (c) a second vector (prey), which may be maintained episomally or integrated into the genome of the host cell, comprising a second nucleic acid coding for a prey protein or part thereof, β9 peptide and a modified IN or modified IC, the second prey vector may further comprise a promoter. In one aspect, the kit further includes (d) a plasmid library encoding second proteins or parts thereof.

[0115] As such, in another embodiment, the present disclosure provides for a method for screening a binding inhibitor between two proteins, the method comprising:

[0116] (a) providing the system according to an embodiment of the present disclosure, the bait protein and the prey protein being proteins known to bind; (b) (i) incubating the bait construct, the prey construct and a compound to be tested in a homogenous liquid phase under conditions that allow the formation of the GP41-1 (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound to be tested under conditions that allow the formation of the GP41 -1 (second incubate);

[0117] (c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and

[0118] (d) measuring a difference in luminescence signal between the first incubate and the second incubate. A greater luminescence signal in the second incubate is indicative that the compound is a binding inhibitor between the bait protein and the prey protein.

[0119] (3) Binding Strength

[0120] SIMPL2 can be used to study the binding strength between two proteins. The luminescence signal generated by two proteins can be measured and quantified using the system of the present disclosure. In one embodiment, the measured luminescence signal is compared to a known control signal in order to determine the binding strength between the two proteins.

[0121] As such, in another embodiment, the present disclosure provides for a method for determining a binding strength between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein), the method comprising: (a) providing a system according to an embodiment of the present disclosure, the bait protein and the prey protein being proteins known to bind;

[0122] (b) (i) incubating the bait construct, the prey construct and a compound known to inhibit the binding between the bait protein and the prey protein in a homogenous liquid phase under conditions that allow the formation of the GP41-1 (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound under conditions that allow the formation of the GP41-1 (second incubate); (c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and

[0123] (d) determining the binding strength between the bait protein and prey protein in the presence of the compound and the binding strength between the bait protein and the prey protein in the absence of the compound based on the luminescence signal of the first incubate and the second incubate.

[0124] Advantages

[0125] The system of the present disclosure presents several advantages over existing techniques to study interactors of membrane proteins: (i) SIMPL2 can be carried out in virtually any cell line due to the availability of prey / bait / reporter vectors for lentivirus generation, which poses the advantage of single copy integration and diminishes overexpression artifacts. Moreover, SIMPL2 is carried out in living cells, thus avoiding signal changes arising from cell lysis or protein purification used in biochemical PPI methods; (ii) SIMPL can be carried out in solution, allowing for a fast, high throughput and quantifiable results; (iii) SIMPL2 can detect subtle changes in interaction patterns, which can be induced / repressed by either drugs, various stimuli or phosphorylation events, in a highly specific manner; (iv) SIMPL2 can be used as a platform for drug discovery, specifically used to screen for novel compounds capable of inhibiting signaling mediated by oncogenic receptors; (v) SIMPL2 can be used in quantitative studies to measure the strength or affinity of PPI; (vi) as splicing occurs in situ, both loss of specific interaction and gain of nonspecific interaction during processing steps, which are common problems for many affinity-based methods such as co-immunoprecipitation and AP-MS, are avoided; and SIMPL can detect PPIs in various cellular compartments.

[0126] In one embodiment, the present disclosure relates to method for detecting the interaction between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein) comprising:

[0127] (a) providing a system according to an embodiment of the present disclosure; (b) incubating the bait construct and the prey construct in a homogenous liquid phase under conditions that allow the formation of the intact protein; and

[0128] (c) adding to the incubate of (b) a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc, whereby detection of the luminescence signal being indicative that the first protein or part thereof and the second protein or part thereof interact.

[0129] In one embodiment of the method for detecting the interaction between a bait protein and a prey protein of the present disclosure, the method further comprises comparing the luminescence signal with the luminescence signal of a control to determine a binding strength between the bait protein and the prey protein.

[0130] In another embodiment, the present disclosure relates to a method for screening a binding inhibitor between to proteins, the method comprising: (a) providing a system according to an embodiment of the present disclosure, the bait protein and the prey protein being proteins known to bind; (b) (i) incubating the bait construct, the prey construct and a compound to be tested in a homogenous liquid phase in the presence under conditions that allow the formation of the intact protein (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound to be tested under conditions that allow the formation of the intact protein (second incubate); (c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and (d) measuring a difference in luminescence signal between the first incubate and the second incubate, wherein greater luminescence signal in the second incubate is indicative that the compound is a binding inhibitor between the bait protein and the prey protein.

[0131] In another embodiment the present disclosure relates to a method for determining a binding strength between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein) in the presence and absence of a compound known to inhibit the binding between the bait protein and the prey protein, the method comprising: (a) providing a system according to an embodiment of the present disclosure wherein the bait protein and the prey protein being proteins known to interact; (b) (i) incubating the bait construct, the prey construct and the compound in a homogenous liquid phase in the presence under conditions that allow the formation of the intact protein (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound under conditions that allow the formation of the intact protein (second incubate); (c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and (d) determining the binding strength between the bait protein and prey protein in the presence of the compound and the binding strength between the bait protein and the prey protein in the absence of the compound based on the luminescence signal of the first incubate and the second incubate.

[0132] In another embodiment, the present disclosure provides for a kit for testing binding between a first protein or part thereof and a second protein or part thereof comprising:

[0133] (a) a first construct comprising the first protein or part thereof and a first handle connected to the first protein or part thereof, the first handle comprising (i) a β10 peptide of a nanoluciferase (NanoLuc) fused to the N-terminal side of a modified N-terminus intein of GP41-1 (IN), or (ii) a β9 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and

[0134] (b) a second construct comprising the second protein or part thereof and a second handle fused to the second protein or part thereof, the second handle comprising a modified C-terminus intein of GP41-1 (IC) fused at the C-terminal side (i) to the β9 peptide of the NanoLuc when the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified IN, or (ii) to the β10 peptide of the NanoLuc when the β9 peptide of the NanoLuc is fused to the N-terminal side of the modified IN; (c) a Δ11S peptide of NanoLuc, and

[0135] (d) a substrate that generates a luminescence signal in the presence of NanoLuc. In one embodiment, the modified IN and modified IC of the kit of the present disclosure comprises a modified IN and a modified IC of any embodiment of the present disclosure.

[0136] In one embodiment of the kit of the present disclosure, the kit comprises a cell line, including mammalian and non-mammalian cells, preferably mammalian cells, expressing a first construct and a second construct according to an embodiment of the present disclosure, a Δ11S peptide of NanoLuc, and a substrate that generates a luminescence signal in the presence of NanoLuc.

[0137] In order to aid in the understanding and preparation of the present disclosure, the following illustrative, non-limiting examples are provided.

[0138] EXAMPLES

[0139] Example 1

[0140] Materials and Methods

[0141] Table 1. Reagents and Tools

[0142]

[0143]

[0144] Molecular cloning and library preparation

[0145] SIMPL2 vectors were generated by inserting (39 or [310 sequence (each containing 11 amino acids) into the original SIMPL vectors using Gibson assembly (Gene Assembly Master mix, New England BioLabs, E2611). In brief, linear DNA fragments, each of which contains a sequence of 30 nucleotides at its 3’ end identical to the 5’ end of another fragment, were created by PCR. The fragments were incubated with the Gibson reaction mixture, which contains 5’ exonuclease, polymerase and ligase. The exonuclease removes nucleotides from the 5’ end of each fragment. The resulted single strand anneals with the corresponding complementary sequence of another fragment. The annellation allows strand extension at each 3’ end by the polymerase. Finally, the nick is sealed by the ligase. Four SIMPL2 vectors with different orientations of GP41-1 split intein and tNLuc tags were used in this study (Fig. 9 and Table 3). All of them contain an attR cassette compatible for LR Gateway cloning. pZY1001: attR-|310-IN. pZY1004:

[0146] [310-IN-attR. pZY1003: IC-|39-attR. pZY1006: attR-IC-|39. These cassettes are under the control of the third generation Tet-On promotor. In addition, all the vectors contain an expression cassette including RPBSA promoter, rtTA, P2A and an antibiotics marker, which is either HygR in pZY1001 and pZY1004, or PuroR in pZY1003 and pZY1006. Similarly, NanoBiT vectors were created: pZY1007: LgBiT-attR; pZY1008: attR-LgBiT; pZY1009: SmBiT-attR; pZY1010: attR-SmBiT (Fig. 9). The above vectors were generated using Gibson assembly. Most cDNAs of target genes were originally obtained from human ORFeome collection or from the Openfreezer collection at Lunenfeld-Tanenbaum Research Institute. Those not in entry clone vectors were cloned into pDONR223 by PCR and BP cloning. Different cDNA fragments were then cloned into SIMPL2 vectors by LR reactions. Some DNA constructs are listed in Table 3. All cDNA sequences were verified using Sanger sequencing.

[0147] CRISPR plasmids were prepared by insert guide sequence into pX330 vector as described46. Guide sequences used in this study are listed as follow:

[0148] CCR5: TAATAATTGATGTCATAGAT (SEQ ID NO: 25)

[0149] ROSA26: GTCGAGTCGCTTCTCGATTA (SEQ ID NO: 26) Temp2 (specific cleavage site in SIMPL2 vector): GATTTCTTCTTGCGCTATCT (SEQ ID NO: 27)

[0150] Cell culture and treatment

[0151] HEK 293T cells were grown in DMEM supplemented with 10% fetal calf serum. For SIMPL2 assay, cells were seeded in 384 well plates with 5,000 - 10,000 cells per well. For Western blot analysis, HEK 293T cells were seeded in 24 well plates and transfected with various plasmids with polyethylenimine (PEI). After 24 hours of transfection, the cells were treated with doxycycline (1 pg / ml) for overnight followed by lysis and Western blot analysis.

[0152] Stable cell preparation

[0153] HEK 293T cells were grown in a 24 well plate. Mix a SIMPL2 plasmid (125 ng), Temp2-pX330 (62.5 ng) and CCR5 / ROSA26-pX330 (62.5 ng) with 0.75 ul PEI in 0.25 ml PBS and incubate for 30 minutes. The DNA / PEI mixture was added in the cell culture. After 24 hours, the cells were trypsinized and reseeded in a 10 cm plate with a medium containing hygromycin or puromycin according to the antibiotics resistant marker in the SIMPL2 plasmid. After colonies appeared, they were individually picked followed by validation of protein expression.

[0154] Western blot analysis and immunoprecipitation

[0155] Cells were lysed in buffer H (Triton X 100 1%, [3-glycerophosphate pH7.3 50 mM, EGTA 1.5 mM, EDTA 1 mM, orthovanadate 0.1 mM, DTT 1 mM supplemented with protease inhibitors (Roche)). After centrifugation at 21,000 x g for 10 min, the supernatants were mixed with Laemmli sample buffer, boiled at 95 °C for 3-5 min and subjected to Western blot analysis. Antibodies used for Western blot analysis were: a-FLAG antibody purchased from Sigma-Aldrich (F1804) with 1:10,000 dilution and a-V5 antibody from Cell Signaling Technology (#13202) with 1:10,000 dilution. Images were acquired using ChemiDoc MP system and its accompanied software Image Lab Touch.

[0156] Specificity of A11S

[0157] Recombinant proteins containing (39, [310, or [310- [39 tNLuc tag C-terminally fused to protein G (C2 domain) and His tag were expressed in BL21-gold (DE3) E. Coli bacteria. The proteins were purified using Nickel Sepharose according to the manufacturer’s instruction. The proteins were diluted at varied concentrations, mixed with A11S protein in a buffer containing Tris pH7.5 20 mM, EDTA 2 mM, NaCI 25 mM, TCEP 0.5 mM, Tween200.1%, albumin 0.05%, and incubated for 3 hours. Furimazine was then added in the mixture followed by luminescence measurements. Final concentration of A11S is 250 pM and furimazine is 25 pM. Kd of A11S with each tag was calculated through fitting the data to a model of specific binding with Hill slope using GraphPad.

[0158] SIMPL2 assay

[0159] Culture medium of tested cells in 384 well plate was aspirated. An aliquot (50 pl) of reaction mixture containing A11 S (0.4 pM) and furimazine (50 pM) in reaction buffer (Tris pH7.5 20 mM, EDTA 1.2 mM, TCEP 0.24 mM, NaCI 30 mM, NP40 0.15%) was added into each well. After 1.5-hour incubation at room temperature, luminescence was read in a Biotek Synergy luminescence microplate reader with integration time of one second.

[0160] SIMPL2 assay for benchmark reference

[0161] HEK293 cells were seeded in 384 well plates with 5,000 cells / well. SIMPL2 plasmids of genes under investigation were transfected into the cells in triplicates using PEI. Two genes (B and P) in SIMPL2 constructs (B-|310-IN and IC-|39-P) were co-transfected to measure the interaction within the corresponding protein pair. In parallel, each individual gene was co-transfected with a corresponding IN / IC handle to roughly measure the expression level. After 24 hours, doxycycline (0.5 pg / ml final concentration) was added for 16 hours incubation. The samples were subject to SIMPL2 assay. Comparison of SIMPL2 interaction signals was performed in three ways as follows. (1) Raw signals. (2) Normalized signals. Here the interaction signals were normalized (divided) by protein expression estimated using IN / IC handle. (3) Data transformed based on RRS distribution. For this, the linearity of SIMPL2 signals (in logarithmic scale) of RRS with their corresponding protein expression, in term of the product of two involved proteins (in logarithmic scale), was verified using simple linear regression. The derived linear equation (y = a + 3x) of this regression was used to transform SIMPL2 signals of both RRS and PRS pairs using the formula Stransformed = Sraw- a - (3SBP. in which Stransformed is the interaction signal after transformation, Sraw is the raw interaction signal, SBP is the product of B and P expressions, and all terms are in logarithmic scale.

[0162] SIMPL2 assay for characterizing PPI modulators

[0163] Stable cells of target protein pair were created using Cas9 CRISPR. To characterize PPI modulators, the cells were seeded in 384 well plate (10,000 cells / well) and incubated for 5 hours. Subsequently, they were treated with doxycycline (1 pg / ml) and tested compound for 16 hours (sotorasib, AG1478, afatinib, selumetinib or RO-5963). In some experiments, the cells were treated with doxycycline and compound such as BI2865 for 5 hours. Average signal of wells without either deoxycycline or compound treatment was used as blank or ‘floor’ signal (SF). The average signal of wells with deoxycycline treatment but without compound treatment was used as maximal interaction or ‘ceiling’ (Sc). SIMPL2 reading of each sample (Ssampie) was normalized according to the formula S*sampie = (Ssampie - SF) / fSc - SF) X 100%, in which S*sampie stands for normalized signal. The normalized data of inhibitor experiments were fitted into a four-parameter inhibitor-response model. For PROTAC characterization, the cells were pretreated with doxycycline (1 pg / ml) for 16 hours followed by ARV-825 treatment for 3 hours in absence or present of MG132 (1 pM).

[0164] NanoBiT assay

[0165] Stable cells of target protein pair were created using Cas9 CRISPR. The cells were grown in 396 well plate with OptiMEM medium containing 5% FBS and doxycycline (0.5 pg / ml) for 16 hours, followed by treatment of rapamycin for 2 hours, BI2865 for 5 hours or ARV-825 for 3 hours. Furimazine was then added to the final concentration of 31.25 pM. After 30 minutes, signals were read in a luminescence microplate reader.

[0166] Statistics

[0167] Curve fitting to various mathematical models, receiver operating characteristic and student’s t test analysis were performed using GraphPad.

[0168] RESULTS

[0169] Incorporation of tNLuc into the SIMPL system Nanoluciferase (NanoLuc) is a small sized enzyme, originally discovered in deep sea shrimp, that is widely used in biochemical research2324. The activity of NanoLuc does not require ATP, and its engineered version uses furimazine as a substrate, producing the most brilliant luminescence among all the known luciferases. This makes it valuable as a highly sensitive reporter in many applications, including various PPI detection tools. The approach of splitting NanoLuc into fragments has also been used to study PPIs, exemplified by the two fragment NanoBiT approach25. Recent studies have further built upon this with the development of a new version of split NanoLuc, called tri-part NanoLuc (tNLuc)21’22, where NanoLuc is split into three fragments: A11 S and two C-terminal 11 amino acid peptides ([39 and [310). In this format, [39 and [310 are separately fused to bait and prey proteins. The luciferase is reconstituted upon bait / prey interaction in presence of A11S, allowing for production of luminescence upon additionof substrate. tNLuc is characterized by reduced background signal and minimized potential to perturb tagged proteins due to the small size of the [39 and [310 peptides24. tNLuc has also been successfully applied in serological assays26-29

[0170] We inserted a [310 tag at the N-terminal side of the IN tag construct, and correspondingly inserted a [39 tag at the C-terminal side of IC tag construct (Fig.5A). It should be noted that a pair of common peptide tags for antibody recognition are still retained in the constructs solely for later validation of the system: a V5 tag in the IN construct and a FLAG tag in the IC construct. It should be understood that the V5 tag and the FLAG tag are not necessary for the study of PPI using SIMPL2. Interaction between the two tagged proteins of interest (‘B’ and ‘P’) induces splicing, ligating B and P together and arranging the [310 and [39 peptides in a tandem position. When A11S is externally added into the reaction system, it will bind the [310- [39 peptide to form a stable and active luciferase complex and produce luminescence signal. A schematic diagram of the new SIMPL2 system is provided in Fig. 1A. To validate our SIMPL2 system, we first characterized tNLuc in the SIMPL setting because the specificity and affinity of A11 S had previously not been strictly assessed. Thus, the binding kinetics of A11 S with different tNLuc tags were determined using recombinant protein G fused to (39 peptide alone, [310 peptide alone, or a tandem |310-|39 peptide (Fig. 1B). No significant binding was detected with [39 or [310 peptide tags alone within the measured concentration range (0.1 - 300 nM). In contrast, the (310-(39 peptide fusion tag resulted a binding curve with a Kd of 11.3 nM (Fig. 1 B), suggesting strong affinity and exclusive specificity of A11S towards only (31O-(39, but not (39 or [310 sequence. Therefore, the tNLuc readout should not be influenced by individually tagged protein partners in the absence of interaction.

[0171] Like our previous version of SIMPL, SIMPL2 involves co-expression of tagged ‘B’ and ‘P’ proteins of interest in cultured mammalian cells to allow the test PPIs to occur in their native environment. The subsequent detection of spliced protein is initiated by addition of a mixture of furimazine and A11S, alongside detergent which permeabilizes the cells to allow A11S uptake. After incubation, the luminescence signal is then recorded (Fig. 1C). We first evaluated the system using the rapamycin-induced dimerization of FKBP1A and FRB (the FKBP rapamycin-binding domain of mTOR) and compared it with the NanoBiT assay (Dixon et al, 2016). For this purpose, a stable SIMPL2 cell line of FRB-[310-IN / IC-(39-FKBP1A and a stable NanoBiT cell line of FRB-LgBiT / FKBP1A-SmBiT were created using CRISPR-Cas9. For the assay, the stable cells were simultaneously treated with rapamycin at different concentrations for two hours followed by measurement of either tNLuc or NanoBiT luminescence. Kinetics analysis revealed very similar EC50 values, 5.0 nM for SIMPL2 assay and 5.2 nM for NanoBiT assay, indicating their comparable performance in this rapamycin induced interaction (Fig. 1 D)

[0172] We also compared SIMPL2 with SIMPL coupled with ELISA in two formats: FRB-IN / IC-FKBP1A (Fig. EV1B) and FRB-IN / FKBP1A-IC (Fig. EV1C) and obtained EC50 values for rapamycin of 5.5 nM and 6.2 nM respectively in SIMPL2 assay. They are about one order of magnitude lower than those obtained using the SIMPL-ELISA platform19, suggesting a significant improvement in the sensitivity. An approximately 5-6-fold improvement was also observed in the signal / background ratio for SIMPL2 versus the previous ELISA-based readout. We therefore conclude that coupling SIMPL with tNLuc substantially improves not only the operability of the assay but also its quantifiability.

[0173] Evaluation of SIMPL2 using reference PPI sets

[0174] True SIMPL signal is potentially confounded by non-specific splicing derived from intrinsic affinity between the IN and IC intein fragments. To develop SIMPL, the GP41-1 split intein used was therefore engineered to minimize this intrinsic affinity to a residual level19. However, if the proteins under study are sufficiently stable, even minimal levels of non-specifically spliced product could accumulate and generate background due to the irreversible nature of splicing. In this case, such non-specific splicing should be proportional to the abundance of the parental proteins. To test this hypothesis, we co-expressed [310-IN-tagged B proteins alongside an ‘IC handle’ construct consisting of only IC-|39, or similarly coexpressed IC-[39-tagged P protein alongside an ‘IN handle’ consisting of only [310-IN peptide (Fig. 5A). The cell lysates were then subjected to Western analysis using anti-FLAG and anti-V5 antibodies to identify total vs. spliced protein products (Fig. 6A and 6B). Here, the splicing signals reflect only the basal interactions driven by IN / IC affinity because the IN and IC handles do not contain any extra protein sequences. As expected, the blots showed that splicing signals were well correlated with total protein B or P expression. This observation suggests it is necessary to distinguish true PPI signal from background signal resulting from varied expression of the involved proteins. On the other hand, the signal produced by IN / IC handles could also be used to estimate protein expression, which can be harnessed to help modify SIMPL2 signal to provide a better measurement of true interactions.

[0175] We next used reference PPIs to evaluate the performance of the SIMPL2 system. The Homo sapiens positive reference set version 1 (hsPRS-v1) was previously created for the purpose of method evaluation42. The hsPRS-v2 (n=60) is an upgraded version with removal of interaction pairs with false annotations or of incorrect DNA sequences30. Since it is well accepted, hsPRS-v2 was used as the PRS in our current study47. The corresponding random reference set (RRS) contains 80 protein pairs with baits and preys selected from hsPRS-v1 but used in combinations computationally determined to have low probability of interaction. In the same experiment, each bait or prey was also tested with corresponding IN / IC handle to obtain an estimate of its expression. Comparing the SIMPL2 results using the PRS vs. RRS revealed a statistically significant difference in the overall raw mean signal between the two datasets (Fig. 2A). Receiver operating characteristic (ROC) analysis (Fig. 2B) produced an area under curve (AUC) value of 0.769. However, overlap in signal between the two reference sets was high (Fig.

[0176] 2A), making it challenging to distinguish more true positives from background.

[0177] We next used scatter plots to compare the SIMPL2 signals obtained with the reference sets to the corresponding expression of proteins B (SB, Fig. 6C) or P (SP, Fig. 6D) estimated by co-expressing B and P proteins with IC or IN handle, respectively. Notably, PRS vs. RRS did show certain separation in this two-dimensional space. This separation was further improved upon plotting against the product of B and P expressions (SBP, or SB+SP since they are in logarithmic scale in the plot, Fig. 2C), suggesting that normalizing SIMPL2 signal based on B and P expression could improve assay performance. Thus, we normalized the SIMPL2 signals by dividing them by the product of the B and P expression values. This normalization made the RRS signals more convergent, allowing better separation of the two sets as shown in a box-whisker plot (Fig. 2D). ROC analysis also demonstrated a mild improvement in AUC (AUC = 0.790) (Fig. 2E).

[0178] We noted that, in the scatter plot of SIMPL2 vs. SB (Fig. 2C), RRS values were distributed in a tight linear pattern (R2= 0.821). This fit confirms a strong relationship between estimated protein expression level and background signal, and more importantly the linear distribution supports the assumption that all RRS pairs follow the same non-specific kinetics. However, the fitted linear regression equation of raw RRS signals presented a slope far different from 1.0 (Fig.2C), and correspondingly normalization using SB did not bring RRS to an even level (Fig. 2 F), suggesting the process deviates from the standard second-order mass action kinetics, a model laying the basis for the normalization approach. We reasoned that the empirical regression equation of RRS (Fig. 2C) may more precisely describe the true behavior of non-specific binding, and therefore it can be used to remove the contribution of non-specific binding to a SIMPL2 signal. Indeed, performing an RRS-based transformation flattened the RRS signal fit in the scatter plot (Fig. 2I) and produced maximal convergence in RRS signal values, thereby allowing for the best separation (Fig. 2G). ROC analysis furthermore demonstrated a significantly improved AUC (0.864) (Fig. 2H). The performance difference among these three data processing methods can be further intuitively seen by comparing the sensitivity values obtained at certain specificity (Fig. 6E), or the specificity values obtained at certain sensitivity (Fig. 6F), derived from the ROC curves (Figs.

[0179] 2B, 2E, 2H): the transformed set showing the best sensitivity and specificity in all conditions.

[0180] Finally, we compared the performance of SMIPL2 with SIMPL-ELISA. Our pervious SIMPL-ELISA analysis used a PRS set containing 88 protein pairs and an RRS set of 88 pairs19. A SIMPL2 assay on the same reference sets presented a ROC curve of transformed data with slightly better AUC than SIMPL-ELISA (0.86 for SIMPL2 vs. 0.81 for SIMPL-ELISA) (Fig. 6G). However, a significant improvement was observed in sensitivity when maintaining a high specificity. For example, and the point of 100% specificity, SIMPL2 allowed a ~50% sensitivity while only ~33% was achieved in SIMPL-ELISA assay. Thus, this framework of RRS-based transformation appears to be an effective approach for enhancing the sensitivity and specificity of SIMPL2 assay.

[0181] SIMPL2 as an efficient tool for identifying PPI modulators

[0182] The operational, quantitative, and modular features of SIMPL2 suggest that it can potentially work as a powerful cell-based platform for identifying PPI modulators. Since stable cells are best suited for this type of application (as their use both simplifies assay setup and reduces variability) we developed cell lines stably expressing target protein pairs of interest in SIMPL2 format. To prevent the occurrence of target PPIs before treatment with PPI modulators, expression of the SIMPL2 constructs was placed under the control of a Tet-on promoter, which can only be induced by addition of doxycycline.

[0183] We first evaluated the performance of SIMPL2 using mutant specific KRAS G12C inhibitor sotorasib (AMG 510), which irreversibly inhibits KRAS G12C mutant through covalent binding, locking it into an inactive conformation and thereby preventing interaction with RAF proteins31. For our study, we created several reporter cell lines stably expressing IC-[39-KRAS (WT or oncogenic KRAS mutants: G12C, G12D, G12R, G12V and Q61H) and the Ras-binding domain (RBD) of RAF1 in the RBD-[310-IN format. Nonrelated cell lines expressing HRAS / RBD or EGFR / SHC1 were used as controls. Expression of the target proteins was induced by doxycycline in the presence of co-treatment with different concentrations of sotorasib and the target interactions were evaluated as SIMPL2 luciferase signals (Fig. 3A). The efficacy of sotorasib on KRAS G12C mutant was verified as a dose-responsive inhibition of KRAS / RBD interaction signal with an IC50 of 74 nM, a value consistent with previous reports31. In contrast, no apparent inhibition of WT KRAS or other mutants was observed within the range of sotorasib concentrations tested, consistent with known sotorasib specificity31. Importantly, no significant inhibition was observed in HRAS / RBD and EGFR / SHC1 cells. SIMPL2 was further compared with NanoBiT using the pan-KRAS inhibitor BI286544(Fig. 7A). The high similarity of the obtained IC50 values (37 nM for SIMPL2 and 24 nM for NanoBiT) demonstrates that the assays have comparable performance in evaluating RAS inhibitors.

[0184] We also evaluated two EGFR tyrosine kinase inhibitors (TKIs), AG1478 (tyrphostin), a prototype EGFR inhibitor that functions through competitive binding to ATP-binding sites32, and afatinib, a second-generation EGFR inhibitor operating through irreversible binding to EGFR33. These molecules indirectly inhibit EGFR PPIs by antagonizing EGFR tyrosine kinase activity, blocking its autophosphorylation and subsequent interaction with downstream signaling molecules such as SHC1. Two reporter cell lines stably expressing EGFR (WT or oncogenic L858R mutant)-[310-IN and IC-[39-SHC1 were created for this study to examine whether these inhibitors exert differential effects on WT and the oncogenic mutant. The EGFR / SHC1 interaction was evaluated as SIMPL2 luciferase signal while treated with different concentrations of the EGFR inhibitors, revealing a distinct dose-responsive inhibition of EGFR / SHCI interaction signal (Fig. 3 B, 30). More specifically, AG1478 inhibited WT EGFR with an ICso of ~100 nM, and L858R mutant with 2.9 nM, while afatinib displayed more potent inhibition, with an ICso of 4.32 nM for WT EGFR and 0.32 nM for L858R mutant (Table 2) which are in line with previously reported profiles45. As expected, neither compound showed any significant inhibitory effect on the KRAS(G12C) / RBD interaction within the tested concentration range (Fig. 3B, 3C, Table 2). Treatment of the same set of cells with selumetinib (AZD6244), a MEK inhibitor, was performed here as a negative control, and showed no effect on the interaction of either EGFR WT or L858R mutant (Table 2). The inhibitory effects of AG1478 and afatinib on WT and L858R EGFR were also verified by Western blot analysis, in which the splicing between EGFR and SHC1 was reduced or abolished by inhibitor treatment in a trend consistent with the SIMPL2 results (Fig. 7B). Collectively, these results demonstrate the ability of SIMPL2 to detect the specific action of inhibitors against their corresponding targets, making it a potentially useful tool for rapid PPI- and enzymatic- inhibitor discovery.

[0185] We next tested the system with molecular glues, which are small molecule compounds that induce target protein interaction. In our current study, we used rapamycin to characterize SIMPL2 (Fig.lD). Notably, the natural product rapamycin is a well-studied molecular glue for FKBP12A / FRB dimerization, and our results support the suitability of molecular glues. However, we also tested another molecular glue, RO-596334, which operates through a mode of action distinct from that of rapamycin. As a p53 inhibitor, RO-5963 induces homo- or hetero-dimerization between MDM2 or MDM4 (MDMX). For our experiments, cells co-expressing MDM2-[310-IN and IC-[39-MDM4 were treated with RO-5963 at different concentrations and then subjected to the SIMPL2 assay (Fig. 4A). The MDM2 / MDM4 interaction was observed at RO-5963 concentrations above 0.3 pM, reaching a peak at 3 pM (Fig. 4A, red line). The interaction was not observed, however, when cells were treated with control compound selumetinib (Fig. 4A, blue line). Control interaction between KRAS G12C and RBD was also not affected upon RO-5963 treatment, demonstrating the ability of SIMPL2 to detect the specificity of RO-5963 (Fig. 4A, black line). Interestingly, the interaction started to decline above 3 pM. This hook effect, however, can be explained by the mode of action of RO-596334. To clarify, the two protein molecules are indeed anchored together through a pair of dimerized RO-5963 molecules. When excess RO-5963 molecules are present, however, the balance of binding reaction is driven toward the production of mono-protein bound RO-5963 molecules. In contrast, rapamycin engages its targets most likely through a sequential pathway, wherein it first binds to FKBP12A followed by the interaction with FRB35Therefore, its binding kinetics display a monophasic change.

[0186] We also examined the ability of SIMPL2 to monitor the activity of PROTACs. These molecules are different from molecular glue, consisting of two active heads (one for binding target and the other for binding E3 ligase) linked through a linear chain, and are specialized for use in directing targets for ubiquitination and degradation by the proteasome. We evaluated SIMPL2 performance for characterizing PROTACs using ARV-825, a well characterized PROTAC that recruits the bromodomain protein BRD4 to the E3 ligase cereblon (CRBN)36. Cells co-expressing [310-IN-CRBN and IC-[39-BRD4 were treated with ARV-825 at different concentrations (Fig. 4B). Since ARV-825 treatment leads to proteasome-mediated BRD4 degradation, a set of cell samples were also co-treated with the proteasome inhibitor MG132. A positive correlation between ARV-825 concentration and BRD4 / CRBN interaction was observed, verifying the activity of ARV-825 (Fig. 4). The same effect was also verified using Western blot analysis (Fig. 8A). The SIMPL2 signals were profoundly enhanced by MG132 treatment, further confirming the in vivo functionality of ARV-825. On the other hand, PROTACs usually produce hook effects at high concentration16which was also captured by the assay especially when MG132 was added. We further compared the capability of NanoBiT in detecting the activity of ARV-825 (Fig. 8B). Performance of SIMPL2 and NanoBiT in measuring PROTAC activity was further compared by plotting the luminescence readings of control and ARV-825 (1000 nM) treated samples obtained from both assays (Fig. 8C). Although ARV-825-induced proximity could be captured by NanoBiT, the signals were low (two orders of magnitude lower than SIMPL2 signals), leading to a low signal / background ratio, 2.6 vs. 7.8 produced by SIMPL2 (Fig. 8C). This result suggests a potential difficulty of NanoBiT in producing statistically significant data in high throughput studies involving weak PPIs or induced proximity, however further work with additional candidates will need to be performed in order to determine the extent of this difference.

[0187] Based on the above results, we therefore conclude that SIMPL2 is well-suited for the sensitive detection and monitoring of both PROTACs and molecular glues, as well as PPI inhibitors.

[0188] The present disclosure describes the d SIMPL2 platform, coupling theSIMPL system to an improved tNLuc readout. SIMPL2 offers advantages over SIMPL. For example, in SIMPL2 reactions occur in the homogenous liquid phase, making the process quick and easy to perform. Additionally, expensive reagents such as labelled antibodies are not required, greatly reducing cost. This makes the SIMPL2 system more labor- and cost-effective (and better suited for use in HTS formats) than the original SIMPL, which relied on Western blotting, ELISA or HTRF readouts.

[0189] One challenge with SIMPL is the non-specific signals caused by low intrinsic affinity between the IN and IC intein fragments used in the tags. Although the affinity has been greatly reduced via targeted protein engineering, it is still possible for fusion product to accumulate to non-negligible levels when using very highly expressed proteins, due to the irreversible nature of intein-mediated splicing. Taking the advantage of the strong quantifiability provided by the SIMPL2 system, however, we were able to set up a framework using this background as an indirect measure of protein expression, which could be used to modify signal and help distinguish true vs non-specific PPIs. This framework was evaluated using benchmark PPI reference sets and provided results favouring signal from the PRS vs. the RRS; specifically, 48% of PPIs within the PRS were identified at a threshold where no RRS interactions were reported (with ROC analysis presenting an AUC value of 0.86). These parameters demonstrate good behavior of SIMPL2 in detecting true PPIs, although they cannot be directly compared to those obtained from other methods30’37’38due to the differences in the RRS employed. It should also be noted that there are a total of eight different combinations of bait / prey tagging orientation possible: B-|310-IN / IC-|39-P, B-|310-IN / P-IC-|39, (310-IN-B / IC-(39-P, (31O-IN-B / P-IC-(39, IC-(39-B / P-(39-IN, IC-(39-B / (39-IN-P, B-IC-(39 / P-(39-IN and B-IC-|39 / |39-I N-P. In this proof-of-technology study, however, we only explored one tagging strategy. Thus, the sensitivity can potentially be further improved in future work via exploration of all possible tagging orientations as demonstrated by previous studies3038Thus, we conclude that SIMPL2 can be readily applied in high throughput PPI discovery by screening IN- or IC-fused open reading frame (ORF) libraries in an array-based format.

[0190] We also validated our SIMPL2 system for use as drug discovery tool. With the advances in our understanding of PPI mechanisms, and the importance of these mechanisms in mediating disease states, it is now more important than ever to translate this knowledge into practical medical applications. To this end, new strategies for manipulating PPIs have opened up avenues for hitting disease targets previously considered undruggable. Indeed, PPI targeted drug discovery has recently become an intensive area of research in both the pharmaceutical industry and academia, and PPI detection methods have become increasingly important players in the early phases of this type of drug discovery. However, although numerous PPI methods are available, many of them are not ideally suited for hit identification and characterization. In particular, traditionally employed biophysical and biochemical techniques are accompanied by certain limitations such as high cost, requirement for labeling, inability to recapitulate in vivo conditions, incompatibility with HTS formats and the requirement for highly specialized equipment and expertise, all of which have the potential to slow the pace of drug discovery.

[0191] In contrast, new drug discovery systems such as SIMPL2 provide promising alternatives to help improve the drug discovery paradigm. For instance, the incorporation of the tNLuc readout in SIMPL2 bestows an improved operability and quantifiability, which enhances the scalability of the system to an HTS test format. We have also shown how the system is capable of monitoring different types of PPI modulators, including inhibitors, molecular glues and PROTACs, all within a live cell format that provides immediate feedback on important parameters like toxicity and permeability. Furthermore, the modular design of SIMPL2 allows it to be rapidly adapted to different types of protein targets, irrespective of their cellular localization.

[0192] The last few years have also seen considerable development and maturation of artificial intelligence (Al) and quantum computing39, and its active participation in drug development4041. The power of Al and quantum computing has the potential to greatly alter drug discovery in the future. However, candidates developed by Al / quantum computing-based drug discovery approaches still need a means of rapid experimental validation which avoids a possible bottleneck step in the discovery process. A highly scalable, labor- and cost-effective system such as SIMPL2 may be suitable for such a role.

[0193] It should be noted that SIMPL2 does not come without limitations, however. There are several important issues that need to be considered especially for PPI modulator studies. Firstly, unlike some assays like NanoBiT which can follow PPI kinetics in a real-time manner, SIMPL can only capture interaction ‘snapshots’. Secondly, the SIMPL reporter mechanism involves an irreversible process which captures only protein association but not dissociation. Once splicing occurs, it can not be broken up by an inhibitor. Thus, it is essential to treat cells with an inhibitor ahead of target protein expression. Thirdly, various PPIs may behave differently in SIMPL. Those with fast association kinetics may rapidly undergo splicing in SIMPL and thereby be unsuitable for inhibitor discovery. Finally, like any cell-based assay, SIMPL can sometimes be confounded by off-target effects on cellular activities such as cell viability, transcription, translation, luciferase activity etc. As such, it is imperative that all experiments be well-designed, employing proper controls and quick orthogonal validation of top candidates, to help distinguish direct and indirect modes of action.

[0194] Overall, the new SIMPL2 system offers a wide range of benefits that should make it a suitable tool for advancing both PPI studies and drug discovery endeavors.

[0195] Table 2. IC50 determined by SIMPL2

[0196]

[0197] Table 3. Various plasmid constructs used in Example 1

[0198]

[0199] SEQUENCE LISTING GP41-1 IN (WT) (SEQ ID NO: 1)

[0200] CLDLKTQVQT PQGMKEISNI QVGDLVLSNT GYNEVLNVFP KSKKKSYKIT LEDGKEIICS EEHLFPTQTG EMNISGGLKE GMCLYVKE GP41-1 IC (WT) (SEQ ID NO: 2)

[0201] MMLKKILKIE ELDERELIDI EVSGNHLFYA NDILTHN GP41-1 IN (C25) (SEQ ID NO: 3)

[0202] CLDLKTQVQT PQGMKEISNI QVGDLVLSNT GYNEVLNVFP KSKKKSYKIT LEDGKEIICS EEHLFPTQTG EMNISGGLKE GMCLYVKEMM LKKILKIEEL GP41-1 IC (C25) (SEQ ID NO: 4)

[0203] DERELIDIEV SGNHLFYAND ILTHN GP41-1 IN (C24) (SEQ ID NO: 5)

[0204] CLDLKTQVQT PQGMKEISNI QVGDLVLSNT GYNEVLNVFP KSKKKSYKIT LEDGKEIICS EEHLFPTQTG EMNISGGLKE GMCLYVKEMM LKKILKIEEL D GP41-1 IC (C24) (SEQ ID NO: 6) ERELIDIEVS GNHLFYANDI LTHN

[0205] GP41-1 IN (C23) (SEQ ID NO: 7)

[0206] CLDLKTQVQT PQGMKEISNI QVGDLVLSNT GYNEVLNVFP KSKKKSYKIT LEDGKEIICS EEHLFPTQTG EMNISGGLKE GMCLYVKEMM LKKILKIEEL DE GP41-1 IN (C23) (SEQ ID NO: 8)

[0207] RELIDIEVSG NHLFYANDIL THN

[0208] A11S nucleotide sequence (SEQ ID NO: 9)

[0209] 1 ATGggccatc atcatcatca tcatcatcat ATGgttttta cacttgaaga ttttgtgggt 61 gattgggaac aaactgctgc atataattta gatcaagttt tagaacaggg tggagtgagt 121 tctcttttac aaaatcttgc agtctcagtt acaccaatac aaagaatagt tagaagtgga 181 gaaaatgcat taaagatcga tatacatgta ataattcctt atgagggact atcagcagac 241 caaatggcac aaattgagga agtattcaaa gttgtatatc cagtagacga tcatcacttt 301 aaagtaatat taccttatgg aactttagta attgatggtg taacaccaaa tatgttaaat 361 tattttggta gaccatacga agggattgca gtttttgatg gaaagaaaat aaccgtaaca 421 ggcactttgt ggaatggaaa taaaataata gatgaaaggt taattacacc tgatTAA A11S amino acid sequence (SEQ ID NO: 10)

[0210] 1 MetGlyHisHisHisHisHisHisHisHisMetValPheThrLeuGluAspPheValGly 21 AspTrpGluGlnThrAlaAlaTyrAsnLeuAspGlnValLeuGluGlnGlyGlyValSer 41 SerLeuLeuGlnAsnLeuAlaValSerValThrProIleGlnArglleValArgSerGly 61 GluAsnAlaLeuLysIleAspIleHisValllelleProTyrGluGlyLeuSerAlaAsp 81 GlnMetAlaGlnlleGluGluValPheLysValValTyrProValAspAspHisHisPhe 101 LysVallleLeuProTyrGlyThrLeuVallleAspGlyValThrProAsnMetLeuAsn 121 TyrPheGlyArgProTyrGluGlylleAlaValPheAspGlyLysLysIleThrValThr 141 GlyThrLeuTrpAsnGlyAsnLys I lei leAspGluArgLeuI leThr ProAsp

[0211] p9 nucleotide sequence (SEQ ID NO: 11)

[0212] 1 GGCTCCATGC TGTTCCGAGT AACCATCAAC AGT

[0213] P9 amino acid sequence (SEQ ID NO: 12)

[0214] 1 GlySerMetLeuPheArgValThrIleAsnSer

[0215] P10 nucleotide sequence (SEQ ID NO: 13)

[0216] 1 GTGAGCGGCT GGCGGCTGTT CAAGAAGATT AGC

[0217] P10 amino acid sequence (SEQ ID NO: 14)

[0218] 1 ValSerGlyTrpArgLeuPheLysLysIleSer

[0219] IN handle nucleotide sequence (SEQ ID NO: 15)

[0220] 1 ATGGGTAAGC CTATCCCTAA CCCTCTCCTC GGTCTCGATT CTACGggtgg cggaggctct 61 GTGAGCGGCT GGCGGCTGTT CAAGAAGATT AGCggaggtg Gatccaccag atccggatac 121 TGCCTGGACC TGAAGACCCA GGTGCAGACC CCTCAGGGCA TGAAGGAGAT CAGCAACATC 181 CAGGTGGGCG ACCTGGTGCT GAGCAACACC GGCTACAACG AGGTGCTGAA CGTGTTCCCC 241 AAGAGCAAGA AGAAGAGCTA CAAGATCACC CTGGAGGACG GCAAGGAGAT CATCTGCAGC 301 GAGGAGCACC TGTTCCCCAC CCAGACCGGC GAGATGAACA TCAGCGGCGG CCTGAAGGAG 361 GGCATGTGCC TGTACGTGAA GGAGATGATG CTGAAGAAGA TCCTGAAGAT CGAGGAGCTG 421 tag

[0221] IN handle amino acid sequence (SEQ ID NO: 16)

[0222] 1 MetGlyLysProIleProAsnProLeuLeuGlyLeuAspSerThrGlyGlyGlyGlySer 21 ValSerGlyTrpArgLeuPheLysLysIleSerGlyGlyGlySerThrArgSerGlyTyr 41 CysLeuAspLeuLysThrGlnValGlnThrProGlnGlyMetLysGluIleSerAsnlle 61 GlnValGlyAspLeuValLeuSerAsnThrGlyTyrAsnGluValLeuAsnValPhePro 81 LysSerLysLysLysSerTyrLysIleThrLeuGluAspGlyLysGluIlelleCysSer 101 GluGluHisLeuPheProThrGlnThrGlyGluMetAsnlleSerGlyGlyLeuLysGlu 121 GlyMetCysLeuTyrValLysGluMetMetLeuLysLysIleLeuLysIleGluGluLeu IC handle nucleotide sequence (SEQ ID NO: 17)

[0223] 1 ATGGACGAGA GAGAGCTGAT CGACATCGAG GTGAGCGGCA ACCACCTGTT CTACGCCAAC 61 GACATCCTGA CCCACAACag ctcctctgac gtgggtggcg gaggctctGG CTCCATGCTG 121 TTCCGAGTAA CCATCAACAG TGAAGCCGCT GCTAAGggag gtggatccga ctacaaagac 181 catgacggtg attataaaga tcatgacatc gattacaagg atgacgatga caagTAA IC handle amino acid sequence (SEQ ID NO: 18)

[0224] 1 MetAspGluArgGluLeuIleAspIleGluValSerGlyAsnHisLeuPheTyrAlaAsn 21 AspIleLeuThrHisAsnSerSerSerAspValGlyGlyGlyGlySerGlySerMetLeu 41 PheArgValThrIleAsnSerGluAlaAlaAlaLysGlyGlyGlySerAspTyrLysAsp 61 HisAspGlyAspTyrLysAspHisAspIleAspTyrLysAspAspAspAspLys

[0225] 9- protein G nucleotide sequence (SEQ ID NO: 19)

[0226] 1 ATGGGCTCCA TGCTGTTCCG AGTAACCATC AACAGTggag gtggatccGG TGGTGGAGGG 61 AGCacctaca aacttGtcAT Taacggtaaa accctgaaag Gtgaaaccac caccgaagct 121 gttgacgctg ctaccgcgga aaaagttttc aaacagtacg ctaacgacaa cggtgttgac 181 ggtgaatgga cctacgacga Cgctaccaaa acctTcaccg taacggaaGG TGGCGGTAGC 241 catcaccacc atcaccacTG A

[0227] 09- protein G amino acid sequence (SEQ ID NO: 20)

[0228] 1 MetGlySerMetLeuPheArgValThrIleAsnSerGlyGlyGlySerGlyGlyGlyGly 21 SerThrTyrLysLeuVallleAsnGlyLysThrLeuLysGlyGluThrThrThrGluAla 41 ValAspAlaAlaThrAlaGluLysValPheLysGlnTyrAlaAsnAspAsnGlyValAsp 61 GlyGluTrpThrTyrAspAspAlaThrLysThrPheThrValThrGluGlyGlyGlySer 81 HisHisHisHisHisHis

[0229] 010-protein G nucleotide sequence (SEQ ID NO: 21)

[0230] 1 ATGGTGAGCG GCTGGCGGCT GTTCAAGAAG ATTAGCggag gtggatccGG TGGTGGAGGG 61 AGCacctaca aacttGtcAT Taacggtaaa accctgaaag gtgaaaccac caccgaagct 121 gttgacgctg ctaccgcgga aaaagttttc aaacagtacg ctaacgacaa cggtgttgac 181 ggtgaatgga cctacgacga cgctaccaaa acctTcaccg taacggaaGG TGGCGGTAGC 241 catcaccacc atcaccacTG A

[0231] 010-protein G amino acid sequence (SEQ ID NO: 22)

[0232] 1 MetValSerGlyTrpArgLeuPheLysLysIleSerGlyGlyGlySerGlyGlyGlyGly 21 SerThrTyrLysLeuVallleAsnGlyLysThrLeuLysGlyGluThrThrThrGluAla 41 ValAspAlaAlaThrAlaGluLysValPheLysGlnTyrAlaAsnAspAsnGlyValAsp 61 GlyGluTrpThrTyrAspAspAlaThrLysThrPheThrValThrGluGlyGlyGlySer 81 HisHisHisHisHisHis

[0233] 010-09-protein G nucleotide sequence (SEQ ID NO: 23)

[0234] 1 ATGGTGAGCG GCTGGCGGCT GTTCAAGAAG ATTAGCggag gtggatccac cagatccgga 61 tacagctcct ctgacgtggg tggcggaggc tctGGCTCCA TGCTGTTCCG AGTAACCATC 121 AACAGTGAAG CCGCTGCTAA Gggaggtgga tccGGTGGTG GAGGGAGCac ctacaaactt 181 GtcATTaacg gtaaaaccct Gaaaggtgaa accaccaccg aagctgttga cgctgctacc 241 gcggaaaaag ttttcaaaca gtacgctaac gacaacggtg ttgacggtga atggacctac 301 Gacgacgcta ccaaaacctT caccgtaacg gaaGGTGGCG GTAGCcatca ccaccatcac 361 cac

[0235] β10-β9-protein G amino acid sequence (SEQ ID NO: 24)

[0236] 1 MetValSerGlyTrpArgLeuPheLysLysIleSerGlyGlyGlySerThrArgSerGly

[0237] 21 TyrSerSerSerAspValGlyGlyGlyGlySerGlySerMe LeuPheArgValThrIle

[0238] 41 AsnSerGluAlaAlaAlaLysGlyGlyGlySerGlyGlyGlyGlySerThrTyrLysLeu

[0239] 61 VallleAsnGlyLysThrLeuLysGlyGluThrThrThrGluAlaValAspAlaAlaThr

[0240] 81 AlaGluLysValPheLysGlnTyrAlaAsnAspAsnGlyValAspGlyGluTrpThrTyr

[0241] 101 AspAspAlaThrLysThrPheThrValThrGluGlyGlyGlySerHisHisHisHisHis

[0242] 121 His

[0243] References

[0244] 1. Yao, Z. et al. A Global Analysis of the Receptor Tyrosine Kinase- Resource A Global Analysis of the Receptor. Mol Cell 65, 347-360 (2017).

[0245] 2. Snider, J. et al. Fundamentals of protein interaction network mapping. Mol Syst Biol 11, 848-848 (2015).

[0246] 3. Rizzolo, K. et al. Features of the Chaperone Cellular Network Revealed through Systematic Interaction Mapping. Cell Rep 20, (2017).

[0247] 4. Lopes, J. P. et al. The role of parkinson’s disease-associated receptor GPR37 in the hippocampus: Functional interplay with the adenosinergic system. J Neurochem 134, (2015).

[0248] 5. Benleulmi-Chaachoua, A. et al. Protein interactome mining defines melatonin MT1 receptors as integral component of presynaptic protein complexes of neurons. J Pineal Res 60, (2016).

[0249] 6. Lim, S. H. et al. CFTR interactome mapping using the mammalian membrane two- hybrid high-throughput screening system. Mol Syst Biol 18, (2022).

[0250] 7. Saraon, P. et al. A drug discovery platform to identify compounds that inhibit EGFR triple mutants. Nat Chem Biol 16, 577-586 (2020).

[0251] 8. Pathmanathan, S., Grozavu, I., Lyakisheva, A. & Stagljar, I. Drugging the undruggable proteins in cancer: A systems biology approach. Current Opinion in Chemical Biology vol. 66 Preprint at https: / / doi. Org / 10.1016 / j.cbpa.2021.07.004 (2022).

[0252] 9. Arkin, M. M. R. & Wells, J. A. Small-molecule inhibitors of protein-protein interactions: Progressing towards the dream. Nature Reviews Drug Discovery vol. 3 301-317 Preprint at https: / / doi.org / 10.1038 / nrd1343 (2004).

[0253] 10. Arkin, M. R., Tang, Y. & Wells, J. A. Small-molecule inhibitors of protein-protein interactions: Progressing toward the reality. Chemistry and Biology vol. 21 1102— 1114 Preprint at https: / / doi. Org / 10.1016 / j.chembiol.2014.09.001 (2014).

[0254] 11. Scott, D. E., Bayly, A. R., Abell, C. & Skidmore, J. Small molecules, big targets: Drug discovery faces the protein-protein interaction challenge. Nature Reviews Drug Discovery vol. 15533-550 Preprint at https: / / doi.org / 10.1038 / nrd.2016.29 (2016). 12. Roberts, A. W., Stilgenbauer, S., Seymour, J. F. & Huang, D. C. S. Venetoclax in Patients with Previously Treated Chronic Lymphocytic Leukemia. Clinical Cancer Research 23, 4527-4534 (2017). 13. Lu, H. etal. Recent advances in the development of protein-protein interactions modulators: mechanisms and clinical trials. Signal Transduction and Targeted Therapy vol. 5 Preprint at https: / / doi.org / 10.1038 / s41392-020-00315-3 (2020).

[0255] 14. Schreiber, S. L. The Rise of Molecular Glues. Cell 184, 3-9 (2021).

[0256] 15. Neklesa, T. K., Winkler, J. D. & Crews, C. M. Targeted protein degradation by PROTACs. Pharmacology and Therapeutics vol. 174138-144 Preprint at https: / / doi. Org / 10.1016 / j.pharmthera.2017.02.027 (2017).

[0257] 16. Pettersson, M. & Crews, C. M. PROteolysis TArgeting Chimeras (PROTACs) — Past, present and future. Drug Discovery Today: Technologies vol. 31 15-27 Preprint at https: / / doi. Org / 10.1016 / j.ddtec.2019.01.002 (2019).

[0258] 17. Liu, X. & Ciulli, A. Proximity- Based Modalities for Biology and Medicine. ACS Central Science vol. 91269-1284 Preprint at https: / / doi.org / 10.1021 / acscentsci.3c00395 (2023).

[0259] 18. Yao, Z., Petschnigg, J., Ketteler, R. & Stagljar, I. Application guide for omics approaches to cell signaling. Nat Chem Biol 11, 387-397 (2015).

[0260] 19. Yao, Z. et al. Split Intein-Mediated Protein Ligation for detecting protein-protein interactions and their inhibition. Nat Commun 11, 2440 (2020).

[0261] 20. Grozavu, I. et al. D154Q Mutation does not Alter KRAS Dimerization. J Mol Biol 434, (2022).

[0262] 21. Ohmuro-Matsuyama, Y. & Ueda, H. Homogeneous Noncompetitive Luminescent Immunodetection of Small Molecules by Ternary Protein Fragment Complementation. Anal Chem 90, 3001-3004 (2018).

[0263] 22. Dixon, A. S., Kim, S. J., Baumgartner, B. K., Krippner, S. & Owen, S. C. A Tri-part Protein Complementation System Using Antibody-Small Peptide Fusions Enables Homogeneous Immunoassays. Sci Rep 7, 8186 (2017).

[0264] 23. Hall, M. P. etal. Engineered luciferase reporter from a deep sea shrimp utilizing a novel imidazopyrazinone substrate. ACS Chem Biol 7, 1848-1857 (2012).

[0265] 24. Oliayi, M., Emamzadeh, R., Rastegar, M. & Nazari, M. Tri-part NanoLucas a new split technology with potential applications in chemical biology: a mini-review.

[0266] Analytical Methods vol. 153924-3931 Preprint at https: / / doi.org / 10.1039 / d3ay00512g (2023).

[0267] 25. Dixon, A. S. etal. NanoLuc Complementation Reporter Optimized for Accurate Measurement of Protein Interactions in Cells. ACS Chem Biol 11, 400-8 (2016). 26. Kim, S. J. et al. Homogeneous surrogate virus neutralization assay to rapidly assess neutralization activity of anti-SARS-CoV-2 antibodies. Nat Commun 13, 3716 (2022).

[0268] 27. Yao, Z. et al. A homogeneous split-luciferase assay for rapid and sensitive detection of anti-SARS CoV-2 antibodies. Nat Commun 12, 1806 (2021).

[0269] 28. Hall, M. P. etal. Toward a Point-of-Need Bioluminescence-Based Immunoassay Utilizing a Complete Shelf-Stable Reagent. Anal Chem 93, 5177-5184 (2021).

[0270] 29. Kim, S. J., Dixon, A. S., Adamovich, P. C., Robinson, P. D. & Owen, S. C.

[0271] Homogeneous Immunoassay Using a Tri-Part Split-Luciferase for Rapid Quantification of Anti-TNF Therapeutic Antibodies. ACS Sens 6, 1807-1814 (2021).

[0272] 30. Choi, S. G. et al. Maximizing binary interactome mapping with a minimal number of assays. Nat Commun 10, (2019).

[0273] 31. Canon, J. etal. The clinical KRAS(G12C) inhibitor AMG 510 drives anti-tumour immunity. Nature 575, 217-223 (2019).

[0274] 32. Yaish, P., Gazit, A., Gilon, C. & Levitzki, A. Blocking of EGF-Dependent Cell Proliferation by EGF Receptor Kinase Inhibitors. Science (1979) 242, 933-935 (1988).

[0275] 33. Li, D. etal. BIBW2992, an irreversible EGFR / HER2 inhibitor highly effective in preclinical lung cancer models. Oncogene 27, 4702-4711 (2008). 34. Graves, B. et al. Activation of the p53 pathway by small-molecule-induced MDM2 and MDMX dimerization. Proc Natl Acad Sci U SA 109, 11788-11793 (2012).

[0276] 35. Banaszynski, L. A., Liu, C. W. & Wandless, T. J. Characterization of the FKBP- rapamycin-FRB ternary complex. J Am Chem Soc 127, 4715-4721 (2005).

[0277] 36. Lu, J. et al. Hijacking the E3 Ubiquitin Ligase Cereblon to Efficiently Target BRD4.

[0278] Chem Biol 22, 755-763 (2015).

[0279] 37. Braun, P. et al. An experimentally derived confidence score for binary protein-protein interactions. Nat Methods 6, 91-97 (2009).

[0280] 38. Trepte, P. et al. LuTHy: a double-readout bioluminescence-based two-hybrid technology for quantitative mapping of protein-protein interactions in mammalian cells. Mol Syst Biol 14, e8071 (2018).

[0281] 39. Vakili, M. G. et al. Quantum Computing-Enhanced Algorithm Unveils Novel Inhibitors for KRAS. ArXiv 2402.08210, (2024).

[0282] 40. Arnold, C. Inside the nascent industry of Al-designed drugs. Nature Medicine vol. 29 1292-1295 Preprint at https: / / doi.org / 10.1038 / s41591-023-02361-0 (2023).

[0283] 41. Marissa Mock, Suzanne Edavettal, Christopher Langmead & Alan Russell (2023) Al can help to speed up drug discovery-but only if we give it the right data Comment. Nature 621: 467-470.

[0284] 42. Venkatesan K, Rual JF, Vazquez A, Stelzl U, Lemmens I, Hirozane-Kishikawa T, Hao T, Zenkner M, Xin X, Goh K II, et al (2009) An empirical framework for binary interactome mapping. Nat Methods 6: 83-90.

[0285] 43. Grozavu I, Stuart S, Lyakisheva A, Yao Z, Pathmanathan S, Ohh M & Stagljar I (2022) D154Q Mutation does not Alter KRAS Dimerization. J Mol Biol 434

[0286] 44. Kim D, Herdeis L, Rudolph D, Zhao Y, Bdttcher J, Vides A, Ayala-Santos Cl, Pourfarjam Y, Cuevas- Navarro A, Xue JY, et al (2023) Pan-KRAS inhibitor disables oncogenic signalling and tumour growth. Nature 619: 160-166.

[0287] 45. Hirano T, Yasuda H, Tani T, Hamamoto J, Oashi A, Ishioka K, Arai D, Nukaga S, Miyawaki M, Kawada I, et al (2015) In vitro modeling to determine mutation specificity of EGFR tyrosine kinase inhibitors against clinically relevant EGFR mutants in non-small-cell lung cancer. Oncotarget 6: 38789-38803.

[0288] 46. Ran FA, Hsu PD, Wright J, Agarwala V, Scott DA & Zhang F (2013) Genome engineering using the CRISPR-Cas9 system. Nat Protocols 8: 2281-2308.

[0289] 47. Trepte P, Seeker C, Olivet J, Blavier J, Kostova S, Maseko SB, Minia I, Silva Ramos E, Cassonnet P, Golusik S, et al. (2024) Al-guided pipeline for protein-protein interaction drug discovery identifies a SARS-CoV-2 inhibitor. Mol Syst Biol 20: 428- 457.

[0290] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0291] It should be understood that the materials, methods, and examples provided here are representative of preferred embodiments, are exemplary, and are not intended as limitations on the scope of the disclosure.

[0292] The disclosure has been described broadly and generically herein. Each of the narrower species and sub-generic groupings falling within the generic disclosure also form part of the disclosure. This includes the generic description of the disclosure with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein.

[0293] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0294] All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety, to the same extent as if each were incorporated by reference individually. In case of conflict, the present specification, including definitions, will control.

[0295] Other embodiments are set forth within the following claims.

Claims

CLAIMSTherefore, what is claimed is:

1. A system for detecting interactions between a first protein or fragment thereof (bait protein) and a second protein or fragment thereof (prey protein), the system comprising two separate constructs:(a) a first construct comprising the bait protein and a first handle connected to the bait protein, the first handle comprising (i) a β10 peptide of a nanoluciferase (NanoLuc) fused to the N-terminal side of a modified N-terminus intein of GP41-1 (IN), or (ii) a β9 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and(b) a second construct comprising the prey protein and a second handle fused to the prey protein, the second handle comprising a modified C-terminus intein of GP41-1 (IC) fused at the C-terminal side (i) to the β9 peptide of the NanoLuc when the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified IN, or (ii) to the β10 peptide of the NanoLuc when the β9 peptide of the NanoLuc is fused to the N-terminal side of the modified IN.

2. The system of claim 1, wherein the first handle comprises the β10 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and the second handle comprises the modified IC fused at the C-terminal side to the β9 peptide of the NanoLuc.

3. The system of claim 1 or claim 2, wherein(i) the modified IC includes amino acid residues at positions 13 to 37 of wild type IC of GP41-1, and the modified IN includes amino acid residues at positions 1 to 88 of the wild type IN of GP41-1 fused to amino acid residues 1 to 12 of wild type IC of GP41 -1 (C25 GP41 -1 split intein), or (ii) the modified IC includes amino acids at positions 14 to 37 of wild type IC of GP41-1 and the modified IN includes amino acid residues at positions 1 to 88 of wild type IN of GP41-1 fused to amino acid residues at positions 1 to 13 of wild type IC of GP41-1 (C24 GP41-1 split intein), or (iii) the modified IC includes amino acids at positions 15 to 37 of wild type IC ofGP41-1 and the modified IN includes amino acid residues at positions 1 to 88 of wild type IN of GP41-1 fused to amino acid residues at positions 1 to 14 of wild type IC of GP41-1 (C23 GP41-1 split intein).

4. The system of claim 1 or claim 2, wherein the modified IN comprises SEQ ID NO: 3 and the modified IC comprises SEQ ID NO: 4, or the modified IN comprises SEQ ID NO: 5 and the modified IC comprises SEQ ID NO: 6, or the modified IN comprises SEQ ID NO: 7 and the modified IC comprises SEQ ID NO: 8.

5. The system according to any one of claims 1 to 4, wherein the β9 peptide comprises SEQ ID NO: 12 and the β10 peptide comprises SEQ ID NO: 14.

6. The system according to any one of claims 2 to 4, wherein the first handle comprises SEQ ID NO: 16 and the second handle comprises SEQ ID NO:

18.

7. The system according to any one of claims 1 to 6, wherein the first handle is connected to the N-terminus of the bait protein.

8. The system according to any one of claims 1 to 6, wherein the first handle is connected to the C-terminus of the bait protein.

9. The system according to any one of claims 1 to 8, wherein the second handle is connected to the N-terminus of the prey protein.

10. The system according to any one of claims 1 to 8, wherein the second handle is connected to the C-terminus of the prey protein.

11. The system according to any one of claims 1 to 10, wherein the first construct and the second construct further comprise a tag, wherein the tag is one or more of a V5 tag, an HA tag, and a 3xFLAG tag.

12. A method for detecting the interaction between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein) comprising:(a) providing the system according to any one of claims 1 to 11;(b) incubating the first construct and the second construct in a homogenous liquid phase under conditions that allow the formation of the GP41-1; and(c) adding to the incubate of (b) a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc, whereby detection of the luminescence signal being indicative that the bait protein or part thereof and the prey protein or part thereof interact.

13. The method of claim 12, wherein the method further comprises comparing the luminescence signal with the luminescence signal of a control to determine a binding strength between the bait protein and the prey protein.

14. A method for screening a binding inhibitor between two proteins, the method comprising:(a) providing the system according to any one of claims 1 to 11, the bait protein and the prey protein being proteins known to bind;(b) (i) incubating the bait construct, the prey construct and a compound to be tested in a homogenous liquid phase under conditions that allow the formation of the GP41-1 (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound to be tested under conditions that allow the formation of the GP41-1 (second incubate);(c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and(d) measuring a difference in luminescence signal between the first incubate and the second incubate,wherein greater luminescence signal in the second incubate is indicative that the compound is a binding inhibitor between the bait protein and the prey protein.

15. A method for determining a binding strength between a first protein or part thereof (bait protein) and a second protein or part thereof (prey protein), the method comprising:(a) providing the system according to any one claims 1 to 11, the bait protein and the prey protein being proteins known to bind;(b) (i) incubating the bait construct, the prey construct and a compound known to inhibit the binding between the bait protein and the prey protein in a homogenous liquid phase under conditions that allow the formation of the GP41-1 (first incubate), and (ii) incubating the bait construct and the prey construct in a homogenous liquid phase in the absence of the compound under conditions that allow the formation of the GP41-1 (second incubate);(c) adding to the first incubate and to the second incubate a Δ11S peptide of NanoLuc and a substrate that generates a luminescence signal in the presence of NanoLuc; and(d) determining the binding strength between the bait protein and prey protein in the presence of the compound and the binding strength between the bait protein and the prey protein in the absence of the compound based on the luminescence signal of the first incubate and the second incubate.

16. The method according to any one of claims 12 to 15, wherein the Δ11S peptide comprises SEQ ID NO: 10.

17. A recombinant polynucleotide, wherein the recombinant polynucleotide comprises SEQ ID NO: 15 or SEQ ID NO: 17, or a recombinant polynucleotide having at least 90% sequence identity to SEQ ID NO: 15 or to SEQ ID NO: 17.

18. A recombinant peptide, wherein the recombinant peptide comprises an amino acid sequence selected from SEQ ID NO: 16 or SEQ ID NO: 18, or a recombinant peptide having at least 90% sequence identity to SEQ ID NO: 16 or to SEQ ID NO: 18.

18. An isolated polynucleotide, wherein the isolated polynucleotide comprises SEQ ID NO: 15 or SEQ ID NO: 17, or an isolated polynucleotide having at least 90% sequence identity to SEQ ID NO: 15 or to SEQ ID NO: 17.

19. An isolated peptide, wherein the isolated peptide comprises an amino acid sequence selected from SEQ ID NO: 16 or SEQ ID NO: 18, or an isolated peptide having at least 90% sequence identity to SEQ ID NO: 16 or to SEQ ID NO: 18.

20. A kit comprising (a) a cell line expressing a first construct and a second construct, (b) a Δ11S peptide of NanoLuc, and (d) a substrate that generates a luminescence signal in the presence of NanoLuc, wherein a) the first construct comprises a first protein and a first handle connected to the first protein, the first handle comprising (i) a β10 peptide of a nanoluciferase (NanoLuc) fused to the N-terminal side of a modified N-terminus intein of GP41-1 (IN), or (ii) a β9 peptide of the NanoLuc fused to the N-terminal side of the modified IN; and(b) a second construct comprising a second protein and a second handle fused to the second protein, the second handle comprising a modified C-terminus intein of GP41-1 (IC) fused at the C-terminal side (i) to the β9 peptide of the NanoLuc when the β10 peptide of the NanoLuc is fused to the N-terminal side of the modified IN, or (ii) to the β10 peptide of the NanoLuc when the β9 peptide of the NanoLuc is fused to the N-terminal side of the modified IN.

21. The kit of claim 20, wherein the modified IN comprises SEQ ID NO: 3 and the modified IC comprises SEQ ID NO: 4, or the modified IN comprises SEQ ID NO: 5 and the modified IC comprises SEQ ID NO: 6, or the modified IN comprises SEQ ID NO: 7 and the modified IC comprises SEQ ID NO: 8.