Method for bioprotac design

The method addresses the challenge of designing bioPROTACs by using a scalable workflow for mRNA-based screening, enabling the identification of functional bioPROTACs and enhancing the efficiency of targeted protein degradation.

WO2025125630A1PCT designated stage expired Publication Date: 2025-06-19MEDIMMUNE LTD

Patent Information

Application Number
PCT/EP2024/086358
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for designing bioPROTACs lack an effective screening technology to systematically examine design features, hindering the identification of productive module combinations for targeted protein degradation.

Method used

A scalable and automatable workflow is established to quickly assemble numerous bioPROTAC domain arrangements using mRNA-based screening, enabling the identification of functional bioPROTACs against new targets.

Benefits of technology

This method allows for the accurate identification of functional bioPROTACs, significantly speeding up the process of finding bioPROTACs against novel targets and demonstrating effectiveness in degrading proteins like c-Myc and K-Ras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000044_0001
    Figure IMGF000044_0001
  • Figure IMGF000045_0001
    Figure IMGF000045_0001
  • Figure IMGF000046_0001
    Figure IMGF000046_0001
Patent Text Reader

Abstract

The present disclosure provides methods for identifying bioPROTACs capable of inducing degradation of target proteins. The disclosure further provides BIOPROTACs which have been generated according to these methods, including bioPROTACs capable of specifically causing degradation of c-Myc or K-Ras.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for BioPROTAC Design

[0002] Field

[0003] Provided herein are methods for identifying bioPROTACs capable of inducing degradation of target proteins. BioPROTACs identified according to these methods may be used therapeutically. Also provided herein are bioPROTACs which have been generated according to these methods, including bioPROTACs capable of specifically causing degradation of c-Myc or K-Ras.

[0004] Background

[0005] The disruption of protein function is central to therapy and dissection of biological phenomena. Breakthrough targeted protein degradation (TPD) technologies permit acute protein elimination, usually by hijacking the cellular ubiquitin-proteasome system (UPS).

[0006] Proteolysis-targeting chimeras (PROTACs) are heterobifunctional drugs that simultaneously recruit a target protein and one of a few amenable E3 ubiquitin ligases to induce ubiquitylation and degradation of the target. They contain two small molecules, one a ligand for the target protein and the other a ligand for the E3 ligase, joined by a linker (Burslem & Crews, Cell 181 : 102-114, 2020).

[0007] PROTACs can outperform their small-molecule drug constituents in potency, duration and selectivity (Verma et al., Molecular Cell 77: 446-460, 2020). Unlike other strategies to deplete a protein target, such as CRISPR and RNAi, PROTACs facilitate rapid and long- lasting protein degradation. Unlike occupancy-driven pharmacology, PROTACs only require target recruitment and are therefore well positioned to broaden the “druggable” target space. Still, their usage is restricted to targets containing ligand-binding pockets due to their small- molecule nature. Despite the presence of more than 600 human E3 ligases, there is a limited repertoire of E3 ligases (VHL, CRBN and lAPs) routinely recruited by PROTACs. However, the evidence of targeted-protein degradation being dependent on the recruited E3 ligase (Lai et al., Angewandte Chemie International Edition 11 : 807-810, 2016) highlights the necessity for expanding the toolbox for broader implications and enhanced selectivity.

[0008] An exciting alternative are biological PROTACs (hereafter bioPROTACs) genetically- encoded protein fusions analogous to PROTACs capable of complete, swift and specific endogenous target destruction. BioPROTACs comprise a target-binding protein fused to either an E3 ligase or a recruiter of an E3 ligase, and thus act in the same way as a chemical PROTAC to bring together a target protein and an E3 ligase to induce ubiquitylation and degradation of the target.

[0009] The biological nature of bioPROTACs enables the recruitment of otherwise non- ligandable targets such as intrinsically disordered proteins or even specific protein variants or post-translationally modified proteins to the vicinity of an unrestricted selection of E3 ligases. BioPROTACs are well suited to expand drug target space, as demonstrated in the destruction of difficult-to-drug oncogenic Ras family proteins (Bery et al., Nature Communications 11 : 3233, 2020).

[0010] However, the implementation of bioPROTACs against more targets suffers from a lack of maturity in the field. Only a few E3 ligases have been surveyed, usually validated against GFP-fused targets instead of endogenous proteins and most proof-of-principle studies have focused on long-lived proteins more sensitive to stability changes. An additional hurdle is that target engagement is not itself predictive of degradation (Moutel et al., eLife 5: e16228, 2016). Compatible bioPROTAC domain partners must be put together and in the right order to elicit target destruction (Bery et al., supra), suggesting that individual module properties (e.g. target-binding affinity) and structural positioning act in concert to govern target ubiquitylation.

[0011] Degrader development programmes currently lack a screening technology which systematically examines bioPROTAC design features to find productive module combinations. The process provided herein meets this need by establishing a scalable, automatable workflow to quickly assemble numerous bioPROTAC domain arrangements and measure their cellular effect in an mRNA-based screening approach, enabling easier generation and identification of functional bioPROTACs against new targets.

[0012] Summary

[0013] Provided herein is an effective method to identify bioPROTACs capable of inducing degradation of a target protein. The method has been validated against the transcription factor c-Myc and the GTPase K-Ras, both targets in cancer therapy, as described in the Examples below. BioPROTACs effective for inducing degradation of these proteins are also provided.

[0014] In a first aspect, provided herein is a method of identifying a bioPROTAC capable of inducing degradation of a target protein in a eukaryotic cell, wherein the bioPROTAC comprises a binding domain and a degradation domain; wherein the binding domain specifically binds the target protein, and the degradation domain is capable of causing ubiquitylation of the target protein; the method comprising:

[0015] (i) providing a first DNA library encoding a library of binding domains;

[0016] (ii) providing a second DNA library encoding a library of degradation domains;

[0017] (iii) assembling a vector library encoding a bioPROTAC library, wherein each vector in the library comprises an expression cassette comprising a gene encoding a bioPROTAC, and each vector is assembled by joining an expression vector backbone, a first DNA molecule from the first DNA library and a second DNA molecule from the second DNA library;

[0018] (iv) producing an mRNA library by in vitro transcription of the genes encoding bioPROTACs;

[0019] (v) transfecting eukaryotic cells which express the target protein with the mRNA library, wherein separate transfections are performed using mRNA encoding each different bioPROTAC; and screening the transfected cells for depletion of the target protein, wherein depletion of the target protein in a transfected cell indicates the bioPROTAC expressed in that cell is capable of inducing degradation of the target protein.

[0020] The bioPROTAC identified by the method may additionally include a linker sequence between the binding domain and the degradation domain. In this case, the method further comprises providing a third DNA library encoding a library of linkers, and in step (iii), each vector is assembled by joining an expression vector backbone, a first DNA molecule from the first DNA library, a second DNA molecule from the second DNA library and a third DNA molecule from the third DNA library.

[0021] In a second aspect, provided herein is a method of treating a disease in a subject, comprising:

[0022] (a) identifying a target protein associated with the disease;

[0023] (b) identifying a bioPROTAC capable of inducing degradation of the target protein according to the method of the first aspect; and

[0024] (c) administering the bioPROTAC or a nucleic acid molecule encoding the bioPROTAC to the subject.

[0025] In a third aspect, provided herein is a bioPROTAC comprising a binding domain and a degradation domain, the degradation domain comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 1-100, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity thereto.

[0026] Preferably, the degradation domain comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-2, 5, 7-8, 11-16, 18-24, 26 or 31-52, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity thereto.

[0027] The bioPROTAC provided according to this aspect may further comprise a linker comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 120, 126, 132 or 133, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity thereto.

[0028] In a fourth aspect, provided herein is a bioPROTAC capable of degrading human c-Myc, comprising a binding domain and a degradation domain, the binding domain comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 101 to 110, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity thereto. The bioPROTAC provided according to this aspect may comprise a degradation domain comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 1-52, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity thereto. The bioPROTAC may further comprise a linker comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 120, 126, 132 or 133, or having at least 80, 85, 90, 95 or 99 % sequence identity thereto.

[0029] In a fifth aspect, provided herein is a bioPROTAC capable of degrading human K-Ras, comprising, from N-terminus to C-terminus:

[0030] (i) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 134, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 134;

[0031] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 120; and;

[0032] (iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 52, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 52.

[0033] In a sixth aspect, provided herein is a nucleic acid molecule encoding the bioPROTAC of the fourth or fifth aspect. Preferably, the nucleic acid molecule is an mRNA molecule or a vector, such as a viral vector.

[0034] In a seventh aspect, provided herein is a pharmaceutical composition comprising a bioPROTAC of the fourth or fifth aspect, or a nucleic acid molecule of the sixth aspect.

[0035] In an eighth aspect, provided herein is a bioPROTAC of the fourth or fifth aspect, or a nucleic acid molecule of the sixth aspect, for use in therapy.

[0036] In a ninth aspect, provided herein is a bioPROTAC of the fourth or fifth aspect, or a nucleic acid molecule of the sixth aspect, for use in treatment of cancer or an inflammatory disease.

[0037] This aspect also provides a method of treating cancer or an inflammatory disease in a subject in need thereof, comprising administering to the subject a bioPROTAC of the fourth or fifth aspect, or a nucleic acid molecule of the sixth aspect.

[0038] This aspect also provides the use of a bioPROTAC of the fourth or fifth aspect, or a nucleic acid molecule of the sixth aspect, in the manufacture of a medicament for treating cancer or an inflammatory disease in a subject.

[0039] In a tenth aspect, provided herein is a method of degrading a protein in a eukaryotic cell, comprising contacting the cell with a nucleic acid molecule or vector encoding a bioPROTAC, wherein:

[0040] (i) the protein is c-Myc, and the bioPROTAC is as defined in the fourth aspect; or (ii) the protein is K-Ras, and the bioPROTAC is as defined in the fifth aspect.

[0041] Detailed Description

[0042] BioPROTACS are an exciting new possibility for targeted therapy, and a useful tool in biological research. At present, bioPROTAC design requires to some extent a trial-and-error approach, in order to identify functional combinations of target binder and degradation domain. However, identification of functional bioPROTACs has been held back by the lack of an effective screening method. The method provided herein enables accurate identification of functional bioPROTACs, which can significantly speed up the identification of bioPROTACs against novel targets. BioPROTACs obtainable from this method, and comprising a degradation domain as exemplified herein, are also provided. bioPROTACs

[0043] A bioPROTAC is a protein construct capable of inducing degradation of a target protein in a eukaryotic cell.

[0044] As defined herein, a bioPROTAC is a fusion protein comprising a binding domain and a degradation domain. These are discussed further below, but in basic terms the binding domain is a domain of the bioPROTAC capable of binding the target protein and the degradation domain is a domain of the bioPROTAC capable of ubiquitylating the target protein, or of recruiting a protein capable of ubiquitylating the target protein. Thus when a bioPROTAC comes into contact with the target protein the binding domain binds the target protein, bringing the target protein into close proximity with the degradation domain. The degradation domain either ubiquitylates the target protein itself, or recruits to the bioPROTAC-target protein complex an E3 ubiquitin ligase which ubiquitylates the target protein. Following ubiquitylation, the target protein is directed to the proteasome for degradation.

[0045] The binding and degradation domain may be arranged in either order, i.e. the binding domain may be N-terminal to the degradation domain or the degradation domain may be N- terminal to the binding domain. As described in the examples, different combinations of binding and degradation domains have different optimal arrangements.

[0046] The bioPROTAC may consist of the binding domain and degradation domain, but commonly comprises one or more additional elements. For example, the bioPROTAC commonly comprises a linker between the binding and degradation domains, as discussed further below. Where the target protein (i.e. the protein targeted for degradation by the bioPROTAC) is located in a sub-cellular compartment or organelle, the bioPROTAC may comprise a targeting sequence directing the bioPROTAC for trafficking into the relevant organelle. For instance, if the target protein is a nuclear protein, the bioPROTAC may contain a nuclear localisation sequence (NLS), or if the target protein is mitochondrial, the bioPROTAC may contain a mitochondrial localisation sequence (MLS). Similarly, the bioPROTAC may contain a sequence which excludes it from particular sub-cellular locations. For example, if it is desired to exclude the bioPROTAC from the nucleus, it may contain a nuclear export signal (NES).

[0047] As noted above, the bioPROTAC is capable of inducing degradation of the target protein in a eukaryotic cell. Since bioPROTACs act by causing ubiquitylation of the target protein, binding of a bioPROTAC to its target protein would not cause degradation of the target protein under all conditions, only where the necessary ubiquitylation machinery is also present, along with a means by which ubiquitylated proteins can be degraded. Degradation of soluble proteins is generally via the proteasome, whereas membrane proteins may be degraded via the lysosomal pathway (Cotton et al., Journal of the American Chemical Society 143: 593-598, 2021). In contexts such as in vitro or in a cellular system lacking ubiquitylation machinery and / or an appropriate degradation pathway the bioPROTAC would not be capable of inducing degradation of the target protein.

[0048] Ubiquitin is expressed in eukaryotic cells, and so the bioPROTAC is capable of inducing degradation of the target protein in eukaryotic cells. Examples of eukaryotic cells in which the bioPROTAC is capable of inducing degradation of the target protein include e.g. yeast and plant cells, but primarily the bioPROTAC is capable of inducing degradation of the target protein in mammalian cells, in particular human cells.

[0049] Whether a bioPROTAC is capable of inducing degradation of a target protein in mammalian cells can be determined using a mammalian cell line. Suitable non-human cell lines which may be used for such a determination include e.g. Chinese hamster ovary (CHO) cells and simian COS cells. Suitable human cells which may be used for this determination include e.g. human embryonic kidney (HEK)-293 cells, HEK-293T cells, HeLa cells, HCT116 cells, etc. Many other suitable human and non-human mammalian cell lines are known to the skilled persons. Cell lines can be acquired from a cell depositary, e.g. the ATCC (VA, USA).

[0050] The cell line may be a modified cell line, e.g. a modified version (also referred to herein as a derivative) of the cell lines set out above, e.g. a modified HCT116 cell line, a modified HEK-293 or HEK-293T cell line or a modified HeLa cell line. In a particular embodiment, the cell line is HCT116, or a modified version thereof. For example, the cell line may be modified to aid determination of whether a bioPROTAC of interest is capable of inducing degradation of its target. To this end the cell line may be modified to express a tagged version of the target protein. The tag may be a reporter sequence which enables detection of the target protein, such as a fluorescent protein (e.g. GFP, YFP, etc.) or a luciferase. Other modifications to the cell line may additionally or alternatively be made, e.g. to express a non-native target protein (such as a viral protein), etc., as desired by the skilled person.

[0051] A bioPROTAC can be considered capable of inducing degradation of a target protein in a cell (e.g. a mammalian, preferably human, cell) if expression of the bioPROTAC in the cell results in a reduction in the level of the target protein in the cell (which may be referred to as depletion of the target protein in the cell). Any level of reduction is encompassed, though clearly the higher the degree of reduction the more effective the bioPROTAC. For example, a bioPROTAC capable of inducing degradation of a target protein in a cell may reduce the level of the target protein by at least 10, 20, 30, 40, 50, 60, 70, 75, 80, 85, 90 or 95 %.

[0052] Whether expression of a bioPROTAC causes a reduction in the level of the target protein in a cell, and if so the degree of reduction, may be determined by comparison to any suitable comparator. For instance, the level of the target protein in a cell line following expression of a bioPROTAC may be compared to the level in the same cell line prior to expression of the bioPROTAC. More preferably, the level of the target protein in a cell line following expression of a bioPROTAC is compared to the level of the target protein in a control cell line. By “control cell line” is meant a cell line which is identical to, and is treated identically to, the cell line expressing the bioPROTAC, except the bioPROTAC itself is not expressed. For example, while the experimental cell line is transfected with a vector encoding the bioPROTAC, the control cell line may be untransfected, mock transfected or transfected with a control vector, e.g. the backbone only of the vector used to encode the bioPROTAC. The skilled person is capable of selecting a suitable control cell line. The process of determining whether expression of the bioPROTAC causes a reduction in the level of the target protein in the cell (compared to a suitable comparator) may be referred to herein as screening the cell for depletion of the target protein.

[0053] To determine whether expression of a bioPROTAC causes a reduction in the level of the target protein in a cell, any suitable technique may be used. A suitable technique allows qualitative or quantitative (preferably quantitative) analysis and comparison of the level of target protein between cell lines. Suitable techniques include immunofluorescence microscopy, luminescence kinetic assay and quantitative immunoblotting. These techniques are well known in the art and demonstrated in the examples below. Preferably, at least two different techniques are used to screen for depletion of a target protein by a bioPROTAC of interest, for example at least two of the techniques set out above. As shown below, some bioPROTACs showed different results in different screens, and so testing bioPROTACs in multiple different screens can increase the reliability of the screening results. For example, a bioPROTAC found to cause a reduction in target protein level in multiple screens is more likely to be truly capable of inducing degradation of the target protein than a bioPROTAC only shown to be effective in a single screen which might have been impacted by an artefact or suchlike. Conversely, performing multiple screens can enable the identification of an effective bioPROTAC which might fail to induce target degradation in one of the screening techniques, e.g. due to an artefact or a lack of compatibility.

[0054] Target Protein

[0055] The target protein is a protein which it is desired to degrade, and thus against which it is desired to design a bioPROTAC. Thus the target protein may be any protein. It may be a natural protein, a synthetic protein or a modified protein, i.e. a protein which occurs in nature but which has been changed, e.g. in its sequence or post-translational modification.

[0056] In some embodiments, the protein is a human protein. Alternatively, the protein may be a non-human protein. The target protein may be a eukaryotic protein or a prokaryotic protein. For example the target protein may be an animal protein, such as a mammalian protein (including a non-human mammalian protein) an insect protein, or a protozoan or protist protein. The target protein may be a bacterial protein, or a viral protein. In some embodiments, the target protein is a protein from a pathogen.

[0057] The target protein may be a protein associated with cancer or chronic disease, such as a protein which is commonly overexpressed or mutated in cancer cells, or a viral oncoprotein. The target protein may be a naturally-occurring mutant protein, containing e.g. a deletion, addition or substitution, which is associated with cancer.

[0058] The target protein is an intracellular protein (e.g. a soluble protein or a membrane- anchored protein) or a transmembrane protein with an intracellular domain. Where the target protein is an intracellular protein, the target protein may be located in any part of the cell, i.e. in any compartment or organelle. For example the target protein may be located in or associated with the cytoplasm, nucleus (as mentioned above) or any other compartment such as the endoplasmic reticulum, Golgi, etc.

[0059] The target protein may be, for example, a receptor, such as a G-protein-couple receptor (GPCR), a ligand-gated ion channel, an enzyme-linked receptor such as a receptor tyrosine kinase, a hormone receptor, a growth factor receptor, a cytokine receptor, etc. The target protein may be a transcription factor; an enzyme, such as a kinase or a GTPase; etc. Any type of protein, with any function, may be targeted.

[0060] Binding Domain

[0061] The binding domain is a domain of the bioPROTAC which binds, or is capable of binding, the target protein. By “domain of the bioPROTAC” herein is meant one of the elements of the bioPROTAC fusion protein. Commonly, each domain of the bioPROTAC forms its own structure, independent of the other element(s), in line with traditional terminology in molecular biology. However, this is not essential.

[0062] The binding domain specifically binds the target protein. By this it is meant that the binding domain binds the target protein with a greater affinity than that with which it binds to other molecules (e.g. with an affinity as described below), or at least most other molecules. Thus, for example, if a binding domain which specifically binds human c-Myc (as exemplified in the Examples below) were contacted with a lysate of human cells, the antibody would bind primarily to c-Myc. In particular, the binding domain may bind to a sequence or configuration present on the target protein, preferably a unique sequence or configuration which is present on the target protein but is not present on other proteins.

[0063] A binding domain which specifically binds a target protein from one species (e.g. humans) does not necessarily bind only to the target protein from that species. A binding domain may cross-react with certain other undefined molecules, or may display a level of non-specific binding when contacted with a mixture of a large number of molecules (such as a cell lysate or suchlike). In particular, a binding domain which specifically binds a human target protein may display cross- reactivity with corresponding proteins from other species. Regardless, the skilled person can readily determine whether a binding domain specifically binds a target protein using standard techniques in the art, e.g. ELISA, Western-blot, surface plasmon resonance (SPR), etc.

[0064] The binding domain may be any type of protein capable of specifically binding to the target protein, with the restriction that it is a single chain protein (e.g. it is not a full antibody). In some embodiments, the binding domain is derived from the antigen binding domain of an antibody, that is to say, the binding domain may be a single chain antibody fragment.

[0065] In some embodiments, the single chain antibody fragment is an scFv (single chain variable fragment). An scFv is a synthetic construct which comprises the variable regions of the heavy and light chains of an antibody joined in a single polypeptide chain. Typically an scFv is produced by recombinantly engineering antibody genes to encode a fusion protein comprising the VHand VLregions of the antibody.

[0066] Generally an scFv includes a peptide linker covalently joining the VHand VLregions, which contributes to the stability of the molecule. The linker may comprise from 1 to 20 amino acids, such as for example 5 to 20, 10 to 20 or 5 to 15 amino acids, 1 , 2, 3 or 4 amino acids, 5, 10 or 15 amino acids, or other intermediate numbers in the range 1 to 20 as convenient. The peptide linker may be formed from any generally convenient amino acid residues, such as glycine and / or serine. One example of a suitable linker is Gly4Ser. Multimers of such linkers may be used, such as for example a dimer, a trimer, a tetramer or a pentamer, e.g. (Gly4Ser)2, (Gly4Ser)3, (Gly4Ser)4or (Gly4Ser)5. However, it is not essential that a linker be present, and the VL region may be linked to the VH region by a peptide bond between amino acids at the termini of the regions.

[0067] In other embodiments the binding domain is a single chain antibody, i.e. an antibody of a type which naturally requires only a single chain to function. Such antibodies include heavy chain antibodies from sharks and camelids, which are antibodies consisting of only heavy chains, i.e. they lack a light chain. Single chain (heavy chain) human antibodies can also be generated by mutating hydrophobic residues V37, G44, L45 and W47 in framework region 2 of the human heavy chain variable region, which naturally form hydrophobic interactions with the light chain (Asaadi et al., Biomarker Research 9: 87, 2021). Mutation of these residues can yield a human heavy chain which functions in the absence of a light chain. For example these residues can be mutated to F37, E44, R45, and G47 or F47. Thus the binding domain may be an antibody chain from a heavy chain antibody.

[0068] The binding domain may be a single domain antibody, also known as a nanobody. A "single domain antibody" (sdAb) is an antibody fragment generally having a molecular weight of 12-15 kDa, consisting of a single monomeric variable domain derived from a heavy chain of an antibody which lacks a light chain. Examples of single domain antibodies include VHH single domain antibodies, which are derived from camelid antibodies, and VNAR single domain antibodies which are derived from antibodies from sharks and other species of cartilaginous fish of the class Chondrichthyes. Alternatively the single domain antibody may be derived from a modified single chain human antibody, as discussed above.

[0069] Antibodies or antibody fragments which recognise the target protein, and thus can be used as or used to generate the binding domain for use herein, can be obtained by any method known in the art. A commercially-available or otherwise prior-existing antibody may be used, or a new antibody may be generated by any known technique, e.g. hybridoma technology or phage display.

[0070] In other embodiments the binding domain is a DARPin (designed ankyrin repeat protein). A DARPin is a binding protein derived from a natural ankyrin protein. Ankyrins consist of a repeating section which is stacked together to form a rigid protein. The DARPins used herein may be as described in WO 2016 / 023898, which is hereby incorporated by reference. A DARPin for use herein comprises at least 2 ankyrin repeats, e.g. at least 3 or 4 ankyrin repeats, e.g. 2, 3, 4 or 5 ankyrin repeats. A DARPin may be generated by any method known in the art, e.g. phage display.

[0071] In other embodiments the binding domain is a Tn3 protein. Tn3 proteins are based on the structure of a type III fibronectin module (Fnlll) and are derived from the third Fnlll domain of human tenascin C. The generation and use of Tn3 proteins is described for example in WO 2009 / 058379, WO 2011 / 130324, WO 2011 / 130328 and Gilbreth et al., Protein Engineering, Design and Selection 27: 411-418, 2014. The Tn3 proteins and the native Fnlll domain from tenascin C are characterized by the same tridimensional structure, namely a beta-sandwich structure with three beta strands (A, B, and E) on one side and four beta strands (C, D, F, and G) on the other side, connected by six loop regions. These loop regions are designated according to the beta- strands connected to the N- and C-terminus of each loop. Accordingly, the AB loop is located between beta strands A and B, the BC loop is located between strands B and C, the CD loop is located between beta strands C and D, the DE loop is located between beta strands D and E, the EF loop is located between beta strands E and F, and the FG loop is located between beta strands F and G. Fnlll domains possess solvent-exposed loops tolerant of randomization, which facilitates the generation of diverse pools of protein scaffolds capable of binding specific targets with high affinity.

[0072] Tn3 proteins can be subjected to directed evolution designed to randomize one or more of the loops which are analogous to the complementarity-determining regions (CDRs) of an antibody variable region. Such a directed evolution approach results in the production of antibody-like binding members with high affinities for targets of interest. Tn3 proteins which specifically bind a particular target protein can be identified by standard techniques in the art, e.g. phage display.

[0073] In yet other embodiments the binding domain is a ligand of the target protein, or a fragment thereof (i.e. a fragment of a ligand of a target protein). As used herein, the term “ligand” refers to a protein, polypeptide or peptide which specifically binds to the target protein as part of its natural, biological function. Thus a ligand may be a natural binding partner of the target protein, e.g. a protein with which the target protein forms a protein- protein interaction.

[0074] Commonly binding of a ligand to its target does not entail the entirety of the ligand interacting with the target, but rather only a part of the ligand is required, e.g. a particular domain or a short peptide within the ligand. In this case, the entire ligand need not be used as the binding domain, but only a part of the ligand including or consisting of the peptide or polypeptide sequence which directly binds the target protein. That is to say a fragment of a ligand may be used. By “fragment” of a ligand is simply meant less than the full length, naturally occurring ligand. A fragment of a ligand may be of any length, so long as it specifically binds the target protein, e.g. the fragment of a ligand may be at least 5, 10, 15, 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450 or 500 amino acids long. A fragment of the ligand may be derived from any part of the ligand, as determined by the part of the ligand which binds the target protein, for example the fragment may be from the N-terminus, or C-terminus of the ligand, or from the internal section of the ligand (i.e. not including either terminus of the ligand). The term “ligand” as used herein includes variants of wild type ligand, i.e. which have a modified amino acid sequence relative to the wild type ligand. Such a variant may have e.g. at least 50, 60, 70, 80, 90, 95 or 99 % sequence identity to the wild type ligand, so long as the variant retains its ability to specifically bind the target protein. By extension, the fragment of a ligand may be a fragment of a ligand which is a variant of the wild type ligand.

[0075] Generally, the ligand is not the target protein itself. That is to say, where the target protein homo-dimerises or homo-multimerises, the binding domain of the bioPROTAC is generally not a full copy of the target protein itself, as this would lead to the bioPROTACs binding to each other. However, a fragment of the protein itself can be used, so long as it (i) binds the target protein, (ii) does not bind itself (to avoid dimerization / multimerization of the bioPROTAC), and (iii) is not recognised by the same E3 ubiquitin ligase as the full length protein (to avoid the bioPROTAC degrading itself).

[0076] In other embodiments the binding domain is an inhibitor of the target protein. An inhibitor of the target protein is any protein, polypeptide or peptide which specifically binds to the target protein and inhibits (i.e. reduces) its activity. An inhibitor as used herein may, upon binding the target protein, reduce its activity by any amount, e.g. at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 75, 80, 85, 90 or 95 %. The inhibitor may be a natural or synthetic inhibitor of the target protein.

[0077] In some cases a fragment of an inhibitor may be used as the binding domain. As for the fragment of a ligand, the fragment of an inhibitor may be of any length so long as it specifically binds the target protein (in some cases a fragment of an inhibitor may bind but not inhibit the activity of the target protein; such a fragment remains useful in the methods of the present disclosure described herein). For example, the fragment of an inhibitor may be at least 5, 10, 15, 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450 or 500 amino acids long. A fragment of an inhibitor may be derived from the N-terminus, C-terminus or internal section of the inhibitor.

[0078] As for the ligand, the inhibitor of the target protein may in some cases be a fragment of the target protein (subject to the same requirements as where a fragment of the target protein is used as a ligand). The Examples below demonstrate the use of the H1S6A / F8Apeptide derived from c-Myc in the production of a bioPROTAC against c-Myc.

[0079] In some cases a binding domain may fall within multiple different types of protein as set out above, e.g. a binding domain may be an antibody fragment and an inhibitor of the target protein, or both a ligand and inhibitor of the target protein.

[0080] Once a potential binding domain for a target protein has been identified, the potential binding domain can be tested to confirm whether it binds the target protein in an intracellular context be expressing the binding domain in a cell line expressing the target protein, and performing immunoprecipitation. If the potential binding domain is pulled down with the target protein, the binding domain is confirmed to bind the target protein in the intracellular environment.

[0081] The binding domain binds the target protein with sufficiently high affinity that the bioPROTAC is able to induce degradation of the target through its degradation domain (discussed below). For example, the binding domain may bind the target protein with an affinity having a KDof less than about 10 pM, 5 pM, 1 pM, 750 nM, 500 nM, 400 nM, 300 nM, 200 nM, 100 nM, 50 nM, 40 nM, 30 nM, 25 nM, 20 nM, 15 nM, 10 nM, 5 nM or 1 nM. The affinity may be measured by any suitable method (e.g. surface plasmon resonance, SPR) under any suitable conditions. For example the affinity may be calculated under “physiological” conditions (i.e. a pH of about 7.4). If the target protein is located in a cellular compartment having a pH higher or lower than 7.4, the affinity may be calculated at the pH of the compartment where the target protein is located.

[0082] Similarly, the binding domain may bind the target protein with a dissociative half-life (ti / 2) of greater than about 1 minute, 1.5 minutes, 1.75 minutes, 2 minutes, 2.5 minutes, 3 minutes, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 25 minutes or 30 minutes as measured using an assay such as SPR under suitable conditions, as discussed above in relation to affinity.

[0083] In some embodiments, the binding domain is modified to avoid or reduce ubiquitylation (that is to say the binding domain may be modified so that ubiquitylation of the binding domain is reduced or avoided).

[0084] The binding domain may be modified to avoid or reduce ubiquitylation by mutation of one or more residues to which ubiquitin may be attached. Such residues may be mutated to residues which cannot by ubiquitylated (or may be deleted). In some embodiments, the binding domain may be modified by one or more amino acid substitutions relative to the original or native sequence, such that ubiquitylation of the binding domain is reduced relative to the original or native sequence.

[0085] The original or native sequence of the binding domain is its sequence before any modification to reduce ubiquitylation. When the binding protein is a natural protein, such as a ligand of the target protein, the native sequence may be the wild type sequence of the protein (or of the fragment of the protein being used). Where the binding protein is a protein which has been generated for the purpose of binding the target protein, such as an antibody, Tn3 or DARPin, the original sequence may be the sequence of the protein which was initially generated, e.g. in a hybridoma or by phage display.

[0086] Alternatively, the original or native sequence may be a sequence which has been modified relative to the wild type or initially-generated sequence for purposes other than altering its ubiquitylation, e.g. a humanised sequence, or a sequence which has been modified to remove or introduce a glycosylation site or a site for any other post-translation modification.

[0087] Modification of the binding domain in this manner may result in the level of ubiquitylation being reduced by any amount in the modified binding domain relative to the original or native sequence. For instance, ubiquitylation of the modified binding domain may be reduced by at least 10, 20, 30, 40, 50, 60, 70, 80 or 90 %. In other cases, modification of the binding domain may remove possible ubiquitylation sites, without actually reducing ubiquitylation of the binding domain in practice. A reduction in ubiquitylation may be measured in any suitable way, for example the number of ubiquitin chains attached to the binding domain in a given time period.

[0088] Most commonly, ubiquitin is attached to proteins destined for degradation via an isopeptide bond between the C-terminal carboxyl group of ubiquitin and a lysine side chain on the target protein. In some cases, ubiquitin is attached to other proteins via an ester bond between the ubiquitin C-terminal carboxyl group and the side chain of a serine or threonine on the target protein, or via a thioester bond between the ubiquitin C-terminal carboxyl group and the side chain of a cysteine on the target protein.

[0089] The binding domain sequence may thus be modified by substitution of one or more lysine, serine, threonine or cysteine residues. Such residues may be substituted for any other residue (except another ubiquitylatable residue, i.e. a lysine, serine, threonine or cysteine residue cannot be substituted for another in this group when the substitution is in order to reduce or avoid ubiquitylation of the binding domain), though in particular embodiments the substitution is a conservative substitution. Conservative amino acid substitutions are described further below.

[0090] Where the binding domain is modified to avoid or reduce ubiquitylation, the binding domain preferably includes one or more lysine mutations, i.e. preferably one or more lysine residues from the original or native sequence is substituted for a different amino acid. Preferably, lysine is mutated to arginine, which is also basic and positively charged but is not ubiquitylatable. Thus the binding domain may comprise one or more lysine to arginine substitutions relative to the native or original sequence, e.g. 1 , 2, 3, 4, 5, 6 or 7 or more lysine to arginine substitutions relative to the native or original sequence.

[0091] Where a cysteine, serine or threonine is mutated relative to the original binding domain sequence in order to reduce its ubiquitylation, the amino acid may be replaced with another uncharged, polar residue such as asparagine or glutamine. The binding domain may comprise any number of such substitutions, e.g. 1 to 10 such substitutions. In some cases where several instances of the same type of residue are substituted, different substitutions may be made at different locations, e.g. if two serine residues are substituted one might be substituted for an asparagine residue and the other for a glutamine residue. The binding domain may be modified by substitution of more than one type of residue, e.g. at least two of a cysteine residue, a lysine residue, a serine residue or a threonine residue may be substituted, as described above.

[0092] Degradation Domain

[0093] The degradation domain is a domain of the bioPROTAC that causes ubiquitylation of the target protein, once it has been bound by the binding domain. Specifically, the degradation domain is capable of causing ubiquitylation of the target protein. That is to say, when the degradation domain is brought into proximity of the target protein under suitable conditions for ubiquitylation to take place (e.g. where ubiquitin is present, such as within a eukaryotic cell as discussed above), the degradation domain causes the target protein to be ubiquitylated.

[0094] In some embodiments, the degradation domain is an E3 ubiquitin ligase (or E3 ligase). In these embodiments, the degradation domain causes ubiquitylation of the target protein by directly ubiquitylating it. Any type of E3 ligase may be used as the degradation domain, including HECT E3 ligases, RING E3 ligases, U-box E3 ligases, RBR E3 ligases and NEL E3 ligases. Where an E3 ligase is used, the E3 ligase may be derived from any species, i.e. it may be a non-mammalian or a mammalian E3 ligase, e.g. a human E3 ligase (humans encode over 600 E3 ligases, any of which can be used herein). In some embodiments a microbial E3 ligase is used, e.g. a bacterial or viral E3 ligase. Certain pathogenic microorganisms, including viruses and bacteria, express E3 ubiquitin ligases which they use to hijack the host cell’s ubiquitin-proteasome system, including Salmonella enterica, Escherichia coll (particular enterohaemorrhagic E. coll (EHEC) and enteropathogenic E. coli (EPEC)), Legionella pneumophila, Shigella flexneri, Herpes simplex virus type 1 (HSV-1), Kaposi’s sarcoma-associated herpesvirus (KSHV) and rotavirus (see Maculins et al., Cell Research 26: 499-510, 2016; and Zhang et al., Frontiers in Immunology 9: 1083, 2018). Any microbial E3 ligase may be used, including a ligase derived from one of the microorganisms listed above.

[0095] In some cases, a fragment of an E3 ubiquitin ligase is used as the degradation domain. A fragment of an E3 ligase can be used as the degradation domain where the entirety of the full-length protein is not required for the enzyme’s ligase activity, i.e. where a fragment of the E3 ligase is sufficient to perform ubiquitylation of a target. Thus where a fragment of an E3 ligase is used, the fragment comprises the E3 ligase active site. Such a fragment may be referred to as an E3 ligase functional region. An E3 ligase functional region can be used where the functional region is formed from a continuous polypeptide chain (i.e. where an E3 ligase contains a functional region formed from non-continuous sections of the protein, it may not be possible to isolate such a functional region as a fragment of the full- length protein). Where a fragment of an E3 ligase is used, the fragment may be of any length, so long as it is functional. In some cases the fragment may be e.g. at least 100, 150, 200, 250, 300, 350, 400, 450 or 500 amino acids long. In other cases the fragment may be shorter, e.g. less than 100, 90, 80, 70, 60 or 50 amino acids long.

[0096] Indeed, in some cases it may be advantageous to use a fragment of an E3 ligase. Removing the natural substrate binding domain can avoid unwanted gain-of-function by the bioPROTAC.

[0097] Where an E3 ligase is used as the degradation domain (or a fragment thereof), it is advantageous if the E3 ligase does not require post-translational modification for its activity.

[0098] In other embodiments, the degradation domain is an E2 ubiquitin-conjugating enzyme (or E2 enzyme). E2 enzymes carry ubiquitin at their active site (after receiving it from an E1 ubiquitin-activating enzyme), and generally cause ubiquitylation of a target protein by binding to a cognate E3 ligase, which in turn binds the target protein, and transfers the ubiquitin molecule from the E2 enzyme to the target protein. Thus an E2 enzyme can be seen to function by recruiting an E3 ligase. Without being bound by theory, however, in some cases an E2 enzyme may be able to directly ubiquitylate a target protein, without the assistance of an E3 ligase.

[0099] Any E2 enzyme may be used as a degradation domain. For example, an E2 enzyme from any species may be used, though generally the E2 enzyme is a mammalian E2 enzyme, preferably a human E2 enzyme. Humans have about 30 E2 enzymes, any one of which can be used herein.

[0100] As for the E3 ligase above, where a full-length E2 enzyme is not required for its functionality (i.e. to transfer or take part in the transfer of an ubiquitin molecule to the target protein), a functional fragment of an E2 enzyme can be used as the degradation domain.

[0101] In other embodiments, the degradation domain is a protein capable of binding, preferably specifically binding, an E3 ligase, and thereby recruiting an E3 ligase to the bioPROTAC. Specific binding is defined above. Such a protein is referred to herein as an E3 binding partner. Any E3 binding partner, of any type, may be used in the bioPROTACs herein.

[0102] An E3 binding partner may be a synthetic protein, i.e. a protein which has been generated for the purpose of binding an E3 protein of interest. Such E3 binding partners include the types described above in the context of the binding domain, e.g. a single chain antibody such as an scFv or single domain antibody, a DARPin or a Tn3. Such synthetic binding proteins may bind any E3 protein of interest, and may be generated as set out above. Degradation domains based on synthetic proteins of these types may be referred to herein as “affinity scaffolds”. Alternatively, an E3 binding partner for use herein may be natural binding partner for an E3 ligase. For example, an E3 binding partner may be a degron. Degrons are short linear motifs bound by E3 ubiquitin ligases which mediate ubiquitylation of substrates. Generally, however, recruitment of an E3 ligase to a degron does not cause ubiquitylation of the degron itself, but at a site a short distance from the degron. Several degron sequences are known in the art, and degron sequences can also be identified using bioinformatics, e.g. Degpred (Hou et al., BMC Biology 20: 162, 2022).

[0103] Where an E3 binding partner is used as the degradation domain, if the full length of the protein is not required for binding to the E3 ligase, a fragment of the binding partner may be used, the fragment being capable of binding the E3 ligase.

[0104] In other embodiments, the degradation domain is a component of an E3 ligase complex. The term “E3 ligase complex” as used herein refers to a multi-protein complex which, as a complex, has E3 ligase activity, but none of the individual components of which have independent E3 ligase activity. In particular, the degradation domain may be a component of a cullin-RING E3 ubiquitin ligase (CRL) complex. CRLs are well known in the art, and are distinct from other types of E3 ligase in that they are E3 ligase complexes, rather than single protein enzymes, and they do not form an E3-ubiquitin intermediate conjugate when transferring ubiquitin from an E2 enzyme to a substrate (target protein).

[0105] CRLs are protein complexes containing 3 or 4 components: all CRLs contain a cullin protein (in mammals, selected from Cull , Cul2, Cul3, Cul4A, Cul4B, Cul5 and Cul7), a substrate receptor and a RING protein, and with the exception of those containing Cul3, CRLs also contain an adaptor protein. The cullin protein forms the central scaffold of the CRL, the C-terminus of which binds the RING protein (Rbx1 or Rbx2) which serves to recruit ubiquitin-bound E2 enzymes. The N-terminus of the cullin protein binds the adaptor protein, which assists the cullin to recruit the substrate receptor. The substrate receptor in turn recruits the ubiquitylation target (substrate) to the CRL complex (Nguyen et al., Cullin-RING E3 Ubiquitin Ligases: Bridges to Destruction. In: Harris, J., Maries- Wright, J. (eds) Macromolecular Protein Complexes. Subcellular Biochemistry, Vol 83., Chapter 12; Springer, 2017).

[0106] The degradation domain may be any CRL component, i.e. a cullin, a substrate receptor, an adaptor protein or a RING protein. In particular, the degradation domain may be a cullin, a substrate receptor or an adaptor protein, preferably a substrate receptor.

[0107] Different cullins form complexes with different binding partners: Cul5 uses Rbx2 as RING protein, whereas all other cullins use Rbx1 . Cull and Cul7 use the Skp1 adaptor protein, Cul2 and Cul5 use EloBC (the elongin B / elongin C complex) and Cul4A and Cul4B use DDB1 . Each adaptor protein / cullin complex interacts with a different family of substrate receptors: Cul1 / Skp1 recruits F-box proteins, Cul2 / EloBC recruits VHL box proteins, Cul4A / DDB1 and Cul4B / DDB1 recruit DCAF proteins, Cul5 / EloBC recruits SOCS box proteins and Cul7 / Skp1 recruits only the specific F-box protein Fbw8 (as far as is known). Cul3 interacts directly with BTB protein substrate receptors, without using a separate adaptor protein. Any such adaptor or substrate receptor protein can be used as the degradation domain herein.

[0108] Where a component of an E3 ligase complex is used as degradation domain, the component acts to recruit the other components of the complex to the bioPROTAC, such that the entire E3 ligase complex forms on the bioPROTAC and ubiquitylates the target protein.

[0109] In some embodiments, the degradation domain is a wild type E3 ligase, E3 ligase complex component, E3 binding partner or E2 enzyme. The term “wild type” here refers to a protein (e.g. E3 ligase) which exists in nature (e.g. a human protein) and has the same sequence as the natural protein, i.e. which is not modified relative to the natural protein. In other embodiments, the degradation domain is a fragment of a wild type E3 ligase, E3 ligase complex component, E3 binding partner or E2 enzyme. That is to say, the degradation domain is a fragment of an E3 ligase, E3 ligase complex component, E3 binding partner or E2 enzyme which exists in nature and is not modified relative to the natural protein, such that the fragment corresponds to an amino acid sequence present in the wild type protein.

[0110] In other embodiments, the degradation domain is a modified E3 ligase, a modified fragment of an E3 ligase, a modified component of an E3 ligase, a modified E3 ligase binding partner or a modified E2 conjugating enzyme, comprising one or more amino acid substitutions, deletions or insertions relative to the initial or wild type sequence (e.g. up to 3, 5, 10, 15 or 20 amino acid substitutions, deletions and / or insertions relative to the initial or wild type sequence). Such a modified degradation domain may have at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to the initial or wild type sequence. As set out above, the wild type sequence of such a protein is the sequence which exists in nature. The “initial” sequence, as used herein, applies to non-natural proteins (scFvs, DARPins, etc.) and refers to the sequence as initially developed or isolated, e.g. from phage display or a hybridoma, prior to any alteration.

[0111] In some embodiments, the degradation domain (e.g. E3 ligase, component of an E3 ligase, E3 ligase binding partner or E2 conjugating enzyme) is modified to avoid or reduce ubiquitylation (that is to say the degradation domain may be modified so that ubiquitylation of the degradation domain is reduced or avoided). Modifications to reduce or avoid ubiquitylation may be as described above in the context of the binding protein. In particular, the degradation domain may comprise at least one substitution of a lysine residue relative to the initial or wild type sequence, preferably at least one lysine to arginine substitutions relative to the initial or wild type sequence. For example, the degradation domain may comprise 1 to 5, 1 to 10, 1 to 15 or 1 to 20 lysine to arginine substitution relative to the initial or wild type sequence, e.g. 1 , 2, 3, 4, 5, 6 or 7 or more lysine to arginine substitutions relative to the initial or wild type sequence.

[0112] Modification of the degradation domain in this manner may result in the level of ubiquitylation being reduced by any amount in the modified degradation domain relative to the initial or wild type sequence. For instance, ubiquitylation of the modified degradation domain may be reduced by at least 10, 20, 30, 40, 50, 60, 70, 80 or 90 %. In other cases, modification of the degradation domain may remove possible ubiquitylation sites, without actually reducing ubiquitylation of the degradation domain in practice. A reduction in ubiquitylation may be measured in the same way as for the binding domain.

[0113] The degradation domain may also or alternatively comprise one or more sequence modifications relative to the initial or wild type sequence for a reason other than reducing ubiquitylation.

[0114] As is well known in the art, different E3 ligases recognise and ubiquitylate different target proteins. Thus the E3 ligase used in a bioPROTAC is capable of inducing ubiquitylation of the target protein recognised by the binding domain. Suitable degradation domains can be identified by screening of degradation domains of interest. For some potential target proteins, E3 ligases which naturally ubiquitylate them will also be known in the art. As shown in the Examples, a functional bioPROTAC also requires a compatible combination of degradation domain and binding domain: any degradation domain which ubiquitylates a given target cannot be combined with any binding domain which recognises that target to yield a functional bioPROTAC. Rather, compatible combinations of degradation domain and binding domain can be identified using methods as described herein. When screening for compatible combinations of degradation and binding domains for a target protein with known native degraders, both native degraders and degraders which do not, or are not known to, natively drive degradation of the target protein may be included. In some cases, non-native degraders of a target protein may yield more effective bioPROTACs.

[0115] Exemplary degradation domains which may be used in the bioPROTACs provided herein and produced according to the method provided herein are set out in SEQ ID NOs: 1 to 100, detailed in Table 1 below.

[0116] The degradation domain used herein may comprise or consist of any one of SEQ ID NOs: 1 to 100, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to any one of SEQ ID NOs: 1 to 100. For example, the degradation domain used herein may comprise any one of SEQ ID NOs: 1-2, 5, 7-8, 11-16, 18-24, 26 or 31-52, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to any one of SEQ ID NOs: 1 to 100. Where the degradation domain is a variant of one of SEQ ID NOs: 1 to 100, the variant retains its ability to induce degradation of a target protein.

[0117] Several of the degradation domains listed in Table 1 contain one or more mutations relative to the wild type protein from which they’re derived (as set out in the table below). Where variants of these degradation domains are used, the mutations may remain. That is to say, where a variant of one of the degradation domains listed in Table 1 is used, the amino acid(s) at the position(s) corresponding to the mutated position(s) listed in the table may be unchanged from those specified in the table.

[0118] The mutations listed in Table 1 are numbered according to the full length wild type, sequences. In a variant of one of the degradation domains set out below, the amino acid at a position corresponding to a mutated position can thus be identified by aligning the variant sequence with the wild type sequence. The amino acid in the variant degradation domain aligned to a particular position in the wild type sequence corresponds to that position.

[0119] Table 1. Exemplary Degradation Domains

[0120]

[0121]

[0122]

[0123] Number ranges indicate the degradation domain is a fragment from those positions in the native protein, e.g. ICPO (115-158) is a fragment of ICPO corresponding to amino acids 115-158 of the original protein. KR numbering indicates the number of lysine to arginine substitutions in the protein, e.g. 7KR indicates the degradation domain comprises 7 lysine to arginine substitutions relative to the native sequence. A225 / A229 are antibody clone numbers.

[0124] Linkers

[0125] The bioPROTACs provided herein or produced according to the method provided herein may comprise a linker. A linker is a non-functional sequence between the binding and degradation domains which serves to separate the two domains.

[0126] It is known that the structural arrangement of bioPROTAC domains affects function, and so the choice of whether to include a linker, and if so what linker to select, is important in generating a functional bioPROTAC. As shown in the Examples below, different combinations of domains work better with different linkers.

[0127] The linker may be of any length, as optimal for the combination of active domains. For example, the linker may be 1 to 200 amino acids in length. In some cases, the linker may be a short linker, e.g. of about 1 to 20, 3 to 20, 5 to 20, 1 to 18, 3 to 18, 5 to 18, 1 to 16, 3 to 16 or 5 to 16 amino acids. In other cases, the linker may be of medium length, e.g. of about 20 to 100, 30 to 100, 40 to 100, 50 to 100, 60 to 100, 20 to 90, 30 to 90, 40 to 90, 50 to 90 or 60 to 90 amino acids. In yet other cases, the linker may a long linker, e.g. of over 100 amino acids, such as over 110, 120, 130, 140, 150, 160, 170, 180, 190 or 200 amino acids, or of about 100 to 200, 110 to 200, 120 to 200, 130 to 200, 140 to 200, 150 to 200, 100 to 190, 100 to 180, 100 to 170, 100 to 160 or 100 to 150 amino acids.

[0128] The linker may be flexible or rigid, or a combination of both (i.e. part flexible and part rigid). A flexible linker is unstructured and is generally composed of small, polar or non-polar amino acid residues (e.g. glycine, serine, threonine and alanine). Rigid linkers on the other hand may adopt a secondary structure (a-helix or (3-sheet). Other rigid linkers are based on the rigidity of proline, and comprise a high proportion of proline residues (e.g. at least 20, 30, 40, 50, 60, 70, 80 or 90 %) proline or a polyproline sequence. Some linkers comprise both rigid and flexible sections, and thus are part flexible and part rigid (referred to herein as rigid / flexible linkers). A flexible linker does not guarantee any particular steric separation of the binding and degradation domains - the degree of separation may range at maximum to the length of the linker, to nothing, due to the linker’s flexibility. A rigid linker or part-rigid on the other hand guarantees some minimum degree of steric separation between binding and degradation domains.

[0129] In some embodiments, the linker is a short, flexible linker, or a medium length flexible linker. In other embodiments, the linker is a short or medium length rigid / flexible linker. In yet other embodiments, the linker is a short or medium length rigid linker. As shown in the Examples below, in our screens short flexible linkers generally yielded superior results to longer, more rigid linkers, and as such short flexible linkers may be preferred, though longer and / or more rigid linkers may yield better results against other targets. Various linker sequences for joining fusion protein constituents are known in the art and can be used in the present bioPROTACs. Flexible linkers include glycine-serine linkers, and these are a preferred type of linker for inclusion in the bioPROTACs provided herein. Glycine-serine linkers are linkers consisting of glycine and / or serine residues, and often include a repeating glycine / serine sequence of 1 to 5 amino acids, for example: [G]n, [S]n, [GS]n, [GG]n, [GGG]n, [GGS]n, [GGGSjn (SEQ ID NO: 113), [GGGGSjn (SEQ ID NO: 114), [GGSGjn (SEQ ID NO: 115), [GSGGjn (SEQ ID NO: 116), [SGGGjn (SEQ ID NO: 117), [SSGGjn (SEQ ID NO: 118), [SSSGjn (SEQ ID NO: 119) and [GGGSGjn (SEQ ID NO: 120). Other flexible linkers include [A]n, [SA]n, [SAGSjn (SEQ ID NO: 121) and [SGASjn (SEQ ID NO: 122), and other such short sequences of glycine, serine and / or alanine residues. In all such linkers, ‘n’ is generally between 1 and 20, preferably between 1 and 5 (though the exact number of ‘n’ may depend on the length of the repeating units, with a higher ‘n’ used for a shorter repeating unit and vice versa. Exemplary longer linkers, made up of repeating units as set out above, include those set out in SEQ ID NOs: 123-126. The bioPROTACs provided herein may comprise or consist of any of the above-listed flexible linkers. Preferred glycine-serine linkers include those of SEQ ID NOs: 120 and 126, which are used as linkers L1 and L2 in the Examples below.

[0130] Several rigid and rigid / flexible linkers are described in Klein et al., Protein Engineering, Design and Selection 27(10): 325-330, 2014, incorporated by reference herein, any of which may be used in the bioPROTACs provided, including SEQ ID NOs: 127-131. Other suitable rigid or rigid / flexible linker sequences are set out in SEQ ID NOs: 132 and 133, which are used as linkers L3 and L4 in the Examples below. The bioPROTACs provided herein may comprise or consist of any of the above-listed rigid or rigid / flexible linkers.

[0131] Variants of the linkers above may also be used, and the bioPROTACs provided herein may comprise a linker comprising or consisting of an amino acid sequence having at least 75, 80, 85, 90, 95 or 99 % identity to any one of SEQ ID NOs: 113-133.

[0132] Identifying Functional bioPROTACs

[0133] As set out above, producing a functional bioPROTAC against a given target requires the selection of a compatible binding domain / degradation domain pair, and optionally linker. At present, functional combinations require empirical identification by screening. Provided herein is an effective screening method, which as shown in the Examples has been successfully used to identify bioPROTACs against two different target proteins, both with therapeutic relevance. The method is thus for identifying a bioPROTAC capable of inducing degradation of a target protein in a eukaryotic cell. The bioPROTAC comprises a binding domain and a degradation domain, and optionally a linker, as defined above.

[0134] Generally, the method described below may be miniaturised (such that all reactions are performed in small volumes) and / or automated, e.g. using pipetting robots, etc. Advantageously, in the method liquids may be dispensed by acoustic dispensing. All or some steps involving dispensing of liquids may be performed by acoustic dispensing (e.g. as opposed to more traditional pipetting methods). Acoustic dispensing enables accurate dispensing of smaller volumes of liquid than is easily achievable by pipetting. Acoustic dispensing may be performed with an acoustic droplet ejector (e.g. the Echo Liquid Handler, Beckman Coulter).

[0135] DNA Libraries

[0136] Step (i) of the method is to provide a first DNA library encoding a library of binding domains, and step (ii) of the method is to provide a second DNA library encoding a library of degradation domains. Thus the method begins with a first DNA library encoding a library of binding domains and a second DNA library encoding a library of degradation domains. Where the bioPROTACs to be generated in the screening method are also to comprise a linker, a third DNA library encoding a library of linkers is also provided.

[0137] The term “DNA library” as used herein refers to a collection of DNA elements (or molecules). Any suitable type of DNA element may be used, e.g. a linear DNA molecule. Commonly a DNA vector is used, such as a plasmid, a phagemid, a cosmid, an artificial chromosome such as a yeast artificial chromosome (YAC), bacterial artificial chromosome (BAC) or P1 -derived artificial chromosome (PAC), a bacteriophage such as lambda phage or M13 phage, or a viral vector. Suitable viral vectors include adeno-associated virus (AAV) vectors, adenovirus vectors, herpes simplex virus vectors, retrovirus vectors, lentivirus vectors, alphavirus vectors, flavivirus vectors, rhabdovirus vectors, measles virus vectors, Newcastle disease virus vectors, poxvirus vectors and picornavirus vectors. Preferably, the DNA libraries provided herein utilise plasmids. Whatever DNA element is used, generally all the elements in a DNA library are of the same type, e.g. all plasmids, and not a mixture of different types.

[0138] Any type of plasmid may be used in a DNA library in the present method. Generally the plasmid is a cloning vector (i.e. a cloning plasmid), for example a pENTR (Gateway), pUC19, pBR322, pBluescript vectors (Stratagene Inc.) or pCR TOPO® plasmid.

[0139] Vectors, such as plasmids, forming the DNA library may be stored in any suitable manner. The plasmids may be stored in solution (generally frozen). Alternatively, the plasmids may be stored within a population of microorganisms, such as bacteria (e.g. E. coll or Bacillus subtilis) or yeast (e.g. Saccharomyces cerevisiae).

[0140] The DNA elements in the DNA libraries each encode a bioPROTAC component (binding domain, degradation domain or linker). Each DNA element in a DNA library (e.g. each plasmid) encodes a single bioPROTAC component, i.e. a single binding domain, a single degradation domain or a single linker. By “encoding” a bioPROTAC component is meant that each DNA element comprises a nucleotide sequence (gene sequence) which codes for a bioPROTAC component. Generally the DNA libraries are made up of cloning vectors and so the bioPROTAC components cannot be expressed from them.

[0141] The gene sequences encoding the bioPROTAC components may be generated by any suitable method and then, if necessary, inserted into a vector, such as a plasmid, to provide an element of a DNA library. For example, sequences encoding a bioPROTAC component can be chemically synthesised or produced by PCR amplification from plasmid or genomic DNA. Where a sequence contains one or more mutations relative to a wild type sequence, these can be introduced by any suitable means, e.g. site-directed mutagenesis. Linear DNA fragments encoding bioPROTAC components can then be inserted into DNA vectors, e.g. plasmids, as discussed further below.

[0142] As noted above, the first DNA library encodes a library of binding domains. This means that a large collection of binding domains is encoded across the first DNA library. As noted above, a single DNA element (e.g. plasmid) encodes only a single bioPROTAC component, and so the first DNA library contains as many different DNA elements as different binding domain sequences. Thus if the first DNA library encodes 30 different binding domains, this means that across all the elements in the DNA library 30 different binding domains are encoded, and there are 30 different elements in the DNA library (e.g. 30 plasmids).

[0143] The first DNA library may encode any number of binding domains, as many as it is desired to screen, e.g. at least 1 , 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 binding domains. For example the first DNA library may encode 1-50, 1-40, 1-30, 1-25, 1-20, 1-15, 1-10, 5-50, 5-40, 5-30, 5-25, 5-20, 10-50, 10-40, 10-30, 10-25 or 10-20 different binding domains. The binding domains are all known, believed or expected to specifically bind the target protein. The binding domains encoded may be of the types described above. Preferably at least two different types of binding domain are encoded by the first DNA library (e.g. at least one scFv and at least one DARPin, etc.).

[0144] The second DNA library encodes a library of degradation domains. Again, the second DNA library may encode any number of degradation domains, as many as it is desired to screen, e.g. at least 1 , 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90,100, 110, 120, 130, 140 or 150 degradation domains. For example, the second DNA library may encode 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-40, 1-30, 1-25, 1-20, 1-15, 1-10, 5-100, 5-90, 5-80, 5-70, 5-60, 5-50, 5-40, 5-30, 5-25, 5-20, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-25, 10-20, 20-100, 20-80, 20-60, 20-50, 20-40, 30-100, 30-80, 30-60, 30-50, 40-100, 40-90, 40-80, 40-60 or 50-100 degradation domains.

[0145] Degradation domains which natively have a continuous amino acid sequence, and are not reliant on post-translational modification to function, are particularly suitable for use in bioPROTACs and thus for inclusion in the library of degradation domains. As noted above, where an E3 ligase is used as a degradation domain, it is advantageous where possible to use a fragment of the natural enzyme lacking the natural substrate binding domain.

[0146] The degradation domains encoded by the second DNA library may be selected based on the identity of the target protein. For any given target, degradation domains encoded by the library may encode E3 ligases which are known to ubiquitylate the target protein and direct it for degradation, components of E3 ligase complexes which are known to ubiquitylate the target protein and direct it for degradation (e.g. a CRL component, particularly a substrate receptor of a CRL known to interact with the target protein), E2 enzymes known to interact with E3 ligases which are known to ubiquitylate the target protein and binding partners of E3 ligases E3 ligases which are known to ubiquitylate the target protein.

[0147] Degradation domains can also be included in the library based on a likelihood or possibility that they will be capable of ubiquitylating the target protein. An important consideration is the sub-cellular localisation of an E3 ligase. It is advantageous if the degradation domain either is, or recruits, an E3 ligase which is naturally located in the same sub-cellular compartment as the target protein, as it is more likely that such an E3 ligase will be capable of ubiquitylating the target protein than an E3 ligase which naturally resides in a different sub-cellular compartment to the target and never naturally contacts it. This is particularly important where the degradation domain acts to recruit an E3 ligase (or other components of an E3 ligase complex to the bioPROTAC), as a bioPROTAC is unlikely to function if it needs to simultaneously bind to two proteins (target protein and E3 ligase) located in different sub-cellular compartments.

[0148] For example, if the target protein is a cytoplasmic protein, it is advantageous to include in the library of degradation domains cytoplasmic E3 ligases, cytoplasmic E2 enzymes, components of cytoplasmic E3 ligase complexes and other proteins capable of recruiting cytoplasmic E3 ligases (e.g. E3 ligase partners). Similarly, if the target protein is a nuclear protein, it is advantageous to include in the library of degradation domains nuclear E3 ligases, nuclear E2 enzymes, components of nuclear E3 ligase complexes and other proteins capable of recruiting nuclear E3 ligases (e.g. E3 ligase partners). The library of degradation domains may contain one or more of the specific degradation domains set out in SEQ ID NOs: 1-100 (and / or variants thereof), which are discussed above. For instance, the IDOL E3 ligase, from which SEQ ID NO: 8 is derived, is naturally located in the cytoplasm and the nucleus, and so can reasonably be included in libraries for producing bioPROTACs against both nuclear and cytoplasmic proteins. MDM2, recognised by the single-domain antibodies of SEQ ID NOs: 13 and 14, is located in the nucleus. The CRL substrate receptors AnkB (SEQ ID NO: 15), FBXO17 (SEQ ID NOs: 21 and 22), LegU1 (SEQ ID NO: 24) and pTrCP (SEQ ID NOs: 29 and 30) all interact with the Skp1 CRL adaptor protein, which localizes to the cytoplasm and nucleus. The CRL substrate receptors CSA (SEQ ID NO: 16), DDB2 (SEQ ID NOs: 17 and 18) and SV5-V (SEQ ID NO: 26), along with DDA1 (SEQ ID NO: 35) all interact with the DDB1 adaptor protein, which localizes to the cytoplasm and nucleus. The CRL substrate receptors E40rf6 (SEQ ID NO: 19), E7 (SEQ ID NO: 20), VHL (SEQ ID NO: 27) and Vif (SEQ ID NO: 28) all interact with the nuclear EloC component of the EloBC adaptor complex. The CRL substrate receptors KEAP1 (SEQ ID NO: 23) and SPOP (SEQ ID NO: 25) interact with Cul3, which is present in the cytoplasm and nucleus. The SelK degron (SEQ ID NOs: 31 and 32) interacts with the nuclear E3 ligase KLHDC2.

[0149] Where the degradation domain acts to recruit an E3 ligase, or components of an E3 ligase complex, it is important that the recruited E3 ligase is expressed by the cell line in which the bioPROTACs produced by the method are to be screened. If the bioPROTAC is being generated with the intention of use in therapy, e.g. for cancer, it is advantageous if the recruited E3 ligase is expressed in a wide range of tissues. It is also advantageous if the recruited E3 ligase is widely expressed in cancer, ideally overexpressed, to ensure sufficient E3 ligase is available in target cells to ubiquitylate the target protein. It is also advantageous if the recruited E3 ligase is essential, to avoid cellular resistance to the bioPROTAC developing. Essential E3 ligases are E3 ligases which are necessary for cellular survival, and may be identified experimentally or from the Online Gene Essentiality (OGEE) database (http: / / oqee.medqenius.info / , Chen et al., Nucleic Acids Research 45(D1): D940-D944, 2017).

[0150] The library of degradation domains may include multiple different types of degradation domain, i.e. at least two of an E3 ligase, an E2 enzyme, a component of an E3 ligase complex and an E3 ligase binding partner.

[0151] The third DNA library (where used) encodes a library of linkers. The third DNA library may encode any number of linkers, e.g. at least 1 , 2, 3, 4, 5, 10, 15 or 20 different linkers. For example the third DNA library may encode 1-20, 1-10 or 1-5 different linkers.

[0152] Any suitable linker sequences may be included in the linker library, such as the linkers discussed above. Preferably the linker library includes linkers of a variety of lengths, e.g. at least one short linker and at least one medium length linker. Preferably the linker library also includes linkers of differing flexibility, e.g. at least one flexible linker and at least one rigid linker, or at least one flexible linker and at least one rigid / flexible linker, or at least one flexible linker, at least one rigid linker and at least one rigid / flexible linker.

[0153] Vector Library

[0154] Step (iii) of the method provided herein entails assembling a vector library encoding a bioPROTAC library. The term “vector library” is used herein to refer specifically to the library of vectors produced in step (iii), which encodes a bioPROTAC library. In general terms, the vector library is a DNA library, but for distinction from the DNA libraries which encode the bioPROTAC components discussed above, is referred to as a vector library. Like the DNA libraries described above, the vector library may be stored in any suitable form, e.g. in solution or in microorganisms.

[0155] The vector library is thus a library of vectors. The vectors in the vector library may be of any type, e.g. a plasmid, a phagemid, a cosmid, an artificial chromosome such as a yeast artificial chromosome (YAC), bacterial artificial chromosome (BAC) or P1 -derived artificial chromosome (PAC), a bacteriophage such as lambda phage or M13 phage, or a viral vector. Although the vectors in the vector library may be of any type, any given vector library contains only a single type of vector. The vectors in the vector library may be used for expression of the bioPROTACs, and so the vectors may be expression vectors. An “expression vector” as used herein is a DNA molecule which can be used for expression of genetic material in an in vitro system. Preferably the vectors in the vector library are plasmids, and the vector library is a plasmid library. Suitable plasmids for use in the vector library include pcDNA3.1 (+), pRSET vectors (Thermo Fisher Scientific) and pET vectors (Merck Millipore).

[0156] The vector library encodes a bioPROTAC library. By analogy to the DNA libraries discussed above, this means that a large collection of bioPROTACs are encoded across the vector library. Each vector encodes only a single bioPROTAC, and thus there are as many different vectors in the vector library as there are bioPROTACs in the bioPROTAC library. The bioPROTAC library is for screening to identify bioPROTACs which induce degradation of the target protein. Each bioPROTAC in the bioPROTAC library comprises a binding domain from the binding domain library (encoded by the first DNA library) and a degradation domain from the degradation domain library (encoded by the second DNA library, and optionally a linker (encoded by the third DNA library). Thus each bioPROTAC in the library is assembled from the components encoded by the first, second and third DNA libraries. Assembly of the vector library from the DNA libraries is discussed below. Since not all bioPROTACs in a bioPROTAC library can be expected to work, generally the bioPROTAC library contains a large number of bioPROTACs, for screening. For example, the bioPROTAC library may comprise at least 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 bioPROTACs, e.g. 50-1000, 50-900, 50-800, 50-700, 50-600, 50-500, 50-400, 50-300, 50-200, 100-1000, 100-900, 100-800, 100-700, 100-600, 100-500, 100-400, 100-300, 100-200, 200-1000, 200-900, 200-800, 200-700, 200-600, 200- 500, 200-400, 400-1000, 400-800, 400-600, 500-1000 or 500-750 different bioPROTACs.

[0157] All the bioPROTACs in a bioPROTAC library are different, i.e. they have different sequences. In some cases a bioPROTAC library may contain bioPROTACs with different components, e.g. some bioPROTACs in a library may contain a linker and others not. Since the arrangement of domains in a bioPROTAC can be important, a bioPROTAC library may contain multiple bioPROTACs with the same domains but arranged differently. For example a bioPROTAC library may comprise two bioPROTACs with the same binding and degradation domains (and linker), but one arranged with the binding domain at the N-terminus and the other arranged with the degradation domain at the N-terminus.

[0158] Each vector in the vector library comprises an expression cassette comprising a gene encoding a bioPROTAC. An “expression cassette” as used herein is a polynucleotide sequence that is capable of effecting transcription of an expression product, i.e. the bioPROTAC. A “coding sequence” is intended to mean a portion of a gene’s polynucleotide sequence that encodes the expression product. In the context of a protein such as a bioPROTAC, this sequence may be referred to as a “protein coding sequence”. The protein coding sequence typically begins at the 5’ end with a start codon and ends at the 3’ end with a stop codon.

[0159] Typically, the expression cassette comprises a promoter operably linked to the bioPROTAC coding sequence. The term “operably linked” includes the situation where a coding sequence and promoter are covalently linked in such a way as to place the expression of the coding sequence under the influence or control of the promoter. Thus, a promoter is operably linked to the bioPROTAC coding sequence if the promoter is capable of effecting transcription of the bioPROTAC coding sequence. The resulting transcript may then be translated into the bioPROTAC protein.

[0160] As set out further below, the bioPROTAC genes in the vector library are transcribed in vitro, and thus the promoter in the expression cassette is a promoter suitable for use in in vitro transcription. Any promoter may be used, depending on the polymerase used for in vitro transcription. Commonly a phage RNA polymerase is used for in vitro transcription, such as the T7, T3 or SP6 RNA polymerase. In particular embodiments, the expression cassette comprises a T7, T3 or SP6 promoter, operably linked to the bioPROTAC gene. In some embodiments, the promoter may be a modified version of a natural promoter, e.g. a modified version of a T7, T3 or SP6 promoter.

[0161] The expression cassette also comprises sequences encoding a 5’ untranslated region (UTR) and a 3’ UTR, which flank the bioPROTAC coding sequence. As is well known in the art, the 5’ and 3’ UTRs are transcribed into mRNA with the coding sequence - the 5’ UTR forms the 5’ end of the transcribed mRNA and the 3’ UTR forms the 3’ end. The UTRs are not translated into protein. Any suitable combination of UTR sequences may be used in the expression cassette in the vectors herein. The 5’ and 3’ UTRs may originate or be derived from a gene or may be synthetic. The 5’ and 3’ UTRs may have the same source or different sources, e.g. the 5’ UTR may be synthetic and the 3’ UTR may originate or be derived from a gene, or vice versa. A UTR originating from a gene refers to a UTR which is identical to a UTR found in nature, in the context of a gene. A UTR derived from a gene refers to a UTR which is based on a UTR found in nature, but which is modified relative to that native (parent UTR). A UTR derived from a gene may have e.g. at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to the parent UTR, and / or may comprise additional sequence elements 5’ and / or 3’ to the parent UTR.

[0162] Where a UTR originates or is derived from a gene, that gene is a eukaryotic gene. The species from which the UTRs originate may depend on the origin of the cell line in which the bioPROTACs are to be expressed and screened. The UTRs may originate or be derived from a mammalian gene, such as a human gene. Suitable UTR sequences are known in the art, see e.g. Thess et al., Molecular Therapy 23(9): 1456-1464, 2015; and De Genst et al., Nature Communications 13: 3018, 2022. Different UTRs may be optimal (in terms of the level of translation attained) for expression of different proteins.

[0163] The 3’ UTR comprises, at its 3’ end, a polyadenosine (polyA) tail. Commonly, lengthy of polyA tails of 100-300 adenosine residues are used in synthetic mRNAs, but as shown in the Examples, the present inventors have found that in the present context a shorter polyA tail is sufficient to provide mRNA stability and efficient translation. The 3’ UTR encoded in the expression cassette may thus comprise a polyA tail which is about 30-100 nucleotides (i.e. adenosine nucleotides) in length e.g. about 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-100, 50-90, 50-80, 50-70, 50-60 60-100, 60- 90, 60-80, 60-70, 70-100, 70-90, 70-80 or 80-100 nucleotides in length. In particular embodiments the expression cassette encodes a 3’ UTR comprising a polyA tail which is about 40-80, 40-60 or 60-80 nucleotides in length, e.g. about 40, about 60 or about 80 nucleotides.

[0164] An expression cassette may further comprise additional elements useful for control of transcription, e.g. one or more enhancer sequences and / or one or more transcription termination (terminator) sequences. As set out in the examples, a vector suitable for use in the vector library can be generated from a commercially available vector by inserting an expression cassette comprising a cloning site, into which a bioPROTAC gene can be inserted / assembled.

[0165] The vectors in the vector library may additionally comprise other standard sequence elements such as a selectable marker, e.g. encoding an antibiotic resistance gene or suchlike.

[0166] The vectors in the vector library are each assembled by joining a vector backbone (i.e. an expression vector backbone), a first DNA molecule from the first DNA library and a second DNA molecule from the second DNA library. Thus the first DNA molecule comprises a gene / gene fragment (or DNA sequence) encoding a binding domain and the second DNA molecule comprises a gene / gene fragment (or DNA sequence) encoding a degradation domain.

[0167] Where a vector is to encode a bioPROTAC comprising a linker, the vector is assembled by joining a vector backbone (i.e. an expression vector backbone), a first DNA molecule from the first DNA library, a second DNA molecule from the second DNA library and a third DNA molecule from the third DNA library. As set out above, in this case the first and second DNA molecules encode a binding domain and a degradation domain, respectively, and the third DNA molecule comprises a DNA sequence encoding a linker.

[0168] Generally, the vector backbone comprises all elements required of the expression vector aside from the sequences of the bioPROTAC components, including all elements of the expression cassette. Where it is desired for the bioPROTACs to be located in a particular cellular compartment upon expression, the vector backbone may encode a trafficking signal which forms part of the bioPROTAC coding sequence. For instance, the examples below describe a vector backbone comprising an NLS.

[0169] As described above, the bioPROTAC components can be arranged in either order, i.e. the gene encoding a bioPROTAC can be assembled with the first DNA molecule, comprising the binding domain sequence, at the 5’ end and the second DNA molecule, encoding the degradation domain sequence, at the 3’ end, such that the encoded bioPROTAC has an N-terminal binding domain and a C-terminal degradation domain. Alternatively, the gene encoding a bioPROTAC can be assembled with the first DNA molecule, comprising the binding domain sequence, at the 3’ end and the second DNA molecule, encoding the degradation domain sequence, at the 5’ end, such that the encoded bioPROTAC has a C-terminal binding domain and an N-terminal degradation domain. If the encoded bioPROTAC contains a linker that is of course located between the binding domain and the degradation domain, so the third DNA molecule is located between the first and second DNA molecules in the bioPROTAC gene. The expression vectors in the vector library are assembled from DNA molecules from the DNA libraries. In some cases, if the DNA libraries contain linear DNA fragments, essentially the entirety of the DNA molecules making up the DNA libraries may be assembled into the vectors. However, if the DNA libraries comprise vectors (e.g. plasmids) comprising sequences encoding the bioPROTAC components, the entire vectors are not used. Rather, the DNA molecules for use in assembling the vectors in step (iii) are obtained from the vectors of the DNA libraries, e.g. by digesting the vectors or by amplifying and isolating the desired sequences.

[0170] The vectors may be assembled by any method known in the art. For instance, the vectors may be assembled by sequential rounds of traditional cloning using restriction enzymes. More preferably, a single step DNA molecule assembly technique is used, for example Gibson assembly or Golden Gate assembly. Kits for Gibson assembly can be obtained from New England Biolabs (USA), see also Gibson et al., Science 329(5987): 52- 56, 2010. Kits for Golden Gate assembly can also be obtained from NEB, and see Engler et al., PLoS ONE 3(11): e3647, 2008.

[0171] Preferably, the vectors are assembled by Golden Gate assembly. That is to say, preferably in step (iii), each vector is assembled by Golden Gate assembly, as discussed further below.

[0172] Golden Gate Assembly

[0173] Golden Gate assembly (also referred to herein as Golden Gate cloning) is a well known technique for assembly of multiple (i.e. more than two) DNA fragments into a single molecule, e.g. plasmid, which unlike traditional restriction enzyme-based methods yields scarless joins between molecules (i.e. no restriction sites are retained in the final assembly).

[0174] The Golden Gate method relies upon the use of type IIS endonucleases. These are a distinct type of restriction endonuclease which have non-palindromic, directional recognition sites, and cleavage sites which are a defined distance (e.g. 1 to 20 nucleotides) and direction outside of their recognition sites. Type IIS endonucleases generally produce sticky ends when cutting DNA, and thus DNA sequences which are to be joined during assembly can be designed such that type IIS cleavage yields complementary overhangs. By using different complementary overhangs between each pair of DNA fragments to be joined, multiple DNA fragments can be assembled in a directed, pre-defined order and orientation. Because the cut sites are outside of the type IIS recognition sites, the assembled DNA product need not include the recognition sequence used to produce the original fragments. Suitable type IIS endonucleases for use in Golden Gate cloning, and which can be used herein, include Bsal, Esp3l, BsmBI and Bbsl. A single type IIS endonuclease is generally used to produce all the components of the assembly.

[0175] Generally Golden Gate assembly has the following steps: digestion, where starting DNA molecules are digested using the type II endonuclease to produce the fragments for assembly; and assembly / ligation, where the fragments assemble based on their complementary overhangs and are ligated by a ligase enzyme. The products can then be transformed into competent cells and isolated based on their selectable marker. Products can then be sequenced to confirm correct assembly.

[0176] In a particular embodiment, the method provided herein comprises steps of:

[0177] (a) obtaining a first gene fragment encoding a binding domain, a second gene fragment encoding a degradation domain and a third gene fragment encoding a linker;

[0178] (b) inserting the first, second and third gene fragments into cloning vector backbones, thereby yielding a first intermediate vector, a second intermediate vector and a third intermediate vector;

[0179] (c) digesting the first, second and third intermediate vectors and an expression vector with a type IIS restriction enzyme; and

[0180] (d) ligating the products of step (c), thereby yielding a vector encoding a bioPROTAC.

[0181] The first, second and third gene fragments are linear DNA molecules and can be obtained by any suitable method. For instance, the gene fragments may be generated by chemical synthesis, by amplification from a template (e.g. a vector or genomic DNA), or by digestion of an existing vector using a suitable restriction enzyme. Amplification can be performed by any method known in the art, most commonly PCR. The gene fragments are each inserted into a vector backbone, yielding intermediate vectors. Both the gene fragments and the intermediate vectors may be seen as forming the first, second and third DNA libraries described above.

[0182] The intermediate vectors may be any type of vector, as set out above. Preferably the intermediate vectors are plasmids.

[0183] In a particular embodiment, the gene fragments are inserted into the cloning vector backbones by topoisomerase-based cloning. Topoisomerase-based cloning (including TOPO™ cloning, Thermo Fisher, USA) utilises DNA topoisomerase I, which has both restriction enzyme and ligase activity. Any suitable topoisomerase I enzyme can be used, for instance the Vaccinia DNA topoisomerase I used in TOPO™ cloning. The topoisomerase I enzyme can thus be used to both linearise the cloning vector and religate it following insertion of the gene fragment. The cloning vector backbone may be provided as a linear DNA molecule with a topoisomerase I enzyme attached (e.g. covalently) to the 3’ phosphates at both ends. This can then readily ligate a DNA fragment with compatible ends into the cloning vector. A TOPO™ kit can be used for this purpose.

[0184] The gene fragments may be sticky-ended or blunt ended, and the cloning vector backbones linearised to match. Suitable sticky ends include single deoxyadenosine nucleotides, which can match to single overhanging thymidine nucleotides on the linearised cloning vector backbone in a process known as TA cloning. Gene fragments with single deoxyadenosine nucleotide overhangs can be generated by amplification with Taq polymerase. Commercial kits are available for all such processes (TOPO™ kits, Thermo Fisher). Suitable cloning vectors for use in this process include the Invitrogen pCR plasmids, such as pCR Blunt ll-TOPO (used in the Examples below), pCR4Blunt-TOPO, pCR2.1- TOPO TA, pCR4-TOPO TA, pCRXL-2-TOPO, etc.

[0185] Intermediate vectors generated by insertion of a gene fragment into a cloning vector backbone can be isolated by transformation into competent cells (e.g. E. coli such as TOP10 cells). Transformants can be isolated using the vector’s selectable marker (e.g. TOPO plasmids carry a kanamycin resistance gene and so successful TOPO transformants can be isolated by kanamycin selection). Isolated intermediate vectors can then be sequenced to confirm the gene fragment has been correctly produced and inserted.

[0186] The intermediate vectors comprise type IIS restriction enzyme recognition sites flanking the gene fragments. These recognition sites may be provided in the cloning vector backbones, or in the gene fragments inserted into the cloning vectors. As described above, the cutting sites in each intermediate vector are designed so that the gene fragments can be directionally assembled during Golden Gate cloning.

[0187] Fragments for Golden Gate cloning are then produced from the intermediate vectors by digestion with a type IIS restriction enzyme. Suitable type IIS restriction enzymes are mentioned above. In some embodiments, the first, second and third intermediate vectors are digested with the type IIS restriction enzyme. In other embodiments, the gene fragments (including the type IIS restriction enzyme recognition and cutting sites) are amplified (e.g. by PCR) and then digested. The gene fragments produced by digestion with the type IIS endonuclease may then be purified, e.g. by a PCR purification step, or if the intermediate vectors are themselves digested with the endonuclease the digestion products may be separated by gel electrophoresis and the desired fragment cut out of and purified from the gel. An expression vector, from which the bioPROTAC is to be expressed, is also digested with the type IIS restriction enzyme, to yield a backbone into which the gene fragments encoding the bioPROTAC components can be inserted.

[0188] Following the digestion step, the products of the digestions (i.e. the vector backbone and the gene fragments encoding the bioPROTAC components) are ligated together using a ligase enzyme, e.g. T3, T4 or T7 DNA ligase. In some cases, the gene fragments are individually produced by separate digestion of the first, second and third intermediate vectors (or amplification products thereof), in which case the expression vector is also separately digested. In this case, the digestion products are mixed and ligated.

[0189] Alternatively, the digestion and ligation reactions may be performed in a single reaction mix. In this case, the intermediate vectors and intact expression vector are mixed with both a type IIS restriction enzyme and a DNA ligase at the same time. The reaction may be performed according to the manufacturer’s instructions, e.g. in a thermocycler using repeated cycles of about 37°C (e.g. for about 5 minutes) followed by about 16°C (e.g. for about 5 minutes).

[0190] Where the digestion and ligation reactions are performed in a single reaction mix, the reaction mix preferably has a low volume, e.g. at most 10, 5 or 2 pl. The minimum volume is determined only by the limits of pipetting technology, e.g. the minimum reaction volume may be about 0.5 or 1 pl. For example, the reaction may be performed in a volume of 1-5, 1-2 or 2-5 pl, e.g. about 1 , 2, 3, 4 or 5 pl. Use of low volumes is enabled by the use of only a small number of components in the reaction, and is particularly compatible with automation of the process.

[0191] In normal cloning methods, ligation products are transformed into competent cells, and transformants selected by growth on selective medium, using a selectable marker present in the vector / plasmid produced. Individual clones can then be sequenced, and products with the correct sequence taken forward. This procedure can be used in the present method - once produced in step (iii), vectors, however assembled, can be isolated and sequenced. However, the Golden Gate assembly method set out above has been found to demonstrate an approximately 99 % accuracy rate in terms of producing plasmids with the desired sequence. This compares to only about 80 % for more traditional methods. The high level of accuracy means that clonal selection is not required.

[0192] Regardless of the method of vector assembly, the ligation products are generally transformed into competent cells, commonly competent bacterial cells, e.g. competent E. coli.

[0193] In some embodiments, where Golden Gate cloning is performed as described above, the ligation products are transformed into competent cells and antibiotic selection is performed in liquid culture. That is to say, rather than being plated out onto an antibiotic plate for clonal selection, the transformed bacteria are put straight into liquid medium and cultured (e.g. overnight). This results in a pooled culture of all successful transformants, rather than a single clone, and speeds up the method by avoiding the time required for clonal selection.

[0194] In some embodiments, where Golden Gate cloning is performed as described above, the gene encoding the bioPROTAC is not sequenced prior to in vitro transcription in step (iv). As noted above, the high accuracy rate for this method means that clonal selection and sequencing is not necessary. Not sequencing the expression vectors produced in step (iii) saves further time compared to standard cloning processes. The present method is thus streamlined to provide a rapid process.

[0195] Vectors may then be isolated, if desired, by culturing transformed bacteria in liquid (including in the context of antibiotic selection), harvesting the cells and isolating the vectors. This can be achieved using a suitable commercial kit, e.g. a miniprep kit (available from e.g. Qiagen, Germany).

[0196] In Vitro Transcription

[0197] Following production of the vector library in step (iii), step (iv) entails producing an mRNA library by in vitro transcription of the genes encoding bioPROTACs. That is to say, the genes encoding bioPROTACs in the vectors produced in step (iii) are transcribed in vitro. The mRNA library is the collection of mRNA molecules produced by transcription of the bioPROTAC genes. The mRNA library may be stored by any suitable means, e.g. frozen in solution.

[0198] In vitro transcription processes are well known in the art and can be performed using a commercial kit, e.g. a HiScribe® kit from New England Biolabs (NEB, USA) or a RiboMAXTM kit from Promega (USA). Transcription is generally performed using a phage RNA polymerase, e.g. the T7 RNA polymerase, SP6 RNA polymerase or T3 RNA polymerase. The polymerase used must recognise the promoter contained in the expression cassette in the vectors.

[0199] The template for the in vitro transcription (IVT) may be the vector itself (i.e. the expression vector produced in step (iii). Alternatively, the expression cassette in the vector may be amplified, yielding an amplification product comprising the expression cassette, and the amplification product used as the in vitro transcription template.

[0200] As noted above, the vectors may be isolated from transformed bacteria, and the IVT template may thus be the purified vector. Similarly, amplification of the expression cassette may be performed using purified vector as template.

[0201] Alternatively, to streamline the process further, no vector purification step may be performed and the liquid culture itself may be used as template for IVT or amplification. That is to say, a small volume of liquid culture may be taken, optionally heated to release DNA, and IVT or amplification then performed. In a particular embodiment, amplification of the expression cassette is performed directly upon liquid cell culture, without an intermediate vector purification step. The amplification may be performed by any suitable method known in the art, preferably PCR. Where an amplification step is performed to generate the template for IVT, the amplification product may be purified (or “cleaned up”) prior to performance of IVT. Purification of the amplification product commonly includes degradation of the DNA template, and of genomic DNA where this is present in the amplification reaction (as is the case where amplification is performed using liquid culture as template). This may be achieved by incubation with Dpnl. A further purification step may be performed to remove other components of the amplification reaction, e.g. salts, nucleotides and DNA polymerase. Such a step may be performed using a standard, commercially available PCR purification kit, e.g. AMPure XP Beads (Beckman Coulter, USA) or a QIAquick PCR purification kit (Qiagen).

[0202] IVT may be performed using an RNA polymerase / kit as set out above. An IVT reaction utilises nucleotides (ATP, GTP, UTP and CTP) which are synthesised into RNA by the RNA polymerase. One or more of the nucleotides used for IVT may be replaced at least in part by a modified nucleotide. That is to say, the IVT reaction may contain one or more modified nucleotides.

[0203] Such a modified nucleotide may be a modified adenosine (A), guanosine (G), uridine (U) or cytidine (C). By “modified” here is meant that the structure of the nucleotide is altered relative to the natural structure. For any given nucleotide, all of (i.e. 100 % of) the nucleotide in the IVT reaction may be modified, all of the nucleotide in the IVT reaction may be unmodified (i.e. have the native structure of the nucleotide), or a certain percentage of the nucleotide in the IVT reaction may be modified, e.g. about 10, 20, 30, 40, 50, 60, 70, 80 or 90 % of the nucleotide in the IVT reaction may be modified. The proportion of a nucleotide which is modified in the IVT reaction can be expected to correlate to the proportion of the instances of the nucleotide in the synthesised mRNA which is modified, e.g. if 10 % of the uridine in the IVT reaction is modified, it would be expected that 10 % of the uridine in the synthesised mRNA would be modified.

[0204] Where the IVT reaction contains a modified nucleotide, the percentage of the nucleotide in the IVT reaction which is modified may be e.g. 5 to 20 %, 5 to 30 %, 5 to 40 %, 5 to 50 %, 5 to 60 %, 5 to 70 %, 5 to 80 %, 5 to 90 %, 5 to 100 %, 10 to 20 %, 10 to 30 %, 10 to 40 %, 10 to 50 %, 10 to 60 %, 10 to 70 %, 10 to 80 %, 10 to 90 %, 10 to 100 %, 20 to 30 %, 20 to 40 %, 20 to 50 %, 20 to 60 %, 20 to 70 %, 20 to 80 %, 20 to 90 %, 20 to 100 %, 40 to 60 %, 40 to 70 %, 40 to 80 %, 40 to 90 %, 40 to 100 %, 60 to 80 %, 60 to 90 %, 60 to 100 %, 80 to 90 % or 80 to 100 %. A modified nucleotide may comprise a modified nucleobase (modified adenine, guanine, cytosine or uracil), a modified ribose sugar or a modified phosphate group.

[0205] The IVT reaction may contain modified versions of 1 , 2, 3 or all 4 natural nucleosides, e.g. modified adenosine and modified cytidine; modified adenosine and modified uridine; modified adenosine and modified guanosine; modified cytidine and modified uridine; modified cytidine and modified guanosine; modified uridine and modified guanosine; modified adenosine, cytidine, and guanosine; modified adenosine, cytidine and uridine; modified adenosine, guanosine and uridine; or modified cytidine, guanosine and uridine.

[0206] Where the IVT reaction contains one or more modified nucleotides, the total proportion of nucleotides in the reaction which are modified may be anything, up to 100 %. For example, the IVT reaction may comprise about 10, 20, 30, 40, 50, 60, 70, 80 or 90 % modified nucleotides. For instance, the IVT reaction may comprise 1 to 10 %, 1 to 20 %, 1 to 25 %, 1 to 50 %, 1 to 60 %, 1 to 70 %, 1 to 80 %, 1 to 90 %, 1 to 95 %, 1 to 100 %, 5 to 10 %, 5 to 20 %, 5 to 25 %, 5 to 50 %, 5 to 60 %, 5 to 70 %, 5 to 80 %, 5 to 90 %, 5 to 95 %, 10 to 20 %, 10 to 25 %, 10 to 50 %, 10 to 60 %, 10 to 70 %, 10 to 80 %, 10 to 90 %, 10 to 95 %, 10 to 100 %, 20 to 25 %, 20 to 50 %, 20 to 60 %, 20 to 70 %, 20 to 80 %, 20 to 90 %, 20 to 95 %, 20 to 100 %, 50 to 60 %, 50 to 70 %, 50 to 80 %, 50 to 90 %, 50 to 95 %, 50 to 100 %, 70 to 80 %, 70 to 90 %, 70 to 95 %, 70 to 100 %, 80 to 90 %, 80 to 95 % or 80 to 100 % modified nucleotides.

[0207] In some embodiments, the IVT reaction does not contain any modified nucleotides, i.e. it contains only the native adenosine, cytidine, guanosine and uridine structures.

[0208] Where a modified nucleotide comprises a modified ribose moiety, the 2' hydroxyl group (OH) can be modified or replaced with a number of different "oxy" or "deoxy" substituents. Examples of "oxy" -2' hydroxyl group modifications include, but are not limited to, alkoxy or aryloxy (-OR, e.g. R = an alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar moiety); polyethyleneglycols (PEG); "locked" nucleic acids (LNA) in which the 2' hydroxyl is connected, e.g. by a methylene bridge, to the 4' carbon of the same ribose sugar; and amino groups (-O-amino, wherein the amino group can be e.g. an alkylamino, dialkylamino, heterocyclylamino, arylamino, diarylamino, heteroarylamino, diheteroaryl amino, ethylene diamine, polyamino or aminoalkoxy group).

[0209] "Deoxy" modifications include hydrogen, amino (e.g. NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diaryl amino, heteroaryl amino, diheteroaryl amino, or amino acid); or the amino group can be attached to the sugar through a linker, wherein the linker comprises one or more of the atoms C, N , and O.

[0210] The ribose moiety may alternatively or further comprise any additional suitable or desired modification at the 2’ OH group or at any other position on the sugar.

[0211] Where a modified nucleotide comprises a modified nucleobase, the nucleobase may be modified at any location by any suitable or desired modification, e.g. by addition of an amino group, a thiol group, an alkyl group, or a halo group.

[0212] Modified nucleotides which may be used herein include: 2-methylthio-N6-(cis- hydroxyisopentenyl)adenosine; 2-methylthio-N6-methyladenosine; 2-methylthio-N6-threonyl carbamoyladenosine; N6 glycinylcarbamoyladenosine; N6-isopentenyladenosine; N6- methyladenosine; N6 threonylcarbamoyladenosine; l,2'-0-dimethyladenosine; 1- methyladenosine; 2' O methyladenosine; 2'-0-ribosyladenosine; 2-methyladenosine; 2- methylthio-N6-isopentenyladenosine; 2-methylthio-N6-hydroxynorvalyl carbamoyladenosine;

[0213] N6-(cis-hydroxyisopentenyl)adenosine; N6,2'-0-dimethyladenosine; N6,N6,2'-O- trimethyladenosine; N6,N6-dimethyladenosine; N6-acetyladenosine; N6- hydroxynorvalylcarbamoyladenosine; N6-methyl-N6-threonylcarbamoyladenosine; 2- methylthio-N6-isopentenyladenosine; 7 deaza-adenosine; N1 -methyladenosine; N6-cis- hydroxy-isopentenyladenosine; a-thio-adenosine; 2-aminoadenosine; 2- (aminopropyl)adenosine; 2-propyladenosine; 2'-amino-2'-deoxyadenosine; 2'-azido-2'- deoxyadenosine; 8 aminoadenosine; 8-hydroxyadenosine; 8 thioadenosine; 8- azidoadenosine; 8-azaadenosine; 7-deaza-8-aza-adenosine; 7 methyladenosine; 1- deazaadenosine; 2'-fluoro-N6-benzoyl-deoxyadenosine; 2'-0-methyl-2-amino-adenosine; 2'- 0-methyl-N6-benzoyl-deoxyadenosine; 2-ethynyladenosine; 2' O trifluoromethyladenosine; 2-azidoadenosine; 2'-ethynyladenosine; 2-bromoadenosine; 2 trifluoromethyladenosine; 2- chloroadenosine; 2'-deoxy-2,2'-difluoroadenosine; 2'-deoxy-2'-mercaptoadenosine; 2'-deoxy- 2'-aminoadenosine; 2'-deoxy-2'-azidoadenosine; 2'-deoxy-2'-bromoadenosine; 2'-deoxy-2'- chloroadenosine; 2'-deoxy-2'-fluoroadenosine; 2'-deoxy-2'-iodoadenosine; 2- fluoroadenosine; 2-iodoadenosine; 2-mercaptoadenosine; 2 methoxyadenosine; 2- methylthioadenosine; 3-deaza-3-bromoadenosine; 3-deaza-3-chloroadenosine; 3-deaza-3- fluoroadenosine; 3-deaza-3-iodoadenosine; 3-deazaadenosine; 4' azidoadenosine; 8- bromoadenosine; 8-trifluoromethyladenosine; 9 deazaadenosine; 2 thiocytidine; 3- methylcytidine; 5 hydroxymethylcytidine; 5-methylcytidine; N4 acetylcytidine; 2'-O- methylcytidine; 5-formyl-2'-0-methylcytidine; lysidine; N4,2'-O-dimethylcytidine; N4-acetyl-2'- O-methylcytidine; N4 methylcytidine; N4,N4-dimethyl-2'-O-methylcytidine; 4-methylcytidine; 5-azacytidine; pseudoisocytidine; a-thio-cytidine; 2' amino-2'-deoxycytidine; 2'-azido-2'- deoxycytidine; 3 deaza-5-azacytidine; 5 propynylcytidine; 5-trifluoromethylcytidine; 5- bromocytidine; 5 iodocytidine; 6 azacytidine; pseudoisocytidine; 2' O-methyl-5-methyl- cytidine; 2-thio-5-methyl-cytidine; 5-methylzebularine; zebularine; (E)-5-(2- bromovinyl)cytidine; N4-benzoyl-2'-fluoro-2'-deoxycytidine; N4-acetyl-2'-fluoro-2'- deoxycytidine; 2'-O-methyl-N4-acetylcytidine; N4 benzoyl-2' O methylcytidine; 2'- ethynylcytidine; 5-trifluoromethyl-2'-deoxycytidine; 2' deoxy-2',2'-difluorocytidine; 2'-deoxy-2'- mercaptocytidine; 5-bromo-2'-deoxycytidine; 2' chloro-2' deoxycytidine; 2'-deoxy-2'- fluorocytidine; 5-iodo-2'-deoxycytidine; 5 (I propynyl)-2'-0-methylcytidine; 3'-C- ethynylcytidine; 4'-azidocytidine; 5 aminoallylcytidine; cyanocytidine; 5-ethynylcytidine; 5 methoxycytidine; N4 aminocytidine; N4-benzoylcytidine; 7-methylguanosine; N2,2'-O- dimethylguanosine; N2-methylguanosine; wyosine; l,2'-0-dimethylguanosine; 1- methylguanosine; 2' O methylguanosine; 7-aminomethyl-7-deazaguanosine; 7-cyano-7- deazaguanosine; archaeosine; methylwyosine; N2,7-dimethylguanosine; N2,N2,2'-O- trimethylguanosine; N2,N2,7-trimethylguanosine; N2,N2-dimethylguanosine; N2,7,2'-O- trimethylguanosine; 6 thioguanosine; 7-deazaguanosine; 8-oxoguanosine; a-thio-guanosine;

[0214] 2'-amino-2'-deoxyguanosine; 2'-azido-2'-deoxyguanosine; 6-O-methylguanosine; 8- aminoguanosine; 8 hydroxyguanosine; 8-thioguanosine; 8-azaguanosine; 7-methyl-6- thioguanosine; 6-thio-7-methylguanosine; 7-deaza-8-azaguanosine; 7-methyl-8- oxoguanosine; N2-isobutyryl-2'-0-methylguanosine; 2'-deoxy-2',2'-difluoroguanosine; 2'- deoxy-2'-chloroguanosine; 2'-deoxy-2'-fluoroguanosine; 8-bromoguanosine; 9- deazaguanosine; 1 -methylinosine; inosine; 1, 2' O dimethylinosine; 2'-0-methylinosine; 7- methylinosine; epoxyqueuosine; galactosyl-queuosine; mannosyl-queuosine; queuosine; 2'- O-methyluridine; 2-thiouridine; 3 methyluridine; 5-carboxymethyluridine; 5-hydroxyuridine; 5- methyluridine; 5 taurinomethyl-2-thiouridine; 5-taurinomethyluridine; dihydrouridine; pseudouridine; 3 (3 amino-3-carboxypropyl)uridine; l-methyl-3-(3-amino-3- carboxypropyl)pseudouridine; 1 -methylpseudouridine; 2'-0-methylpseudouridine; 2-thio-2'- O-methyluridine; N3 methyl-2' O methyluridine; 3-methylpseudouridine; 4-thiouridine; 5 carboxyhydroxymethyluridine; 5-methyl-2'-O-methyluridine; 5,6-dihydrouridine; 5 aminomethyl-2-thiouridine; 5 carbamoylmethyl-2'-0-methyluridine; 5 carbamoylmethyluridine; 5 carboxymethylaminomethyl-2'-0-methyluridine; 5 carboxymethylaminomethyl-2-thiouridine; 5-carboxymethylaminomethyluridine; 5- methoxycarbonylmethyl-2'-0-methyluridine; 5-methoxycarbonylmethyl-2-thiouridine; 5- methoxycarbonylmethyluridine; 5-methoxyuridine; 5-methyl-2-thiouridine; 5- methylaminomethyl-2-selenouridine; 5 methylaminomethyl-2-thiouridine; 5- methylaminomethyluridine; 5-methyldihydrouridine; uridine-5-oxyacetic acid; methyluridine- 5-oxyacetic acid; 5 (isopentenylaminomethyl)uridine; 5-propynyluridine; a-thio-uridine; 2'- deoxyuridine; 2’ deoxy-2' fluorouridine; 2'-amino-2'-deoxyuridine; 2'-azido-2'-deoxyuridine; 4 thiopseudouridine; 5-(aminopropyl)uridine; 5-methyl-4-thiouridine; 5 (trifluoromethyl)uridine;

[0215] 5-(3-aminopropyl)uridine; 5-aminoallyluridine; 5-bromouridine; 5-iodouridine; 5-chlorouridine; 5-fluorouridine; 6-azauridine; 3-deazauridine; 2-thio-6-azauridine and 2-thiopseudouridine.

[0216] Where the IVT reaction comprises a modified version of a particular nucleotide, it may contain one or more, e.g. 2, 3 or 4 different modified versions of that nucleotide. In particular embodiments, modified uridine is used in the in vitro transcription reaction, such that the resulting mRNA comprises modified uridine. In a preferred embodiment, 5-methoxyuridine (5moU) is used. The structure of 5moU is set forth in Formula I: Up to 20, 40, 60, 80 or 100 % 5moU may be used in the IVT reaction. That is to say, up to 20, 40, 60, 80 or 100 % of the total uridine used in the IVT reaction may be 5moU. For example, 20-40, 20-60, 20-80, 20-100, 40-60, 60-80, 40-100, 60-80, 60-100 or 80-100 % 5moU may be used in the IVT reaction.

[0217] As is well known in the art, naturally-occurring mRNA comprises a 5’ cap structure. Accordingly, the IVT step herein generally includes capping of the produced mRNA molecule. The capping may be performed co-transcriptionally (at the same time as the mRNA is synthesised) or post-transcriptionally (i.e. after the mRNA has been synthesised. Co-transcriptional capping is thus performed within the IVT reaction, whereas post- transcription capping is commonly performed after and separately to the IVT reaction. As shown in the Examples, co-transcriptional capping appears more effective (in terms of maximising protein expression from the synthesised mRNA) and so is generally preferred.

[0218] In particular embodiments, the cap has a Cap1 structure. The Cap1 structure comprises an N7-methylguanosine connected to the 5' nucleotide of the mRNA molecule through a 5' to 5' triphosphate linkage and a methyl group at the 2’0 position on the first residue of the mRNA. The Cap1 structure may be referred to as an m7GpppNm cap (where Nm refers to any 2’-O-methylated nucleotide). This structure may alternatively be presented as m7Gppp(2’Om)N. A Cap1 structure is shown in Formula II below (showing guanosine as an exemplary first residue):

[0219] The Cap1 structure may in particular have the structure m7GpppAm (i.e. wherein the first nucleotide in the mRNA molecule is adenosine).

[0220] In other embodiments, the mRNA molecule provided herein comprises a 5’ cap having a modified Cap1 structure. A modified Cap1 structure is defined herein as a Cap1 structure (i.e. comprising the m7GpppNm structure) having an additional structural modification to the methylguanosine cap itself, the 5’ nucleotide of the mRNA molecule or the triphosphate linker between them. For example, the 5’ cap may have a Cap1 structure in which the N7-methylguanosine cap comprises an additional methyl group at the 3’0 position. Such a modified Cap1 may be referred to as an m7(3’Om)GpppNm cap. In particular embodiments, the mRNA molecule comprises a modified cap with the structure m7(3’Om)GpppAm (i.e. wherein the first nucleotide residue of the mRNA molecule is adenosine). Another example of a modified Cap1 structure is a cap having the structure m7(3’Om)Gpppm6(2’Om)A. Such a cap is based on the m7(3’Om)GpppAm cap described above, further comprising an additional methyl group at position 6 of the adenosine residue at the 5’ terminus of the mRNA molecule.

[0221] Such caps can be applied to mRNA molecules by any suitable process, e.g. chemically or enzymatically. For example, a Cap1 structure can be applied to mRNA co- transcriptionally using the CleanCap® Reagent AG (TriLink, CA, USA). Modified Cap1 structures as described above can be applied to mRNA using the CleanCap® Reagent AG (3’0Me) and the CleanCap® Reagent M6, respectively.

[0222] Thus in particular embodiments, the IVT reaction comprises co-transcriptional capping of the mRNA with a cap having a Cap1 structure. This may be achieved by adding the reagent m7G(5')ppp(5')(2'OMeA)pG (CleanCap® Reagent AG) to the IVT reaction. The reagent m7G(5')ppp(5')(2'OMeA)pG has the structure set forth in Formula II:

[0223] Formula II:

[0224] Such a reagent may be used according to the manufacturer’s instructions, e.g. it may be added in equimolar amount to the NTPs.

[0225] At the end of the IVT reaction, the DNA template may be degraded, e.g. by addition of a DNase. The IVT reaction may be performed in a small reaction volume, e.g. at most 10, 5 or 2 pl. The minimum volume is determined only by the limits of pipetting technology, e.g. the minimum reaction volume may be about 0.5 or 1 pl. For example, the reaction may be performed in a volume of 1-5, 1-2 or 2-5 pl, e.g. about 1 , 2, 3, 4 or 5 pl. Use of small reaction volumes is particularly compatible with automation of the process.

[0226] Transfection

[0227] The mRNA produced by the IVT reaction is then used in step (v) to transfect eukaryotic cells. Prior to transfection the mRNA may be purified, or cleaned up, e.g. using a commercial RNA purification kit such as the Monarch® RNA Cleanup Kit (NEB). RNA purification removes debris (e.g. degraded DNA) and other impurities such as salts and leftover NTPs. However, RNA purification results in a loss of yield, and as shown in the Examples below the present inventors have found that a purification step may not be necessary. Accordingly, performance of an RNA purification step prior to use of the RNA for transfection is optional. While performing a purification step results in loss of RNA yield, it enables quantification of the amount / concentration of RNA produced by IVT.

[0228] The mRNA molecules encoding the bioPROTACs are transfected into the eukaryotic cells in which the bioPROTACs are to be screened. Any suitable eukaryotic cell may be used, as discussed above. Considerations for selection of the eukaryotic cell type for use in screening include whether it expresses the target protein, whether it expresses E3 ligases and / or E3 ligase complexes which can be recruited by degradation domains used in the bioPROTACs, ease of transfection, relevance to therapeutic indication for which the bioPROTACs are of interest, etc. As set out above, in particular embodiments the eukaryotic cells used for screening are mammalian cells, e.g. human cells. The cells may be from a cell line, e.g. a mammalian cell line such as a human cell line. Suitable cell lines for screening are described above.

[0229] The eukaryotic cells may be transfected using any technique known in the art, e.g. DEAE-dextran transfection, polybrene-mediated transfection, electroporation, microinjection, biolistic transfection, optical transfection or lipofection. Lipofection is a preferred technique, in particular using lipofectamine (a mixture of 2'-(1",2"-dioleoyloxypropyldimethyl-ammonium bromide)-N-ethyl-6-amidospermine tetratrifluoroacetic acid salt (DOSPA) and 1 ,2-dioleoyl- sn-glycero-3-phosphatidylethanolamine (DOPE)).

[0230] In some embodiments, the concentration of the mRNA produced by IVT is not determined, at least for most IVT reactions performed in the method. In this case, concentrations may be estimated based on the concentrations of a small, dedicated subset of IVT reactions. To avoid small deviations in mRNA concentration between samples affecting the level of protein expression in the transfected cells, transfection may be performed with a saturating amount of mRNA (as estimated as set out above). By a saturating amount of mRNA is meant an amount of mRNA which yields the maximum possible amount of expression of the encoded protein. That is to say, sufficient mRNA is applied to the cells that even if the amount of mRNA is increased, the level of protein expression does not increase any further. For example, using GFP expression as a proxy, the inventors have found that in a well of a 384-well plate seeded with 8000 HOT 116 cells, protein expression is not increased when the amount of mRNA is increased beyond 50 ng. If this set up is selected, the cells may be transfected with at least 50 ng mRNA.

[0231] Saturating amounts of mRNA for other experimental set-ups can be easily determined by the skilled person, by testing expression of a fluorescent protein such as GFP in cells transfected with a series of increasing amounts of mRNA. Once a maximal level of fluorescence is reached, this indicates that the amount of mRNA added is now saturating.

[0232] Each separate mRNA molecule produced by the IVT step is separately transfected into the eukaryotic cells, such that each group of transfected cells expresses a single bioPROTAC, and each bioPROTAC in the library is expressed by a single group of transfected cells.

[0233] Screening

[0234] The final step of the method, step (vi) is screening of the transfected cells for depletion of the target protein. Depletion of the target protein in a particular transfected cell indicates that the bioPROTAC expressed in that cell is capable of inducing degradation of the target protein. Screening may be performed as set out above, using the assays mentioned above.

[0235] Identifying bioPROTACs for Treatment

[0236] In one aspect, provided herein is a method of treating a disease in a subject, comprising:

[0237] (a) identifying a target protein associated with the disease;

[0238] (b) identifying a bioPROTAC capable of inducing degradation of the target protein according to the method described above; and

[0239] (c) administering the bioPROTAC or a nucleic acid molecule encoding the bioPROTAC to the subject.

[0240] A target protein associated with a disease is a protein which is known or believed to contribute to the pathogenesis or severity of a disease, such that degradation of the protein is expected to result in the disease being cured or its symptoms being attenuated. Examples of such proteins include cancer-associated proteins, as described above. Such a protein may be identified in a given subject by diagnosing that subject with a particular disease. Several diseases or conditions are known to be associated with particular proteins. Alternatively, in some cases (particularly cancer), proteins associated with a disease suffered by a particular subject may require analysis of the disease in that subject. For example a biopsy may be taken from a tumour and cancer cells analysed to identify proteins associated with the tumour’s malignant state. Such analysis may be performed by e.g. histochemistry or gene expression profiling (e.g. by RNA-Seq).

[0241] A bioPROTAC capable of inducing degradation of the target protein may be identified according to the method described above. In some cases, a method may be performed to identify a bioPROTAC against a particular target protein once that target protein has been identified in the subject. In other cases, particularly for target proteins known to be associated with diseases, bioPROTACs which recognise the target protein may be identified without reference to a specific subject. Accordingly, steps (a) and (b) of this method can be performed in either order.

[0242] Once a suitable bioPROTAC with which to treat the subject has been identified, the bioPROTAC or a nucleic acid molecule encoding the bioPROTAC is administered to the subject. Administration of bioPROTACs to subjects is described below.

[0243] BioPROTACs which Degrade c-Myc

[0244] Provided herein are bioPROTACs which are capable of inducing degradation of human c-Myc (also referred to as MYC), a known oncoprotein with the UniProt accession number P01106. The bioPROTACs may also be capable of inducing degradation of MYC from other species, particularly closely related species to humans.

[0245] The bioPROTACs are as set out above, i.e. they comprise a binding domain, a degradation domain and optionally a linker. Since MYC is a nuclear protein, the bioPROTAC also generally comprises an NLS.

[0246] The binding domain comprises or consists of an amino acid sequence as set forth in any one of SEQ ID NOs: 101 to 110, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to any one of SEQ ID NOs: 101 to 110. SEQ ID NOs: 101 to 110 are proteins which specifically bind human MYC. Where the binding domain is a variant of one of SEQ ID NOs: 101 to 110, the variant retains its ability to specifically bind human MYC.

[0247] The degradation domain is capable of ubiquitylating human MYC, and may comprise or consist of an amino acid sequence as set forth in any one of SEQ ID NOs: 1-100, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to any one of SEQ ID NOs: 1-100. Variants of these degradation domains are discussed above.

[0248] Where the bioPROTAC comprises a linker, the linker may be as described above. For example the linker may comprise or consist of an amino acid sequence as set forth in any one of SEQ ID NOs: 113 to 133, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to any one of SEQ ID NOs: 113 to 133. In particular embodiments the linker may comprise or consist of an amino acid sequence as set out in any one of SEQ ID NOs: 120, 126, 132 or 133.

[0249] Determination of whether a particular bioPROTAC is capable of inducing degradation of MYC may be performed as demonstrated in the Examples below.

[0250] In a particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0251] (i) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 108, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 108, wherein the amino acid at the position corresponding to position 78 of SEQ ID NO: 108 is not lysine, preferably wherein it is arginine;

[0252] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 120; and

[0253] (iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 40, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 40, wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is not threonine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is not lysine, preferably wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is alanine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is arginine.

[0254] Such a bioPROTAC may in particular comprise or consist of the amino acid sequence of SEQ ID NO: 138, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 138. Alternatively, such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 146 or 147, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 146 or 147.

[0255] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0256] (i) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 105, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 105, wherein the amino acids at the positions corresponding to positions 5, 6, 57, 90, 123, 133 and 156 of SEQ ID NO: 105 are not lysine, preferably wherein they are arginine;

[0257] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 120; and

[0258] (iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 25, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 25.

[0259] Such a bioPROTAC may in particular comprise or consist of the amino acid sequence of SEQ ID NO: 139, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 139. Alternatively, such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 153, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 153.

[0260] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0261] (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 10, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 10;

[0262] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 126, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 126; and

[0263] (iii) a binding domain comprising the amino acid sequence set forth in SEQ ID

[0264] NO: 101 , or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 101.

[0265] Such a bioPROTAC may in particular comprise or consist of the amino acid sequence of SEQ ID NO: 140, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 140.

[0266] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0267] (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 10, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 10;

[0268] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 120; and

[0269] (iii) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 101 or SEQ ID NO: 108, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 101 or SEQ ID NO: 108, wherein in the variant of SEQ ID NO: 108 the amino acid at the position corresponding to position 78 of SEQ ID NO: 108 is not lysine, preferably wherein it is arginine.

[0270] Such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 141 , SEQ ID NO: 142 or SEQ ID NO: 143, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 141 , SEQ ID NO: 142 or SEQ ID NO: 143.

[0271] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0272] (i) a binding domain comprising the amino acid sequence set forth in SEQ ID

[0273] NO: 106, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 106, wherein the amino acid at the position corresponding to position 78 of SEQ ID NO: 106 is not lysine, preferably wherein it is arginine;

[0274] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 126, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 126; and

[0275] (iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 40, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 40, wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is not threonine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is not lysine, preferably wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is alanine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is arginine.

[0276] Such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 144 or SEQ ID NO: 145, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 144 or SEQ ID NO: 145.

[0277] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0278] (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 7, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 7;

[0279] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 133, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 133; and

[0280] (iii) a binding domain comprising the amino acid sequence set forth in SEQ ID

[0281] NO: 109, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 109, wherein the amino acids at the positions corresponding to positions 45, 67 and 78 of SEQ ID NO: 109 are not lysine, preferably wherein they are arginine.

[0282] Such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 148, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 148.

[0283] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0284] (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 7, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 7;

[0285] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 126 or SEQ ID NO: 133, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 126 or SEQ ID NO: 133; and

[0286] (iii) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 101 , or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 101.

[0287] Such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 149 or SEQ ID NO: 150, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 149 or SEQ ID NO: 150

[0288] In another particular embodiment, the bioPROTAC capable of inducing degradation of MYC comprises, from N-terminus to C-terminus:

[0289] (i) a binding domain comprising the amino acid sequence set forth in SEQ ID

[0290] NO: 101 , or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 101 ;

[0291] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120 or SEQ ID NO: 126, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 120 or SEQ ID NO: 126; and

[0292] (iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 7, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 7.

[0293] Such a bioPROTAC may comprise or consist of the amino acid sequence of SEQ ID NO: 151 or SEQ ID NO: 152, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 151 or SEQ ID NO: 152.

[0294] As set out above, corresponding amino acid positions can be identified by sequence alignment, i.e. the position in a protein of interest which corresponds to a specified position in a reference protein is the position which aligns to the specified position when the sequences of the protein of interest and the reference protein are aligned. For example, the position in a variant of SEQ ID NO: 105 corresponding to position 5 of SEQ ID NO: 105 is the position which corresponds to (or aligns to) position 5 of SEQ ID NO: 105 when the variant sequence is aligned to SEQ ID NO: 105.

[0295] BioPROTACs which Degrade K-Ras

[0296] Provided herein are bioPROTACs which are capable of inducing degradation of human K-Ras (another oncoprotein, with the UniProt accession no. P01116). The bioPROTACs may also be capable of inducing degradation of K-Ras from other species, particularly closely related species to humans.

[0297] The bioPROTACs are as set out above, i.e. they comprise a binding domain, a degradation domain and a linker. The bioPROTAC comprises, from N-terminus to C-terminus:

[0298] (i) a binding domain comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 134, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 134. SEQ ID NO: 134 is a DARPin which specifically binds human K-Ras. Where the binding domain is a variant of SEQ ID NO: 134, the variant retains its ability to specifically bind human K-Ras;

[0299] (ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 120; and

[0300] (iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 82, or a variant thereof having at least 70, 75, 80, 85, 90, 95 or 99 % sequence identity to SEQ ID NO: 82. Variants of degradations domains, including that of SEQ ID NO: 82, are discussed above.

[0301] Such a bioPROTAC may in particular comprise or consist of the amino acid sequence of SEQ ID NO: 154, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99 % identity to SEQ ID NO: 154.

[0302] Nucleic Acid Molecules and Vectors

[0303] Also provided herein is a nucleic acid molecule encoding a bioPROTAC capable of inducing degradation of human MYC or human K-Ras, as described above. The nucleic acid molecule provided herein may be an isolated nucleic acid molecule. The nucleic acid molecule provided herein may be a DNA or RNA molecule, particularly an mRNA molecule. The nucleic acid molecule provided herein may be single-stranded or double-stranded. The nucleic acid molecule provided herein may be a linear nucleic acid molecule or a circular nucleic acid molecule. The nucleic acid molecule may be provided in the context of an expression vector, such as an expression vector as described below. Alternatively, the nucleic acid molecule may be provided in the context of a linear construct for cloning into a vector, or in the context of a cloning vector.

[0304] The nucleic acid molecule may be a chemical derivative of DNA or RNA, such as a molecule having a radioactive isotope or a chemical adduct such as a fluorophore, chromophore or biotin (“label”). Thus the nucleic acid may comprise modified nucleotides, as described above. The nucleic acid molecule may have a modified backbone, e.g. a phosphorothioate, phosphoroselenate, borano phosphates, hydrogen phosphonates or phosphoroamidate backbone.

[0305] In some embodiments, the nucleic acid molecule is a vector or is comprised within a vector. The vector may be any type of vector, as described above. The vector may be a cloning vector or an expression vector. The vector may be a bacterial or prokaryotic vector, or it may be a eukaryotic vector, particularly a mammalian vector. For example, the vector may be a bacterial cloning vector, e.g. an E. coli cloning vector, such as described above.

[0306] Alternatively, the vector may be an expression vector, which comprises an expression cassette as described above. In particular, the expression vector may be a mammalian expression vector.

[0307] In particular embodiments, the expression vector is a viral vector, that it is to say a non-pathogenic virus which can be used for delivery of genetic material to a target cell. The viral vector may be suitable for use in gene therapy, in which case it is suitable for administration to a subject (particularly a human subject) and in vivo delivery of its genome to target cells. In any event, the viral vector is generally replication deficient.

[0308] Suitable viral vectors for use herein include adeno-associated virus (AAV) vectors, adenovirus vectors, herpes simplex virus vectors, retrovirus vectors, lentivirus vectors, alphavirus vectors, flavivirus vectors, rhabdovirus vectors, measles virus vectors, Newcastle disease virus vectors, poxvirus vectors and picornavirus vectors.

[0309] In preferred embodiments, the expression vector is an adeno-associated virus (AAV) vector. AAVs are one of the most actively investigated gene therapy vehicles and are characterized by excellent safety profile and high efficiency of transduction in a broad range of target tissues. The use of AAVs as a vector for gene therapy is described in, for example, Naso et al., BioDrugs 31 (4): 317-334, 2017 and Colella et al., Molecular Therapy - Methods and Clinical Development 8: 87-104, 2018. Various AAV serotypes, including AAV1 , AAV3, AAV4, AAV5, AAV6, AAV6.2, AAV6.2FF, AAV8, AAV 8.2, AAV9, and AAV rh10 and pseudotyped AAV such as AAV2 / 8, AAV2 / 5 and AAV2 / 6 can be used. Further examples of serotypes and their isolation are described in Srivastava, Current Opinion in Virology 21 : 75- 80, 2016. The AAV particle is a small (25 nm) virus from the Parvoviridae family, and it is composed of a non-enveloped icosahedral capsid (protein shell) that contains a linear single-stranded DNA genome of around 4.8 kb. The AAV genome encodes several protein products, namely, four non-structural Rep proteins, three capsid proteins (VP1-3), and the assembly-activating protein (AAP). The AAV genes are flanked by two AAV-specific palindromic inverted terminal repeats (ITRs).

[0310] The AAV vector may be engineered, for example in order to improve its function. Examples of AAVs that have been engineered for clinical gene therapy are described in Kotterman and Schaffer, Nature Reviews Genetics 15(7): 445-51 , 2014.

[0311] Therapeutic Methods

[0312] The bioPROTACs described above capable of inducing degradation of MYC or K-Ras are also provided for use in therapy. The nucleic acids and vectors described above which encode these bioPROTACs are also provided for use in therapy. The term “therapy” as used herein encompasses both treatment (including palliative treatment) and prophylaxis of a disease or disorder.

[0313] The term “treatment,” as used herein in the context of treating a disease or disorder, pertains generally to treatment and therapy of a human, or alternatively of a non-human animal, in which some desired therapeutic effect is achieved, for example the inhibition of the progress of the disease or disorder, and includes a reduction in the rate of progress, a halt in the rate of progress, regression of the disease or disorder, amelioration of the disease or disorder, and cure of the disease or disorder.

[0314] “Prophylaxis” in the context of the present specification should not be understood to circumscribe complete success i.e. complete protection or complete prevention. Rather prophylaxis in the present context refers to a measure which is administered in advance of detection of a symptomatic condition with the aim of preserving health by helping to delay, mitigate or avoid that particular condition.

[0315] The individual treated according to the therapeutic methods described herein may be referred to as a subject or (normally when referring to a human subject) a patient.

[0316] The therapy may comprise administering a bioPROTAC as specified to a subject, i.e. administering a bioPROTAC protein. As described above, bioPROTACs operate intracellularly, so in this instance the bioPROTAC may comprise a feature or be administered in a manner which enables cellular entry. For example, the bioPROTAC may comprise a trafficking signal which directs it into a target cell. Alternatively, a bioPROTAC could be delivered to a target cell in a delivery vehicle, e.g. encapsulated in a lipid nanoparticle (LNP). In other embodiments, the therapy described herein utilises an in vivo expressed biologic (IVEB). An IVEB is a biologic agent (i.e. a peptide or protein, such as a bioPROTAC) which is produced within the subject being treated. In other words, rather than being delivered to the subject in protein form, a nucleic acid molecule encoding the IVEB is administered to the subject (optionally in the context of a vector), who expresses the IVEB from the nucleic acid molecule.

[0317] Where the bioPROTAC is expressed as an IVEB, the administered nucleic acid molecule may be a DNA or RNA molecule encoding the bioPROTAC. For example an mRNA encoding the bioPROTAC may be used. Delivery of mRNA molecules can be effected using any suitable delivery vehicle, e.g. an LNR Alternatively the administered nucleic acid molecule may be a DNA vector, such as a plasmid, which may be delivered in a similar manner to an mRNA, e.g. using an LNP

[0318] Another suitable means of nucleic acid molecule delivery is a viral vector. Suitable viral vectors are described above and in Bulcha et al., Signal Transduction and Targeted Therapy 6: 53, 2021.

[0319] When used in therapy, the bioPROTAC or nucleic acid molecule or vector encoding the bioPROTAC is administered in the context of a pharmaceutical composition comprising (i) the bioPROTAC or nucleic acid molecule or vector encoding the bioPROTAC, and (ii) a pharmaceutically acceptable carrier or diluent. The term “pharmaceutically acceptable,” as used herein, pertains to compounds, ingredients, materials, compositions, etc., which are, within the scope of sound medical judgment, suitable for use in contact with the tissues of the subject in question (e.g., human) without excessive toxicity, irritation, allergic response, or other problem or complication, commensurate with a reasonable benefit / risk ratio. Each carrier, diluent, excipient, etc. must also be “acceptable” in the sense of being compatible with the other ingredients of the formulation.

[0320] In particular embodiments, the therapy may be for cancer. The bioPROTACs described herein, and the nucleic acid molecules and vectors encoding them, may be used to treat any form of cancer. That is to say, provided herein is a bioPROTAC as described above, or a nucleic acid molecule or vector encoding such a bioPROTAC, for treatment of cancer. This aspect may be seen as providing a method of treating cancer in a subject in need thereof, comprising administering to the subject a bioPROTAC as defined herein, or a nucleic acid molecule or vector encoding such a bioPROTAC. This aspect may alternatively be seen as providing the use of bioPROTAC as defined herein, or a nucleic acid molecule or vector encoding such a bioPROTAC, in the manufacture of a medicament for treating cancer.

[0321] In particular, the bioPROTACs which target MYC may be used to treat cancers which express MYC, particularly cancers which overexpress or constitutively express MYC. MYC is known to be aberrantly expressed in over 70 % of human cancers and so targeting of MYC has broad relevance in cancer therapy. Cancers which express or overexpress MYC can be identified by any suitable means, e.g. histochemistry or RT-PCR.

[0322] Cancers which may be treated using the bioPROTACs which target MYC include lung cancer, pancreatic cancer, oesophageal cancer, breast cancer, liver cancer, colorectal cancer, prostate cancer and gastric cancer. BioPROTACs which target MYC may also be used to treat blood cancers, e.g. acute myeloid leukaemia, acute lymphoblastic leukaemia and non-Hodgkin lymphomas, e.g. Burkitt lymphoma.

[0323] The bioPROTACs which target K-Ras may in particular be used to treat cancers which express K-Ras, particularly cancers which overexpress K-Ras. Cancers which comprise mutations in K-Ras can also be treated with the bioPROTACs provided herein, in particular cancers which express K-Ras comprising one or more driver mutations, particularly mutations which cause K-Ras to become constitutively active or which prevent its deactivation.

[0324] Cancers which may be treated using the bioPROTACs which target K-Ras include lung cancer, pancreatic cancer, oesophageal cancer, breast cancer, liver cancer, colorectal cancer, prostate cancer and gastric cancer.

[0325] The cancer treated using the bioPROTACs herein may be at any stage, i.e. stage I, stage II, stage III or stage IV. Thus the cancer may be metastatic cancer.

[0326] In other embodiments, the therapy may be for an inflammatory disease. The bioPROTACs described herein, and the nucleic acid molecules and vectors encoding them, may be used to treat any inflammatory disease. That is to say, provided herein is a bioPROTAC as described above, or a nucleic acid molecule or vector encoding such a bioPROTAC, for treatment of an inflammatory disease. This aspect may be seen as providing a method of treating an inflammatory disease in a subject in need thereof, comprising administering to the subject a bioPROTAC as defined herein, or a nucleic acid molecule or vector encoding such a bioPROTAC. This aspect may alternatively be seen as providing the use of bioPROTAC as defined herein, or a nucleic acid molecule or vector encoding such a bioPROTAC, in the manufacture of a medicament for treating an inflammatory disease. MYC and K-Ras are known to be associated with inflammation (Kortlever et al., Cell 171(6): 1301-1315, 2017) and thus targeting of these proteins may successfully ameliorate an inflammatory disease.

[0327] Inflammatory diseases which may be treated with the bioPROTACs provided herein include autoimmune diseases, such as rheumatoid arthritis, multiple sclerosis, systemic lupus erythematosus, Addison’s disease, Grave’s disease, scleroderma, polymyositosis, diabetes, autoimmune uveoretinitis, ulcerative colitis, pemphigus vulgaris, inflammatory bowel disease, autoimmune thyroiditis, uveitis, Behget’s disease, Sjogren’s syndrome and psoriasis.

[0328] Methods of Degrading Proteins

[0329] Also provided herein is a method of degrading a protein in a eukaryotic cell, comprising contacting the cell with a nucleic acid molecule or vector encoding a bioPROTAC which is capable of inducing degradation of the protein.

[0330] In some embodiments, the method is for degrading MYC and the bioPROTAC is a bioPROTAC capable of inducing degradation of MYC, as described above.

[0331] In other embodiments, the method is for degrading K-Ras and the bioPROTAC is a bioPROTAC capable of inducing degradation of K-Ras, as described above.

[0332] The eukaryotic cell may be any type of eukaryotic cell, as described above. In particular embodiments, the eukaryotic cell is a human cell.

[0333] The cell may be in vivo, i.e. the method may include administering the nucleic acid molecule or vector to a subject, as discussed above.

[0334] Alternatively, the cell may be in vitro or ex vivo, e.g. the cell may be from a cell line or may be a primary cell isolated from a subject. In this instance, the cell may be transfected or transduced with the nucleic acid molecule or vector, as described above.

[0335] Seguence Identity and Alterations

[0336] Sequence identity is commonly defined with reference to the algorithm GAP (Wisconsin GCG package, Accelerys Inc, San Diego USA). GAP uses the Needleman and Wunsch algorithm to align two complete sequences, maximising the number of matches and minimising the number of gaps. Generally, default parameters are used, with a gap creation penalty equalling 12 and a gap extension penalty equalling 4. Use of GAP may be preferred but other algorithms may be used, e.g. BLAST (which uses the method of Altschul et al. (1990)), FASTA (which uses the method of Pearson and Lipman (1988)), or the Smith- Waterman algorithm (Smith and Waterman (1981)), or the TBLASTN program, of Altschul et al. (1990) supra, generally employing default parameters. In particular, the psi-Blast algorithm may be used.

[0337] Where the disclosure makes reference to a particular amino acid sequence having at least 90 % sequence identity to a reference amino acid sequence, this includes the amino acid sequence having 90 %, 91 %, 92 %, 93 %, 94 %, 95 %, 96 %, 97 %, 98 %, 99 % and 100 % sequence identity to the reference amino acid sequence (including as rounded to the nearest integer percentage). The term “sequence alterations” or “sequence modifications” as used herein, and equivalent terms, are intended to encompass the substitution, deletion and / or insertion of an amino acid residue. Thus, a protein containing one or more amino acid sequence alterations compared to a reference sequence contains one or more substitutions, one or more deletions and / or one or more insertions of an amino acid residue as compared to the reference sequence. The terms “amino acid mutation” is herein used interchangeably with “sequence alteration”, unless the context clearly identifies otherwise.

[0338] In some embodiments in which one or more amino acids are substituted with another amino acid, the substitutions may be conservative substitutions, for example according to the following Table. In some embodiments, amino acids in the same block in the middle column are substituted, i.e. a non-polar amino acid is substituted for another non-polar amino acid for example. In some embodiments, amino acids in the same line in the rightmost column are substituted, i.e. G is substituted for A or P for example.

[0339] In some embodiments, substitution(s) may be functionally conservative. That is, in some embodiments the substitution may not affect (or may not substantially affect) one or more functional properties (e.g. binding affinity) of the protein comprising the substitution as compared to the equivalent unsubstituted protein.

[0340] ***

[0341] The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the present disclosure in diverse forms thereof.

[0342] While the present disclosure has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art. Accordingly, the exemplary embodiments of the present disclosure set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the present disclosure.

[0343] For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations.

[0344] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0345] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0346] It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%.

[0347] Figure Legends

[0348] Figure 1 shows the measured growth of HCT116 cells across a period of 6h to 54h after transfection with mRNA or plasmid DNA, or in untransfected cells for comparison. For mRNA transfection, two distinct transfection reagents were used: either Lipofectamine™ RNAiMAX or Lipofectamine™ 2000. Plasmid DNA was transfected with Lipofectamine™ 2000. Live cells were imaged every 2h by brightfield microscopy, and the percentage of surface occupied by cells (confluency) was quantified.

[0349] Figure 2 shows flow cytometry histograms depicting the GFP intensity 22 hr post transfection of live HCT116 cells transfected with mRNAs encoding for GFP. The different histograms are from cells transfected with identical mRNAs except in their polyA tail sizes: 20, 40, 60 or 80 adenosine nucleosides (A20, A40, A60 and A80, respectively), and are compared to the untransfected control, a.u. = arbitrary units. Figure 3 shows flow cytometry histograms depicting the GFP intensity of live HCT116 cells transfected with mRNAs encoding GFP, with various capping mechanisms. The mRNAs used for transfection of the different samples have identical sequences, except that the transcription start site may be modified based on capping requirements. (A) Comparison of enzymatic capping using the Vaccinia Capping System (VCS) and co-transcriptional capping with CleanCap AG, both generating a Cap-1 structure. (B) Comparison of co-transcriptional capping using different reagents: Anti-Reverse Cap Analog (ARCA,

[0350] 3'-O-Me-m7G(5')ppp(5')G), CleanCap AG (m7G(5')ppp(5')(2'OMeA)pG), m7G(5')ppp(5')G (NEB S1404) or m7G(5’)ppp(5’)A (NEB S1405). Note that polyA length in these examples is 60 residues.

[0351] Figure 4 shows the live-cell imaging results of HCT116 cells transfected with mRNAs encoding GFP. Between samples, only the ratio of 5-methoxyuridine-5’-triphosphate (5-moU) to UTP is changed. (A) The ratio of cell area which is positive for GFP was used as a metric for transfection efficiency (i.e. fraction of cells which have been successfully transfected, 1.0 = all cells have been transfected). (B) The average integrated intensity among GFP-positive cells (in arbitrary units) is used as a metric for GFP expression levels. NT = not transfected.

[0352] Figure 5 shows the live-cell imaging results of HCT116 cells transfected with mRNAs encoding GFP, whereby identical mRNA produced in 2 μL or 10 μL reactions is compared. (A) The ratio of cell area which is positive for GFP was used as a metric for transfection efficiency. (B) The average integrated intensity among GFP-positive cells (in arbitrary units) is used as a metric for GFP expression levels. NT = not transfected.

[0353] Figure 6 shows the live-cell imaging results of HCT116 cells transfected with purified and unpurified mRNA encoding GFP. mRNA synthesis is performed in otherwise identical conditions. NT = not transfected.

[0354] Figure 7 shows the live-cell imaging results of HCT116 cells transfected with mRNA encoding GFP. mRNA synthesis differs only in the origin of the PCR template, whether it was first purified from bacteria, or whether PCR was performed directly on a culture of bacteria containing the plasmid. (A) The ratio of cell area which is positive for GFP was used as a metric for transfection efficiency. (B) The average integrated intensity among GFP-positive cells (in arbitrary units) is used as a metric for GFP expression levels. NT = not transfected.

[0355] Figure 8 shows the results of FLAG immunoprecipitation (IP) experiments following transient transfection of plasmids encoding the indicated FLAG-Halo-tagged MYC binders or nonbinding DARPin (DPctrl). Input is 5 % of cell lysate. Vinculin is used as a control for loading and purity.

[0356] Figure 9 shows the measured bioPROTAC-HA expression in single cells by IF for Halo-L1 controls in the MYC bioPROTAC screening (example of N = 3). The entire cell population is included, and black lines indicate the median abundance, a.u. = arbitrary units.

[0357] Figure 10 shows nuclear MYC abundance (IF) and area under the curve (kinetic assay) fold changes (FC) in log2scale (left) or as z-scores (right) in the primary screening (N=1) for (DPA / H discovery). bioPROTACs were delivered as mRNA into HCT116 cells, and FCs are relative to controls where the degradation domain is replaced by a Halo tag. Symbols designate the MYC binder: DP07 (•), DP20(o), VH12(«) and VH15(n). Shaded areas depict the detection limit (IF) or most extensive MYC downregulation (kinetic assay), determined by 100 pg / mL cycloheximide treatment.

[0358] Figure 11 shows confirmatory screening results (average of 3 independent experiments) for the DPA / H discovery screening, with bioPROTACs containing DP07(®), DP20(o), VH12(«) or VH15(n).Volcano plot for 17 h end-point immunofluorescence (left) or 48 h kinetic data (middle) incorporate average Log2FC nuclear MYC or area under the curve for bioPROTAC- expressing samples vs matching control samples. Significance was evaluated by false discovery rate (FDR)-controlled t-tests. FC and q-value thresholds are represented by dotted lines, and shaded areas indicate the detection limit (IF) or most extensive MYC downregulation attainable (kinetic assay), determined by cycloheximide treatment. Venn diagrams (right) depict top-performing constructs (Log2FC < -2, q-value < 0.05 for IF readout and Log2FC < -1 , q-value < 0.05 in kinetic assay).

[0359] Figure 12 shows MYC abundance effect dependence on construct expression (normalized to mock and the average of controls, respectively), measured by IF in the confirmatory screening (average of N=3) when HA-tagged bioPROTACs (•) or controls (o) are expressed. Pearson correlation tests were performed and the resulting correlation coefficient (r) are indicated, p-value: **** < 0.0001 , ns = non-significant (> 0.05).

[0360] Figure 13 shows (A) confirmatory screening results (average of 3 independent experiments) for H1 discovery screening with bioPROTACs containing H1 (>) or the concomitant validation & optimization run with unmodified (•) or K>R mutant (o) bioPROTACs. Volcano plots for 17 h end-point immunofluorescence (left) or 48 h kinetic data (middle) incorporate average Log2FC nuclear MYC or area under the curve for bioPROTAC-expressing samples vs matching control samples. Significance was evaluated by false discovery rate (FDR)- controlled t-tests. FC and q-value thresholds are represented by dotted lines, and shaded areas indicate the detection limit (IF) or most extensive MYC downregulation attainable (kinetic assay), determined by cycloheximide treatment. Venn diagrams (right) depict top- performing constructs (Log2FC < -2, q-value < 0.05 for IF readout and Log2FC < -1 , q-value < 0.05 in kinetic assay); (B-C) fold change (FC) in l_og2scale (left) or z-score (right) correlation between MYC abundance assay outputs for the follow-up primary MYC bioPROTAC screening (N=1) in HCT116 cells transfected with mRNA. (B) shows H1 discovery primary screening data, whereby FCs are relative to controls lacking a functional degradation domain. (C) shows results from the optimization & validation screening subset (primary screening), whereby FCs are calculated vs controls unable to recruit MYC. Unmodified (•), K>R mutant (o) or bioPROTACs containing further ICPO truncations (■) were expressed. Shaded areas depict MYC downregulation by cycloheximide, taken as the background signal.

[0361] Figure 14 shows the impact of KR substitutions on bioPROTAC abundance and effect: (A) heatmap of validation & optimization confirmatory screening constructs (values are the average of N=3), depicting nuclear bioPROTAC-HA intensity normalized to mean of controls (“HA”), nuclear MYC FC in IF (“MYC FC”) and the area under the curve FC from the kinetic assay (“AUC FC”). Black dots highlight nonsignificant values (q-value > 0.05 in FDR- corrected t-tests). #K>R: number of Lys to Arg substitutions. Numbers in x / y format represent KR substitutions in the degradation domain and in the target-binding domain, respectively, and a single number means that all mutations are in the target-binding domain; (B, C) immunoblotting following HCT116 mRNA-based expression of bioPROTACs bearing the indicated number of K>R substitutions in the ICPO degradation domain (DD) or binders (VH15 or DP20). Note that for immunoblotting experiments, screening workflow simplifications were not employed, and the same mRNA amount was transfected in all samples. A representative example of three independent experiments is shown. Numbers left of HA blots indicate the molecular weight of protein markers. Vinculin is used as loading control and anti-HA detects bioPROTACs.

[0362] Figure 15 shows composition preferences for selected MYC-targeting bioPROTACs. Immunoblotting following HCT116 mRNA-based expression of bioPROTACs (HA-tagged) with varying compositions to investigate (A) DD-binder relationships (constructs have the flexible 16AA linker) and (B) ICPO positioning and the contribution of the different linkers for MYC downregulation. For these experiments, mRNA was re-synthesized without the screening workflow simplifications and the same amount transfected in all samples. A representative example of three independent experiments is shown. Linkers: L1 , flexible 5AA; L2, flexible 16AA; L3, rigid 65AA and L4, rigid+flexible 81 AA. Numbers left of HA blots indicate the molecular weight of protein markers (note that the H1 peptide is not detected due to its small size). Vinculin is used as loading control and anti-HA detects bioPROTACs.

[0363] Figure 16 shows that shortlisted bioPROTACs require concomitant MYC binding and E3 ligase function to eliminate MYC and elicit MYC-specific downstream proteome changes: (A) IF (top, scale bar: 50 pm) and kinetic (bottom) MYC abundance readouts for one of N=3 upon expression of the indicated hits from the validation and optimization screening (left and centre with nonbinding control) and H1 discovery screening (right with no E3 recruitment control). Time-course data depicts normalized relative light units (RLU) in samples (black lines) and matching controls (grey lines). Data is the mean (solid lines) ± standard deviation (SD, dotted lines) of technical replicates in a representative experiment; (B) hit validation by immunoblotting (representative of N=3) after 17 h mRNA-based expression of constructs containing (+), lacking (-) or bearing inactive (i) domains, or split by a T2Aautocleavable peptide (+W+) with a resulting N-terminal FLAG tag remnant. The H1 peptide cannot be detected due to small size. RNAi targeting CUL3 was used 48 h in advance to prevent E3 ligase assembly. Vinculin and GAPDH serve as loading controls; (C) volcano plots depicting proteome changes upon mRNA-mediated expression of the bioPROTACs vs the corresponding unfunctionalized binder controls. Significant changes of protein levels (Log2FC < -0.5 or > 0.5 and FDR-adjusted p-values (q-values) < 0.05) from 3 independent experiments are depicted for MYC transcriptional targets (dark grey dot) or unrelated proteins (black dot); (light grey dot) represents proteins with non-significant changes in levels. All expressed constructs contain a C-terminal HA tag and NLS.

[0364] Figure 17 shows validation of mRNA-based bioPROTAC-HA expression and MYC downregulation in the HCT116 cell lysates analysed by mass spectrometry (one example of N=3). Controls include VHctrl, DPctrl and PepCtrl as non-MYC binding moieties or lack a degradation domain. Numbers left of the HA blot indicate the molecular weight of protein markers. Vinculin is used as loading control and anti-HA detects bioPROTACs.

[0365] Figure 18 shows proteome changes caused by anti-MYC bioPROTACs. Proteins significantly downregulated by the bioPROTACs vs unfunctionalized binder control (q-value < 0.05 and Log2FC < -0.5) were categorized as (from left to right in the bars) known MYC targets, MYC interactors, off-targets of the degradation domain and unknown relation to MYC. Figure 19 shows MAX downregulation in the presence of MYC-targeting bioPROTACs: (A) immunoblotting after expression of the indicated constructs by mRNA transfection for 17 h; (B) similar analysis following induction of bioPROTAC-T2A-mCherry constructs with 1 pg / mL doxycycline in stable cell lines for the indicated time-points. Representative data of 3 independent experiments are shown. Numbers left of HA blots indicate the molecular weight of protein markers. Ectopically-expressed constructs were detected with anti-HAtag antibody and vinculin was used as loading control.

[0366] Figure 20 shows bioPROTAC-mediated MYC destruction hinders cell proliferation in stable cell lines: (A) Live-cell imaging for monitoring of cell confluence (% cell-covered area, top) and mCherry reporter expression (average integrated intensity per cell, bottom) of the indicated HCT116 doxycycline-inducible cell lines expressing construct-T2A-mCherry fusions or inhibitor-treated unmodified cells (60 μM 10058-F4). For inducible cell lines, strong colours designate induction of the indicated construct, as opposed to non-induced conditions represented by fainter colours. Continuous lines are the mean of 3 technical replicates for a representative biological replicate (of N=3). Dotted lines depict the error (SD); (B) cell growth was further addressed via a CellTiter Gio 72 h end-point assay in the same conditions. Values obtained following induction with doxycycline or treatment with 10058-F4 are depicted relative to untreated controls (no doxycycline or DMSO, respectively). Sample key: parental cells (first bar), cells expressing dox-inducible binder-only controls (second, fourth and sixt bar) or bioPROTACs (third, fifth and seventh bar) treated with doxycycline, or parental cells treated with 60 μM 10058-F4 (eight bar). Shown is the mean of 3 biological replicates (columns) ±SD (error bars). Statistical analysis was performed using a one-way ANOVA followed by Bonferroni multiple comparisons test. Multiplicity-adjusted p-values: “ < 0.01 , *** < 0.001 , **** < 0.0001 , ns = non-significant (> 0.05); (C) construct induction and MYC downregulation were monitored in the same cell lines for the indicated time-points by immunoblotting. Numbers left of the HA blot are the molecular weight of protein markers. Upper bands are uncleaved construct-T2A-mCherry fusions. Vinculin was used as a loading control and the expression of constructs was detected with anti-HA antibody.

[0367] Figure 21 shows the effect of bioPROTAC expression on KRAS levels: (A) Western blot, HA indicates bioPROTAC and vinculin used as loading control; (B) densitometry analysis of immunoblotting results, normalised to KRAS levels in mock transfected HCT116 cells (left- hand panel, all bioPROTACs except those with SPOP DD; right-hand panel only SPSB1 bioPROTAC (+SPSB1 bioPROTAC with control DARPin). ** =p = 0.008 Examples

[0368] Example 1 - Development of mRNA-based screening method

[0369] Results and Discussion

[0370] Unlike high-throughput screening of small-molecule drugs, which are first synthesized and then individually delivered to the cells as arrays, the polypeptide nature of bioPROTACs means that the plasma membrane is a significant barrier to intracellular delivery. An alternative method requires intracellular delivery as nucleic acids;upon translation, the effector bioPROTAC proteins are produced within the cells where they can elicit the desired function. The two options we considered for this purpose due to scalability of both production and intracellular delivery were the transient transfection of plasmid DNA or of mRNA, via lipofection. We identified mRNA delivery as offering significant advantages for our screening strategy, namely 1) a very high transfection efficiency by lipofection, close to 100% in certain cellular systems (Avci-Adali et al., J Biol Eng 8: 8, 2014), thus constituting an attractive option where confounding effects of non-transfected cells are essentially absent; 2) short time (30-60 min) from delivery to expression of the encoded gene without the cell-to-cell heterogeneity of expression onset observed in plasmid DNA transfection (Andreev et al., Gene 578(1): 1-6, 2016). Thus, mRNA-based delivery by lipofection allows an almost immediate start of the cellular effect across the entire cell population. An additional important consideration is the unwanted inherent cellular toxicity of nucleic acid lipofection, which we addressed in Figure 1. As shown, live-cell imaging measurements of HCT116 cell growth indicated plasmid-based transient transfection using Lipofectamine™ 2000 as being largely cytotoxic, with cells ceasing to proliferate for the duration of the assay. On the contrary, not only was in v / fro-transcribed mRNA transfection using the same reagent markedly less toxic, a different reagent, Lipofectamine™ RNAiMAX, was seemingly innocuous and was therefore chosen as the method of choice for this and the ensuing examples.

[0371] We proceeded by optimizing the in vitro transcription reactions, using GFP expression as a reporter. Efficient mRNA production requires a linear DNA template, which we generated by PCR amplification from plasmid DNA. During the amplification step, a polyA (polyadenosine) tail was added to enhance intracellular mRNA translation and stability. We tested the effect of varying polyA tail lengths in GFP expression following mRNA transfection (Figure 2), and observed substantial improvements in expression with increased polyA length. However, this improvement was progressively lower, with 80 adenosine residues having only a moderate advantage over 60 residues. We therefore chose not to investigate even longer polyA lengths. Next, we focused on optimizing mRNA capping, a process which is also important for mRNA translation and stability. First, enzymatic post-transcriptional and co-transcriptional generation of the natural Cap-1 structure were compared, and we observed substantially higher GFP expression for the latter setup (Figure 3A). This observation was followed by testing different commercially available cap analogs for co- transcriptional capping, with maximal GFP expression observed for the CleanCap AG reagent (Figure 3B). Finally, because mRNA modifications can reduce its immunogenicity and boost stability and expression (Kariko et al., Mol Ther 16(11): 1833-1840, 2008), we tested the inclusion of a modified nucleoside, 5-Methoxy-UTP (5moU), in its ability to enhance GFP expression. Due to potential consequences in mRNA stability, GFP expression was monitored by live-cell imaging to address both short-term and long-term consequences of the used modification (Figure 4). We found more GFP-expressing cells at the later time- points when UTP was replaced with 5moU (Figure 4A), irrespectively of the amount of 5moU used, arguing for a positive effect of the modification on mRNA stability. Interestingly, cells expressed higher levels of GFP at earlier time-points with lower amounts of 5moU, with maximal effect at 20% (Figure 4B). Using this low concentration of 5moU also translates into a lower cost per reaction, an advantage for implementation of this method at scale. Note that with this experiment, we could also identify a time window where expression from mRNA is maximal in HCT116 cells.

[0372] We then focused in miniaturizing the reactions and reducing steps of the production workflow to enable mRNA synthesis in high-throughput. We started by simply reducing the in vitro transcription reaction volume, which did not cause a significant decrease in GFP expression (Figure 5). Typically, mRNA synthesized by in vitro transcription is purified before transfection, to eliminate reaction components such as unincorporated NTPs and DNase- digested DNA. We hypothesized that this purification step may not be essential for mRNA transfection, and performed a side-by-side comparison of purified and unpurified mRNA (Figure 6). Indeed, we observed identical GFP expression in the two settings, arguing that this laborious step could be excluded when desired. Likewise, plasmid DNA is typically purified before amplification of the linear template by PCR, and we therefore tested whether this step could be eliminated. We thus compared the standard workflow, whereby plasmid DNA is purified from lysed E. coli before PCR, with a workflow where intact bacteria (containing the plasmid) are used directly while in culture medium for PCR. Surprisingly, we found that not only could we successfully perform PCR in these settings, but the resulting mRNA elicited GFP expression similar to the standard conditions (Figure 7).

[0373] Furthermore, in setting up of the screening platform used in Example 2, we implemented lab automation instrumentation to replace manual steps and increase throughput. For example, acoustic droplet dispensing was used for molecular cloning and transfections instead of manual pipetting, and PCR products purified using an automation station. We additionally miniaturized molecular cloning and implemented protocol simplifications as described (Oling et al., ACS Synth Biol 11(7): 2229-2237, 2022). This included the elimination of plating after bacterial transformation due to the high fidelity of this step.

[0374] Methods

[0375] PCR from a plasmid encoding a T7-5’UTR-GFP-3’UTR cassette was performed using the Phusion® High-Fidelity DNA Polymerase (NEB) following manufacturer’s instructions, using oligonucleotide primers which bind the template DNA upstream of the T7 promoter (forward primer) and downstream of the 3’ UTR (reverse primer). The reverse primer’s sequence contained a 5’ extension of 80 thymine bases (unless otherwise noted), which results in the same number of adenine bases in the forward strand following PCR. in vitro transcription reactions were performed using the HiScribe® T7 high yield RNA synthesis kit (NEB) as per manufacturer instructions, with the following variations: 1) reagent volumes were proportionally reduced to 10 μL or 2 μL reactions, 2) 40% or otherwise indicated amounts of UTP were replaced with 5-methoxyuridine-5’-triphosphate (5-moUTP, TriLink), 3) co- transcriptional capping was performed by including CleanCap® AG (TriLink) in equimolar amount to NTPs. Where specified, co-transcriptional capping was instead performed with other cap analogs: ARCA (NEB S1411), m7G(5')ppp(5’)G (NEB S1404) or m7G(5')ppp(5’)A (NEB S1405). mRNAwas purified using a MEGAclear Transcription Clean-Up kit (Invitrogen). If post-transcriptional capping was performed, in vitro transcription was performed in absence of any cap analog, mRNAwas purified as described and capping performed using the Vaccinia Capping System (NEB) and Cap 2'-O-Methyltransferase (NEB) to generate a Cap-1 structure.

[0376] For live-cell imaging experiments, 8000 HCT116 cells prepared in 40 μL DMEM (Gibco) supplemented with 10% FBS were transfected in each well of a 384-well CellCarrier Ultra plate (PerkinElmer) with 10 uL Opti-MEM (Gibco) containing 50 ng mRNA and 0.15 μL Lipofectamine™ 2000 (Invitrogen). Cells were subsequently imaged for brightfield and green fluorescence every 2h on an Incucyte S3 live-cell analysis system (Sartorius). For flow cytometry measurements, a proportional amount of cells were instead transfected in 12-well plates using Lipofectamine™ RNAiMAX. Cells were harvested by incubating with Versene solution (Gibco) for 30 min at 37°C and analysed on a BD LSRFortessa™ cytometer.

[0377] Example 2 - Identification of MYC degraders Results and Discussion

[0378] The screening method developed in Example 1 was implemented to identify bioPROTACs capable of inducing degradation of human MYC. MYC is an oncogenic driver upregulated in approximately 50 % of human cancers. MYC’s inactivation triggers tumour regression, underpinning a clear therapeutic opportunity, but its intrinsically disordered nature has proved to be a difficulty in traditional drug discovery (Whitfield & Soucek, Journal of Cell Biology 220(8): 6202103090, 2021). MYC’s very short half-life (-30 min) has raised scepticism about its amenability for targeted degradation (Whitfield & Soucek, supra). MYC was selected as proof of concept for the screening method as a target to challenge our strategy.

[0379] To develop MYC degraders, we first established a DNA fragment library of bioPROTAC modules (degradation domains (DDs), binding domains and linkers). For our library, we selected a representative panel of DDs covering a wide variety of E3 ligase families and recruitment strategies to identify optimal attributes. 11 DDs were sourced from the literature (e.g. SPOP and TRIM21) and the others were new designs, chosen based on their suitability to be employed in bioPROTACs (continuous polypeptide; associated with protein degradation) and chance of the cognate E3 ligase encountering oncogenic MYC (nuclear localization; widespread expression across tissues and cancer types of the recruited E3 ligase). We partly centred on bacterial / viral proteins to leverage their qualities in UPS hijacking. The DDs in the initial panel are set out in SEQ ID NOs: 1-11 and 13-39.

[0380] For MYC recruitment, we sourced two VH single-domain antibody fragments from the literature (Zeng et al., Journal of Immunological Methods 426: 140-143, 2015), CMYCVH-12-321 and CMYCVH-15-321 (hereafter VH 12 (SEQ ID NO: 103) and VH15 (SEQ ID NO: 110), respectively), and the DARPin DP20 (SEQ ID NO: 102), previously generated in-house by phage display. All binders were originally selected as standalone inhibitors. Intracellular binder-MYC engagement was confirmed by co-immunoprecipitation in HCT116 cells (Figure 8).

[0381] Finally, since structural arrangements influence bioPROTAC action, we included 5 and 16 amino acid (AA) glycine-serine flexible linkers, a long 65AA rigid linker (Klein et al., supra) and an 81AA rigid / flexible combination (named L1 to L4, SEQ ID NOs: 120, 126, 132 and 133, respectively).

[0382] Using the methods developed in Example 1 , a library of 720 bioPROTACs was assembled from the panels of DDs, binding domains and linkers, bioPROTAC genes were transcribed in vitro and the resulting mRNA transfected into cell lines. A nuclear localization signal (NLS) was included in all constructs to ensure co-localization with MYC. Strikingly, almost 100 % of cells transfected with the bioPROTAC constructs exhibited bioPROTAC expression when using this method (Figure 9). Endogenous MYC abundance changes were measured by end-point immunofluorescence, or kinetically monitored by luminescence in cells expressing NanoLuc- tagged MYC[T58A], The T58A mutation increases MYC half-life to -80 min (Bahram et al., Blood 95(6): 2104-2110, 2000) and was chosen to identify degraders that struggle with targeting endogenous MYC.

[0383] The bioPROTAC library was first assessed in primary screening (Figure 10), and a subset (91) was selected for subsequent validation, based on a reasonable effect (z-score < 1 .5 in either assay, avoiding L1 / L2 and L3 / L4 linker redundancy) or for design feature dissection.

[0384] We identified 7 strong bioPROTACs that achieved the required thresholds for both immunofluorescence and kinetic analyses (q-value < 0.05, l_og2fold change (FC) < -2 in immunofluorescence / Log2FC < -1 in the kinetic assay) (Figure 11). Of these, 4 hits contained the newly-designed HSV-1 ICP0 DD. Some bioPROTACs yielded disparate results in the two assays, which may be due to the different turnover of the T58A mutant and / or the saturation of the E3 ligase by increased abundance of MYC in the cell line, an interference of NanoLuc in target ubiquitylation, or due to degradation-independent effects.

[0385] Notably, we observed a significant correlation between bioPROTAC abundance and effect (Pearson’s r = -0.63, p < 0.0001) (Figure 12). Given that E3 ligases can be co- depleted during TPD (Clift et al., Cell 171 (7): 1692-1706, 2017), we hypothesized that some potent bioPROTACs are destabilized by self-ubiquitylation. We followed up with a “validation and optimization” screening run where 1) rationally-designed lysine-arginine substitutions (K>R) were introduced in hits and low abundant constructs in an attempt to boost their abundance (SEQ ID NOs: 12 and 40-52) and 2) non-MYC-binding controls were included (SEQ ID NOs: 111 ,-112 and 155).

[0386] In parallel, we established a new lead generation strategy using the original DD / Linker DNA fragment library coupled to another MYC binder, the H1S6A’F8A(hereafter H1) peptide (Giorello et al., Cancer Research 58(16): 3654-3659, 1998) , to further increase compositional diversity. Primary and confirmatory runs were conducted as above, yielding 12 further strong degraders common to both assays (Figure 13).

[0387] An attempt to reduce ICP0 domain length failed. K>R substitutions enhanced 12 bioPROTAC abundance and effect, though most changes were subtle except for the dramatic improvement of DP20-L1-SPOP upon seven K>R substitutions in the binding domain (Figure 14).

[0388] We further extracted favourable bioPROTAC features from our datasets and corroborated specific characteristics by immunoblotting. We noted that 1) the strongest DDs incorporate RING-family E3 ligases (particularly ICP0); 2) module relationships are intricate and unpredictable, as demonstrated by DD-specific binder preferences (Figure 15A), alluding to a key contribution of specific structural arrangements; 3) all investigated binders can assemble proficient degraders if combined with the right DD, thus finding the right domain pair supersedes affinity properties; 4) enhanced target degradation can be achieved by changing linkers or orientations, as seen for ICPO (Figure 15B); and 5) composition changes can also alter bioPROTAC efficacy by influencing the bioPROTAC’s own abundance.

[0389] In total, as detailed above, 19 strong Myc-degraders were identified: VH151KR-L1- ICP01KR(SEQ ID NO: 138), DP207KR-L1-SPOP (SEQ ID NO: 139), TRIM21-L2-H1 (SEQ ID NO: 140), TRIM21-L1-H1 (SEQ ID NO: 141), TRIM21-L1-VH151KR(SEQ ID NO: 142), TRIM21-L1-VH153KR(SEQ ID NO: 143), VH123KR-L2-ICP01KR(SEQ ID NO: 144), VH121KR- l 2-ICP01KR(SEQ ID NO: 145), VH153KR-L1-ICP01KR(SEQ ID NO: 146), VH15-L1-ICP01KR(SEQ ID NO: 147), ICP0-L4-VH153KR(SEQ ID NO: 148), ICP0-L4-H1 (SEQ ID NO: 149), ICP0-L2-H1 (SEQ ID NO: 150), H1-L1-ICP0 (SEQ ID NO: 151), H1-L2-ICP0 (SEQ ID NO: 152) and DP207KR-L1-SPQP3KR(SEQ ID NO: 153).

[0390] Of these high-performing bioPROTACs, three with unique domains were chosen for further study (Figure 16A): VH151KR-L1-ICP01KR(SEQ ID NO: 138), DP207KR-L1-SPOP (SEQ ID NO: 139) and TRIM21-L2-H1 (SEQ ID NO: 140). The SPOP-based degrader was chosen as an example of E3 indirect recruitment which underperforms in the engineered kinetic assay yet excels in endogenous MYC destruction. Densitometry analysis following immunoblotting revealed MYC levels down to 3.55±1.03%, 6.01 ±4.23% and 4.86±2.53% (mean±standard deviation of 3 biological replicates, relative to mock), respectively, upon bioPROTAC expression (Figure 16B).

[0391] In the case of DP207KR-L1-SPOP, we additionally depleted the recruited E3 ligase’s subunit CUL3 to substantiate our findings. H1 can itself elicit mild MYC downregulation, which is expected given that its stand-alone role in preventing MYC heterodimerization with its interacting partner MAX (Giorello et al., supra) modulates MYC protein levels (Mathsyaraja et al., Genes & Development 33: 1252-1264, 2019). Finally, we demonstrate that bioPROTACs cannot operate when the functional domains are split by the T2A autocleavable peptide, supporting a mechanism whereby the selected bioPROTACs potently induce endogenous MYC degradation by bringing the target and an active E3 ligase into close proximity.

[0392] Next, we confirmed the biological effect of these bioPROTACs by measuring global proteome changes by mass spectrometry in mRNA-transfected cells. Though MYC itself couldn’t be reliably quantified via LC-MS / MS, we observed widespread significant proteome changes (Log2FC < -0.5 or >0.5, q-value < 0.05) induced by MYC-targeting bioPROTACs (Figure 16C). Accordingly, MYC downregulation was confirmed by immunoblotting (Figure 17). We also observed that some MYC interactors were co-degraded (Figure 18). Accordingly, MAX was depleted in the presence of the MYC-targeting bioPROTACs after mRNA transfection or in doxycycline (dox)-inducible cell lines (Figure 19), albeit DP207KR-L1-SPOP was slow-acting. Consistent with MYC functions, we observed a marked inhibition of cell proliferation upon bioPROTAC expression in these cell lines (Figure 20), including for unfunctionalized DP207KRwhich by design inhibits MYC. This proliferation delay was comparable to treatment with the MYC inhibitor 10058-F4 in the case of DP207KR-L1-SPOP, but substantially more pronounced for VH151KR-L1-ICP01KRand TRIM21-L2-H1. Overall, these data show that VH151KR-L1-ICP01KRexcels in on-pathway specificity and induces strong MYC destruction and downstream effects, while also affecting the MYC interactome. DP207KR-L1-SPOP is an adequate candidate when interactor co- depletion is to be avoided.

[0393] Our study demonstrates that bioPROTACs can be used to very efficiently modulate the high-turnover MYC, and also provides a roadmap to destroy virtually any endogenous protein with high specificity in short time, without the need for prior cell line engineering or protein tagging. We uncover several properties of biological degraders, and identify ICPO as a new, specific and strong DD. To our knowledge, this high-throughput development platform represents the most extensive exploration of bioPROTAC composition to date against any target. The combinatorial approach rapidly uncovers productive pairings between DDs and binders, which are currently unpredictable and could easily be missed by a small-scale manual screen.

[0394] Methods

[0395] Selection and engineering of degradation domains

[0396] Selected degradation domains originate from human, bacterial or viral E3 ligases or intermediate proteins capable of recruiting human E3 ligases. Domains were either retrieved from other studies or newly developed, in all cases extracted from proteins matching the following criteria: 1) associated with protein degradation, 2) functional domains well understood (biochemically and / or the structure has been determined and is available in the protein data bank (PDB)), and can easily be reformatted as an individual bioPROTAC module (ideally uninterrupted sequence and not dependent on post-translational modifications). If the domain is an intermediate, the indirectly-recruited E3 ligase is 3) present in the nucleus (to co-localize with MYC), 4) abundant in a wide range of tissues (to ensure wide applicability) and 5) widely expressed in cancer, ideally overexpressed, for applicability and to ensure enough supply to degrade overexpressed MYC (sources: Uniprot, www.uniprot.org; Human Protein Atlas, www.proteinatlas.org; and GEPIA, gepia.cancer- pku.cn). Additionally, we favoured intermediates where the recruited partner is found to be essential in > 50 % of the studies shown in the online gene essentiality (OGEE) database (v3.ogee.info / # / home), to avoid cellular resistance to the bioPROTACs, though this was not a requirement. We considered an E3 ligase essential if the direct interacting partner is reported as an essential gene in > 80 % of the studies included in the OGEE database and not essential if in < 20 % of the studies (otherwise labelled “conditional”).

[0397] For new designs, we compiled a non-exhaustive list of proteins matching the above criteria after manual literature searches, supported by UniProt curated information. We prioritized covering a diverse range of module characteristics in the assembled DD list, for instance by exploiting E3 ligases of various families and / or utilizing structurally and functionally distinct moieties. To circumvent deleterious gain-of-function, most natural substrate-binding sequences were removed or mutated (e.g. T67A in ICPO, see Chaurushiya et al., Molecular Cell 46: 79-90, 2012), guided by literature observations and UniProt annotations, unless the functional regions cannot be fully separated and a resulting gain-of- function is deemed not harmful. We also explored the benefit of positioning a functional group at either the N- or C- terminus relative to the MYC binder when there is evidence that natural substrates may bind in either end. In addition, we tested two truncations of the CRL adaptors pTrCP, FBXO17 and DDB2 in search of a minimal functional domain and also included HUWE1 constructs with and without its autoinhibitory sequence as an attempt of mitigating self-ubiquitylation and degradation. Undesired trafficking sequences were also removed, such as membrane-targeting or nuclear exclusion signals. See Table 1 .

[0398] Affinity protein generation a-MYC DARPin 20 (WO 2016 / 023898) was isolated from a phage display library ahead of this work as previously described (Guillard et al., Nature Communications 8: 16111 , 2017), by selection against GST-MYC bHLH-LZ (AA 330-439; Novus Biologicals). Deselection against GST-Ubiquitin was also performed to eliminate GST-binding molecules. Validation of antigen-DARPin binding was performed by immunoassays as previously described (Guillard et al., supra). MDM2 VHHs were generated by Hybrigenics Services SAS, Paris, France via a yeast two-hybrid screening against LexA-MDM2 (AA 1-188). As non-MYC binding scaffold controls, we used DARPin E3_5 (Kohl et al., Biophysics and Computational Biology 100(4): 1700-1705), a VH targeting hen egg white lysozyme (Rouet et al., Journal of Biological Chemistry 290: 11905-11917, 2015) and an a-helical peptide from the DARPin backbone. Sequences of the binders used are set out in SEQ ID NOs: 101-110.

[0399] Single insert cloning and mutagenesis Individual DNA fragments were generated by gene synthesis (Twist Bioscience, GenScript or Invitrogen GeneArt) or PCR, except fragments smaller than 80 bp which were obtained by annealing of two complementary oligonucleotides by heating at 95°C and cooling down to 22°C at a rate of 0.1 °C / s. Non-human sequences were subjected to codon optimization (GenScript). DNA fragments encoding the degradation domains, linkers, MYC-binding domains and controls were cloned into the pCR Blunt ll-TOPO vector using the Zero Blunt TOPO PCR Cloning Kit (Thermo Fisher) following the manufacturer’s instructions. cDNA encoding human MYC was subcloned into pFN31K / pFC32K vectors (Promega) for fusion with NanoLuc luciferase (expression controlled by the hPGK promoter), DNA encoding FLAG-tagged DP20, DARPin E3_5, VH12 and VH15 was subcloned into the pFC14K vector (Promega), DNA encoding shortlisted bioPROTAC-HA-NLS constructs and controls was subcloned into a modified ODIn-inv-neo mammalian expression plasmid (Lundin et al., Nature Communications 11 :4903, 2020) containing T2A-mCherry. The Golden Gate destination (GG-dest) vector was generated by integration of a cassette comprising the T7 promoter with AG initiator sequence, UTRs (Mandal & Rossi, Nature Protocols 8: 568-582, 2013), the Kozak consensus sequence, the ccdB killer gene plus a chloramphenicol resistance gene flanked by Esp3l recognition sequences, HA tag and MYC NLS (Ray et al., Bioconjugate Chemistry 26: 1004-1007, 2015) sequences onto a pcDNA3.1 (+) vector (Invitrogen).

[0400] MYC Thr 58 was mutated to Ala (T58A) following a “round-the-horn” PCR-based protocol. Briefly, after PCR the template vector was digested with Dpnl (New England Biolabs, NEB) for 4 h at 37°C, purified using the QIAquick PCR purification kit (Qiagen), blunt ends were phosphorylated using PNK (NEB) in T4 ligase buffer for 20 min at 37°C. Following heat-inactivation of PNK at 75°C for 10 min, ends were ligated using T4 ligase (NEB) by incubating for 2 h at room temperature and overnight at 16°C. All other mutants were obtained by gene synthesis. PCR was performed using Phusion High-Fidelity Master Mix (NEB) following vendor instructions. Decisions to mutate lysine residues in degradation domains were based on conservation with orthologs (which may include mouse, rat, fruit fly, cow, pig, zebrafish, western clawed frog and African clawed frog), conservation in human paralogs, presence within critical functional domains and surface exposure (inferred from the positioning of a lysine’s side chain in available 3D structures). Mutagenesis of MYC-binding domains relied on the conservation and surface exposure of residues from similar affinity proteins. Sequence and structural information was obtained from UniProt and PDB. All sequences were analysed using Geneious Prime 2020.

[0401] Automated multi-fragment cloning and in vitro transcription Arrayed unique combinations of library DNA fragments encoding the degradation domains (or Halo tag as negative control), linkers and MYC-binding domains (or non-binding controls), flanked by Esp3l recognition sequences were pre-cloned into the pCR Blunt II- TOPO vector, and individually assembled by a Golden Gate cloning protocol modified from as described in Example 1 which resulted in > 98 % correct clones in preliminary control reactions. All resulting fusions contain a C-terminal HA tag and NLS sequence. Overhang sequences ATGG, GCCG, GAAG and GGAT were chosen based on T4 ligase fidelity information, using the Ligase Fidelity Viewer (NEB; ggtools.neb.com / viewset / run.cgi). 2 μL reactions containing 1 fmol of DNA fragments (3 fragments per reaction), 1 fmol GG-dest vector, 0.1 μL Esp3l FastDigest Enzyme (Thermo Fisher), 0.1 μL T4 DNA ligase

[0402] 2 x 106U / mL (NEB) and 0.2 μL 10x T4 ligase buffer (NEB) were subjected to 30 incubation cycles alternating between 37°C for 5 min and 16°C for 5 min, followed by 30 min at 37°C and 15 min at 75°C. 0.5 μL of a master mix containing 0.125 μL Plasmid-Safe ATP- Dependent DNase (Lucigen / Epicentre) and 0.125 μLATP 25 mM in 1x T4 DNA ligase buffer (NEB) were then added to each reaction and incubated for 1 h at 37°C. High efficiency chemically competent DH5a E. coli (NEB) were grown in liquid 2x TY medium overnight at 37°C following transformation, after which 0.5 μL of the liquid cultures were directly used for amplification of the T7-5’UTR-bioPROTAC-HA-NLS-3’UTR-polyA cassette by PCR.

[0403] Following incubation with Dpnl for 1 h at 37°C, products were bound to AMPure XP beads (Beckman Coulter), washed with 70 % ethanol and eluted with RNase-free water. Subsequently, arrayed in vitro transcription reactions were performed using the HiScribe T7 high yield RNA synthesis kit (NEB) as per manufacturer instructions, with the following variations: 1) reagent volumes were proportionally reduced for 2 μL reactions, 2) 20 % of UTP was replaced with 5-methoxyuridine-5’-triphosphate (5-moUTP, TriLink) and 3) co- transcriptional capping was performed by including CleanCap AG (TriLink) in equimolar amount to NTPs.

[0404] Template DNA was digested with DNase I (NEB). DNA quality control was performed by electrophoresis and sanger sequencing of a small fraction of PCR products and mRNA integrity and yield was examined using a Bioanalyzer RNA 6000 Nano assay (Agilent) for randomly-selected constructs. mRNA was not purified prior to screening, and the mRNA concentration of screening library elements was extrapolated based on the analysed reaction subset and transfected in excess of previously-determined saturating amounts (we have determined that for GFP expression, transfection with > 50 ng mRNA does not increase protein expression levels, i.e. the fluctuation of mRNA yields should not be reflected in bioPROTAC protein abundance following mRNA transfection).

[0405] For follow-up experiments, plasmid DNA from individual clones was purified and sequenced and PCR reactions were column-purified using a QIAquick PCR purification kit. §RNAwas synthesized in 10 μL reactions containing 40 % 5-moUTP (to further evade cellular innate immune responses and stabilize mRNA), purified using a MEGAclear Transcription Clean-Up kit (Invitrogen) or RNeasy 96 Kit (Qiagen) according to vendor instructions and quantified with a Lunatic nucleic acid quantification system (Unchained labs) to allow transfections with equal amounts of mRNA. Instrumentation for screening: purified DNA (plasmids and PCR products) was dispensed by acoustic droplet ejection using an Echo 550 liquid handler (Labcyte). For automated cloning and in vitro transcription, shared reaction components (except GG-dest vector) were transferred as a master mix using a Certus Flex liquid dispenser (Fritz Gyger). Other dispensing steps during screening were performed using a Hamilton Microlab Star liquid handling system (Hamilton). DNA sequences were analyzed using Geneious Prime 2020.

[0406] Cell culture and cell line generation

[0407] All cell culture reagents are from Gibco, Thermo Fisher unless otherwise indicated. HCT116 cells were purchased from ATCC and cultured in DMEM supplemented with 10 % FBS and 1 % penicillin-streptomycin at 37°C and 5 % CO2. In absence of a C02-controlled environment during kinetic luminescence experiments, Leibovitz’s L-15 medium was used instead of DMEM. To establish modified cell lines, cells were transfected using FuGene HD (Promega) following manufacturer’s instructions. ObLiGaRe Doxycycline Inducible (ODin) cell lines (Lundin et al., supra) were generated by co-transfecting bioPROTAC or control- encoding ODIn-inv-neo and ZFN-AAVS1 plasmids at a ratio of 1 :2 and selecting with 0.5 mg / mL geneticin-containing growth media. All cells were tested and confirmed to be mycoplasma-free before use.

[0408] Immunofluorescence microscopy

[0409] For immunofluorescence (IF) experiments, unmodified HCT116 cells were transfected with mRNA using Lipofectamine RNAiMAX (Invitrogen), which resulted in > 99 % transfected cells when tested with GFP, as follows: ≥ 50 ng bioPROTAC-encoding mRNAwas added via acoustic dispensing to the wells of 384-well plates containing Opti-MEM (Gibco) in triplicate, onto which 0.15 μL Lipofectamine RNAiMAX in Opti-MEM was added followed by a 15-30 min incubation at room temperature and the seeding of 8000 cells per well. 17 h post- transfection, cells were incubated in 4 % formaldehyde for 15 min, washed 3x with PBS and incubated in blocking buffer (3 % BSA, 0.1 % Triton X-100 in PBS) for 1 h.

[0410] Fixed cells were incubated overnight with anti-MYC antibody Y69 (Abeam) and anti- HAtag antibody 16B12 (Abeam and Enzo Life Sciences), both diluted 1 :5000 in blocking buffer, followed by 3 washes with PBS and incubation with donkey anti-rabbit IgG / Alexa Fluor 488 1 :500, goat anti-mouse IgG / Alexa Fluor 568 1 :500 and Hoechst 2 pg / mL (all from Thermo Fisher) in blocking buffer for 1 h and then 3 washes with PBS. 4 images per well were acquired by a CV7000 spinning disk confocal microscope (Yokogawa) using a 20x objective. Raw images were processed using the Columbus 2.9.1.532 software (PerkinElmer) for quantification of nuclear MYC and HA (bioPROTAC) abundance for all cells in the field of view. Raw values per cell are an average of the nuclear pixel intensity for the respective fluorescence channel. The values brought forward per well are the median of the cell population. In each experiment, three technical replicates were measured (these are from distinct samples and not repeated measurements of the same sample). Up to three independent experiments were performed. Treatment with 100 pg / mL of the protein synthesis inhibitor cycloheximide (Sigma Aldrich) was used as a proxy for background fluorescence in absence of MYC and provides a measure of the detection limit, based on the estimation of a negligible residual MYC of 5.8x10-9% after 17 h treatment (considering a half-life of 30 min and an exponential decay of its abundance). Figure panels were constructed using Imaged 1.53c and Adobe Illustrator 2020 (version 24.2.1). Prior to this study, we confirmed the high specificity of the anti-MYC Y69 antibody by siRNA-mediated downregulation of MYC and that the anti-HA antibody does not yield detectable signal when HA-tagged proteins are not expressed.

[0411] Luminescence kinetic assay

[0412] HCT116 MYC[T58A]-NanoLuc-expressing cells were transfected in triplicate following the protocol indicated for the IF experiments, except that Nano-Gio Vivazine live cell NanoLuc substrate (Promega) was pre-mixed with cells prior to seeding. Readings were taken every 30 min for 48 h using an EnVision plate reader (PerkinElmer) with ultrasensitive luminescence detection. In each experiment, three technical replicates were measured (these are from distinct samples and not repeated measurements of the same sample). Up to three independent experiments were performed. Time resolution was lower in primary screens due to practical reasons. 100 pg / mL cycloheximide treatment was used to estimate the maximum attainable MYC downregulation, based on the assumption that its decay in complete absence of protein synthesis (80 min half-life (Bahram et al., supra)) is faster than induced protein degradation following mRNA transfection. Residual MYC[T58A] is 3.8 x 10-4% 24 h post-transfection considering an 80 min half-life.

[0413] Screening data processing

[0414] For the IF readout, the median of nuclear MYC fluorescent intensity measurements for the cell population was normalized to the plate-specific mock control to eliminate plate-to-plate variations and signal disparities across biological replicates. Nuclear HA staining (bioPROTACs) lacks a universal standard reference, and thus the data was not transformed except for the visualization of bioPROTAC expression relative to the average of controls in Figures 12 and 14. HA measurements were therefore not corrected for plate-to-plate variations. In the kinetic assay, luminescence relative light unit (RLU) values were normalized to the respective mock control data points to eliminate time-point- and plate- specific variations. Kinetic datasets were then normalized to the first measured data point, and the area under each curve was calculated following the trapezoidal rule. Values for technical replicates were averaged and nuclear MYC fluorescence intensity (in IF) or area under the curve (kinetic) fold changes (FC) were calculated as the MYC abundance ratio in the presence of bioPROTACs vs specific innocuous control fusion.

[0415] For discovery-type screenings, controls contain the linker and MYC binder, but lack a degradation domain (replaced by the Halo tag) and for the validation & optimization screening, controls use target-binding scaffolds which do not recruit MYC. In the primary screenings, FCs were standardized as z-scores to evaluate performance relative to the dataset samples, also providing a dimension to compare both readouts. We considered that values with z-score < -1.5 are sufficiently away from the mean of the dataset to call for further scrutiny in a confirmatory screening. The probability that values in a given normal distribution fit this criterion is 6.7 % (see https: / / www.calculator.net / z-score-calculator.html). Selected hits were carried forward for confirmation by performing two additional independent experiments (i.e. final N=3).

[0416] Wells with median HA fluorescence intensity values (in IF) lower than the average of mock samples’ medians + 3x standard deviation (SD) were considered expression-, transfection- or staining-negative and were excluded on grounds of evidence of technical error. Accordingly, MYC abundance readings in the IF and kinetic assays from constructs not detected in any technical replicate were excluded, thereby resulting in some cases with 2 usable biological replicates. Where an innocuous control fusion was not detected in primary screenings, the most similar one was chosen to calculate FCs of the cognate bioPROTACs (does not apply to confirmatory runs). Statistical analysis of MYC abundance and area under the curve measurements were performed vs sample-specific controls using multiple unpaired two-tailed t-tests corrected for false discovery rate (FDR) with the two-stage Benjamini, Krieger and Yekutieli method. We assumed a normal distribution of the biological replicate values, and that a given sample has the same SD as its respective control.

[0417] Reported are Log2FC, difference between mean values and associated standard error, t ratio (difference between means divided by the standard error of the difference), degrees of freedom and q-values (FDR-adjusted p-values). Discoveries were assigned based on a 5 % FDR threshold, and discoveries with Log2FC < -2 (IF) or < -1 (Kinetic assay) were considered strong hits. This arbitrary classification deems a 75 % reduction in endogenous MYC levels (endpoint IF), or 50 % reduction in area under the curve across 48 h (kinetic assay), as a very pronounced reduction in MYC abundance (albeit without guarantee of sufficient MYC reduction to elicit downstream effects). BioPROTAC expression and MYC abundance correlations were evaluated by Pearson correlation tests, with reported correlation coefficient (r) and two-sided p-value. Basic calculations and data handling were performed in MS Excel (v2102) and all statistical analyses were performed using the GraphPad Prism 9 software.

[0418] Cell growth experiments

[0419] Cell proliferation was assessed in biological triplicate by seeding 2000 cells of HCT 116 ODin cell lines or unmodified HCT116 cells 24 h before bioPROTAC induction with 1 pg / mL doxycycline or treatment with 60 μM 10058-F4 MYC inhibitor (Kuser-Abali et al., Epigenetics 9: 634-643, 2014) or DMSO (all reagents from Sigma Aldrich). Live cells were imaged every 2 h for 72 h on an Incucyte S3 system (Sartorius) using default settings. Cell confluence and average mCherry reporter fluorescence (used as a surrogate for bioPROTAC expression) were extracted as specified in the provided software (version 2019A). 72 h end-point assays were performed using the CellTiter-Glo 2.0 cell viability assay (Promega) following manufacturer instructions, and measured in an EnVision plate reader. Statistical analysis was performed by an ordinary one-way ANOVA followed by Bonferroni’s multiple comparisons test (we assumed a normal distribution of the biological replicate values) using the GraphPad Prism 9 software. Means were considered significantly different if the multiplicity-adjusted p-value was < 0.05.

[0420] Immunoprecipitation

[0421] 4 x 106HCT116 cells were transfected with 17 pg pFC14K encoding FLAG / Halo-tagged affinity proteins using FugeneHD as per manufacturer’s instructions. 2 days after transfection, cells were recovered and incubated for 20 min in lysis buffer (1 % NP40, 50 mM Tris-HCI pH 7.5, 150 mM NaCI, 1 mM EGTA, 1 mM EDTA, 10 mM glycerophosphate, 50 mM sodium fluoride, 0.27 M sucrose, 5 mM sodium pyrophosphate, 1 mM sodium orthovanadate) supplemented with complete Mini EDTA-free protease inhibitor (Roche). Cell debris were removed by centrifugation and protein content quantified with the Pierce BCA Protein Assay Kit (Thermo Scientific) following vendor instructions. 12.5 pg cell lysate was diluted in reducing Laemmli sample buffer (Bio-Rad; reducing agent added as instructed) and incubated at 95°C for 5 min. 5 μL anti-FLAG M2 magnetic beads (Sigma) were pre- washed with lysis buffer and incubated overnight at 4°C with 250 pg cell lysate. Beads were subsequently washed Ix with lysis buffer and 2x with high salt wash buffer (1 % NP40, 50 mM Tris-HCI pH 7.5, 300 mM NaCI, 1 mM EGTA, 1 mM EDTA, 10 mM glycerophosphate, 50 mM sodium fluoride, 0.27 M sucrose, 5 mM sodium pyrophosphate, 1 mM sodium orthovanadate) supplemented with protease inhibitors. Captured proteins were eluted in 1x reducing Laemmli sample buffer and analysed by immunoblotting. Experiments were performed in > 3 biological replicates.

[0422] Immunoblottinq

[0423] In mRNA transfection experiments, 3.2 x 105HCT116 cells were transfected with 1 pg mRNA using Lipofectamine RNAiMAX and harvested with 1x reducing Laemmli buffer after 17 h. Where indicated, Lipofectamine RNAiMAX-mediated transfection of 20 nM Silencer Select CUL3 siRNAs (IDs S16048, S16049, S531239, S531241) or Silencer Select Negative Control No. 2 siRNA (all from Thermo Fisher) was performed 48 h ahead of mRNA transfection. Constructs in ODin system cell lines were induced as previously specified. For all experiments, samples were prepared by lysing cells directly in 1x reducing Laemmli buffer, except in immunoprecipitation experiments (see above). SDS-PAGE and immunoblotting were performed using Criterion Cell and Criterion Blotter systems (Bio-Rad), as instructed by the manufacturer for Tris / Glycine buffers.

[0424] PVDF membranes were incubated with blocking buffer (5 % milk, 0.05 % Tween in PBS) for 1 h and overnight with one of the following antibodies diluted 1 :1000 in blocking buffer (unless otherwise stated): Anti-MYC Y69 (Abeam), anti-HA 16B12 (Enzo Life Sciences) diluted 1 :2000, anti-FLAG M2-Peroxidase (Sigma), anti-Vinculin E1 E9V (Cell Signaling Technology, CST), anti-GAPDH 14C10 (Sigma), anti-MAX (Bethyl, cat. no. A302- 866A-T) and anti-CUL3 (Bethyl, cat. no. A301-109A-T). Vinculin and GAPDH were used as loading controls. Where applicable, washed membranes were incubated for 1 h with the respective anti-mouse IgG or anti-rabbit IgG secondary antibodies conjugated to HRP (CST) diluted 1 :2500 in blocking buffer.

[0425] Detection was performed using Immobilon Crescendo Western HRP substrate (Millipore) in a ChemiDoc imaging system (Bio-Rad) using the Image Lab 6.0.1 software (Bio-Rad). When necessary, HRP was inactivated with 0.1 % Sodium Azide (Sigma) in PBS and membranes re-incubated with primary and secondary antibodies as instructed. Band quantification by densitometry was performed using Image Lab, after which MYC integrated intensity values were normalized to those of vinculin and provided as a percentage relative to the mock control. Figure panels were constructed using Adobe Illustrator 2020 (version 24.2.1) and Adobe Photoshop CC 2019. Experiments were performed in biological triplicate.

[0426] Proteomics sample preparation and data acquisition

[0427] 6 x 106HCT116 cells were transfected with 9 pg mRNA using Lipofectamine RNAiMAX, using different plate configurations in each biological replicate to mitigate plate-to- plate effects. 17 h post-transfection, cells were harvested and washed with ice-cold PBS. Cell lysates from 9 different conditions (VH151KR-L1-ICP01KR, VHCtrl-L1-ICP01KR, VH151KR-L1 , DP207KR-L1-SPOP, DPCtrl-L1-SPOP, DP207KR-L1 , TRIM21-L2-H1S6A / F8A, TRIM21-L2-PepCtrl, and L2-H1S6A / F8A) each with 3 biological replicates were prepared in S-Trap lysis buffer (5 % SDS, 50 mM triethylammonium bicarbonate (TEAB) buffer, pH 7.55). Cells were also collected for parallel immunoblot analysis. Samples were solubilized using a Retsch mill (MM400) bead beater for 2 min at frequency 30. Protein concentration was measured using Pierce gold BCA kit. 50 pg of protein lysates were digested using micro S-Trap method (ProtiFi) according to the manufacturer’s instructions. Briefly, proteins were reduced using 20 mM tris (2-carboxyethyl) phosphine for 15 min at 60°C, alkylated using 80 mM iodoacetamide for 1 h at room temperature, and digested on the micro S-Trap cartridge using mass spectrometry grade trypsin / lys-C (Promega) for 2 h at 47°C. The digested peptides were eluted with 50 mM TEAB buffer, followed by 0.2 % formic acid (FA) in water, and 50 / 50 acetonitrile / water with 0.2 % FA. The eluted peptides were dried in a speedvac and reconstituted in 0.15 % FA in water.

[0428] For spectral library construction, pooled digested samples (10 % peptides from each sample) were fractionated based on high-pH reversed-phase separation using a Dionex RS LC system (Thermo Scientific) (Wang et al., Proteomics 11 : 2019-2026, 2011). The HPLC system was installed with XBridge peptide BEH C18, 130A, 3.5 pm column (Waters). The samples were run at 0.1 mL per minute flow rate with acetonitrile gradient (mobile phase A: 10 mM ammonium bicarbonate, mobile phase B: acetonitrile) for 40 min. A total of 96 fractions were collected and concatenated to a final set of 24 fractions.

[0429] LC-MS / MS analyses were conducted on a timsTOF Pro mass spectrometer (Bruker) coupled with a nanoElute LC-system and nano-electrospray ion source (CaptiveSpray Source, Bruker). Samples were loaded onto a 15 cm x 75 pm, 1.9 pm ReproSil, C18 column (PepSep) using an oven temperature of 50°C. The peptides were eluted at a flow rate of 500 nL / min over a total 51 min gradient, from 4 to 24 % solvent B (36 min), 24 to 36 % solvent B (7 min), 36 to 64 % solvent B (5 min), 22 and 64 to 98 % solvent B (3 min). Solvent A was composed of 0.15 % FA in water, and solvent B was composed of 0.15 % FA in acetonitrile.

[0430] Data-dependent acquisition (DDA) was performed in PASEF mode with 6 PASEF scans at a duty cycle close to 100 %. MS acquisition recorded spectra from 100-1600 m / z and ion mobility was scanned from 0.85-1.30 Vs / cm2over a ramp time of 100 ms. The total duty cycle time was 1.15 s. The collision energy was linearly increased from 27 to 45 as a function of ion mobility. An active exclusion of 0.4 min was applied to precursors that reach to target intensity of 20,000 units. Data-independent acquisition (DIA)-PASEF mode was performed with a scheme that consists of 2 rows of 32 windows (8 PASEF scans per row and 4 steps per PASEF scan) with a 25 m / z isolation width (Meier et al., Nature Communications 12: 1185, 2021). The mass scan range was from 100 to 1700 m / z and ion mobility was scanned from 0.57-1.47 Vs / cm2over a ramp time of 100 ms. The collision energy was ramped linearly from 20 to 52 as a function of mobility.

[0431] Proteomics data analysis and integration

[0432] To generate a comprehensive spectral library for the DIA analysis, we first analyzed 24 concatenated fractions from high-pH reversed-phase fractionation in DDA mode, followed by acquiring 27 individual samples in DDA mode with the same LC gradient. In total 51 DDA acquisition raw files were analyzed via Spectronaut with Pulsar search engine (SN14.10.201222) to build the library. The Uniprot human proteome database was used (UP000005640, 96,797 entries) and the search parameters were set as default but included an additional deamidation (NQ) in variable modifications. DIA files were processed via Spectronaut using the default settings with precursor and protein FDR cut-off set to 0.01 , data filtering set to 0.5 percentile Q-value with global imputing, and normalization strategy set to global normalization on a median.

[0433] Peptides containing missing intensities in one or more of the samples and duplicated instances were filtered. The resulting peptide intensities were normalised using the median approach and batch-corrected with the use of the plate information using in-house and the data.table libraries of the R programming language (https: / / www.r-project.org / ). To find differentially expressed proteins, a statistical analysis was carried out using the Bioconductor library limma (Ritchie et al., Nucleic Acids Research 43: e47, 2015). FDR-adjusted p-values (q-values) using the Benjamini and Hochberg’s method are presented, and discoveries were assigned based on a 5 % FDR threshold and Log2FC < -0.5 or > 0.5. Volcano plots were generated in GraphPad Prism 9 and Venn diagrams using Venny 2.1 (bioinfogp.cnb.csic.es / tools / venny / ).

[0434] The lists of MYC transcriptional targets were obtained by extracting SLAM-seq, RNA- seq and expression microarray information from Muhar et al. (Muhar et al., Science 360: 800-805, 2018) (SLAM-seq after auxin-induced degradation of MYC-AID in HCT116 cells), Lorenzin et al. (Lorenzin et al., Elife 5: e15161 , 2016) (RNA-seq in U2OS cells treated with MYC siRNA) and Satoh et al. (Satoh et al., PNAS 114: E7697-E7706, 2017) (expression microarray in HCT116 cells treated with 2 siRNAs targeting MYC). Data integration with the proteomics dataset and gene identifier conversion were performed using the data.table and g:Profiler libraries of the R programming language. We considered a gene to be differentially expressed at the transcript level if its adjusted p-value < 0.05 and Log2FC < -0.5 or > 0.5 upon MYC modulation. When comparing the protein expression signatures to the reported mRNA signatures, the direction of change in the mass spectrometry measurements must be the same as in the transcriptomics experiments. MYC protein- protein interactions information was retrieved from the BioGRID database (thebiogrid.org).

[0435] Example 3 - Identification of HuR and K-ras degraders

[0436] To validate the approach taken in Example 2 to identify MYC degraders, it was repeated in an attempt to identify degraders of human antigen R (HuR). A VHH against HuR was identified and used for the screen encompassing a panel of -70 E3 ligase fusions (SEQ ID NOs: 1 , 2, 4-6, 8-10, 16-18, 21-23, 25, 27, 29, 30-32, 34, 35 and 53-100). Subsequently, top hits were confirmed via Western blotting against HuR. Following identification of E3 ligase- HuR VHH fusions that degraded HuR, screening against an additional target was undertaken. A single KRAS binder, DARPin K19 (Bery et al., Nature Communications 10: 2607, 2019) was used for this screen, along with the following E3 ligases: KLHL21 (1-240) (SEQ ID NO: 67), TRIM31 (1-343) (SEQ ID NO: 87), TNFRSF1 A (356-455) (SEQ ID NO: 84), SPSB1 (232-273) (SEQ ID NO: 82) and SPOP (167-374) (SEQ ID NO: 25). BioPROTACS were generated containing DARPin K19 with each of these E3 ligases, in each case joined by linker L1 (SEQ ID NO: 120).

[0437] BioPROTAC constructs were made and in vitro transcribed as described above, and the mRNA transfected into HCT116 cells. KRAS degradation was assessed by immunoblotting (Figure 21A). Of these, only K19-L1-SPSB1 was found to successfully degrade KRAS. The effect of bioPROTAC expression on KRAS levels was quantified by densitometry, showing that K19-L1-SPSB1 reduced KRAS levels by more than 50 % (Figure 21 B).

[0438] Immunoblotting was performed using mouse anti-human KRAS antibody 2C1 (Absolute Biotech, catalogue number LS-C175665), rabbit anti-HAtag antibody C29F4 (Cell Signalling Technology (CST), catalogue number 3724) and the anti-vinculin antibody used in Example 2; anti-mouse IgG or anti-rabbit IgG secondary antibodies conjugated to HRP (CST) were used for detection. All primary antibodies were used at 1 :1000 dilution and secondary antibodies at 1 :2000 dilution.

[0439] Sequences

[0440] Upper case sequences are protein sequences. Lower case sequences are nucleic acid sequences.

[0441] SEQ ID NO: 1 - HUWE1 (3810-4374)

[0442] MDVDQPSPSAQDTQSIASDGTPQGEKEKEERPPELPLLSEQLSLDELWDMLGECLKELEE

[0443] SHDQHAVLVLQPAVEAFFLVHATERESKPPVRDTRESQLAHIKDEPPPLSPAPLTPATPSSLD PFFSREPSSMHISSSLPPDTQKFLRFAETHRTVLNQILRQSTTHLADGPFAVLVDYIRVLDFD

[0444] VKRKYFRQELERLDEGLRKEDMAVHVRRDHVFEDSYRELHRKSPEEMKNRLYIVFEGEEG QDAGGLLREWYMIISREMFNPMYALFRTSPGDRVTYTINPSSHCNPNHLSYFKFVGRIVAKA VYDN RLLECYFTRSFYKH I LG KSVRYTDMESEDYH FYQG LVYLLENDVSTLG YDLTFSTEVQ EFGVCEVRDLKPNGANILVTEENKKEYVHLVCQMRMTGAIRKQLAAFLEGFYEIIPKRLISIFT

[0445] EQELELLISGLPTIDIDDLKSNTEYHKYQSNSIQIQWFWRALRSFDQADRAKFLQFVTGTSKV PLQGFAALEGMNGIQKFQIHRDDRSTDRLPSAHTCFNQLDLPAYESFEKLRHMLLLAIQECS

[0446] EGFGLA

[0447] SEQ ID NO: 2 - HUWE1 (3993-4374)

[0448] DFDVKRKYFRQELERLDEGLRKEDMAVHVRRDHVFEDSYRELHRKSPEEMKNRLYIVFEG EEGQDAGGLLREWYMIISREMFNPMYALFRTSPGDRVTYTINPSSHCNPNHLSYFKFVGRI

[0449] VAKAVYDNRLLECYFTRSFYKHILGKSVRYTDMESEDYHFYQGLVYLLENDVSTLGYDLTFS TEVQEFGVCEVRDLKPNGANILVTEENKKEYVHLVCQMRMTGAIRKQLAAFLEGFYEIIPKR

[0450] LISIFTEQELELLISGLPTIDIDDLKSNTEYHKYQSNSIQIQWFWRALRSFDQADRAKFLQFVT GTSKVPLQGFAALEGMNGIQKFQIHRDDRSTDRLPSAHTCFNQLDLPAYESFEKLRHMLLLA

[0451] IQECSEGFGLA

[0452] SEQ ID NO: 3 - lpaH9.8 (254-545)

[0453] LADAVTAWFPENKQSDVSQIWHAFEHEEHANTFSAFLDRLSDTVSARNTSGFREQVAAWL EKLSASAELRQQSFAVAADATESCEDRVALTWNNLRKTLLVHQASEGLFDNDTGALLSLGR EMFRLEILEDIARDKVRTLHFVDEIEVYLAFQTMLAEKLQLSTAVKEMRFYGVSGVTANDLRT

[0454] AEAMVRSREENEFTDWFSLWGPWHAVLKRTEADRWAQAEEQKYEMLENEYPQRVADRLK ASGLSGDADAEREAGAQVMRETEQQIYRQLTDEVLALRLPENGSQLHHS

[0455] SEQ ID NO: 4 - NEDD4 (937-1319)

[0456] YSRDYKRKYEFFRRKLKKQNDIPNKFEMKLRRATVLEDSYRRIMGVKRADFLKARLWIEFDG EKGLDYGGVAREWFFLISKEMFNPYYGLFEYSATDNYTLQINPNSGLCNEDHLSYFKFIGRV AGMAVYHGKLLDGFFIRPFYKMMLHKPITLHDMESVDSEYYNSLRWILENDPTELDLRFIIDE ELFGQTHQHELKNGGSEIWTNKNKKEYIYLVIQWRFVNRIQKQMAAFKEGFFELIPQDLIKIF DENELELLMCGLGDVDVNDWREHTKYKNGYSANHQVIQWFWKAVLMMDSEKRIRLLQFVT GTSRVPMNGFAELYGSNGPQSFTVEQWGTPEKLPRAHTCFNRLDLPPYESFEELWDKLQM

[0457] AIENTQGFDGVD SEQ ID NO: 5 - UBR5 (2462-2799)

[0458] LDLGLVDSSEKVQQENRKRHGSSRSWDMDLDDTDDGDDNAPLFYQPGKRGFYTPRPGK

[0459] NTEARLNCFRNIGRILGLCLLQNELCPITLNRHVIKVLLGRKVNWHDFAFFDPVMYESLRQLI

[0460] LASQSSDADAVFSAMDLAFAIDLCKEEGGGQVELIPNGVNIPVTPQNVYEYVRKYAEHRMLV

[0461] VAEQPLHAMRKGLLDVLPKNSLEDLTAEDFRLLVNGCGEVNVQMLISFTSFNDESGENAEKL

[0462] LQFKRWFWSIVEKMSMTERQDLVYFWTSSPSLPASEEGFQPMPSITIRPPDDQHLPTANTCI SRLYVPLYSSKQILKQKLLLAIKTKNFGFV

[0463] SEQ ID NO: 6 - PRKN (221-465)

[0464] ETSVALHLIATNSRNITCITCTDVRSPVLVFQCNSRHVICLDCFHLYCVTRLNDRQFVHDPQL

[0465] GYSLPCVAGCPNSLIKELHHFRILGEEQYNRYQQYGAEECVLQMGGVLCPRPGCGAGLLP

[0466] EPDQRKVTCEGGNGLGCGFAFCRECKEAYHEGECSAVFEASGTTTQAYRVDERAAEQAR

[0467] WEAASKETIKKTTKPCPRCHVPVEKNGGCMHMKCPQPQCRLEWCWNCGCEWNRVCMG DHWFDV

[0468] SEQ ID NO: 7 - ICPO (1-241)

[0469] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0470] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA

[0471] VCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNAKLVYLIVGVTPSGSFSTIPIVN

[0472] DPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDDDDLDD

[0473] SEQ ID NO: 8 - IDOL (280-445)

[0474] VTSAVMMQYSRDLKGHLASLFLNENINLGKKYVFDIKRTSKEVYDHARRALYNAGWDLVSR

[0475] NNQSPSHSPLKSSESSMNCSSCEGLSCQQTRVLQEKLRKLKEAMLCMVCCEEEINSTFCP

[0476] CGHTVCCESCAAQLQSCPVCRSRVEHVQHVYLPTHTSLLNLTVI

[0477] SEQ ID NO: 9 - LNX1 (1-167)

[0478] MNQPESANDPEPLCAVCGQAHSLEENHFYSYPEEVDDDLICHICLQALLDPLDTPCGHTYC

[0479] TLCLTNFLVEKDFCPMDRKPLVLQHCKKSSILVNKLLNKLLVTCPFREHCTQVLQRCDLEHH

[0480] FQTSCKGASHYGLTKDRKRRSQDGCPDGCASLTATAPSPEVSAA

[0481] SEQ ID NO: 10 - TRIM21 (1-286)

[0482] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL

[0483] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK

[0484] HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHAE

[0485] FVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSALEL

[0486] LQEVIIVLERSESWNLKDLDITSPELRSVCHVPGLKKMLRTCA

[0487] SEQ ID NO: 11 - LubX (1-130)

[0488] MGYRIEMATRNPFDIDHKSKYLREAALEANLSHPETTPTMLTCPIDSGFLKDPVITPEGFVYN

[0489] KSSILKWLETKKEDPQSRKPLTAKDLQPFPELLIIVNRFVETQTNYEKLKNRLVQNARVAARQ KEYT

[0490] SEQ ID NO: 12 - gTrCP (1-261 / 11 KR)

[0491] MDPAEAVLQERALRFMCSMPRSLWLGCSSLADSMPSLRCLYNPGTGALTAFQNSSEREDC NNGEPPRRIIPERNSLRQTYNSCARLCLNQETVCLASTAMRTENCVARTRLANGTSSMIVPR QRRLSASYEKEKELCVKYFEQWSESDQVEFVEHLISQMCHYQHGHINSYLKPMLQRDFITA LPARGLDHIAENILSYLDARSLCAAELVCKEWYRVTSDGMLWKKLIERMVRTDSLWRGLAER

[0492] RGWGQYLFKNRPPDGN

[0493] SEQ ID NO: 13 - VHH a-MDM2 (A225)

[0494] MAEVQLQASGGGFVQPGGSLRLSCAASGFTSKNDSMGWFRQAPGKEREFVSAISESHDG AVYYADSVKGRFTISRDNSKNTVYLQMNSLRAEDTATYYCAWQRGDSPEWAIWMWYWGQ

[0495] GTQVTVSS

[0496] SEQ ID NO: 14 - VHH a-MDM2 (A229)

[0497] MAEVQLQASGGGFVQPGGSLRLSCAASGDTSKIEIMGWFRQAPGKEREFVSAISRSETMF NYYADSVKGRFTISRDNSKNTVYLQMNSLRAEDTATYYCAAEKAHGIWWYPPLWNYWGQG

[0498] TQVTVSS

[0499] SEQ ID NO: 15 - AnkB / LeqAU13 (1-53)

[0500] MKKNFFSDLPEETIVNTLSFLKANTLARIAQTCQFFNRLANDKHLELHQLRQQ

[0501] SEQ ID NO: 16 - CSA (1-32)

[0502] M LG FLSARQTG LEDPLRLRRAESTRRVLG LEL

[0503] SEQ ID NO: 17 - DDB2 (1-115)

[0504] MAPKKRPETQKTSEIVLRPRNKRSRSPLELEPEAKKLCAKGSGPSRRCDSDCLWVGLAGP

[0505] QILPPCRSIVRTLHQHKLGRASWPSVQQGLQQSFLHTLDSYRILQKAAPFDRRAT

[0506] SEQ ID NO: 18 - DDB2

[0507] MAPKKRPETQKTSEIVLRPRNKRSRSPLELEPEAKKLCAKGSGPSRRCDSDCLWVGLAGP

[0508] QILPPCRSIVRTLHQHKLGRASWPSVQQGLQQSFLHTLDSYRILQKAAPFDRRATSLAWHP

[0509] THPSTVAVGSKGGDIMLWNFGIKDKPTFIKGIGAGGSITGLKFNPLNTNQFYASSMEGTTRL QDFKGNILRVFASSDTINIWFCSLDVSASSRMWTGDNVGNVILLNMDGKELWNLRMHKKK

[0510] VTHVALNPCCDWFLATASVDQTVKIWDLRQVRGKASFLYSLPHRHPVNAACFSPDGARLLT TDQKSEIRVYSASQWDCPLGLIPHPHRHFQHLTPIKAAWHPRYNLIWGRYPDPNFKSCTPY

[0511] ELRTIDVFDGNSGKMMCQLYDPESSGISSLNEFNPMGDTLASAMGYHILIWSQEEARTRK

[0512] SEQ ID NO: 19 - E4orf6 (1-139)

[0513] MQRDRRYRYRLAPYNKYQLPPCEEQSKATLSTSENSLWPECNSLTLHNVSEVRGIPSCVGF TVLQEWPIPWDMILTDYEMFILKKYMSVCMCCATINVEVTQLLHGHERWLIHCHCQRPGSL QCMSAGMLLGRWFKMAV

[0514] SEQ ID NO: 20 - E7

[0515] MHGDTPTLHEYMLDLQPETTDLYSYQQLNDSSEEEDEIDGPAGQAEPDRAHYNIVTFCCKC DSTLRLCVQSTHVDIRTLEDLLMGTLGIVCPICSQKP

[0516] SEQ ID NO: 21 - FBXO17 (1-98)

[0517] MGARLSRRRLPADPSLALDALPPELLVQVLSHVPPRSLVTRCRPVCRAWRDIVDGPTVWLL QLARDRSAEGRALYAVAQRCLPSNEDKEEFPLCALAR

[0518] SEQ ID NO: 22 - FBXO17 MGARLSRRRLPADPSLALDALPPELLVQVLSHVPPRSLVTRCRPVCRAWRDIVDGPTVWLL

[0519] QLARDRSAEGRALYAVAQRCLPSNEDKEEFPLCALARYCLRAPFGRNLIFNSCGEQGFRG

[0520] WEVEHGGNGWAIEKNLTPVPGAPSQTCFVTSFEWCSKRQLVDLVMEGVWQELLDSAQIEI CVADWWGARENCGCVYQLRVRLLDVYEKEWKFSASPDPVLQWTERGCRQVSHVFTNFG KGIRYVSFEQYGRDVSSWVGHYGALVTHSSVRVRIRLS

[0521] SEQ ID NO: 23 - KEAP1 (1-300)

[0522] MQPDPRPSGAGACCRFLPLQSQCPEGAGDAVMYASTECKAEVTPSQHGNRTFSYTLEDH

[0523] TKQAFGIMNELRLSQQLCDVTLQVKYQDAPAAQFMAHKWLASSSPVFKAMFTNGLREQG

[0524] MEWSIEGIHPKVMERLIEFAYTASISMGEKCVLHVMNGAVMYQIDSWRACSDFLVQQLDP

[0525] SNAIGIANFAEQIGCVELHQRAREYIYMHFGEVAKQEEFFNLSHCQLVTLISRDDLNVRCESE

[0526] VFHACINWVKYDCEQRRFYVQALLRAVRCHSLTPNFLQMQLQKCEILQSDSRCKDY

[0527] SEQ ID NO: 24 - LeqU1 (1-87)

[0528] MKAKYDPTKPGLQKLPPEIKVMILEFLDAKSKLALSQTNYGWRDLILDRPEYTKEITNTLFRL

[0529] DKKRHRQAIAQMMSGRVTASSMAK

[0530] SEQ ID NO: 25 - SPOP (167-374)

[0531] SVNISGQNTMNMVKVPECRLADELGGLWENSRFTDCCLCVAGQEFQAHKAILAARSPVFS AMFEHEMEESKKNRVEINDVEPEVFKEMMCFIYTGKAPNLDKMADDLLAAADKYALERLKV MCEDALCSNLSVENAAEILILADLHSADQLKTQAVDFINYHASDVLETSGWKSMWSHPHL.V

[0532] AEAYRSLASAQCPFLGPPRKRLKQS

[0533] SEQ ID NO: 26 - SV5-V

[0534] MDPTDLSFSPDEINKLIETGLNTVEYFTSQQVTGTSSLGKNTIPPGVTGLLTNAAEAKIQEST

[0535] NHQKGSVGGGAKPKKPRPKIAIVPADDKTVPGKPIPRPLLGLDSTPSTQTVLDLSGKTLPSG

[0536] SYKGVKLAKFGKENLMTRFIEEPRENPIATSSPIDFKRGRDTGGFHAREYSIGWVGDEVKVT

[0537] EWCNPSCSPITAAARRFECTCHQCPVTCSECERDT

[0538] SEQ ID NO: 27 - VHL

[0539] MPRRAENWDEAEVGAEEAGVEEYGPEEDGGEESGAEESGPEESGPEELGAEEEMEVGR

[0540] PRPVLRSVNSREPSQVIFCNRSPRWLPVWLNFDGEPQPYPTLPPGTGRRIHSYRGHLWLF

[0541] RDAGTHDGLLVNQTELFVPSLNVDGQPIFANITLPVYTLKERCLQWRSLVKPENYRRLDIVR

[0542] SLYEDLEDHPNVQKDLERLTQERIAHQRMGD

[0543] SEQ ID NO: 28 - Vif (80-156)

[0544] HLGQGVSIEWRKKRYSTQVDPDLADQLIHLHYFDCFSESAIRNTILGRIVSPRCEYQAGHNK

[0545] VGSLQYLALAALIKP

[0546] SEQ ID NO: 29 - BTrCP (1-261)

[0547] MDPAEAVLQEKALKFMCSMPRSLWLGCSSLADSMPSLRCLYNPGTGALTAFQNSSEREDC

[0548] NNGEPPRKIIPEKNSLRQTYNSCARLCLNQETVCLASTAMKTENCVAKTKLANGTSSMIVPK

[0549] QRKLSASYEKEKELCVKYFEQWSESDQVEFVEHLISQMCHYQHGHINSYLKPMLQRDFITA LPARGLDHIAENILSYLDAKSLCAAELVCKEWYRVTSDGMLWKKLIERMVRTDSLWRGLAER RGWGQYLFKNKPPDGN MLQRDFITALPARGLDHIAENILSYLDAKSLCAAELVCKEWYRVTSDGMLWKKLIERMVRTD

[0550] SLWRGLAERRGWGQYLFKNKPPDGNAPPNSFYRALYPKIIQDIETIESNWRCGRHSL

[0551] SEQ ID NO: 31 - SelK (79-91)

[0552] HLRGPSPPPMAGG

[0553] SEQ ID NO: 32 - SelK (87-91)

[0554] PMAGG

[0555] SEQ ID NO: 33 - Vpu (49-58)

[0556] AEDSGNESEG

[0557] SEQ ID NO: 34 - CKS1 (1-74)

[0558] MSHKQIYYSDKYDDEEFEYRHVMLPKDIAKLVPKTHLMSESEWRNLGVQQSQGWVHYMIH

[0559] EPEPHILLFRRPLP

[0560] SEQ ID NO: 35 - DDA1

[0561] MADFLKGLPVYNKSNFSRFHADSVCKASNRRPSVYLPTREYPSEQIIVTEKTNILLRYLHQQ

[0562] WDKKNAAKKRDQEQVELEGESSAPPRKVARTDSPDMHEDT

[0563] SEQ ID NO: 36 - E6 (8-158)

[0564] MFQDPQERPRKLPQLCTELQTTIHDIILECVYCKQQLLRREVYDFAERDLCIVYRDGNPYAV

[0565] CDKCLKFYSKISEYRHYSYSLYGTTLEQQYNKPLSDLLIRCINCQKPLSPEEKQRHLDKKQR

[0566] FHNIRGRWTGRCMSCSRSSRTRRETQL

[0567] SEQ ID NO: 37 - Vpr

[0568] MEQAPEDQGPQREPYNEWTLELLEELKSEAVRHFPRGGGSGGGGSGGGGSGDTWAGVE

[0569] AIIRILQQLLFIHFRIGCRHSRIGVTRQRRARNGASRS

[0570] SEQ ID NO: 38 - ICPO (1-158)

[0571] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0572] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA

[0573] VCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNA

[0574] SEQ ID NO: 39 - ICPO (115-158)

[0575] MVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNA

[0576] SEQ ID NO: 40 - ICPO (1-241 / 1 KR)

[0577] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0578] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA

[0579] VCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNARLVYLIVGVTPSGSFSTIPIV

[0580] NDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDDDDLD

[0581] D

[0582] SEQ ID NO: 41 - ICPO (115-241 / 1 KR)

[0583] MVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNARLVYLIVGVTPSGSFS

[0584] TIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDD

[0585] DDLDD

[0586] BLANK PAGE

[0587] SEQ ID NO: 42 - HUWE1 (3993-4374 / 5KR)

[0588] DFDVKRRYFRQELERLDEGLRREDMAVHVRRDHVFEDSYRELHRRSPEEMKNRLYIVFEG EEGQDAGGLLREWYMIISREMFNPMYALFRTSPGDRVTYTINPSSHCNPNHLSYFKFVGRI

[0589] VAKAVYDNRLLECYFTRSFYKHILGKSVRYTDMESEDYHFYQGLVYLLENDVSTLGYDLTFS TEVQEFGVCEVRDLKPNGANILVTEENKREYVHLVCQMRMTGAIRRQLAAFLEGFYEIIPKR

[0590] LISIFTEQELELLISGLPTIDIDDLKSNTEYHKYQSNSIQIQWFWRALRSFDQADRAKFLQFVT GTSKVPLQGFAALEGMNGIQKFQIHRDDRSTDRLPSAHTCFNQLDLPAYESFEKLRHMLLLA IQECSEGFGLA

[0591] SEQ ID NO: 43 - IDOL (280-445 / 9KR)

[0592] VTSAVMMQYSRDLRGHLASLFLNENINLGRRYVFDIRRTSREVYDHARRALYNAGWDLVS RNNQSPSHSPLRSSESSMNCSSCEGLSCQQTRVLQERLRRLREAMLCMVCCEEEINSTFC PCGHTVCCESCAAQLQSCPVCRSRVEHVQHVYLPTHTSLLNLTVI

[0593] SEQ ID NO: 44 - NEDD4 (937- 1319 / 16KR)

[0594] YSRDYRRKYEFFRRRLRKQNDIPNRFEMRLRRATVLEDSYRRIMGVRRADFLKARLWIEFD GERGLDYGGVAREWFFLISKEMFNPYYGLFEYSATDNYTLQINPNSGLCNEDHLSYFRFIGR VAGMAVYHGKLLDGFFIRPFYKMMLHKPITLHDMESVDSEYYNSLRWILENDPTELDLRFIID EELFGQTHQHELKNGGSEIWTNRNKREYIYLVIQWRFVNRIQRQMAAFREGFFELIPQDLIK IFDENELELLMCGLGDVDVNDWREHTRYKNGYSANHQVIQWFWRAVLMMDSERRIRLLQF

[0595] VTGTSRVPMNGFAELYGSNGPQSFTVEQWGTPERLPRAHTCFNRLDLPPYESFEELWDKL QMAIENTQGFDGVD

[0596] SEQ ID NO: 45 - PRKN (221-465 / 7KR)

[0597] ETSVALHLIATNSRNITCITCTDVRSPVLVFQCNSRHVICLDCFHLYCVTRLNDRQFVHDPQL GYSLPCVAGCPNSLIRELHHFRILGEEQYNRYQQYGAEECVLQMGGVLCPRPGCGAGLLP EPDQRKVTCEGGNGLGCGFAFCRECREAYHEGECSAVFEASGTTTQAYRVDERAAEQAR

[0598] WEAASRETIRRTTKPCPRCHVPVERNGGCMHMRCPQPQCRLEWCWNCGCEWNRVCMG DHWFDV

[0599] SEQ ID NO: 46 - TRIM21 (1-286 / 16KR)

[0600] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL RNLRPNRQLANMVNNLREISQEAREGTQGERCAVHGERLHLFCERDGRALCWVCAQSRR HRDHAMVPLEEAAQEYQEKLQVALGELRRRQELAERLEVEIAIRRADWKRTVETQRSRIHA

[0601] EFVQQRNFLVEEEQRQLQELERDEREQLRILGEREARLAQQSQALQELISELDRRCHSSAL ELLQEVIIVLERSESWNLRDLDITSPELRSVCHVPGLKRMLRTCA

[0602] SEQ ID NO: 47 - UBR5 (2462-2799 / 7KR)

[0603] LDLGLVDSSEKVQQENRKRHGSSRSWDMDLDDTDDGDDNAPLFYQPGKRGFYTPRPGR NTEARLNCFRNIGRILGLCLLQNELCPITLNRHVIKVLLGRKVNWHDFAFFDPVMYESLRQLI LASQSSDADAVFSAMDLAFAIDLCKEEGGGQVELIPNGVNIPVTPQNVYEYVRRYAEHRMLV VAEQPLHAMRRGLLDVLPRNSLEDLTAEDFRLLVNGCGEVNVQMLISFTSFNDESGENAEK LLQFKRWFWSIVERMSMTERQDLVYFWTSSPSLPASEEGFQPMPSITIRPPDDQHLPTANT

[0604] CISRLYVPLYSSKQILRQKLLLAIKTRNFGFV SEQ ID NO: 48 - SPOP (167-374 / 3KR)

[0605] SVNISGQNTMNMVKVPECRLADELGGLWENSRFTDCCLCVAGQEFQAHKAILAARSPVFS

[0606] AMFEHEMEESKKNRVEINDVEPEVFKEMMCFIYTGRAPNLDKMADDLLAAADKYALERLKV

[0607] MCEDALCSNLSVENAAEILILADLHSADQLKTQAVDFINYHASDVLETSGWKSMWSHPHLV

[0608] AEAYRSLASAQCPFLGPPRRRLRQS

[0609] SEQ ID NO: 49 - VHH Q-MDM2 (A225) (2KR)

[0610] MAEVQLQASGGGFVQPGGSLRLSCAASGFTSRNDSMGWFRQAPGKEREFVSAISESHDG

[0611] AVYYADSVKGRFTISRDNSRNTVYLQMNSLRAEDTATYYCAWQRGDSPEWAIWMWYWGQ

[0612] GTQVTVSS

[0613] SEQ ID NO: 50 - VHH a-MDM2 (A225) (4KR)

[0614] MAEVQLQASGGGFVQPGGSLRLSCAASGFTSRNDSMGWFRQAPGREREFVSAISESHDG

[0615] AVYYADSVRGRFTISRDNSRNTVYLQMNSLRAEDTATYYCAWQRGDSPEWAIWMWYWGQ

[0616] GTQVTVSS

[0617] SEQ ID NO: 51 - VHH a-MDM2 (A229) (2KR)

[0618] MAEVQLQASGGGFVQPGGSLRLSCAASGDTSRIEIMGWFRQAPGKEREFVSAISRSETMF

[0619] NYYADSVKGRFTISRDNSKNTVYLQMNSLRAEDTATYYCAAERAHGIWWYPPLWNYWGQ

[0620] GTQVTVSS

[0621] SEQ ID NO: 52 - VHH a-MDM2 (A229) (5KR)

[0622] MAEVQLQASGGGFVQPGGSLRLSCAASGDTSRIEIMGWFRQAPGREREFVSAISRSETMF

[0623] NYYADSVRGRFTISRDNSRNTVYLQMNSLRAEDTATYYCAAERAHGIWWYPPLWNYWGQ

[0624] GTQVTVSS

[0625] SEQ ID NO: 53 -

[0626] RMKGQEFVDEIQGRYPHLLEQLLSTSDTTGEENADPPIIHFGPGESSSEDAVMMNTPWKS

[0627] ALEMGFNRDLVKQTVQSKILTTGENYKTVNDIVSALLNAEDEKREEEKEKQAEEMASDDLSL

[0628] IRKNRMALFQQLTCVLPILDNLLKANVINKQEHDIIKQKTQIPLQARELIDTILVKGNAAANIFKN

[0629] CLKEIDSTLYKNLFVDKNMKYIPTEDVSGLSLEEQLRRLQEERTCKVCMDKEVSWFIPCGH

[0630] LVVCQECAPSLRKCPICRGIIKGTVRTFLS

[0631] SEQ ID NO: 54 - BIRC4 (337-497)

[0632] EYINNIHLTHSLEECLVRTTEKTPSLTRRIDDTIFQNPMVQEAIRMGFSFKDIKKIMEEKIQISG

[0633] SNYKSLEVLVADLVNAQKDSMQDESSQTSLQKEISTEEQLRRLQEEKLCKICMDRNIAIVFVP

[0634] CGHLVTCKQCAEAVDKCPMCYTVITFKQKIFMS

[0635] SEQ ID NO: 55 - BIRC7 (53-298)

[0636] GQILGQLRPLTEEEEEEGAGATLSRGPAFPGMGSEELRLASFYDWPLTAEVPPELLAAAGFF

[0637] HTGHQDKVRCFFCYGGLQSWKRGDDPWTEHAKWFPSCQFLLRSKGRDFVHSVQETHSQ

[0638] LLGSWDPWEEPEDAAPVAPSVPASGYPELPTPRREVQSESAQEPGGVSPAEAQRAWWVL

[0639] EPPGARDVEAQLRRLQEERTCKVCLDRAVSIVFVPCGHLVCAECAPGLQLCPICRAPVRSR

[0640] VRTFLS

[0641] SEQ ID NO: 56 - Cbl (358-436) QDHIKVTQEQYELYCEMGSTFQLCKICAENDKDVKIEPCGHLMCTSCLTSWQESEGQGCPF

[0642] CRCEIKGTEPIWDPFDP

[0643] SEQ ID NO: 57 - FBXO1 (28-303)

[0644] RNLTILSLPEDVLFHILKWLSVEDILAVRAVHSQLKDLVDNHASVWACASFQELWPSPGNLKL FERAAEKGNFEAAVKLGIAYLYNEGLSVSDEARAEVNGLKASRFFSLAERLNVGAAPFIWL.FI RPPWSVSGSCCKAWHESLRAECQLQRTHKASILHCLGRVLSLFEDEEKQQQAHDLFEEA

[0645] AHQGCLTSSYLLWESDRRTDVSDPGRCLHSFRKLRDYAAKGCWEAQLSLAKACANANQLG LEVRASSEIVCQLFQASQAVSKQQVFSVQK

[0646] SEQ ID NO: 58 - DCAF11 (1-390)

[0647] MGSRNSSSAGSGSGDPSEGLPRRGAGLRRSEEEEEEDEDVDLAQVLAYLLRRGQVRLVQ GGGAANLQFIQALLDSEEENDRAWDGRLGDRYNPPVDATPDTRELEFNEIKTQVELATGQL

[0648] GLRRAAQKHSFPRMLHQRERGLCHRGSFSLGEQSRVISHFLPNDLGFTDSYSQKAFCGIY

[0649] SKDGQIFMSACQDQTIRLYDCRYGRFRKFKSIKARDVGWSVLDVAFTPDGNHFLYSSWSDY IHICNIYGEGDTHTALDLRPDERRFAVFSIAVSSDGREVLGGANDGCLYVFDREQNRRTLQIE SHEDDVNAVAFADISSQILFSGGDDAICKVWDRRTMREDDPKPVGALAGHQDGITFIDSKGD

[0650] ARYLISNSKDQTIKLWDIRRFSSR

[0651] SEQ ID NO: 59 - DCAF11 (365-546)

[0652] G DARYLI SN SKDQTI KLWD I RRFSSREG M EASRQAATQQ N WDYRWQQ VPKKAWRKLKLPG DSSLMTYRGHGVLHTLIRCRFSPIHSTGQQFIYSGCSTGKVWYDLLSGHIVKKLTNHKACV RDVSWHPFEEKIVSSSWDGNLRLWQYRQAEYFQDDMPESEECASAPAPVPQSSTPFSSP Q

[0653] SEQ ID NO: 60 - DCAF15 (32-474)

[0654] RREHVLKQLERVKISGQLSPRLFRKLPPRVCVSLKNIVDEDFLYAGHIFLGFSKCGRYVLSYT SSSGDDDFSFYIYHLYWWEFNVHSKLKLVRQVRLFQDEEIYSDLYLTVCEWPSDASKVIVFG FNTRSANGMLMNMMMMSDENHRDIYVSTVAVPPPGRCAACQDASRAHPGDPNAQCLRH GFMLHTKYQWYPFPTFQPAFQLKKDQWLLNTSYSLVACAVSVHSAGDRSFCQILYDHSTC

[0655] PLAPASPPEPQSPELPPALPSFCPEAAPARSSGSPEPSPAIAKAKEFVADIFRRAKEAKGGV PEEARPALCPGPSGSRCRAHSEPLALCGETAPRDSPPASEAPASEPGYVNYTKLYYVLESG EGTEPEDELEDDKISLPFWTDLRGRNLRPMRERTAVQGQYLTVEQLTLDFEYVINEVIRHD

[0656] ATWGHQFCSFSDY

[0657] SEQ ID NO: 61 - DCAF16

[0658] MGPRNPSPDHLSESESEEEENISYLNESSGEEWDSSEEEDSMVPNLSPLESLAWQVKCLL

[0659] KYSTTWKPLNPNSWLYHAKLLDPSTPVHILREIGLRLSHCSHCVPKLEPIPEWPPLASCGVP

[0660] PFQKPLTSPSRLSRDHATLNGALQFATKQLSRTLSRATPIPEYLKQIPNSCVSGCCCGWLTK TVKETTRTEPINTTYSYTDFQKAVNKLLTASL

[0661] SEQ ID NO: 62 - DDB1 (391-709)

[0662] RNGIGIHEHASIDLPGIKGLWPLRSDPNRETDDTLVLSFVGQTRVLMLNGEEVEETELMGFV

[0663] DDQQTFFCGNVAHQQLIQITSASVRLVSQEPKALVSEWKEPQAKNISVASCNSSQWVAVG

[0664] RALYYLQIHPQELRQISHTEMEHEVACLDITPLGDSNGLSPLCAIGLWTDISARILKLPSFELL

[0665] HKEMLGGEIIPRSILMTTFESSHYLLCALGDGALFYFGLNIETGLLSDRKKVTLGTQPTVLRTF

[0666] RSLSTTNVFACSDRPTVIYSSNHKLVFSNVNLKEVNYMCPLNSDGYPDSLALANNSTLTIGTI DEIQK SEQ ID NO: 63 - DTX1 (410-620)

[0667] DCTICMERLVTASGYEGVLRHKGVRPELVGRLGRCGHMYHLLCLVAMYSNGNKDGSLQCP

[0668] TCKAIYGEKTGTQPPGKMEFHLIPHSLPGFPDTQTIRIVYDIPTGIQGPEHPNPGKKFTARGF

[0669] PRHCYLPNNEKGRKVLRLLITAWERRLIFTIGTSNTTGESDTWWNEIHHKTEFGSNLTGHG

[0670] YPDASYLDNVLAELTAQGVSEAAAKA

[0671] SEQ ID NO: 64 - FBXQ40 (570-624)

[0672] QNSLTSLPLEILKYIAGFLDSVSLAQLSQVSVLMRNICATLLQERGMVLLQWKKK

[0673] SEQ ID NO: 65 - FBXW5 (1-50)

[0674] MDEGGTPLLPDSLVYQIFLSLGPADVLAAGLVCRQWQAVSRDEFLWREQF

[0675] SEQ ID NO: 66 - FBXW5 (82-540)

[0676] CVEVQTLREHTDQVLHLSFSHSGYQFASCSKDCTVKIWSNDLTISLLHSADMRPYNWSYTQ

[0677] FSQFNKDDSLLLASGVFLGPHNSSSGEIAVISLDSFALLSRVRNKPYDVFGCWLTETSLISGN

[0678] LHRIGDITSCSVLWLNNAFQDVESENVNWKRLFKIQNLNASTVRTVMVADCSRFDSPDLLL

[0679] EAGDPATSPCRIFDLGSDNEEWAGPAPAHAKEGLRHFLDRVLEGRAQPQLSERMLETKVA

[0680] ELLAQGHTKPPERSATGAKSKYLIFTTGCLTYSPHQIGIKQILPHQMTTAGPVLGEGRGSDAF

[0681] FDALDHVIDIHGHIIGMGLSPDNRYLYVNSRAWPNGAWADPMQPPPIAEEIDLLVFDLKTMR

[0682] EVRRALRAHRAYTPNDECFFIFLDVSRDFVASGAEDRHGYIWDRHYNICLARLRHEDWNS

[0683] WFSPQEQELLLTASDDATIKAWRS

[0684] SEQ ID NO: 67 - KLHL21 (1-240)

[0685] MERPAPLAVLPFSDPAHALSLLRGLSQLRAERKFLDVTLEAAGGRDFPAHRAVLAAASPYFR

[0686] AMFAGQLRESRAERVRLHGVPPDMLQLLLDFSYTGRVAVSGDNAEPLLRAADLL.QFPAVKE

[0687] ACGAFLQQQLDLANCLDMQDFAEAFSCSGLASAAQRFILRHVGELGAEQLERLPLARLLRY

[0688] LRDDGLCVPKEEAAYQLALRWVRADPPRRAAHWPQLLEAVRLPFVRRFYLLAHVEA

[0689] SEQ ID NO: 68 - MARCH5 (1-80)

[0690] MPDQALQQMLDRSCWVCFATDEDDRTAEWVRPCRCRGSTKWVHQACLQRWVDEKQRG

[0691] NSTARVACPQCNAEYLIVFPKLG

[0692] SEQ ID NO: 69 - MARCH7 (543-616)

[0693] DSEEEEGDLCRICQMAAASSSNLLIEPCKCTGSLQYVHQDCMKKWLQAKINSGSSLEAVTT

[0694] CELCKEKLELNLE

[0695] SEQ ID NO: 70 - MIB1 (819-1006)

[0696] CMVCSDMKRDTLFGPCGHIATCSLCSPRVKKCLICKEQVQSRTKIEECWCSDKKAAVLFQ

[0697] PCGHMCACENCANLMKKCVQCRAWERRVPFIMCCGGKSSEDATDDISSGNIPVLQKDKD

[0698] NTNVNADVQKLQQQLQDIKEQTMCPVCLDRLKNMIFLCGHGTCQLCGDRMSECPICRKAIE

[0699] RRILLY

[0700] SEQ ID NO: 71 - RFPL2 (1-167)

[0701] MEVAELGFPETAVSQSRICLCAVLCGHWDFADMMVIRSLSLIRLEGVEGRDPVGGGNL.TNK

[0702] RPSCAPSPQDLSAQWKQLEDRGASSRRVDMAALFQEASSCPVCSDYLEKPMSLECGCAV

[0703] CLKCINSLQKEPHGEDLLCCCSSMVSRKNKIRRNRQLERLASHIKEL SEQ ID NO: 72 - RNF114 (1-68)

[0704] MAAQQRDCGGAAQLAGPAAEADPLGRFTCPVCLEVYEKPVQVPCGHVFCSACLQECLKP KKKPVCGVC

[0705] SEQ ID NO: 73 - RNF126 (228-311)

[0706] ECPVCKDDYALGERVRQLPCNHLFHDGCIVPWLEQHDSCPVCRKSLTGQNTATNPPGLTG

[0707] VSFSSSSSSSSSSSPSNENATSNS

[0708] SEQ ID NO: 74 - RNF149 (265-324)

[0709] DAENCAVCIENFKVKDIIRILPCKHIFHRICIDPWLLDHRTCPMCKLDVIKALGYWGEPG

[0710] SEQ ID NO: 75 - RNF181 (75-118)

[0711] LKCPVCLLEFEEEETAIEMPCHHLFHSSCILPWLSKTNSCPLCRY

[0712] SEQ ID NO: 76 - RNF187 (1-54)

[0713] MALPAGPAEAACALCQRAPREPVRADCGHRFCRACWRFWAEEDGPFPCPECAD

[0714] SEQ ID NO: 77 - SIAH1 (5-76)

[0715] TATALPTGTSKCPPSQRVPALTGTTASNNDLASLFECPVCFDYVLPPILQCQSGHLVCSNCR PKLTCCPTCR

[0716] SEQ ID NO: 78 - SIAH1 (5-282)

[0717] TATALPTGTSKCPPSQRVPALTGTTASNNDLASLFECPVCFDYVLPPILQCQSGHLVCSNCR PKLTCCPTCRGPLGSIRNLAMEKVANSVLFPCKYASSGCEITLPHTEKADHEELCEFRPYSC PCPGASCKWQGSLDAVMPHLMHQHKSITTLQGEDIVFLATDINLPGAVDWVMMQSCFGFH FMLVLEKQEKYDGHQQFFAIVQLIGTRKQAENFAYRLELNGHRRRLTWEATPRSIHEGIATAI

[0718] MNSDCLVFDTSIAQLFAENGNLGINVTISMC

[0719] SEQ ID NO: 79 - SKP1 (64-163)

[0720] HHKDDPPPPEDDENKEKRTDDIPVWDQEFLKVDQGTLFELILAANYLDIKGLLDVTCKTVAN

[0721] MIKGKTPEEIRKTFNIKNDFTEEEEAQVRKENQWCEEK

[0722] SEQ ID NO: 80 - SKP2 (2-150)

[0723] HRKHLQEIPDLSSNVATSFTWGWDSSKTSELLSGMGVSALEKEEPDSENIPQELLSNLGHP ESPPRKRLKSKGSDKDFVIVRRPKLNRENFPGVSWDSLPDELLLGIFSCLCLPELLKVSGVC KRWYRLASDESLWQTLDLTGKNLHPD

[0724] SEQ ID NO: 81 - SOCS3 (177-225)

[0725] VLSRPLSSNVATLQHLCRKTVNGHLDSYEKVTQLPGPIREFLDQYDAPL

[0726] SEQ ID NO: 82 - SPSB1 (232-273)

[0727] PEPLPLMDLCRRSVRLALGRERLGEIHTLPLPASLKAYLLYQ

[0728] SEQ ID NO: 83 - CHIP (128-303)

[0729] RLNFGDDIPSALRIAKKKRWNSIEERRIHQESELHSYLSRLIAAERERELEECQRNHEGDED DSHVRAQQACIEAKHDKYMADMDELFSQVDEKRKKRDIPDYLCGKISFELMREPCITPSGIT YDRKDIEEHLQRVGHFDPVTRSPLTQEQLIPNLAMKEVIDAFISENGWVEDY SEQ ID NO: 84 - TNFRSF1A (356-455)

[0730] PATLYAWENVPPLRWKEFVRRLGLSDHEIDRLELQNGRCLREAQYSMLATWRRRTPRREA

[0731] TLELLGRVLRDMDLLGCLEDIEEALCGPAALPPAPSLLR

[0732] SEQ ID NO: 85 - TRPC4AP (1-383)

[0733] MAAAPVAAGSGAGRGRRSAATVAAWGGWGGRPRPGNILLQLRQGQLTGRGLVRAVQFTE

[0734] TFLTERDKQSKWSGIPQLLLKLHTTSHLHSDFVECQNILKEISPLLSMEAMAFVTEERKLTQE

[0735] TTYPNTYIFDLFGGVDLLVEILMRPTISIRGQKLKISDEMSKDCLSILYNTCVCTEGVTKRLAE

[0736] KNDFVIFLFTLMTSKKTFLQTATLIEDILGVKKEMIRLDEVPNLSSLVSNFDQQQLANFCRILAV

[0737] TISEMDTGNDDKHTLLAKNAQQKKSLSLGPSAAEINQAALLSIPGFVERLCKLATRKVSEST GTASFLQELEEWYTWLDNALVLDALMRVANEESEHNQASIVFPPPGASEENGLPHTSARTQ LPQSMKIMH

[0738] SEQ ID NO: 86 - TRIM6 (1-281)

[0739] MTSPVLVDIREEVTCPICLELLTEPLSIDCGHSFCQACITPNGRESVIGQEGERSCPVCQTSY

[0740] QPGNLRPNRHLANIVRRLREWLGPGKQLKAVLCADHGEKLQLFCQEDGKVICWLCERSQ

[0741] EHRGHHTFLVEEVAQEYQEKFQESLKKLKNEEQEAEKLTAFIREKKTSWKNQMEPERCRIQ TEFNQLRNILDRVEQRELKKLEQEEKKGLRIIEEAENDLVHQTQSLRELISDLERRCQGSTM ELLQDVSDVTERSEFWTLRKPEALPTKLRSMFRAP

[0742] SEQ ID NO: 87 - TRIM31 (1-343)

[0743] MASGQFVNKLQEEVICPICLDILQKPVTIDCGHNFCLKCITQIGETSCGFFKCPLCKTSVRKN

[0744] AIRFNSLLRNLVEKIQALQASEVQSKRKEATCPRHQEMFHYFCEDDGKFLCFVCRESKDHK

[0745] SHNVSLIEEAAQNYQGQIQEQIQVLQQKEKETVQVKAQGVHRVDVFTDQVEHEKQRILTEF

[0746] ELLHQVLEEEKNFLLSRIYWLGHEGTEAGKHYVASTEPQLNDLKKLVDSLKTKQNMPPRQL

[0747] LEDIKWLCRSEEFQFLNPTPVPLELEKKLSEAKSRHDSITGSLKKFKDQLQADRKKDENRF

[0748] FKSMNKNDMKSWGLLQKNNHKMNKTSEPGSSSAGG

[0749] SEQ ID NO: 88 - TRIM22 (1-256)

[0750] MDFSVKVDIEKEVTCPICLELLTEPLSLDCGHSFCQACITAKIKESVIISRGESSCPVCQTRFQ

[0751] PGNLRPNRHLANIVERVKEVKMSPQEGQKRDVCEHHGKKLQIFCKEDGKVICWVCELSQE

[0752] HQGHQTFRINEWKECQEKLQVALQRLIKEDQEAEKLEDDIRQERTAWKNYIQIERQKILKGF

[0753] NEMRVILDNEEQRELQKLEEGEVNVLDNLAAATDQLVQQRQDASTLISDLQRRLRGSSVEM LQDVIDVM

[0754] SEQ ID NO: 89 - TRIM24 (1-360)

[0755] MEVAVEKAVAAAAAASAAASGGPSAAPSGENEAESRQGPDSERGGEAARLNLLDTCAVCH

[0756] QNIQSRAPKLLPCLHSFCQRCLPAPQRYLMLPAPMLGSAETPPPVPAPGSPVSGSSPFATQ

[0757] VGVIRCPVCSQECAERHIIDNFFVKDTTEVPSSTVEKSNQVCTSCEDNAEANGFCVECVEW

[0758] LCKTCIRAHQRVKFTKDHTVRQKEEVSPEAVGVTSQRPVFCPFHKKEQLKLYCETCDKLTC

[0759] RDCQLLEHKEHRYQFIEEAFQNQKVIIDTLITKLMEKTKYIKFTGNQIQNRIIEVNQNQKQVEQ

[0760] DIKVAIFTLMVEINKKGKALLHQLESLAKDHRMKLMQQQQEVAGLSKQLEHVM

[0761] SEQ ID NO: 90 - TRIM48 (1-148)

[0762] MNSGISQVFQRELTCPICMNYFIDPVTIDCGHSFCRPCFYLNWQDIPILTQCFECIKTIQQRNL

[0763] KTNIRLKKMASLARKASLWLFLSSEEQMCGIHRETKKMFCEVDRSLLCLLCSSSQEHRYHR

[0764] HCPAEWAAEEHWEKLLKKMQSLW SEQ ID NO: 91 - TRIM63 (1-269)

[0765] MDYKSSLIQDGNPMENLEKQLICPICLEMFTKPWILPCQHNLCRKCANDIFQAANPYWTSR

[0766] GSSVSMSGGRFRCPTCRHEVIMDRHGVYGLQRNLLVENIIDIYKQECSSRPLQKGSHPMCK EHEDEKINIYCLTCEVPTCSMCKVFGIHKACEVAPLQSVFQGQKTELNNCISMLVAGNDRVQ

[0767] TIITQLEDSRRVTKENSHQVKEELSQKFDTLYAILDEKKSELLQRITQEQEKKLSFIEALIQQYQ EQLDKSTKLVETAIQSLDE

[0768] SEQ ID NO: 92 - UHRF1 (724-793)

[0769] CICCQELVFRPITTVCQHNVCKDCLDRSFRAQVFSCPACRYDLGRSYAMQVNQPLQTVLNQ LFPGYGNGR

[0770] SEQ ID NO: 93 - WDR76 (311-626)

[0771] VTTGPIFSMALHPSETRTLVAVGAKFGQVGLCDLTQQPKEDGVYVFHPHSQPVSCLYFSPA

[0772] NPAHILSLSYDGTLRCGDFSRAIFEEVYRNERSSFSSFDFLAEDASTLIVGHWDGNMSLVDR

[0773] RTPGTSYEKLTSSSMGKIRTVHVHPVHRQYFITAGLRDTHIYDARRLNSRRSQPLISLTEHTK SIASAYFSPLTGNRWTTCADCNLRIFDSSCISSKIPLLTTIRHNTFTGRWLTRFQAMWDPKQ EDCVIVGSMAHPRRVEIFHETGKRVHSFGGEYLVSVCSINAMHPTRYILAGGNSSGKIHVFM

[0774] NEKSC

[0775] SEQ ID NO: 94 - TRIM21 (1-255)

[0776] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL

[0777] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK

[0778] HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHAE

[0779] FVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSALEL LQEVIIVLERSE

[0780] SEQ ID NO: 95 - UBR2 (1108-1755)

[0781] CILCQEEQEVKVESRAMVLAAFVQRSTVLSKNRSKFIQDPEKYDPLFMHPDLSCGTHTSSC

[0782] GHIMHAHCWQRYFDSVQAKEQRRQQRLRLHTSYDVENGEFLCPLCECLSNTVIPLLLPPR NIFNNRLNFSDQPNLTQWIRTISQQIKALQFLRKEESTPNNASTKNSENVDELQLPEGFRPD FRPKIPYSESIKEMLTTFGTATYKVGLKVHPNEEDPRVPIMCWGSCAYTIQSIERILSDEDKPL

[0783] FGPLPCRLDDCLRSLTRFAAAHWTVASVSWQGHFCKLFASLVPNDSHEELPCILDIDMFHL LVGLVLAFPALQCQDFSGISLGTGDLHIFHLVTMAHIIQILLTSCTEENGMDQENPPCEEESAV

[0784] LALYKTLHQYTGSALKEIPSGWHLWRSVRAGIMPFLKCSALFFHYLNGVPSPPDIQVPGTSH FEHLCSYLSLPNNLICLFQENSEIMNSLIESWCRNSEVKRYLEGERDAIRYPRESNKLINLPE

[0785] DYSSLINQASNFSCPKSGGDKSRAPTLCLVCGSLLCSQSYCCQTELEGEDVGACTAHTYSC GSGVGIFLRVRECQVLFLAGKTKGCFYSPPYLDDYGETDQGLRRGNPLHLCKERFKKIQKL WHQHSVTEEIGHAQEANQTLVGIDWQHL

[0786] SEQ ID NO: 96 - VHL (Y98N)

[0787] MPRRAENWDEAEVGAEEAGVEEYGPEEDGGEESGAEESGPEESGPEELGAEEEMEAGR PRPVLRSVNSREPSQVIFCNRSPRWLPVWLNFDGEPQPNPTLPPGTGRRIHSYRGHLWLF

[0788] RDAGTHDGLLVNQTELFVPSLNVDGQPIFANITLPVYTLKERCLQWRSLVKPENYRRLDIVR SLYEDLEDHPNVQKDLERLTQERIAHQRMGD

[0789] SEQ ID NO: 97 - VHL (152-213)

[0790] TLPVYTLKERCLQWRSLVKPENYRRLDIVRSLYEDLEDHPNVQKDLERLTQERIAHQRMGD SEQ ID NO: 98 - CRBN (1-317)

[0791] MAGEGDQQDAAHNMGNHLPLLPAESEEEDEMEVEDQDSKEAKKPNIINFDTSLPTSHTYL

[0792] GADMEEFHGRTLHDDDSCQVIPVLPQVMMILIPGQTLPLQLFHPQEVSMVRNLIQKDRTFAV

[0793] LAYSNVQEREAQFGTTAEIYAYREEQDFGIEIVKVKAIGRQRFKVLELRTQSDGIQQAKVQIL PECVLPSTMSAVQLESLNKCQIFPSKPVSREDQCSYKWWQKYQKRKFHCANLTSWPRWL YSLYDAETLMDRIKKQLREWDENLKDDSLPSNPIDFSYRVAACLPIDDVLRIQLLKIGSAIQRL

[0794] RCELDIMNK

[0795] SEQ ID NO: 99 - MDM2 (436-491)

[0796] EPCVICQGRPKNGCIVHGKTGHLMACFTCAKKLKKRNKPCPVCRQPIQMIVLTYFP

[0797] SEQ ID NO: 100 - UBE2D1

[0798] MALKRIQKELSDLQRDPPAHCSAGPVGDDLFHWQATIMGPPDSAYQGGVFFLTVHFPTDYP FKPPKIAFTTKIYHPNINSNGSICLDILRSQWSPALTVSKVLLSICSLLCDPNPDDPLVPDIAQIY KSDKEKYNRHAREWTQKYAM

[0799] SEQ ID NO: 101 - HI

[0800] NELKRAFAALRDQI

[0801] SEQ ID NO: 102 - DP20

[0802] MDLGKKLLEAARAGQDDEVRILMANGADVNAADDWGNTPLHLAALEGHLEIVEVLLKYGAD VNAQDLYGTTPLHLAAWVGHLEIVEVLLKNGADVNAMDWGETPLHLAAEMGHLEIVEVLLK HGADVNAQDKFGKTAFDISIDNGNEDLAEILQKL

[0803] SEQ ID NO: 103 - VH12

[0804] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSRWGMSWVRQAPGKGLEWVSYISHDGTF IYYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCARGIIPRDLVGRLLLFDYWGQGTL VTVSS

[0805] SEQ ID NO: 104 - DP20 (4KR)

[0806] MDLGKKLLEAARAGQDDEVRILMANGADVNAADDWGNTPLHLAALEGHLEIVEVLLRYGAD VNAQDLYGTTPLHLAAWVGHLEIVEVLLRNGADVNAMDWGETPLHLAAEMGHLEIVEVLLR HGADVNAQDRFGKTAFDISIDNGNEDLAEILQKL

[0807] SEQ ID NO: 105 - DP20 (7KR)

[0808] MDLGRRLLEAARAGQDDEVRILMANGADVNAADDWGNTPLHLAALEGHLEIVEVLLRYGA DVNAQDLYGTTPLHLAAWVGHLEIVEVLLRNGADVNAMDWGETPLHLAAEMGHLEIVEVLL

[0809] RHGADVNAQDRFGKTAFDISIDNGNEDLAEILQRL

[0810] SEQ ID NO: 106 - VH12 (1 KR)

[0811] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSRWGMSWVRQAPGKGLEWVSYISHDGTF IYYADSVKGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGIIPRDLVGRLLLFDYWGQGTL VTVSS SEQ ID NO: 107 - VH12 (3KR)

[0812] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSRWGMSWVRQAPGRGLEWVSYISHDGTF

[0813] IYYADSVRGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGIIPRDLVGRLLLFDYWGQGTL

[0814] VTVSS

[0815] SEQ ID NO: 108 - VH15 (1 KR)

[0816] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGKGLEWVSYISRDARVI

[0817] YYADSVKGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYWG

[0818] QGTLVTVSS

[0819] SEQ ID NO: 109 - VH 15 (3KR)

[0820] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGRGLEWVSYISRDARVI

[0821] YYADSVRGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYWG

[0822] QGTLVTVSS

[0823] SEQ ID NO: 110 - VH15

[0824] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGKGLEWVSYISRDARVI

[0825] YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYWG

[0826] QGTLVTVSS

[0827] SEQ ID NO: 111 - Non-binding human VH

[0828] MAEVQLLESGGGLVQPGGSLRLSCAASGFRFDAEDMGWVRQAPGKGLEWVSSIYGPSGS

[0829] TYYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCAKYTSPPQNHGFDYWGQGTLVT

[0830] VSS

[0831] SEQ ID NO: 112 -- Non-binding helical peptide

[0832] GQDDEVRILMANGA

[0833] SEQ ID NO: 113 - Linker

[0834] GGGS

[0835] SEQ ID NO: 114 - Linker

[0836] GGGGS

[0837] SEQ ID NO: 115 - Linker

[0838] GGSG

[0839] SEQ ID NO: 116 - Linker

[0840] GSGG

[0841] SEQ ID NO: 117 - Linker

[0842] SGGG

[0843] SEQ ID NO: 118 - Linker

[0844] SSGG

[0845] SEQ ID NO: 119 - Linker SSSG

[0846] SEQ ID NO: 120 - Linker (L1)

[0847] GGGSG

[0848] SEQ ID NO: 121 - Linker

[0849] SAGS

[0850] SEQ ID NO: 122 - Linker

[0851] SGAS

[0852] SEQ ID NO: 123 - Linker

[0853] TGGGGSGGGGS

[0854] SEQ ID NO: 124 - Linker

[0855] GGGGSGGGGS

[0856] SEQ ID NO: 125 - Linker

[0857] SGGSSGSSGGS

[0858] SEQ ID NO: 126 - Linker (L2)

[0859] GGGGSGGGGSGGGGSG

[0860] SEQ ID NO: 127 - Linker

[0861] AGSGGSGGSGGSPVPSTPPTNSSSTPPTPSPSPVPSTPPTNSSSTPPTPSPSPVPSTPPT NSSSTPPTPSPSAS

[0862] SEQ ID NO: 128 - Linker

[0863] AGSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSPVPSTPPTNSSSTPPTPSPSPVPSTPPTP SPSTPPTPSPSAS

[0864] SEQ ID NO: 129 - Linker

[0865] AGSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSGGSGNSSGSGGSPVPSTPPTPSPSTPP TPSPSAS

[0866] SEQ ID NO: 130 - Linker

[0867] AGSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSPVPSTPPTPSPSTPPTPSPSPVPSTPPTP SPSTPPTPSPSAS

[0868] SEQ ID NO: 131 - Linker

[0869] AGSGNSSGSGGSGGSGNSSGSGGSPVPSTPPTPSPSTPPTPSPSAS

[0870] SEQ ID NO: 132 - Linker (L3)

[0871] GSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSGGSGNSSGSGGSPVPSTPPTPSPSTPPT PSPSAS

[0872] SEQ ID NO: 133 - Linker (L4) GSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSGGSGNSSGSGGSPVPSTPPTPSPSTPPT

[0873] PSPSASGGGGSGGGGSGGGGSG

[0874] SEQ ID NO: 134 - DARPin-K19

[0875] MDLGKKLLEAARAGQDDEVRILMANGADVNASDRWGWTPLHLAAWWGHLEIVEVLLKRG

[0876] ADVSAADLHGQSPLHLAAMVGHLEIVEVLLKYGADVNAKDTMGATPLHLAARSGHLEIVEEL

[0877] LKNGADMNAQDKFGKTTFDISTDNGNEDLAEILQKL

[0878] SEQ ID NO: 135 - ICPO (1 KR) inactive mutant

[0879] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0880] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA VCTDEIAPHLRCDTFPCMHRFCIPCMETWMQLRNTCPLCNARLVYLIVGVTPSGSFSTIPIV

[0881] NDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDDDDLD D

[0882] SEQ ID NO: 136 - TRIM21 inactive mutant

[0883] MASAARLTMMWREVTCPICLRPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCAQRFLL

[0884] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHAE FVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSALEL LQEVIIVLERSESWNLKDLDITSPELRSVCHVPGLKKMLRTCA

[0885] SEQ ID NO: 137 - SPOP inactive mutant

[0886] SVNISGQNTMNMVKVPECRLADELGGLWENSRFTDCCLCVAGQEFQAHKAILAARSPVFS AMFEHEMEESKKNRVEINDVEPEVFKEMMCFIYTGKAPNLDKMADDLLAAADKYALERLKV MCEDALCSNLSVENYHASDVLETSGWKSMWSHPHLVAEAYRSLASAQCPFLGPPRKRLK

[0887] QS

[0888] SEQ ID NO: 138 - VH151KR-L1-ICP01KR

[0889] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGKGLEWVSYISRDARVI

[0890] YYADSVKGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYWG QGTLVTVSSGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEV GGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAP PREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNARLVYLI VGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPT

[0891] HPEPTTDEDDDDLDD

[0892] SEQ ID NO: 139 - DP207KR-L1-SPQP

[0893] MDLGRRLLEAARAGQDDEVRILMANGADVNAADDWGNTPLHLAALEGHLEIVEVLLRYGA DVNAQDLYGTTPLHLAAWVGHLEIVEVLLRNGADVNAMDWGETPLHLAAEMGHLEIVEVLL

[0894] RHGADVNAQDRFGKTAFDISIDNGNEDLAEILQRLGGG SG SVN I SGQ NTM N M VKVP ECRLA DELGGLWENSRFTDCCLCVAGQEFQAHKAILAARSPVFSAMFEHEMEESKKNRVEINDVEP EVFKEMMCFIYTGKAPNLDKMADDLLAAADKYALERLKVMCEDALCSNLSVENAAEILILADL HSADQLKTQAVDFINYHASDVLETSGWKSMWSHPHLVAEAYRSLASAQCPFLGPPRKRLK QS SEQ ID NO: 140 - TRIM21-L2-H1

[0895] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL

[0896] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK

[0897] HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHA

[0898] EFVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSAL

[0899] ELLQEVIIVLERSESWNLKDLDITSPELRSVCHVPGLKKMLRTCAGGGGSGGGGSGGGGS GNELKRAFAALRDQI

[0900] SEQ ID NO: 141 - TRIM21-L1-H1

[0901] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL

[0902] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK

[0903] HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHA

[0904] EFVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSAL ELLQEVIIVLERSESWNLKDLDITSPELRSVCHVPGLKKMLRTCAG G G SG N ELKRAFAALRD QI

[0905] SEQ ID NO: 142 - TRIM21-L1-VH151KR

[0906] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL

[0907] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK

[0908] HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHA

[0909] EFVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSAL ELLQEVIIVLERSESWNLKDLDITSPELRSVCHVPGLKKMLRTCAG G G SG M AEVQ LLESGG

[0910] GLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGKGLEWVSYISRDARVIYYADSVKGRFTI

[0911] SRDNSRNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYWGQGTLVTVSS

[0912] SEQ ID NO: 143 - TRIM21-L1-VH153KR

[0913] MASAARLTMMWEEVTCPICLDPFVEPVSIECGHSFCQECISQVGKGGGSVCPVCRQRFLL

[0914] KNLRPNRQLANMVNNLKEISQEAREGTQGERCAVHGERLHLFCEKDGKALCWVCAQSRK

[0915] HRDHAMVPLEEAAQEYQEKLQVALGELRRKQELAEKLEVEIAIKRADWKKTVETQKSRIHA

[0916] EFVQQKNFLVEEEQRQLQELEKDEREQLRILGEKEAKLAQQSQALQELISELDRRCHSSAL ELLQEVIIVLERSESWNLKDLDITSPELRSVCHVPGLKKMLRTCAG G G SG M AEVQ LLESGG

[0917] GLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGRGLEWVSYISRDARVIYYADSVRGRFT

[0918] ISRDNSRNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYWGQGTLVTVSS

[0919] SEQ ID NO: 144 - VH123KR-L2-ICP01KR

[0920] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSRWGMSWVRQAPGRGLEWVSYISHDGTF IYYADSVRGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGIIPRDLVGRLLLFDYWGQGT LVTVSSGGGGSGGGGSGGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPD

[0921] SSDSEAETEVGGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREE DPGSCGGAPPREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTC

[0922] PLCNARLVYLIVGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLG GHTVRALSPTHPEPTTDEDDDDLDD

[0923] SEQ ID NO: 145 - VH121KR-L2-ICP01KR

[0924] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSRWGMSWVRQAPGKGLEWVSYISHDGTF IYYADSVKGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGIIPRDLVGRLLLFDYWGQGT LVTVSSGGGGSGGGGSGGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPD

[0925] SSDSEAETEVGGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREE DPGSCGGAPPREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTC PLCNARLVYLIVGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLG GHTVRALSPTHPEPTTDEDDDDLDD

[0926] SEQ ID NO: 146 - VH153KR-L1-ICP01KR

[0927] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGRGLEWVSYISRDARV IYYADSVRGRFTISRDNSRNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYW

[0928] GQGTLyTySSGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAET

[0929] EVGGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGG

[0930] APPREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNARLV

[0931] YLIVGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALS PTHPEPTTDEDDDDLDD

[0932] SEQ ID NO: 147 - VH15-L1-ICP01KR

[0933] MAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNWSMLWVRQAPGKGLEWVSYISRDARV IYYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCARGALQHPLISNECVQFHFDYW GQGTLVTVSSGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAET

[0934] EVGGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGG

[0935] APPREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNARLV

[0936] YLIVGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALS PTHPEPTTDEDDDDLDD

[0937] SEQ ID NO: 148 - ICP0-L4-VH153KR

[0938] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0939] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA

[0940] VCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNAKLVYLIVGVTPSGSFSTIPIV

[0941] NDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDDDDL

[0942] DDGSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSGGSGNSSGSGGSPVPSTPPTPSPSTP

[0943] PTPSPSASGGGGSGGGGSGGGGSGMAEVQLLESGGGLVQPGGSLRLSCAASGFTFSNW SMLWVRQAPGRGLEWVSYISRDARVIYYADSVRGRFTISRDNSRNTLYLQMNSLRAEDTAV YYCARGALQHPLISNECVQFHFDYWGQGTLVTVSS

[0944] SEQ ID NO: 149 - ICP0-L4-H1

[0945] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0946] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA

[0947] VCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNAKLVYLIVGVTPSGSFSTIPIVN

[0948] DPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDDDDLDD GSGGSGGSGGSPVPSTPPTPSPSTPPTPSPSGGSGNSSGSGGSPVPSTPPTPSPSTPPT PSPSASGGGGSGGGGSGGGGSGNELKRAFAALRDQI

[0949] SEQ ID NO: 150 - ICP0-L2-H1

[0950] MEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSEAETEVGGRGDADHHDDDS

[0951] ASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSCGGAPPREDGGSDEGDVCA

[0952] VCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNAKLVYLIVGVTPSGSFSTIPIVN

[0953] DPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVRALSPTHPEPTTDEDDDDLDD GGGGSGGGGSGGGGSGNELKRAFAALRDQI

[0954] SEQ ID NO: 151 - H1-L1-ICP0 NELKRAFAALRDQIGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFPCDRDLPDSSDSE AETEVGGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSGGSPPREEDPGSC GGAPPREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTWMQLRNTCPLCNA KLVYLIVGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRFAPRYLTLGGHTVR ALSPTHPEPTTDEDDDDLDD

[0955] SEQ ID NO: 152 - H1-L2-ICP0

[0956] NELKRAFAALRDQIGGGGSGGGGSGGGGSGMEPRPGASTRRPEGRPQREPAPDVWVFP

[0957] CDRDLPDSSDSEAETEVGGRGDADHHDDDSASEADSTDAELFETGLLGPQGVDGGAVSG GSPPREEDPGSCGGAPPREDGGSDEGDVCAVCTDEIAPHLRCDTFPCMHRFCIPCMKTW MQLRNTCPLCNAKLVYLIVGVTPSGSFSTIPIVNDPQTRMEAEEAVRAGTAVDFIWTGNQRF APRYLTLGGHTVRALSPTHPEPTTDEDDDDLDD

[0958] SEQ ID NO: 153 - DP207KR-L1-SPQP3KR

[0959] MDLGRRLLEAARAGQDDEVRILMANGADVNAADDWGNTPLHLAALEGHLEIVEVLLRYGA DVNAQDLYGTTPLHLAAWVGHLEIVEVLLRNGADVNAMDWGETPLHLAAEMGHLEIVEVLL RHGADVNAQDRFGKTAFDISIDNGNEDLAEILQRLGGG SG SVN I SGQ NTM N M VKVP ECRLA DELGGLWENSRFTDCCLCVAGQEFQAHKAILAARSPVFSAMFEHEMEESKKNRVEINDVEP EVFKEMMCFIYTGRAPNLDKMADDLLAAADKYALERLKVMCEDALCSNLSVENAAEILILAD

[0960] LHSADQLKTQAVDFINYHASDVLETSGWKSMWSHPHLVAEAYRSLASAQCPFLGPPRRRL RQS

[0961] SEQ ID NO: 154 - DARPIn-K19-L1-SPSB1

[0962] MDLGKKLLEAARAGQDDEVRILMANGADVNASDRWGWTPLHLAAWWGHLEIVEVLLKRG ADVSAADLHGQSPLHLAAMVGHLEIVEVLLKYGADVNAKDTMGATPLHLAARSGHLEIVEEL LKNGADMNAQDKFGKTTFDISTDNGNEDLAEILQKLGGGSGPEPLPLMDLCRRSVRLALG R ERLGEIHTLPLPASLKAYLLYQ

[0963] SEQ ID NO: 155 - Non-binding DARPin

[0964] MDLGKKLLEAARAGQDDEVRILMANGADVNATDNDGYTPLHLAASNGHLEIVEVLLKNGAD VNASDLTGITPLHLAAATGHLEIVEVLLKHGADVNAYDNDGHTPLHLAAKYGHLEIVEVLLKH GADVNAQDKFGKTAFDISIDNGNEDLAEILQ

Claims

Ciaims1 . A method of identifying a bioPROTAC capable of inducing degradation of a target protein, or a fragment thereof, in a eukaryotic cell, wherein the bioPROTAC comprises a binding domain and a degradation domain; wherein the binding domain specifically binds the target protein, and the degradation domain is capable of causing ubiquitylation of the target protein; the method comprising:(i) providing a first DNA library encoding a library of binding domains;(ii) providing a second DNA library encoding a library of degradation domains;(iii) assembling a vector library encoding a bioPROTAC library, wherein each vector in the library comprises an expression cassette comprising a gene encoding a bioPROTAC, and each vector is assembled by joining an expression vector backbone, a first DNA molecule from the first DNA library and a second DNA molecule from the second DNA library;(iv) producing an mRNA library by in vitro transcription of the genes encoding bioPROTACs;(v) transfecting eukaryotic cells which express the target protein with the mRNA library, wherein separate transfections are performed using mRNA encoding each different bioPROTAC; and(vi) screening the transfected cells for depletion of the target protein, wherein depletion of the target protein in a transfected cell indicates the bioPROTAC expressed in that cell is capable of inducing degradation of the target protein.

2. The method of claim 1 , wherein the binding domain is a DARPin; a single chain antibody fragment, preferably an scFv or a single domain antibody; or a ligand or inhibitor of the target protein or a fragment thereof.

3. The method of claim 2, wherein the binding domain sequence is modified to comprise one or more lysine to arginine substitutions relative to the original or native sequence.

4. The method of any one of claims 1 to 3, wherein the degradation domain is an E3 ubiquitin ligase, a component of an E3 ubiquitin ligase complex, an E3 ligase binding partner or an E2 conjugating enzyme, or a fragment thereof.

5. The method of claim 4, wherein the degradation domain is a wild-type E3 ligase, component of an E3 ligase complex, E3 ligase binding partner or E2 conjugating enzyme, or fragment thereof.

6. The method of claim 4, wherein the degradation domain is a modified E3 ligase, component of an E3 ligase, E3 ligase binding partner or E2 conjugating enzyme, or fragment thereof, comprising at least one lysine to arginine substitution relative to the initial or wild- type sequence.

7. The method of any one of claims 4 to 6, wherein the degradation domain is an E3 ligase, a component of an E3 ligase complex or an E3 ligase binding partner, or fragment thereof, comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 1-100, or a variant thereof having at least 80 % sequence identity thereto.

8. The method of any one of claims 1 to 7, wherein the bioPROTAC further comprises a linker joining the binding domain and the degradation domain; wherein the method further comprises providing a third DNA library encoding a library of linkers; and in step (iii), each vector is assembled by joining an expression vector backbone, a first DNA molecule from the first DNA library, a second DNA molecule from the second DNA library and a third DNA molecule from the third DNA library.

9. The method of claim 8, wherein the linker is a glycine-serine linker.

10. The method of claim 8 or 9, wherein the linker comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 120, 126, 132 or 133, or a variant thereof having at least 80 % sequence identity thereto.11 . The method of any one of claims 1 to 10, wherein the expression cassette encodes a 3’ UTR comprising a polyA tail 40-80 nucleotides in length.

12. The method of any one of claims 1 to 11 , wherein in step (iii) each vector is assembled by Golden Gate cloning.

13. The method of claim 12, comprising:(a) obtaining a first gene fragment encoding a binding domain, a second gene fragment encoding a degradation domain and a third gene fragment encoding a linker;(b) inserting the first, second and third gene fragments into cloning vector backbones, thereby yielding a first intermediate vector, a second intermediate vector and a third intermediate vector;(c) digesting the first, second and third intermediate vectors, or amplification products thereof, and an expression vector with a type IIS restriction enzyme; and(d) ligating the products of step (c), thereby yielding a vector encoding a bioPROTAC.

14. The method of claim 13, wherein the gene fragments are inserted into the cloning vector backbones by topoisomerase-based cloning.

15. The method of claim 13 or 14, wherein the digestion and ligation reactions are performed in a single reaction mix in a volume of at most 5 pl, preferably at most 2 pl.

16. The method of any one of claims 13 to 15, wherein the product of the ligation reaction of step (d) is transformed into competent bacteria, and antibiotic selection of transformants is performed in liquid culture.

17. The method of any one of claims 13 to 16, wherein sequencing of the gene encoding the bioPROTAC is not performed prior to in vitro transcription.

18. The method of any one of claims 1 to 17, wherein in step (iv) the template for in vitro transcription is the vector encoding the bioPROTAC.

19. The method of any one of claims 1 to 17, wherein following step (iii) the expression cassette comprising the gene encoding the bioPROTAC is amplified, yielding an amplification product comprising the expression cassette, and the amplification product is the template for in vitro transcription in step (iv).

20. The method of claim 19, wherein amplification is performed using the liquid culture as the amplification template.21 . The method of claim 19 or 20, wherein the amplification product is purified prior to in vitro transcription.

22. The method of any one of claims 1 to 21 , wherein up to 40 % of the total uridine used for in vitro transcription is 5-methoxyuridine (5moU), preferably wherein 20-40 % 5moU is used.

23. The method of any one of claims 1 to 22, wherein the in vitro transcription comprises co-transcriptional capping of the mRNA.

24. The method of any one of claims 1 to 23, wherein the in vitro transcription comprises capping of the mRNA with a cap having a Cap1 structure.

25. The method of claim 23 or 24, wherein the in vitro transcription comprises co- transcriptional capping of the mRNA with a cap having a Cap1 structure.

26. The method of claim 25, wherein co-transcriptional capping is performed by adding m7G(5’)ppp(5’)(2’OMeA)pG to the in vitro transcription reaction.

27. The method of any one of claims 1 to 26, wherein the in vitro transcription reaction is performed in a reaction volume of at most 5 pl, preferably 2 pl.

28. The method of any one of claims 1 to 27, wherein the mRNA is not purified prior to transfection of the eukaryotic cells.

29. The method of any one of claims 1 to 28, wherein the eukaryotic cells are from a mammalian cell line, preferably a human cell line.

30. The method of claim 29, wherein the human cell line is HOT 116, or a derivative thereof.31 . The method of any one of claims 1 to 30, wherein transfection is performed with a saturating amount of mRNA.

32. The method of any one of claims 1 to 31 , wherein the screening for depletion of the target protein is performed by immunofluorescence microscopy, luminescence kinetic assay and / or quantitative immunoblotting, optionally wherein the screening comprises at least two different screening techniques.

33. The method of claims 1 to 32, wherein liquids are dispensed by acoustic dispensing.

34. A method of treating a disease in a subject, comprising:(a) identifying a target protein associated with the disease;(b) identifying a bioPROTAC capable of inducing degradation of the target protein according to the method of any one of claims 1 to 33; and(c) administering the bioPROTAC or a nucleic acid molecule encoding the bioPROTAC to the subject.

35. A bioPROTAC comprising a binding domain and a degradation domain, the degradation domain comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 1-100, or a variant thereof having at least 80 % sequence identity thereto; preferably wherein the degradation domain comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-2, 5, 7-8, 11-16, 18-24, 26 or 31-52, or a variant thereof having at least 80 % sequence identity thereto; optionally wherein the bioPROTAC further comprises a linker comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 120, 126, 132 or 133, or a variant thereof having at least 80 % sequence identity thereto.

36. A bioPROTAC capable of inducing degradation of human c-Myc, comprising a binding domain and a degradation domain, the binding domain comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 101 to 110, or a variant thereof having at least 80 % sequence identity thereto; optionally wherein the degradation domain comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-52, or a variant thereof having at least 80 % sequence identity thereto; and / or wherein the bioPROTAC further comprises a linker comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 120, 126, 132 or 133, or having at least 80 % sequence identity thereto.

37. The bioPROTAC of claim 36, comprising, from N-terminus to C-terminus:(a) (i) a binding domain comprising the amino acid sequence set forth in SEQ IDNO: 108, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 108, wherein the amino acid at the position corresponding to position 78 of SEQ ID NO: 108 is not lysine, preferably wherein it is arginine;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 120; and(iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 40,or a variant thereof having at least 80 % identity to SEQ ID NO: 40, wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is not threonine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is not lysine, preferably wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is alanine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is arginine;(b) (i) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 105, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 105, wherein the amino acids at the positions corresponding to positions 5, 6, 57, 90, 123, 133 and 156 of SEQ ID NO: 105 are not lysine, preferably wherein they are arginine;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 120; and(iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 25, or a variant thereof having at least 80 % identity to SEQ ID NO: 25;(c) (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 10, or a variant thereof having at least 80 % identity to SEQ ID NO: 10;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 126, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 126; and(iii) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 101 , or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 101 ;(d) (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 10, or a variant thereof having at least 80 % identity to SEQ ID NO: 10;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 120; and(iii) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 101 or SEQ ID NO: 108, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 101 or SEQ ID NO: 108, wherein in the variant of SEQ ID NO: 108 the amino acid at the position corresponding to position 78 of SEQ ID NO: 108 is not lysine, preferably wherein it is arginine;(e) (i) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 106, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 106, wherein the amino acid at the position corresponding to position 78 of SEQ ID NO: 106 is not lysine, preferably wherein it is arginine;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 126, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 126; and(iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 40, or a variant thereof having at least 80 % identity to SEQ ID NO: 40, wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is not threonine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is not lysine, preferably wherein the amino acid at the position corresponding to position 67 of SEQ ID NO: 40 is alanine and the amino acid at the position corresponding to position 159 of SEQ ID NO: 40 is arginine;(f) (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 7, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 7;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 133, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 133; and(iii) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 109, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 109, wherein the amino acids at the positions corresponding to positions 45, 67 and 78 of SEQ ID NO: 108 are not lysine, preferably wherein they are arginine;(g) (i) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 7, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 7;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 126 or SEQ ID NO: 133, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 126 or SEQ ID NO: 133; and(iii) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 101 , or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 101 ; or(h) (i) a binding domain comprising the amino acid sequence set forth in SEQ ID NO: 101 , or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 101 ;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120 or SEQ ID NO: 126, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 120 or SEQ ID NO: 126; and(iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 7, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 7.

38. The bioPROTAC of claim 37, comprising the amino acid sequence set forth in any one of SEQ ID NOs: 138-153, or a variant thereof having at least 80 % sequence identity thereto.

39. A bioPROTAC capable of inducing degradation of human K-Ras, comprising, from N-terminus to C-terminus:(i) a binding domain comprising the amino acid sequence set forth in SEQ IDNO: 134, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 134;(ii) a linker comprising the amino acid sequence set forth in SEQ ID NO: 120, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 120; and;(iii) a degradation domain comprising the amino acid sequence set forth in SEQ ID NO: 82, or a variant thereof having at least 80 % sequence identity to SEQ ID NO: 82.

40. The bioPROTAC of claim 39, comprising the amino acid sequence set forth in SEQ ID NO: 154, or a variant thereof having at least 80 % sequence identity thereto.41 . A nucleic acid molecule encoding the bioPROTAC of any one of claims 36 to 40, optionally wherein the nucleic acid molecule is mRNA.

42. A vector comprising the nucleic acid molecule of claim 41 .

43. The vector of claim 42, wherein the vector is a viral vector, preferably an adeno- associated virus (AAV).

44. The bioPROTAC of any one of claims 36 to 40, the nucleic acid molecule of claim 41 or the vector of claim 42 or 43 for use in therapy.

45. The bioPROTAC of any one of claims 36 to 40, the nucleic acid molecule of claim 41 or the vector of claim 42 or 43 for use in the treatment of cancer or an inflammatory disease.

46. A method of treating cancer in a subject in need thereof, comprising administering to the subject a bioPROTAC as defined in any one of claims 36 to 40, a nucleic acid molecule as defined in claim 41 or a vector as defined in claim 42 or 43.

47. A method of degrading a protein in a eukaryotic cell, comprising contacting the cell with a nucleic acid molecule or vector encoding a bioPROTAC, wherein:(i) the protein is c-Myc, and the bioPROTAC is as defined in any one of claims 36 to 38; or(ii) the protein is K-Ras, and the bioPROTAC is as defined in claim 39 or 40.

Citation Information

Patent Citations

  • Protein scaffolds

    WO2009058379A2

  • Fibronectin type iii domain-based multimeric scaffolds

    WO2011130324A1

  • Trail r2-specific multimeric scaffolds

    WO2011130328A1

  • Intracellular antigen binding

    WO2016023898A2

  • Bio-PROTAC artificial protein targeting UBE2C

    CN114057861A

Cited By

  • Methods of treating a ras related disease or disorder

    WO2026015790A1

  • Methods of treating a ras related disease or disorder

    WO2026015796A1

  • Methods of treating a ras related disease or disorder

    WO2026015801A1

  • Use of ras inhibitor for treating pancreatic cancer

    WO2026015825A1

  • Ras inhibitors

    WO2026050446A1