cancer

A polypeptide derived from the CR1 domain of MYC is used to selectively target and degrade MYC in cancer cells, addressing the lack of effective MYC-targeting therapies by leveraging its self-interaction properties and ubiquitin ligase fusion for cancer treatment.

WO2026115272A1PCT designated stage Publication Date: 2026-06-04IMPERIAL COLLEGE INNVOATIONS LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
IMPERIAL COLLEGE INNVOATIONS LTD
Filing Date
2025-11-28
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Current therapeutic approaches lack effective methods to target the oncogenic transcription factor c-MYC due to its intrinsically disordered nature and lack of binding pockets for small molecule compounds, necessitating improved strategies to inhibit MYC activity in cancer treatment.

Method used

Development of a polypeptide derived from the Compaction Region 1 (CR1) domain of MYC, which can self-interact in trans, allowing for intracellular binding and degradation of MYC through fusion with a ubiquitin ligase, such as a PROTAC, to selectively target and degrade overexpressed MYC in cancer cells.

Benefits of technology

The CR1 domain provides a novel targeting moiety for MYC, enabling selective degradation of MYC in cancer cells, offering a promising therapeutic tool for cancer treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025052613_04062026_PF_FP_ABST
    Figure GB2025052613_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to cancer therapy, and in particular, to novel polypeptides derived from the oncogenic transcription factor c-MYC (MYC), conjugates thereof, pharmaceutical compositions comprising such polypeptides and / or mimics thereof, and their uses in therapies and methods for treating, preventing or ameliorating cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cancer

[0002] The present invention relates to cancer therapy, and in particular, to novel polypeptides derived from the oncogenic transcription factor c-MYC (MYC), conjugates thereof, pharmaceutical compositions comprising such polypeptides and / or mimics thereof, and their uses in therapies and methods for treating, preventing or ameliorating cancer.

[0003] The oncogenic transcription factor c-MYC (commonly abbreviated as MYC) interacts with a myriad of cellular proteins that are involved in regulation of gene expression, RNA processing, ribosome biogenesis, mitosis, and DNA damage detection / replication [2-4]. The cancer-driving aspects of MYC are emphasized by the fact that MYC is overexpressed in the majority of human cancers, thus accounting for millions of deaths worldwide every year [5, 6]. Inhibition of excessive MYC activity halts tumorigenesis and induces tumour regression [7, 8]. For these reasons, MYC has been recognized as a compelling target for cancer therapies [9-11].

[0004] Although MYC's major function as a gene-specific transcription factor (GSTF) is firmly established, its precise mode of action is still controversial and may reveal itself to be more complex compared to other human GSTFs. From a structural point of view, MYC is a medium-sized GSTF consisting of an N-terminal intrinsically disordered region and a DNA-binding domain (DBD) capable of directing MYC to various genomic target loci in a sequence-specific manner (after heterodimerizing with MAX; see Fig. 1). In congruence with this model, MYC directly regulates the expression of ~730 target genes

[0013] by acting as a transcriptional activator

[0014] or repressor [15-17]. MYC also appears to act as a more general "transcriptional amplifier" that enhances the expression levels of gene-expression programs already established through the activities of other GSTFs [18-21]. Furthermore, MYC is not restricted to regulating mRNA transcription by the RNA polymerase (RNAP) II machinery, but also stimulates transcription by the RNAP I and RNAP III transcription systems

[0022] .

[0005] MYC can therefore be considered one of the most universal and unusually versatile GSTFs controlling expression of the human genome. Nevertheless, despite all this knowledge regarding the functional effects of MYC on modulating the gene expression machinery in both normal and cancer cells, there is a surprising lack of understanding of the parts of MYC that are responsible for conducting all these activities. A major proportion of MYC, apart from its DNA-binding domain, is intrinsically disordered and does not contain binding pockets for small molecule compounds used in conventional therapeutic drug screens. There is, therefore, a considerable clinical need for improved therapeutic approaches for targeting MYC and treating cancer.

[0006] The present invention arises from the inventors' work in attempting to overcome the problems associated with the prior art.

[0007] The inventors' work focused on mapping the position of transcriptional activation domains (ADs) of MYC and assessing their conformational accessibility based on insights from MD simulations and mutagenesis of the intrinsically disordered portion of MYC (MYC1-143;

[0025] ). The inventors' systematic computational / experimental approaches show that, although intrinsically disordered, some regions in MYC surprisingly retain the ability of forming longer-lasting structures that are capable of intramolecular contacts. As shown in Figures 1 and 2, the inventors derived a high-resolution map of the locations of three independent transcriptional activation domains (AD-1, AD-2, and AD-3) based on assays is yeast and human cells: AD-1 is permanently functionally accessible, but the in vivo activities of AD-2 and AD-3 are negatively controlled by the presence of nearby " Compaction Regions" (CR1 and CR2, respectively). Surprisingly, the inventors found that CR1 is able to self-interact in trans (i.e. an intermolecular interaction between two copies of MYC). The CR1 trans interaction properties allow this motif to act as a "warhead" to degrade intracellular MYC when, for example, fused with a ubiquitin ligase to generate a proteolysis-targeting chimera (PROTAC).

[0008] Therefore, in view of the above, the inventors have demonstrated for the first time that the CR1 domain represents a novel targeting moiety (or "warhead") that can be used to bind intracellular MYC.

[0009] Hence, in a first aspect of the invention, there is provided an isolated polypeptide, derivative or analogue thereof, comprising a Compaction Region 1 (CR1) domain of the oncogenic transcription factor c-MYC (MYC), or a fragment or variant thereof.

[0010] In some embodiments, the polypeptide, derivative or analogue thereof, according to the present invention comprises the CR1 domain of the native MYC protein, or a fragment or variant thereof.

[0011] References to a "native MYC", "native MYC protein", "native MYC sequence", and the like, refer to the human MYC sequence having accession number CAA25015 in the NCBI database. In one embodiment, therefore, the oncogenic transcription factor c-MYC (MYC) protein may comprise the sequence represented in single letter amino acid code herein as SEQ ID NO: 1, as follows:

[0012] MPLNVSFTNR NYDLDYDSVQ PYFYCDEEEN FYQQQQQSEL QPPAPSEDIW KKFELLPTPP LSPSRRSGLC SPSYVAVTPF SLRGDNDGGG GSFSTADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG SPNPARGHSV CSTSSLYLQD LSAAASECID PSWFPYPLN DSSSPKSCAS QDSSAFSPSS DSLLSSTESS PQGSPEPLVL HEETPPTTSS DSEEEQEDEE EIDWSVEKR QAPGKRSESG SPSAGGHSKP PHSPLVLKRC HVSTHQHNYA APPSTRKDYP AAKRVKLDSV RVLRQISNNR KCTSPRSSDT EENVKRRTHN VLERQRRNEL KRSFFALRDQ IPELENNEKA PKWILKKAT AYILSVQAEE QKLISEEDLL RKRREQLKHK LEQLRNSCA

[0013] [SEQ ID NO: 1]

[0014] As shown in Figure 2B, the inventors have modelled the three-dimensional structure of MYC by computational molecular dynamics simulation and used local compaction plot (LCP) analysis to identify two domains that are conformationally unusually confined, referred to herein as " Compaction Region 1" (CR-1) and " Compaction Region 2" (CR- 2), respectively. The CR1 domain (CR1) corresponds to residues 91-160 of a native MYC protein sequence as shown in SEQ ID NO: 1.

[0015] Accordingly, in some embodiments, the polypeptide, derivative or analogue thereof, may have an amino acid sequence consisting of, or comprising, residues 91-160 of a native MYC protein.

[0016] Thus, in some embodiments, the CR1 domain has the protein sequence represented herein as SEQ ID NO: 2, as follows:

[0017] GSFSTADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0018] [SEQ ID NO: 2]

[0019] Accordingly, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially set out in SEQ ID No: 2, or a fragment or variant thereof.

[0020] The term "derivative or analogue thereof" can mean a peptide within which amino acid residues are replaced by residues (whether natural amino acids, non-natural amino acids or amino acid mimics) with similar side chains or peptide backbone properties. Additionally, the terminals of such peptides may be protected by N- and / or C-terminal protecting groups with similar properties to acetyl or amide groups. Derivatives and analogues of peptides according to the invention may also include a sequence which modifies the intracellular location of the polypeptide. This includes, for example, adding nuclear localization sequences, mitochondria-targeting sequences, or cellpenetrating motifs.

[0021] Derivatives and analogues of peptides according to the invention may also include those that increase the peptide's half-life in vivo. For example, a derivative or analogue of the peptides of the invention may include peptoid and retropeptoid derivatives of the peptides, peptide-peptoid hybrids and D-amino acid derivatives of the peptides.

[0022] Peptoids, or poly-N-substituted glycines, are a class of peptidomimetics whose side chains are appended to the nitrogen atom of the peptide backbone, rather than to the alpha-carbons, as they are in amino acids. Peptoid derivatives of the peptides of the invention may be readily designed from knowledge of the structure of the peptide. Retropeptoids (in which all amino acids are replaced by peptoid residues in reversed order) are also suitable derivatives in accordance with the invention. A retropeptoid is expected to bind in the opposite direction in the ligand-binding groove, as compared to a peptide or peptoid-peptide hybrid containing one peptoid residue. As a result, the side chains of the peptoid residues are able to point in the same direction as the side chains in the original peptide.

[0023] In some embodiments, the polypeptide, derivative or analogue thereof may have an amino acid sequence consisting of residues MYC 91-160, and may further comprise up to about 100 additional amino acids. For example, in some embodiments, the polypeptide, derivative or analogue thereof may have an amino acid sequence consisting of residues 91-160 of a native MYC protein plus, at most, an additional 50, 40, 30, 20, 10 or 5 amino acids N-terminal to this sequence, i.e. SEQ ID No: 2. In addition, or alternatively, in some embodiments, the polypeptide, derivative or analogue thereof may have an amino acid sequence consisting of residues 91-160 of a native MYC protein plus, at most, an additional 50, 40, 30, 20, 10 or 5 amino acids C-terminal to this sequence, i.e. SEQ ID No: 2.

[0024] The inventors have surprisingly found that the CR1 domain of the MYC protein is able to self-interact in trans, such that two separate copies of MYC are able to form an intermolecular interaction via the CR1 domains. Thus, the polypeptide, derivative or analogue thereof consisting of, or comprising, the CR1 domain is capable of binding to other endogenous MYC proteins present in cells. In some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of the native MYC protein, or a fragment or variant thereof, is capable of binding to endogenous MYC.

[0025] Various tests may be used to determine whether the isolated polypeptide, derivative or analogue thereof comprising at least a portion of the CR1 domain is able to bind to endogenous MYC for the purposes of the present disclosure. For example, the yeast two-hybrid studies exemplified in Example 2 of the present disclosure which was used to detect interactions between two polypeptide portions derived from the same protein, i.e. by using a range of MYC fragments of different sizes.

[0026] The inventors have also surprisingly found that the self-interaction properties of the CR1 domain were not limited to MYC91-160 (full-length CR1 domain).

[0027] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, a fragment of the CRT domain of a native MYC protein.

[0028] The term "fragment thereof" refers to a portion or derivative of the polypeptide sequence that is smaller in size than the full-length native protein, for example, comprising fewer amino acids and thus having a lower molecular weight. The reduction of amino acids may be achieved by removal of residues from the C- and / or N-terminal of the CRT peptide, or may be achieved by deletion of one or more amino acids from internal parts of the peptide sequence.

[0029] Thus, a "fragment of the CRT domain" can mean the CRT domain of the native MYC protein (i.e. residues 91-160 of a native MYC protein) is reduced in size by the removal of amino acids. The reduction of amino acids may be achieved by removal of residues from either the C- or N-terminus of the CRT domain of the native MYC protein, or may be achieved by deletion of one or more amino acids from within the core of the CRT domain of the native MYC protein.

[0030] In some embodiments, the polypeptide, derivative or analogue thereof may comprise a fragment of the CRT domain of the native MYC protein having a length, for example, based on the number of amino acid residues, that is greater than 70%, 75%, 80%, 90%, 95%, 96%, 97%, 98% or 99% of the full-length CRT domain of the native MYC protein. Typically, the polypeptide, derivative or analogue thereof may comprise a fragment of the CRT domain of the native MYC protein having a length that is greater than 80%, or greater than 85%, more typically greater than 90%, of the full-length CRT domain of the native MYC protein. In some embodiments, the polypeptide, derivative or analogue thereof comprises at least 10, 20, 30, 40, 50 or 60 amino acids of the CR1 domain of a native MYC protein. Typically, the polypeptide, derivative or analogue thereof comprises at least 50, more typically at least 60 amino acids of the CR1 domain of a native MYC protein. In some embodiments, the polypeptide, derivative or analogue thereof comprises 65 amino acids of the CR1 domain of a native MYC protein.

[0031] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 91-160 of a native MYC protein, wherein at least the last 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid residues from the N-terminus are absent, deleted or removed.

[0032] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 92-160 of a native MYC protein represented herein as SEQ ID NO: 3, as follows:

[0033] -SFSTADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0034] [SEQ ID NO: 3]

[0035] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially as set out in SEQ ID No: 3, or a fragment or variant thereof.

[0036] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 93-160 of a native MYC protein represented herein as SEQ ID NO: 4, as follows:

[0037] — FSTADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0038] [SEQ ID NO: 4]

[0039] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially as set out in SEQ ID No: 4, or a fragment or variant thereof.

[0040] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 94-160 of a native MYC protein represented herein as SEQ ID NO: 5, as follows: - STADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0041] [SEQ ID NO: 5]

[0042] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially as set out in SEQ ID No: 5, or a fragment or variant thereof.

[0043] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 95-160 of a native MYC protein represented herein as SEQ ID NO: 6, as follows:

[0044] - TADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0045] [SEQ ID NO: 6]

[0046] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially as set out in SEQ ID No: 6, or a fragment or variant thereof.

[0047] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 96-160 of a native MYC protein represented herein as SEQ ID NO: 7, as follows:

[0048] - ADQLE MVTELLGGDM VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0049] [SEQ ID NO: 7]

[0050] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially as set out in SEQ ID No: 7, or a fragment or variant thereof.

[0051] In addition, or alternatively, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 91-160 of the native MYC protein, wherein at least the last 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid residues from the C-terminus are absent, deleted or removed.

[0052] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, a single portion of residues 91-160 of a native MYC protein, or a combination of a plurality of portions of residues 91-160 of a native MYC protein. Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 91-160 of a native MYC protein, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid residues from internal parts of the peptide sequence are absent, deleted or removed. In some embodiments, some or all of the amino acid residues that are absent, deleted or removed within residues 91-160 of a native MYC protein occur consecutively. In some embodiments, some or all of the amino acid residues that are absent, deleted or removed within residues 91-160 of a native MYC protein occur non-consecutively.

[0053] In some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, residues 91-160 (D101-110) of a native MYC protein represented herein as SEQ ID NO: 8, as follows:

[0054] GSFSTADQLE - VNQSFICDPD DETFIKNIII QDCMWSGFSA AAKLVSEKLA SYQAARKDSG

[0055] [SEQ ID NO: 8]

[0056] Thus, in some embodiments, the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially as set out in SEQ ID No: 8, or a fragment or variant thereof.

[0057] Surprisingly, the inventors have found that the trans interaction property of MYC91-160 (CR1 domain) is abolished by single alanine substitutions in a select number of distinct residues, including D109, Q113, E122, K126, N127, 1129, 1130, D132, M134, F138, L149 and Y152 (see, for example, Fig. 3C-D). These residues within the primary amino acid sequence clearly represent key positions that play an important role in the self-interaction properties of the CR1 domain.

[0058] Thus, in some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein does not comprise substitutions in one or more residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152. Typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein does not comprise substitutions in at least three, four or five residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152. More typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein does not comprise substitutions in at least six, seven, eight or nine residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152. Yet typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein does not comprise substitutions in at least 10, 11 or 12 residues selected from D109, Q113, E122, K126, N127, 1129, 1130, D132, M134, F138, L149 and / or Y152.

[0059] Alternatively, in some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a conservative substitution in one or more residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152. Typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a conservative substitution in at least three, four or five residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152. More typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a conservative substitution in at least six, seven, eight or nine residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152. Yet typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a conservative substitution in at least 10, 11 or 12 residues selected from D109, Q113, E122, K126, N127, I129, I130, D132, M134, F138, L149 and / or Y152.

[0060] The term "conservative substitution" means replacing an amino acid in a polypeptide with another amino acid having similar biochemical characteristics, e.g. substituting one hydrophobic amino acid for another hydrophobic amino acid. Such substitutions are not likely to change the shape of the polypeptide chain.

[0061] For example, as shown in Fig. 3E, the inventors carried out a complete substitution series of the D132 position with all 19 alternative naturally occurring amino acids and found that only a substitution by another acidic amino acid (D132-E) supports the selfinteraction property fully, while Q, N, H and R substitutions provide detectable, but only very partial functionality.

[0062] In some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a conservative substitution in D132 residue. In some embodiments, the D132 residue is substituted with E, Q, N, H and R. Typically, the D132 residue is substituted with E.

[0063] The inventors have also advantageously found that the trans interaction property of MYC91-160 (CR1 domain) is significantly improved by single alanine (A) substitutions in a select number of distinct residues, including C133, S136 and S139 (see, for example, Fig. 3C-D). Accordingly, substitutions in one or more residues selected from C133, S136 and S139 surprisingly increase the self-interaction properties of the CR1 domain (i.e. binding affinity and / or specificity).

[0064] Thus, in some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises substitutions in one or more residues selected from C133, S136 and S139. Typically, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises substitutions in at least one, two or all three residues selected from C133, S136 and S139.

[0065] In some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a substitution in C133 residue. In some embodiments, the C133 residue is substituted with A, G, I, L, M or V. In some embodiments, the C133 residue is substituted with A

[0066] In some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a substitution in S136 residue. In some embodiments, the S136 residue is substituted with A, G, I, L, M or V. In some embodiments, the S136 residue is substituted with A.

[0067] In some embodiments, the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a substitution in S139 residue. In some embodiments, the S139 residue is substituted with A, G, I, L, M or V. In some embodiments, the S139 residue is substituted with A.

[0068] The inventors' surprising finding that the CR1 domain of MYC is an intrinsically disordered region of MYC that essentially interacts with itself has identified the CR1 domain as a novel peptide drug and / or a targeting moiety for endogenous MYC, providing a promising tool both in diagnostic and therapeutic applications.

[0069] Thus, in a second aspect of the invention, there is provided a conjugate comprising the polypeptide, derivative or analogue thereof according to the first aspect, and a payload molecule.

[0070] In some embodiments, the payload molecule is an enzyme. In some embodiments, the enzyme may be selected from an oxidoreductase, transferase, hydrolase, lyase, isomerase and / or ligase. In some embodiments, the enzyme is a ligase. Recently, PROTAC (Proteolysis-Targeting Chimeras) degraders have emerged as promising anti-cancer drugs through utilization of the cancer cells own protein destruction machinery to selectively degrade essential tumour drivers. MYC is known to be overexpressed in cancer cells. However, MYC is an intrinsically disordered protein and, therefore, does not contain binding pockets for small molecule compounds used in conventional therapeutic drug screen. The inventors have shown for the first time that the CR1 domain can be fused with an ubiquitin ligase as a PROTAC which is capable of degrading extremely high overexpressed MYC in cells (referring to Figure 5).

[0071] Thus, in some embodiments, the ligase is an E2- and / or E3-ubiquitin ligase.

[0072] The term "ubiquitin ligase" refers to a family of proteins that facilitate the transfer of ubiquitin to a specific substrate protein, targeting the substrate protein for degradation.

[0073] In some embodiments, the ubiquitin ligase is an E2 ubiquitin ligase. In some embodiments, the E2 ubiquitin ligase may be selected from E2A, E2B, E2C, E2D1, E2D2, E2D3, E2D4, E2E1, E2E2, E2E3, E2F, E2G1, E2G2, E2H, E2I, E2J1, E2J2, E2K, E2L3, E2L6 and / or E2M. In some embodiments, the E2 ubiquitin ligase is E2B. In some embodiments, the E2 ubiquitin ligase is E2D1.

[0074] In some embodiments, the ubiquitin ligase is an E3 ubiquitin ligase. In some embodiments, the E3 ubiquitin ligase may be selected from PTrCP, FBW7, SKP2, VHL, SPOP, CRBN, DDB2, SOCS2, ASB1 and / or CHIP. In some embodiments, the E3 ubiquitin ligase is SPOP. In some embodiments, the E3 ubiquitin ligase is SOCS2. In some embodiments, the E3 ubiquitin ligase is VHL.

[0075] In some embodiments, the enzyme is a fragment of the full-length protein that retains the specific binding properties and / or catalytic activity of the native protein.

[0076] In some embodiments, the enzyme is a fragment of SPOP. In some embodiments, the fragment of SPOP is the C-terminal sequence of SPOP (SPOP167-374) represented herein as SEQ ID NO: 9, as follows:

[0077] SVNI SGQNTMNMVK VPECRLADELGGLWENSRFTDCCLCVAGQEFQAHKAI LAARS PVFSAMFEHEMEESKKNRVEINDVEPEVFKEMMCF I YTGKAPNLDKMADDLLAAADKYALERLKVMCEDALCSNLSVENAAEILILADLHSADQLKTQAVDFINYHASDVLETSGWKSMWSHPHL VAEA YRS LAS AQC P FL G P P RK RL K Q S

[0078] [SEQ ID NO: 9] Therefore, the conjugate comprises an amino acid sequence substantially set out as SEQ ID No:9, or a fragment or variant thereof.

[0079] In some embodiments, the payload molecule may be a protein involved in autophagy. In some embodiments, the protein may be involved in the formation of autophagosomes. In some embodiments, the protein may be an autophagy-related protein. In some embodiments, the autophagy-related protein may be a member of the autophagy-related protein 8 (Atg8) protein family. In some embodiments, the autophagy-related protein may be microtubule-associated protein 1 light chain 3 (LC3).

[0080] In some embodiments, LC3 may be represented herein as SEQ ID NO: 22, as follows:

[0081] MPSEKTFKQR RTFEQRVEDV RLIREQHPTK I PVI IERYKG EKQLPVLDKT KFLVPDHVNM SELIKIIRRR LQLNANQAFF LLVNGHSMVS VSTPISEVYE SEKDEDGFLY MVYASQETFG MKLSV

[0082] [SEQ ID NO: 22]

[0083] Therefore, the conjugate comprises an amino acid sequence substantially set out as SEQ ID No: 22, or a fragment or variant thereof.

[0084] In some embodiments, the conjugate may comprise a linker by which the polypeptide, derivative or analogue thereof according to the first aspect is conjugated to the payload molecule.

[0085] The term "linker" used herein means a chemical moiety comprising a chain of atoms that covalently attaches the polypeptide fragment of MYC to the payload. Linkers have the general property that they consist of variable lengths / combinations of amino acids that provide flexibility (e.g. glycine) and maintain solubility (e.g. serine, threonine). The inventors have surprisingly shown that conjugation with payloads using linkers containing alternating glycine and serine and / or threonine residues in lengths typically varying between 3 to 15 residues does not affect the ability of the polypeptide of the first aspect to bind to endogenous MYC, nor the activity of the payload molecule (Example 4). Thus, in some embodiments, the conjugate is capable of binding to endogenous MYC for targeted delivery of a payload.

[0086] In some embodiments, the linker comprises (GST)n, (GTGSG)n and / or (GSG)n, wherein n is an integer of at least one. In some embodiments, n may be 2, 3, 4, or 5. Typically, therefore, the conjugate of the present invention has the structure substantially as represented herein as SEQ ID NO: 10, as follows:

[0087] MYPYDVPDYAGYPYDVPDYAGSTGSFSTADQLEMVTELLGGDMVNQSFICDPDDETFIKNIIIQDCMWSGFSAAAKLVSEKLASYQAARKD SGGTGSGEQKLISEEDLGSGSVNI SGQNTMNMVK VPECRLADELGGLWENSRFTDCCLCVAGQEFQAH KAI LAARS PVFSAMFEHEMEESK KNRVEINDVEPEVFKEMMCFI YTGKAPNLDKMADDLLAAADKYALERLKVMCEDALCSNLSVENAAEILILADLHSADQLKTQAVDFINYH ASDVLETSGWKSMWSHPHLVAEAYRSLASAQCPFLGPPRKRLKQS

[0088] [SEQ ID NO: 10]

[0089] Therefore, the conjugate comprises an amino acid sequence substantially set out as SEQ ID No: 10, or a fragment or variant thereof.

[0090] In some embodiments, the conjugate of the present invention has the structure substantially as represented herein as SEQ ID NO: 11, as follows:

[0091] MGSFSTADQLEMVTELLGGDMVNQSFICDPDDETFIKNIIIQDCMWSGFSAAAKLVSEKLASYQAARKDSGSVNISGQNTMNMVKVPECRL ADELGGLWENSRFTDCCLCVAGQEFQAHKAI LAARS PVFSAMFEHEMEESKKNRVEINDVEPEVFKEMMCFI YTGKAPNLDKMADDLLAAA DKYALERLKVMCEDALCSNLSVENAAEILILADLHSADQLKTQAVBFINYHASDVLETSGWKSMWSHPHLVAEAYRSLASAQCPFLGPPR KRLKQS

[0092] [SEQ ID NO: 11]

[0093] Therefore, the conjugate comprises an amino acid sequence substantially set out as SEQ ID No: 11, or a fragment or variant thereof.

[0094] In some embodiments, the conjugate of the present invention has the structure substantially as represented herein as SEQ ID NO: 23, as follows:

[0095] MYPYDVPDYAGYPYDVPDYAGSTGSFSTADQLEMVTELLGGDMVNQSFICDPDDETFIKNIIIQDCMWSGFSAAAKLVSEKLASYQAARKD SGGTGSGEQKLISEEDLGSGMPSEKTFKQRRTFEQRVEDVRLIREQHPTKIPVIIERYKGEKQLPVLDKTKFLVPDHVNMSELIKIIRRRL QLNANQAFFLLVNGHSMVSVSTPI SEVYESEKDEDGFLYMVYASQETFGMKLSV

[0096] [SEQ ID NO: 23]

[0097] Therefore, the conjugate comprises an amino acid sequence substantially set out as SEQ ID No: 23, or a fragment or variant thereof.

[0098] In a third aspect of the invention, there is provided a nucleic acid sequence encoding the polypeptide, derivative or analogue thereof according to the first aspect, or the conjugate according to the second aspect.

[0099] In some embodiments, the nucleic acid may comprise DNA or RNA. In some embodiments, the nucleic acid may be single-stranded or double-stranded. In some embodiments, the nucleic acid is synthetic. In some embodiments, the nucleic acid may comprise natural, non-natural and / or modified nucleotides. In an embodiment, the nucleic acid comprises DNA. In another embodiment, the nucleic acid comprises RNA. In some embodiments, the RNA is messenger RNA (mRNA). In some embodiments, the RNA is synthetic mRNA.

[0100] In some embodiments, the CR1 domain of the first aspect is encoded by the DNA sequence represented herein as SEQ ID NO: 12, as follows:

[0101] GGTAGCTTTAGCACCGCAGATCAGCTGGAGATGGTGACAGAACTGCTGGGCGGCGATATGGTGAACCAGAGCTTCATCTGTGACCCGGACG ACGAAACCTTTATTAAAAATATTATTATCCAGGATTGCATGTGGAGCGGCTTCAGCGCAGCAGCAAAACTGGTGAGCGAGAAACTGGCAAG CTATCAGGCCGCACGTAAAGATAGCGGC

[0102] [SEQ ID NO: 12]

[0103] Therefore, the nucleic acid comprises a nucleotide sequence substantially set out as SEQ ID No: 12, or a fragment or variant thereof.

[0104] In some embodiments, the CR1 domain of the first aspect is encoded by the RNA sequence represented herein as SEQ ID NO: 13, as follows:

[0105] GGUAGCUUUAGCACCGCAGAUCAGCUGGAGAUGGUGACAGAACUGCUGGGCGGCGAUAUGGUGAACCAGAGCUUCAUCUGUGACCCGGACG ACGAAACCUUUAUUAAAAAUAUUAUUAUCCAGGAUUGCAUGUGGAGCGGCUUCAGCGCAGCAGCAAAACUGGUGAGCGAGAAACUGGCAAG CUAUCAGGCCGCACGUAAAGAUAGCGGC

[0106] [SEQ ID NO: 13]

[0107] Therefore, the nucleic acid comprises a nucleotide sequence substantially set out as SEQ ID No: 13, or a fragment or variant thereof.

[0108] In some embodiments, the C-terminal sequence of SPOP (SPOP167-374) is encoded by the DNA sequence represented herein as SEQ ID NO: 14, as follows:

[0109] AGCGTGAACATCAGCGGGCAGAACACCATGAACATGGTGAAGGTGCCCGAGTGCAGACTGGCCGACGAGCTGGGCGGCCTGTGGGAGAACA GCAGATTCACCGACTGCTGCCTGTGCGTGGCCGGCCAAGAGTTCCAAGCCCACAAGGCCATCCTGGCCGCTAGAAGCCCCGTGTTCAGCGC CATGTTCGAGCACGAGATGGAGGAGAGCAAGAAGAACAGAGTGGAGATCAACGACGTGGAGCCCGAGGTGTTCAAGGAGATGATGTGCTTC ATCTACACCGGCAAGGCCCCCAACCTGGACAAGATGGCCGACGACCTGCTGGCCGCCGCCGACAAGTACGCCCTGGAGAGACTGAAGGTGA TGTGCGAGGACGCCCTGTGCAGCAACCTGAGCGTGGAGAACGCCGCCGAGATCCTGATCCTGGCCGACCTGCACAGCGCCGATCAGCTGAA GACCCAAGCCGTGGACTTCATCAACTACCACGCTAGCGACGTGCTGGAGACAAGCGGCTGGAAGAGCATGGTGGTGAGCCACCCCCACCTG GTGGCCGAGGCCTACAGAAGCCTGGCTAGCGCTCAGTGCCCCTTCCTGGGCCCCCCTAGAAAGAGACTGAAGCAGAGC

[0110] [SEQ ID NO: 14] Therefore, the nucleic acid comprises a nucleotide sequence substantially set out as SEQ ID No: 14, or a fragment or variant thereof.

[0111] In some embodiments, the C-terminal sequence of SPOP (SPOP167-374) is encoded by the RNA sequence represented herein as SEQ ID NO: 15, as follows:

[0112] AGCGUGAACAUCAGCGGGCAGAACACCAUGAACAUGGUGAAGGUGCCCGAGUGCAGACUGGCCGACGAGCUGGGCGGCCUGUGGGAGAACA GCAGAUUCACCGACUGCUGCCUGUGCGUGGCCGGCCAAGAGUUCCAAGCCCACAAGGCCAUCCUGGCCGCUAGAAGCCCCGUGUUCAGCGC CAUGUUCGAGCACGAGAUGGAGGAGAGCAAGAAGAACAGAGUGGAGAUCAACGACGUGGAGCCCGAGGUGUUCAAGGAGAUGAUGUGCUUC AUCUACACCGGCAAGGCCCCCAACCUGGACAAGAUGGCCGACGACCUGCUGGCCGCCGCCGACAAGUACGCCCUGGAGAGACUGAAGGUGA UGUGCGAGGACGCCCUGUGCAGCAACCUGAGCGUGGAGAACGCCGCCGAGAUCCUGAUCCUGGCCGACCUGCACAGCGCCGAUCAGCUGAA GACCCAAGCCGUGGACUUCAUCAACUACCACGCUAGCGACGUGCUGGAGACAAGCGGCUGGAAGAGCAUGGUGGUGAGCCACCCCCACCUG GUGGCCGAGGCCUACAGAAGCCUGGCUAGCGCUCAGUGCCCCUUCCUGGGCCCCCCUAGAAAGAGACUGAAGCAGAGC

[0113] [SEQ ID NO:15]

[0114] Therefore, the nucleic acid comprises a nucleotide sequence substantially set out as SEQ ID No: 15, or a fragment or variant thereof.

[0115] In some embodiments, the LC3 is encoded by the DNA sequence represented herein as SEQ ID NO: 24, as follows:

[0116] ATGCCTAGCGAGAAGACCTTCAAGCAGAGAAGAACCTTCGAGCAGAGAGTGGAGGACGTGAGACTGATCAGAGAGCAGCACCCCACCAAGA TCCCCGTGATCATCGAGAGATACAAGGGCGAGAAGCAGCTGCCCGTGCTGGACAAGACCAAGTTCCTGGTGCCCGACCACGTGAACATGAG CGAGCTGATCAAGATCATCAGAAGAAGACTGCAGCTGAACGCCAACCAAGCCTTCTTCCTGCTGGTGAACGGCCACAGCATGGTGAGCGTG AGCACCCCCATCAGCGAGGTGTACGAGAGCGAGAAGGACGAGGACGGCTTCCTGTACATGGTGTACGCTAGCCAAGAGACCTTCGGCATGA AGC GAGCG G

[0117] [SEQ ID NO: 24]

[0118] Therefore, the nucleic acid comprises a nucleotide sequence substantially set out as SEQ ID No:24, or a fragment or variant thereof.

[0119] In some embodiments, the LC3 is encoded by the RNA sequence represented herein as SEQ ID NO: 25, as follows:

[0120] AUGCCUAGCGAGAAGACCUUCAAGCAGAGAAGAACCUUCGAGCAGAGAGUGGAGGACGUGAGACUGAUCAGAGAGCAGCACCCCACCAAGA UCCCCGUGAUCAUCGAGAGAUACAAGGGCGAGAAGCAGCUGCCCGUGCUGGACAAGACCAAGUUCCUGGUGCCCGACCACGUGAACAUGAG CGAGCUGAUCAAGAUCAUCAGAAGAAGACUGCAGCUGAACGCCAACCAAGCCUUCUUCCUGCUGGUGAACGGCCACAGCAUGGUGAGCGUG AGCACCCCCAUCAGCGAGGUGUACGAGAGCGAGAAGGACGAGGACGGCUUCCUGUACAUGGUGUACGCUAGCCAAGAGACCUUCGGCAUGA AGCUGAGCGUG

[0121] [SEQ ID NO: 25]

[0122] Therefore, the nucleic acid comprises a nucleotide sequence substantially set out as SEQ ID No: 25, or a fragment or variant thereof. In some embodiments, the conjugate of the present invention is encoded by the DNA sequence represented herein as SEQ ID NO: 16, as follows:

[0123] ATGGGTAGCTTTAGCACCGCAGATCAGCTGGAGATGGTGACAGAACTGCTGGGCGGCGATATGGTGAACCAGAGCTTCATCTGTGACCCGG ACGACGAAACCTTTATTAAAAATATTATTATCCAGGATTGCATGTGGAGCGGCTTCAGCGCAGCAGCAAAACTGGTGAGCGAGAAACTGGC AAGCTATCAGGCCGCACGTAAAGATAGCGGCAGCGTGAACATCAGCGGGCAGAACACCATGAACATGGTGAAGGTGCCCGAGTGCAGACTG GCCGACGAGCTGGGCGGCCTGTGGGAGAACAGCAGATTCACCGACTGCTGCCTGTGCGTGGCCGGCCAAGAGTTCCAAGCCCACAAGGCCA TCCTGGCCGCTAGAAGCCCCGTGTTCAGCGCCATGTTCGAGCACGAGATGGAGGAGAGCAAGAAGAACAGAGTGGAGATCAACGACGTGGA GCCCGAGGTGTTCAAGGAGATGATGTGCTTCATCTACACCGGCAAGGCCCCCAACCTGGACAAGATGGCCGACGACCTGCTGGCCGCCGCC GACAAGTACGCCCTGGAGAGACTGAAGGTGATGTGCGAGGACGCCCTGTGCAGCAACCTGAGCGTGGAGAACGCCGCCGAGATCCTGATCC TGGCCGACCTGCACAGCGCCGATCAGCTGAAGACCCAAGCCGTGGACTTCATCAACTACCACGCTAGCGACGTGCTGGAGACAAGCGGCTG GAAGAGCATGGTGGTGAGCCACCCCCACCTGGTGGCCGAGGCCTACAGAAGCCTGGCTAGCGCTCAGTGCCCCTTCCTGGGCCCCCCTAGA AAGAGAC T GAAGCAGAG CTAA

[0124] [SEQ ID NO: 16]

[0125] In some embodiments, the conjugate of the present invention is encoded by the RNA sequence represented herein as SEQ ID NO: 17, as follows:

[0126] AUGGGUAGCUUUAGCACCGCAGAUCAGCUGGAGAUGGUGACAGAACUGCUGGGCGGCGAUAUGGUGAACCAGAGCUUCAUCUGUGACCCGG ACGACGAAACCUUUAUUAAAAAUAUUAUUAUCCAGGAUUGCAUGUGGAGCGGCUUCAGCGCAGCAGCAAAACUGGUGAGCGAGAAACUGGC AAGCUAUCAGGCCGCACGUAAAGAUAGCGGCAGCGUGAACAUCAGCGGGCAGAACACCAUGAACAUGGUGAAGGUGCCCGAGUGCAGACUG GCCGACGAGCUGGGCGGCCUGUGGGAGAACAGCAGAUUCACCGACUGCUGCCUGUGCGUGGCCGGCCAAGAGUUCCAAGCCCACAAGGCCA UCCUGGCCGCUAGAAGCCCCGUGUUCAGCGCCAUGUUCGAGCACGAGAUGGAGGAGAGCAAGAAGAACAGAGUGGAGAUCAACGACGUGGA GCCCGAGGUGUUCAAGGAGAUGAUGUGCUUCAUCUACACCGGCAAGGCCCCCAACCUGGACAAGAUGGCCGACGACCUGCUGGCCGCCGCC GACAAGUACGCCCUGGAGAGACUGAAGGUGAUGUGCGAGGACGCCCUGUGCAGCAACCUGAGCGUGGAGAACGCCGCCGAGAUCCUGAUCC UGGCCGACCUGCACAGCGCCGAUCAGCUGAAGACCCAAGCCGUGGACUUCAUCAACUACCACGCUAGCGACGUGCUGGAGACAAGCGGCUG GAAGAGCAUGGUGGUGAGCCACCCCCACCUGGUGGCCGAGGCCUACAGAAGCCUGGCUAGCGCUCAGUGCCCCUUCCUGGGCCCCCCUAGA AAGAGAC U GAAGCAGAG CUAA

[0127] [SEQ ID NO: 17]

[0128] In some embodiments, the conjugate of the present invention is encoded by the DNA sequence represented herein as SEQ ID NO: 18, as follows:

[0129] ATGTACCCTTACGACGTGCCCGACTACGCCGGGTACCCAGATCAGCTGGAGATGGTGACAGAACTGCTGGGCGGCGATATGGTGAACCAGA GCTTCATCTGTGACCCGGACGACGAAACCTTTATTAAAAATATTATTATCCAGGATTGCATGTGGAGCGGCTTCAGCGCAGCAGCAAAACT GGTGAGCGAGAAACTGGCAAGCTATCAGGCCGCACGTAAAGATAGCGGCGGTACCGGCAGCGGCGAGCAGAAGCTGATCAGCGAGGAGGAC CTGGGCAGCGGCAGCGTGAACATCAGCGGGCAGAACACCATGAACATGGTGAAGGTGCCCGAGTGCAGACTGGCCGACGAGCTGGGCGGCC TGTGGGAGAACAGCAGATTCACCGACTGCTGCCTGTGCGTGGCCGGCCAAGAGTTCCAAGCCCACAAGGCCATCCTGGCCGCTAGAAGCCC CGTGTTCAGCGCCATGTTCGAGCACGAGATGGAGGAGAGCAAGAAGAACAGAGTGGAGATCAACGACGTGGAGCCCGAGGTGTTCAAGGAG ATGATGTGCTTCATCTACACCGGCAAGGCCCCCAACCTGGACAAGATGGCCGACGACCTGCTGGCCGCCGCCGACAAGTACGCCCTGGAGA GACTGAAGGTGATGTGCGAGGACGCCCTGTGCAGCAACCTGAGCGTGGAGAACGCCGCCGAGATCCTGATCCTGGCCGACCTGCACAGCGC CGATCAGCTGAAGACCCAAGCCGTGGACTTCATCAACTACCACGCTAGCGACGTGCTGGAGACAAGCGGCTGGAAGAGCATGGTGGTGAGC CACCCCCACCTGGTGGCCGAGGCCTACAGAAGCCTGGCTAGCGCTCAGTGCCCCTTCCTGGGCCCCCCTAGAAAGAGACTGAAGCAGAGCT AA

[0130] [SEQ ID NO: 18] In some embodiments, the conjugate of the present invention is encoded by the RNA sequence represented herein as SEQ ID NO: 19, as follows:

[0131] AUGUACCCUUACGACGUGCCCGACUACGCCGGGUACCCAGAUCAGCUGGAGAUGGUGACAGAACUGCUGGGCGGCGAUAUGGUGAACCAGA GCUUCAUCUGUGACCCGGACGACGAAACCUUUAUUAAAAAUAUUAUUAUCCAGGAUUGCAUGUGGAGCGGCUUCAGCGCAGCAGCAAAACU GGUGAGCGAGAAACUGGCAAGCUAUCAGGCCGCACGUAAAGAUAGCGGCGGUACCGGCAGCGGCGAGCAGAAGCUGAUCAGCGAGGAGGAC CUGGGCAGCGGCAGCGUGAACAUCAGCGGGCAGAACACCAUGAACAUGGUGAAGGUGCCCGAGUGCAGACUGGCCGACGAGCUGGGCGGCC UGUGGGAGAACAGCAGAUUCACCGACUGCUGCCUGUGCGUGGCCGGCCAAGAGUUCCAAGCCCACAAGGCCAUCCUGGCCGCUAGAAGCCC CGUGUUCAGCGCCAUGUUCGAGCACGAGAUGGAGGAGAGCAAGAAGAACAGAGUGGAGAUCAACGACGUGGAGCCCGAGGUGUUCAAGGAG AUGAUGUGCUUCAUCUACACCGGCAAGGCCCCCAACCUGGACAAGAUGGCCGACGACCUGCUGGCCGCCGCCGACAAGUACGCCCUGGAGA GACUGAAGGUGAUGUGCGAGGACGCCCUGUGCAGCAACCUGAGCGUGGAGAACGCCGCCGAGAUCCUGAUCCUGGCCGACCUGCACAGCGC CGAUCAGCUGAAGACCCAAGCCGUGGACUUCAUCAACUACCACGCUAGCGACGUGCUGGAGACAAGCGGCUGGAAGAGCAUGGUGGUGAGC CACCCCCACCUGGUGGCCGAGGCCUACAGAAGCCUGGCUAGCGCUCAGUGCCCCUUCCUGGGCCCCCCUAGAAAGAGACUGAAGCAGAGCU AA

[0132] [SEQ ID NO: 19]

[0133] In some embodiments, the conjugate of the present invention is encoded by the DNA sequence represented herein as SEQ ID NO: 20, which includes the DNA sequence of SEQ ID NO: 18 and parts of the pCDNA3 expression plasmid (for example, the CMV enhancer and promoter module), as follows:

[0134] GACGGATCGGGAGATCTCCCGATCCCCTATGGTGCACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAAGCCAGTATCTGCTCCCTGC TTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAACAAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTT AGGGTTAGGCGTTTTGCGCTGCTTCGCGATGTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTA CGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCG CCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACT GCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCC AGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGGTTTTGGCAGTACAT CAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAA CGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTC TCTGGCTAACTAGAGAACCCACTGCTTACTGGCTTATCGAAATTAATACGACTCACTATAGGGAGACCCAAGCTTGGTACCGAGCTCGGAT CTCCCACCATGTACCCTTACGACGTGCCCGACTACGCCGGGTACCCTTACGACGTGCCCGACTACGCCGGATCCACCGGTAGCTTTAGCAC CGCAGATCAGCTGGAGATGGTGACAGAACTGCTGGGCGGCGATATGGTGAACCAGAGCTTCATCTGTGACCCGGACGACGAAACCTTTATT AAAAATATTATTATCCAGGATTGCATGTGGAGCGGCTTCAGCGCAGCAGCAAAACTGGTGAGCGAGAAACTGGCAAGCTATCAGGCCGCAC GTAAAGATAGCGGCGGTACCGGCAGCGGCGAGCAGAAGCTGATCAGCGAGGAGGACCTGGGCAGCGGCAGCGTGAACATCAGCGGGCAGAA CACCATGAACATGGTGAAGGTGCCCGAGTGCAGACTGGCCGACGAGCTGGGCGGCCTGTGGGAGAACAGCAGATTCACCGACTGCTGCCTG TGCGTGGCCGGCCAAGAGTTCCAAGCCCACAAGGCCATCCTGGCCGCTAGAAGCCCCGTGTTCAGCGCCATGTTCGAGCACGAGATGGAGG AGAGCAAGAAGAACAGAGTGGAGATCAACGACGTGGAGCCCGAGGTGTTCAAGGAGATGATGTGCTTCATCTACACCGGCAAGGCCCCCAA CCTGGACAAGATGGCCGACGACCTGCTGGCCGCCGCCGACAAGTACGCCCTGGAGAGACTGAAGGTGATGTGCGAGGACGCCCTGTGCAGC AACCTGAGCGTGGAGAACGCCGCCGAGATCCTGATCCTGGCCGACCTGCACAGCGCCGATCAGCTGAAGACCCAAGCCGTGGACTTCATCA ACTACCACGCTAGCGACGTGCTGGAGACAAGCGGCTGGAAGAGCATGGTGGTGAGCCACCCCCACCTGGTGGCCGAGGCCTACAGAAGCCT GGCTAGCGCTCAGTGCCCCTTCCTGGGCCCCCCTAGAAAGAGACTGAAGCAGAGCTAAGAATTCTGCAGATATCCATCACACTGGCGGCCG CTCGAGCATGCATCTAGAGGGCCCTATTCTATAGTGTCACCTAAATGCTAGAGCTCGCTGATCAGCCTCGACTGTGCCTTCTAGTTGCCAG CCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCAT CGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGC TGGGGATGCGGTGGGCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCTAGGGGGTATCCCCACGCGCCCTGTAGCGGCGCATTA AGCGCGGCGGGTGTGGTGGTTACGCGCAGCGTGACCGCTACACTTGCCAGCGCCCTAGCGCCCGCTCCTTTCGCTTTCTTCCCTTCCTTTC TCGCCACGTTCGCCGGCTTTCCCCGTCAAGCTCTAAATCGGGGGCTCCCTTTAGGGTTCCGATTTAGTGCTTTACGGCACCTCGACCCCAA AAAACTTGATTAGGGTGATGGTTCACGTAGTGGGCCATCGCCCTGATAGAC

[0135] [SEQ ID NO: 20]

[0136] In some embodiments, the conjugate of the present invention is encoded by the RNA sequence represented herein as SEQ ID NO: 21, as follows:

[0137] GGGAGACCCAAGCUUGGUACCGAGCUCGGAUCUCCCACCAUGUACCCUUACGACGUGCCCGACUACGCCGGGUACCCUUACGACGUGCCCG ACUACGCCGGAUCCACCGGUAGCUUUAGCACCGCAGAUCAGCUGGAGAUGGUGACAGAACUGCUGGGCGGCGAUAUGGUGAACCAGAGCUU CAUCUGUGACCCGGACGACGAAACCUUUAUUAAAAAUAUUAUUAUCCAGGAUUGCAUGUGGAGCGGCUUCAGCGCAGCAGCAAAACUGGUG AGCGAGAAACUGGCAAGCUAUCAGGCCGCACGUAAAGAUAGCGGCGGUACCGGCAGCGGCGAGCAGAAGCUGAUCAGCGAGGAGGACCUGG GCAGCGGCAGCGUGAACAUCAGCGGGCAGAACACCAUGAACAUGGUGAAGGUGCCCGAGUGCAGACUGGCCGACGAGCUGGGCGGCCUGUG GGAGAACAGCAGAUUCACCGACUGCUGCCUGUGCGUGGCCGGCCAAGAGUUCCAAGCCCACAAGGCCAUCCUGGCCGCUAGAAGCCCCGUG UUCAGCGCCAUGUUCGAGCACGAGAUGGAGGAGAGCAAGAAGAACAGAGUGGAGAUCAACGACGUGGAGCCCGAGGUGUUCAAGGAGAUGA UGUGCUUCAUCUACACCGGCAAGGCCCCCAACCUGGACAAGAUGGCCGACGACCUGCUGGCCGCCGCCGACAAGUACGCCCUGGAGAGACU GAAGGUGAUGUGCGAGGACGCCCUGUGCAGCAACCUGAGCGUGGAGAACGCCGCCGAGAUCCUGAUCCUGGCCGACCUGCACAGCGCCGAU CAGCUGAAGACCCAAGCCGUGGACUUCAUCAACUACCACGCUAGCGACGUGCUGGAGACAAGCGGCUGGAAGAGCAUGGUGGUGAGCCACC CCCACCUGGUGGCCGAGGCCUACAGAAGCCUGGCUAGCGCUCAGUGCCCCUUCCUGGGCCCCCCUAGAAAGAGACUGAAGCAGAGCUAAG

[0138] [SEQ ID NO: 21]

[0139] In some embodiments, the conjugate of the present invention is encoded by the DNA sequence represented herein as SEQ ID NO: 26, as follows:

[0140] ATGTACCCTTACGACGTGCCCGACTACGCCGGGTACCCTTACGACGTGCCCGACTACGCCGGATCCACCGGTAGCTTTAGCACCGCAGATC AGCTGGAGATGGTGACAGAACTGCTGGGCGGCGATATGGTGAACCAGAGCTTCATCTGTGACCCGGACGACGAAACCTTTATTAAAAATAT TATTATCCAGGATTGCATGTGGAGCGGCTTCAGCGCAGCAGCAAAACTGGTGAGCGAGAAACTGGCAAGCTATCAGGCCGCACGTAAAGAT AGCGGCGGTACCGGCAGCGGCGAGCAGAAGCTGATCAGCGAGGAGGACCTGGGCAGCGGCATGCCTAGCGAGAAGACCTTCAAGCAGAGAA GAACCTTCGAGCAGAGAGTGGAGGACGTGAGACTGATCAGAGAGCAGCACCCCACCAAGATCCCCGTGATCATCGAGAGATACAAGGGCGA GAAGCAGCTGCCCGTGCTGGACAAGACCAAGTTCCTGGTGCCCGACCACGTGAACATGAGCGAGCTGATCAAGATCATCAGAAGAAGACTG CAGCTGAACGCCAACCAAGCCTTCTTCCTGCTGGTGAACGGCCACAGCATGGTGAGCGTGAGCACCCCCATCAGCGAGGTGTACGAGAGCG AGAAGGACGAGGACGGCTTCCTGTACATGGTGTACGCTAGCCAAGAGACCTTCGGCATGAAGCTGAGCGTGTAA

[0141] [SEQ ID NO: 26]

[0142] In some embodiments, the conjugate of the present invention is encoded by the RNA sequence represented herein as SEQ ID NO: 27, as follows:

[0143] AUGUACCCUUACGACGUGCCCGACUACGCCGGGUACCCUUACGACGUGCCCGACUACGCCGGAUCCACCGGUAGCUUUAGCACCGCAGAUC AGCUGGAGAUGGUGACAGAACUGCUGGGCGGCGAUAUGGUGAACCAGAGCUUCAUCUGUGACCCGGACGACGAAACCUUUAUUAAAAAUAU UAUUAUCCAGGAUUGCAUGUGGAGCGGCUUCAGCGCAGCAGCAAAACUGGUGAGCGAGAAACUGGCAAGCUAUCAGGCCGCACGUAAAGAU AGCGGCGGUACCGGCAGCGGCGAGCAGAAGCUGAUCAGCGAGGAGGACCUGGGCAGCGGCAUGCCUAGCGAGAAGACCUUCAAGCAGAGAA GAACCUUCGAGCAGAGAGUGGAGGACGUGAGACUGAUCAGAGAGCAGCACCCCACCAAGAUCCCCGUGAUCAUCGAGAGAUACAAGGGCGA GAAGCAGCUGCCCGUGCUGGACAAGACCAAGUUCCUGGUGCCCGACCACGUGAACAUGAGCGAGCUGAUCAAGAUCAUCAGAAGAAGACUG CAGCUGAACGCCAACCAAGCCUUCUUCCUGCUGGUGAACGGCCACAGCAUGGUGAGCGUGAGCACCCCCAUCAGCGAGGUGUACGAGAGCG AGAAGGACGAGGACGGCUUCCUGUACAUGGUGUACGCUAGCCAAGAGACCUUCGGCAUGAAGCUGAGCGUGUAA

[0144] [SEQ ID NO: 27] Therefore, the nucleic acid may comprise a nucleotide sequence substantially set out in any one of SEQ ID No: 16-21 or 26-27, or a fragment or variant thereof.

[0145] In a fourth aspect of the invention, there is provided an expression cassette comprising the nucleic acid sequence according to the third aspect.

[0146] The nucleic acid sequences of the invention are preferably harboured in a recombinant vector, for example a recombinant vector for delivery into a host cell of interest to enable production of the polypeptide, derivative or analogue thereof according to the first aspect or conjugate according to the second aspect.

[0147] Accordingly, in a fifth aspect of the invention, there is provided a recombinant vector comprising the expression cassette according to the fourth aspect.

[0148] The vector of the fifth aspect encoding the polypeptide or conjugate may, for example, be a plasmid, cosmid or phage and / or be a viral vector. Such recombinant vectors are highly useful in the delivery systems of the invention for transforming cells with the nucleotide sequences.

[0149] The nucleic acid molecule may (but not necessarily) be one, which becomes incorporated in the DNA of the host cell. Undifferentiated cells may be stably transformed leading to the production of genetically modified daughter cells (in which case regulation of expression in the subject may be required e.g. with specific transcription factors or gene activators). Alternatively, the delivery system may be designed to favour unstable or transient transformation of differentiated cells. When this is the case, regulation of expression may be less important because expression of the DNA molecule will stop when the transformed cells die or stop expressing the protein.

[0150] Alternatively, the delivery system may provide the nucleic acid molecule to the host cell without it being incorporated in a vector. For instance, the nucleic acid molecule may be incorporated within a liposome, lipid nanoparticle or virus particle (as shown in Example 5).

[0151] Accordingly, in a sixth aspect of the invention, there is provided a lipid nanoparticle comprising the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, or the recombinant vector according to the fifth aspect.

[0152] Alternatively, a "naked" nucleic acid molecule, for example mRNA, such as a synthetic mRNA, may be inserted into a host cell by a suitable means e.g. direct endocytotic uptake.

[0153] The inventors believe that the polypeptides described herein may be used both diagnostically and therapeutically, for example to detect or treat cancer.

[0154] Thus, in a seventh aspect of the invention, there is provided the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect, for use in therapy or diagnosis.

[0155] In an eighth aspect of the invention, there is provided the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect, for use in treating, ameliorating or preventing cancer.

[0156] In a ninth aspect of the invention, there is provided a method of treating, ameliorating or preventing cancer in a subject, the method comprising, administering to a subject in need of such treatment, a therapeutically effective amount of the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect.

[0157] In some embodiments, the cancer may be a blood cancer. The cancer may be Leukaemia, lymphoma or myeloma.

[0158] In some embodiments, the cancer may be a solid tumour or solid cancer. The cancer may be bowel cancer, brain cancer, breast cancer, endometrial cancer, gastric cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, or skin cancer. In some embodiments, the cancer is liver cancer.

[0159] It will be appreciated that the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect (herein collectively referred to as the active agents), may be used in a medicament which may be used in a monotherapy (i.e. use of the active agents alone), for treating, ameliorating, or preventing cancer. Alternatively, the active agents may be used as an adjunct to, or in combination with, known therapies for treating, ameliorating, or preventing cancer.

[0160] In one embodiment, the active agents may be used in combination with a drug that damages DNA. Examples of DNA-damaging compounds used in the treatment of cancer include, for example, cisplatin, carboplatin, oxaliplatin, methotrexate, doxorubicin and daunorubicin. Thus, in one embodiment, the drug that damages DNA is cisplatin, carboplatin, oxaliplatin, methotrexate, doxorubicin and / or daunorubicin.

[0161] In one embodiment, the active agents may be used in combination with drugs that block checkpoint proteins from binding with their partner proteins and thus allowing immune cells to target cancer cells. Accordingly, the active agents may be used in combination a checkpoint inhibitor. The checkpoint inhibitor may be a programmed cell death protein 1 (PD-1) inhibitor, a programmed death-ligand 1 (PD-L1) inhibitor and / or a cytotoxic T-lymphocyte-associated protein 4 (CTLA-4) inhibitor. In one embodiment, the checkpoint inhibitor is a PD-1 inhibitor. Typically, the PD-1 inhibitor is Pembrolizumab. In one embodiment, the checkpoint inhibitor is a PD-L1 inhibitor. Typically, the PD-L1 inhibitor is avelumab.

[0162] Alternatively, or additionally, the active agents may be used in combination with ionising radiation that damages DNA.

[0163] Medicaments comprising the active agents described herein may be used in a number of ways. Compositions comprising the active agents may be administered by inhalation (e.g. intranasally). Compositions may also be formulated for topical use. For instance, creams or ointments may be applied to the skin. The active agents may be combined in compositions having a number of different forms depending, in particular, on the manner in which the composition is to be used. Thus, for example, the composition may be in the form of a powder, tablet, capsule, liquid, ointment, cream, gel, hydrogel, aerosol, spray, micellar solution, transdermal patch, liposome suspension or any other suitable form that may be administered to a person or animal in need of treatment. It will be appreciated that the vehicle of medicaments according to the invention should be one which is well-tolerated by the subject to whom it is given.

[0164] The active agents may also be incorporated within a slow- or delayed-release device. Such devices may, for example, be inserted on or under the skin, and the medicament may be released over weeks or even months. The device may be located at least adjacent the treatment site. Such devices may be particularly advantageous when long-term treatment with the active agents used according to the invention is required and which would normally require frequent administration (e.g. at least daily injection).

[0165] The active agents according to the invention may be administered to a subject by injection into the blood stream or directly into a site requiring treatment, for example into a cancerous tumour (e.g. breast cancer) or into the blood stream adjacent thereto. Injections may be intravenous (bolus or infusion) or subcutaneous (bolus or infusion), intradermal (bolus or infusion) or intramuscular (bolus or infusion).

[0166] In some embodiments, the active agents are administered orally. Accordingly, the active agents may be contained within a composition that may, for example, be ingested orally in the form of a tablet, capsule or liquid.

[0167] It will be appreciated that the amount of the active agents that is required is determined by its biological activity and bioavailability, which in turn depends on the mode of administration, the physiochemical properties of the inhibitor, and whether it is being used as a monotherapy, or in a combined therapy. Optimal dosages to be administered may be determined by those skilled in the art, and will vary with the particular active agents in use, the strength of the pharmaceutical composition, the mode of administration, and the advancement of the cancer. Additional factors depending on the particular subject being treated will result in a need to adjust dosages, including subject age, weight, gender, diet, and time of administration. The active agents may be administered before, during or after onset of the cancer to be treated. Daily doses may be given as a single administration. However, typically the active agents is given two or more times during a day, and most preferably twice a day.

[0168] Generally, a daily dose of between O. Olpg / kg of body weight and 500mg / kg of body weight of active agents according to the invention may be used for treating, ameliorating, or preventing cancer. More preferably, the daily dose is between O. Olmg / kg of body weight and 400mg / kg of body weight, more preferably between O.lmg / kg and 200mg / kg body weight, and most preferably between approximately Img / kg and lOOmg / kg body weight.

[0169] A patient receiving treatment may take a first dose upon waking and then a second dose in the evening (if on a two dose regime) or at 3- or 4-hourly intervals thereafter. Alternatively, a slow release device may be used to provide optimal doses of the conjugate according to the invention to a patient without the need to administer repeated doses.

[0170] Known procedures, such as those conventionally employed by the pharmaceutical industry (e.g. in vivo experimentation, clinical trials, etc.), may be used to form specific formulations comprising the conjugate according to the invention and precise therapeutic regimes (such as daily doses of the conjugate and the frequency of administration). The inventors believe that they are the first to describe a pharmaceutical composition for treating cancer, based on the use of the active agents described herein, in particular the CR1 domain of MYC.

[0171] Hence, in a tenth aspect of the invention, there is provided a pharmaceutical composition comprising the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect, and a pharmaceutically acceptable vehicle.

[0172] The pharmaceutical composition can be used in the therapeutic amelioration, prevention or treatment in a subject of cancer.

[0173] The invention also provides, in an eleventh aspect, a process for making the pharmaceutical composition according to the ninth aspect, the process comprising contacting a therapeutically effective amount of the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect, and a pharmaceutically acceptable vehicle.

[0174] A "subject" may be a vertebrate, mammal, or domestic animal. Hence, the conjugate, inhibitor, compositions and medicaments according to the invention may be used to treat any mammal, for example livestock (e.g. a horse), pets, or may be used in other veterinary applications. Typically, however, the subject is a human being.

[0175] A "therapeutically effective amount" of the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect, is any amount which, when administered to a subject, is the amount of drug that is needed to treat or prevent the cancer.

[0176] For example, the therapeutically effective amount of the polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect, used may be from about 0.01 mg to about 800 mg, and typically from about 0.01 mg to about 500 mg. It is considered that the amount of polypeptide, derivative or analogue thereof according to the first aspect, the conjugate according to the second aspect, the nucleic acid sequence according to the third aspect, the expression cassette according to the fourth aspect, the vector according to the fifth aspect or the lipid nanoparticle according to the sixth aspect is an amount from about 0.1 mg to about 250 mg, or from about 0.1 mg to about 20 mg.

[0177] A "pharmaceutically acceptable vehicle" as referred to herein, is any known compound or combination of known compounds that are known to those skilled in the art to be useful in formulating pharmaceutical compositions.

[0178] In some embodiments, the pharmaceutically acceptable vehicle may be a solid, and the composition may be in the form of a powder or tablet. A solid pharmaceutically acceptable vehicle may include one or more substances which may also act as flavouring agents, lubricants, solubilisers, suspending agents, dyes, fillers, glidants, compression aids, inert binders, sweeteners, preservatives, dyes, coatings, or tabletdisintegrating agents. The vehicle may also be an encapsulating material. In powders, the vehicle is a finely divided solid that is in admixture with the finely divided active agents (i.e. the conjugate) according to the invention. In tablets, the conjugate may be mixed with a vehicle having the necessary compression properties in suitable proportions and compacted in the shape and size desired. The powders and tablets preferably contain up to 99% of the conjugate. Suitable solid vehicles include, for example calcium phosphate, magnesium stearate, talc, sugars, lactose, dextrin, starch, gelatin, cellulose, polyvinylpyrrolidine, low melting waxes and ion exchange resins. In another embodiment, the pharmaceutical vehicle may be a gel and the composition may be in the form of a cream or the like.

[0179] However, the pharmaceutical vehicle may be a liquid, and the pharmaceutical composition is in the form of a solution. Liquid vehicles are used in preparing solutions, suspensions, emulsions, syrups, elixirs and pressurized compositions. The conjugate according to the invention may be dissolved or suspended in a pharmaceutically acceptable liquid vehicle such as water, an organic solvent, a mixture of both or pharmaceutically acceptable oils or fats. The liquid vehicle can contain other suitable pharmaceutical additives such as solubilisers, emulsifiers, buffers, preservatives, sweeteners, flavouring agents, suspending agents, thickening agents, colours, viscosity regulators, stabilizers or osmo-regulators. Suitable examples of liquid vehicles for oral and parenteral administration include water (partially containing additives as above, e.g. cellulose derivatives, preferably sodium carboxymethyl cellulose solution), alcohols (including monohydric alcohols and polyhydric alcohols, e.g. glycols) and their derivatives, and oils (e.g. fractionated coconut oil and arachis oil). For parenteral administration, the vehicle can also be an oily ester such as ethyl oleate and isopropyl myristate. Sterile liquid vehicles are useful in sterile liquid form compositions for parenteral administration. The liquid vehicle for pressurized compositions can be a halogenated hydrocarbon or other pharmaceutically acceptable propellant.

[0180] Liquid pharmaceutical compositions, which are sterile solutions or suspensions, can be utilized by, for example, intramuscular, intrathecal, epidural, intraperitoneal, intravenous and particularly subcutaneous injection. The conjugate may be prepared as a sterile solid composition that may be dissolved or suspended at the time of administration using sterile water, saline, or other appropriate sterile injectable medium.

[0181] The active agents and compositions of the invention may be administered in the form of a sterile solution or suspension containing other solutes or suspending agents (for example, enough saline or glucose to make the solution isotonic), bile salts, acacia, gelatin, sorbitan monoleate, polysorbate 80 (oleate esters of sorbitol and its anhydrides copolymerized with ethylene oxide) and the like. The active agents used according to the invention can also be administered orally either in liquid or solid composition form. Compositions suitable for oral administration include solid forms, such as pills, capsules, granules, tablets, and powders, and liquid forms, such as solutions, syrups, elixirs, and suspensions. Forms useful for parenteral administration include sterile solutions, emulsions, and suspensions.

[0182] It will be appreciated that the invention extends to any nucleic acid or peptide or variant, derivative or analogue thereof, which comprises substantially the amino acid or nucleic acid sequences of any of the sequences referred to herein, including variants thereof. The terms "substantially the amino acid / nucleotide / peptide sequence" and "variant", can be a sequence that has at least 40% sequence identity with the amino acid / nucleotide / peptide sequences of any one of the sequences referred to herein, for example 40% identity with any of the sequences identified herein.

[0183] Amino acid / polynucleotide / polypeptide sequences with a sequence identity which is greater than 65%, more typically greater than 70%, even more typically greater than 75%, and still more typically greater than 80% sequence identity to any of the sequences referred to are also envisaged. Typically, the amino acid / polynucleotide / polypeptide sequence has at least 85% identity with any of the sequences referred to, more typically at least 90% identity, even more typically at least 92% identity, even more typically at least 95% identity, even more typically at least 97% identity, even more typically at least 98% identity and, most typically at least 99% identity with any of the sequences referred to herein.

[0184] The skilled technician will appreciate how to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences. In order to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences, an alignment of the two sequences must first be prepared, followed by calculation of the sequence identity value. The percentage identity for two sequences may take different values depending on:- (i) the method used to align the sequences, for example, ClustalW, BLAST, FASTA, Smith-Waterman (implemented in different programs), or structural alignment from 3D comparison; and (ii) the parameters used by the alignment method, for example, local vs global alignment, the pair-score matrix used (e.g. BLOSUM62, PAM250, Gonnet etc.), and gap-penalty, e.g. functional form and constants.

[0185] Having made the alignment, there are many different ways of calculating percentage identity between the two sequences. For example, one may divide the number of identities by: (i) the length of shortest sequence; (ii) the length of alignment; (iii) the mean length of sequence; (iv) the number of non-gap positions; or (v) the number of equivalenced positions excluding overhangs. Furthermore, it will be appreciated that percentage identity is also strongly length dependent. Therefore, the shorter a pair of sequences is, the higher the sequence identity one may expect to occur by chance.

[0186] Hence, it will be appreciated that the accurate alignment of protein or DNA sequences is a complex process. The popular multiple alignment program ClustalW (Thompson et al., 1994, Nucleic Acids Research, 22, 4673-4680; Thompson et al., 1997, Nucleic Acids Research, 24, 4876-4882) is a preferred way for generating multiple alignments of proteins or DNA in accordance with the invention. Suitable parameters for ClustalW may be as follows: For DNA alignments: Gap Open Penalty = 15.0, Gap Extension Penalty = 6.66, and Matrix = Identity. For protein alignments: Gap Open Penalty = 10.0, Gap Extension Penalty = 0.2, and Matrix = Gonnet. For DNA and Protein alignments: ENDGAP = -1, and GAPDIST = 4. Those skilled in the art will be aware that it may be necessary to vary these and other parameters for optimal sequence alignment.

[0187] Typically, calculation of percentage identities between two amino acid / polynucleotide / polypeptide sequences may then be calculated from such an alignment as (N / T)*100, where N is the number of positions at which the sequences share an identical residue, and T is the total number of positions compared including gaps and either including or excluding overhangs. Typically, overhangs are included in the calculation. Hence, a most suitable method for calculating percentage identity between two sequences comprises (i) preparing a sequence alignment using the ClustalW program using a suitable set of parameters, for example, as set out above; and (ii) inserting the values of N and T into the following formula:- Sequence Identity = (N / T)*100. Alternative methods for identifying similar sequences will be known to those skilled in the art. For example, a substantially similar nucleotide sequence will be encoded by a sequence which hybridizes to DNA sequences or their complements under stringent conditions. By stringent conditions, the inventors mean the nucleotide hybridises to filter-bound DNA or RNA in 3x sodium chloride / sodium citrate (SSC) at approximately 45°C followed by at least one wash in 0.2x SSC / 0.1% SDS at approximately 20-65°C. Alternatively, a substantially similar polypeptide may differ by at least 1, but less than 5, 10, 20, 50 or 100 amino acids from any of the sequences identified herein.

[0188] Due to the degeneracy of the genetic code, it is clear that any nucleic acid sequence described herein could be varied or changed without substantially affecting the sequence of the protein encoded thereby, to provide a functional variant thereof. Suitable nucleotide variants are those having a sequence altered by the substitution of different codons that encode the same amino acid within the sequence, thus producing a silent (synonymous) change. Other suitable variants are those having homologous nucleotide sequences but comprising all, or portions of, sequence, which are altered by the substitution of different codons that encode an amino acid with a side chain of similar biophysical properties to the amino acid it substitutes, to produce a conservative change. For example, small non-polar, hydrophobic amino acids include glycine, alanine, leucine, isoleucine, valine, proline, and methionine. Large non-polar, hydrophobic amino acids include phenylalanine, tryptophan and tyrosine. The polar neutral amino acids include serine, threonine, cysteine, asparagine and glutamine. The positively charged (basic) amino acids include lysine, arginine and histidine. The negatively charged (acidic) amino acids include aspartic acid and glutamic acid. It will therefore be appreciated which amino acids may be replaced with an amino acid having similar biophysical properties, and the skilled technician will know the nucleotide sequences encoding these amino acids.

[0189] All features described herein (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.

[0190] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example, to the accompanying Figures, in which:- Figure 1 shows structure and activation domain (AD) mapping in MYC. A. Illustration of the sliding window strategy for high-resolution mapping for mapping the positions of transcriptional activation domains (ADs) in the intrinsically disordered portion of MYC in yeast and human cells. MYC1'352is represented by 50 amino acid portions with start- and endpoints spaced in 10 residue intervals (for example, MYC1'50, MYC11'60, MYC21'70, etc.). B. Principles of the reporter gene assays. In both the yeast and human systems, a fusion protein consisting of a portion of MYC (purple; also see A) is fused in frame to the DNA-binding domain of yeast transcription factor GAL4 (grey). Upon transfection into the Y2HGold yeast strain, the transcriptional activity of the fusion protein results in read-outs from three separate reporter genes (o-galactosidase, Ade-2 and His-3). The binding of the GAL4-MYC fusion protein in human cells is detected by Lucia luciferase activity. Secreted alkaline phosphatase (SEAP, pink) also encoded by the expression plasmid was used to standardize slight variations in transfection efficiencies between constructs. C. Results of the 50 amino acid scanning window strategy in yeast. The constructs expressed in strain Y2H-Gold only enable growth (upper panel). Constructs containing MYC1'50, MYC81'130and MYC151'200show the highest signals on both panels, indicating the presence of three distinct ADs. D.

[0191] Results of the 50 amino acid scanning window strategy in human H460 cells.

[0192] Figure 2 shows computational molecular dynamics of MYC. A. Overview of MYC Structure. The MYC primary protein structure is shown schematically as a light blue box. The positions of evolutionarily conserved MYC-Boxes (MB-0 to MB-4) are shown in dark blue, the C-terminal basic Helix-Loop-Helix (bHLH) motif in violet. The molecular structure depicts the bHLH crystal structure of MYC (blue) bound sequence specifically with its heterodimerization partner MYC-associated factor X (MAX; red) to a DNA fragment containing an E-Box hexanucleotide binding motif (PDB# 1NKP; [1]). The N-terminal 80% of the protein are intrinsically disordered (dark grey dotted double-headed arrow). Higher order structural analysis of MYC1-352. The graphs are aligned and stacked on top of each other to aid cross-comparisons. B. Local compaction plot (LCP) analysis

[0012] with a window size of 40 amino acids, emphasizing medium-range, regional folding patterns. While some parts of MYC1-352 emerge as extended and conformationally heterogeneous (especially MYC270-352), other regions (including MYC1-50, MYC90-120, MYC200-220 and MYC250-260) are arranged in a compact (exemplified by short distances [<20 A]) and highly populated fold (dense black lines following a similar path). See

[0012] for further explanation of this particular visualization method. C. Summary of o-helical propensity from 12 x 1 ps aMD simulations. The blue bars are the median values and the grey error bars are based on the standard deviation calculated from each simulation separately. The central region of MYC (spanning residues MYC90-220) appears most highly structured in terms of o-helical propensity. Underneath the plots, a schematic diagram of the primary sequence of MYC1-352 shows locations of the evolutionarily conserved MYC-boxes. C. Three representative examples of conformations of MYC71-150 (final frames from aMD_no4, aMD_no5 and aMD_no6 after 1 ps simulation time). Residue 71 is shown as a van der Waals representation in light blue, and residues of MB-2 (MYC129-143) in purple. Compaction Region #1 (CR#1) is highlighted in dashed red circles. Note the consistent compaction, but variable folding pattern and changing position of CR#1 and MB-2 relative to each other.

[0193] Figure 3 shows Yeast-2-Hybrid (Y2H) Alanine-Scanning Mutagenesis of the SelfInteraction Properties of MYC91-160. A. Primary amino acid sequence of CR-1 (MYC91-160). The MB-2 motif is underlined. B. All residues - except alanines and glycines - in MYC91-160 were individually substituted by alanine residues to test their contribution to the self-interaction properties. Growth of yeast colonies on nutrient agar deficient of tryptophan and leucine (-WL) tests for successful transformation of the two plasmids required for the Y2H experiment. C. Growth of yeast colonies on nutrient agar deficient of tryptophan, leucine, adenine and histidine (-WLAH, bottom) tests for successful Y2H interaction. Colony size and cell density are direct indicators of the strength of protein- protein interaction between the fusion partners. The identity and position of the mutated residue is shown beneath the corresponding colony. In the bottom right position, wildtype (WT) and negative control colonies are shown. All constructs were negative in an autoactivation assay (ROJW, data not shown). D.

[0194] Quantiation of the mutant phenotypes based on growth under selective conditions. Alanine substitutions of residues K126, N127 and D132 cause the strongest disruption. E. Systematic substitutions of D132 with all 19 alternative amino acids. Only D and E support the interaction fully. F. Molecular dynamics model of the homodimeric structure interaction between two copies of MYC91-160. The polypeptide strand of one copy is labelled in magenta, the other copy in turquoise. Key residues are visualized as van der Waals representations and color-coded. G. A comparison of growth density after 24 hours for MYC91-160, MYC96-160 and MYC91-160(delta 101-110) compared to a vector only control.

[0195] Figure 4 shows combinatorial activation domain (AD) knockout in MYC. A. Schematic diagram showing the N-terminal portion of MYC with all seven combinatorial knockouts of the three ADs (ADKOs) by targeted alanine substitutions. Red crosses indicate the presence of inactivating alanine substitution sets. The final version, MYC AD-1KO AD-2KO AD-3KO is not expected to contain any activation activity. B. In vivo assay of the various the full-length MYC ADKO constructs in yeast. All reporter gene activities are shown relative to the MYC11-60 standard (column 2).

[0196] Figure 5 shows an embodiment of an anti-MYC bioPROTAC. A. Schematic diagram of the CR-l-bioPROTACs containing either a functional or mutant version of the SPOP E3167-374 ligase. B. Schematic diagram of CR-l-bioPROTAC action. CR-1 binds in trans to MYC proteins in cells. This results in recruitment of E2 ligase and ubiquitination of the target protein. The ubiquitinated protein is then recognized by proteosomes and degraded. C. Western blotting showing degradation of increasing amounts of MYC in the presence of CR-l-bioPROTACs containing functional SPOP, but no degradation with SPOPmut (note the small differences in the sizes due to the deletion in SPOPmut). Histone H3 is detected as a loading control. E. and F. Assaying the effect of the CR-l-bioPROTAC on transcription of a reporter genes. See text for more detail.

[0197] Figure 6 shows efficient mRNA delivery in primary humanised liver cancer model (A) or B-cell lymphomas (B).

[0198] Figure 7 shows a reporter gene assay to measure MYC-destruction efficiency of various anti-MYC bioPROTAC variants. A. Schematic principle of the MYC-sensor detection systems. The MYC-Producer plasmid drives the constitutive expression of full-length MYC. The MYC-Sensor plasmid contains four MYC-binding consensus sequences (E-boxes, red) to control expression of the Lucia reporter gene from a minimal promoter. After transfection of both plasmids into cells, MYC produced from the producer plasmid will bind to the E-boxes and stimulate the expression of secretable Lucia luciferase. B. Schematic diagram of the bioPROTAC variants tested. Each variant contains the anti-MYC CR1 warhead (light blue) fused via a linker domain (grey) to various types of E2-, or E3-ligases, or the LC3 protein from the autophagesomal pathway (as indicated). These constructs are co-transfected with the MYC-Producer and MYC Sensor. The efficacy of MYC degradation is measured by the luciferase reporter. C. In vivo quantitation of bioPROTAC efficacy to degrade intracellular MYC.

[0199] Figure 8 shows the phenotypic effect of expression of the CR1-SPO and CR1-SPOPmut bioPROTACs on HeLa cell viability. A. Schematic representation of the AAVOne (Serotype2) constructs. The bioPROTAC (CR1-SPOP or CRl-SPOPmut) is expressed from a constitutive CMV promoter. This transcript also contains a downstream Internal Ribosome Entry Sies (IRES; green) to produce a red fluorescent protein (dTomato, red) from the same transcript. This set-up allows cells that have been transduced successfully with this viral construct to be identified through their red fluorescence. The WPRE and poly(A) elements (light blue and orange, respectively) help with stabilization and processing of the transcript. B. HeLa cells were transduced with AAV2-SPOP (left image) or AAV2-SPOPmut (right image). Only cells expressing the inactive SPOPmut bioPROTAC are detected in fluorescence microscopy because cells taking up the CR1-SPOP version cannot survive due to the intracellular degradation of MYC on which HeLa cells are strongly dependent. C. Flow cytometry quantitation of HeLa cells transduced with CR1-SPOP (blue curve) or CRl-SPOPmut (green curve).

[0200] Examples

[0201] The oncoprotein MYC ('MYC') is a transcription factor that regulates the expression of target genes required for cell proliferation. MYC is overexpressed in more than 2 / 3 of all human cancers and its overexpression is critically relevant for survival of cancer cells due to the key role of MYC in cancer biology. A major proportion of MYC, apart from its DNA-binding domain, is intrinsically disordered. Based on exhaustive mutagenesis approaches, the inventors have derived a high-resolution map of the locations of three independent transcriptional activation domains (AD-1, AD-2, and AD3) based on assays is yeast and human cells. AD-1 is permanently functionally accessible, but the in vivo activities of AD-2 and AD-3 are negatively controlled by the presence of nearby " Compaction Regions" (CR1 and CR2, respectively). Based on results from two-hybrid assays, CR1 is also able to self-interact in trans. The trans interaction, involving two identical copies of CR1, is highly specific as shown by mutations in key residues, such as aspartate (D)132, that abolish it completely.

[0202] Furthermore, the CR1 trans interaction properties allow this motif to act as a warhead to degrade intracellular MYC when fused with SPOP E3 ligase as a bioPROTAC. Overall, the inventors propose that the granular intrinsic disorder of MYC constitutes a new regulatory phenomenon that modulates the physiological functions of MYC in different cellular contexts.

[0203] Materials & Methods

[0204] Atomistic molecular dynamics simulations of MYC and trajectory analyses

[0205] MD simulations were executed with the AMBER package as described previously

[0027] . tLEaP (AmberTools) was used to parameterise and prepare an initially unfolded structure comprising MYC1-352 for simulation

[0028] . The molecule was embedded in a TIP4P-D water box

[0029] (with a minimum of 25 A of distance between the protein structures and the solvent box border) and the ion concentration adjusted to a final concentration of 150 mM NaCI. The resulting files were minimized and used to run a conventional molecular dynamics (cMD) simulation using pmemd.cuda in mixed SPFP precision mode [30, 31]. For accelerated MD (aMD) simulations, total potential (EPTOT) and dihedral (DI-HED) energy values were extracted from the cMD simulation after 25 ns and applied to adjust acceleration parameters based on the recommended alpha-value (alpha= 0.2)

[0032] . The simulations (1 psec each for 12 independent simulations [aMD_nol-aMD_nol2]) were per-formed at 310 K, at 1 atmosphere of pressure.

[0206] AD Mapping and two-hybrid interaction assays in yeast

[0207] All constructs were derived from a codon-optimized version encoding the intrinsically disordered N-terminus of human MYC (MYC1-345, Figure SI). The desired portions of this construct were PCR amplified and cloned into the Ndel site of the yeast two-hybrid vectors pGBKT7 (and, if applicable, pGADT7) (Takara Bio). The in vivo transactivation potentials of the plasmids encoding GAL4-DBD fusion proteins in the yeast strain Y2HGold were assayed qualitatively by growth on solid medium lacking tryptophane (for selecting presence of the plasmid) and adenine and histidine (to select for transcriptional activity), and quantitatively by measuring the growth rate in liquid medium lacking these compounds. The cell density was measured after 36 hours growth in microplates at 600 nm (with shaking at 30°C). Two hybrid-interactions were measured in the same way using combinations of pGBKT7 and pGADT7 constructs containing portions of MYC1-345 in media lacking tryptophane, leucine (for selecting presence of both plasmids), adenine and histidine (to select for two-hybrid interactions).

[0208] Assay of AD activities in human cell culture

[0209] Various wildtype and mutated MYC fragments were joined C-terminally in frame to the yeast GAL4 DNA-binding domain and expressed constitutively from an expression plasmid (pBIND-SEAP; see below) from the ADH promoter. Plasmids used in the transfection assays was verified to be supercoiled monomers to ensure comparability

[0064] . The transfection efficiency was normalized by quantitating expression of secreted alkaline phosphatase (SEAP) expressed via a constitutive SV40 promoter on the same vector backbone. pBIND-SEAP was constructed by replacing the Renilla reniformis luciferase internal control with the open reading frame of SEAP (Addgene) in the Stul-Clal fragment of pBIND (Promega). For biological replicates, independent batches of cells were cultured and transfected in 48-well plates with independently purified plasmids. Luciferase activity was quantitated using a coelenterazine-based luminescence detection (QUANTI-Luc Gold, InvivoGen). Results from each construct in the technical replicates were averaged to provide a single data point for the biological replicates.

[0210] The transactivation activities of the GAL4-MYC fusion proteins were assayed by cotransfection with plasmid pG5-Lucia. pG5-Lucia was derived from pG5luc (Promega) by replacing the luciferase coding region with a codon-optimized secreted version of Lucia luciferase (InvivoGen). pBIND-SEAP and pG5-Lucia were transfected in an equimolar ratio. The combination of pBIND-SEAP and pG5-Lucia allows direct sampling of both SEAP (internal control) and Lucia luciferase (AD function) activity from the cell culture supernatant.

[0211] MYC-Sensor assay

[0212] To create a MYC-responsive reporter plasmid (" MYC Sensor 4x-5"), the GAL4 binding sites located with a KpnI-Hindlll fragment in plasmid pG5luc (Promega E2440) were replaced by a tandem array of four consensus MYC target sequences (E-boxes) interspersed with random sequence spacers of five nucleotides (ggtaccgagtttctagacggCACGTGCACATCACGTGGTACTCACGTGCTATCCACGTGaagacgc tagcggggggctataaaagggggtgggggcgttcgtcctcactctagatctgcgatctaagtaagctt - SEQ ID NO: 28; E-boxes shown in bold, Kpnl and Hindlll sites shown in italics, respectively). The presence of MYC stimulates transcription from the minimal TK promoter upstream of a secretable Lucia luciferase. HEK293-MYCKO cells were transfected (Fugene HD, Promega E2311) with equivalent amounts of pCDNA3-HA2 MYC (Addgene #74164) in combination with either pCDNA3-HA2CR1-SPOP167’374, pCDNA3-HA2CR1-SPOP167’374mut, pCDNA3-HA2CR1-VHL152'213, pCDNA3-HA2CR1-Socs2143'198, pCDNA3-HA2CR1-E2D11'147, pCDNA3-HA2CR1-E2B1'152, or pCDNA3-HA2CR1-LC31'125.

[0213] Creation of the AAVOne serotpye 2 CRl-SPOP(mut)-IRES-dTomato construct

[0214] The Adenovirus-Associated-Virus vector is designed to produce a transcript in transduced mammalian cells that encodes the open reading frames of CR1-SPOP (or SPOPmut) and the red fluroescent protein dTomato from a single transcript. The translation of the dTomato open reading frame is enable by including an Internal Ribosome Entry Sequence (IRES). For this task, we employed the single plasmid " AAVOne" system (Addgene #230930; https: / / aavnergene.com / aav-manufacturing / aavone-single-plasmid-production-system / ). Infectious AAV particles were produced by transfecting the CR1-SPOP (or SPOPmut) AAVOne plasmids into HEK293T cells. After 72 hours the cells and supernatant were harvested and lysed by three freeze-thaw cycles. After determining the titer of the lysate was determined with dPCR (Qiacuity probe mix for detection of the WPRE elements in viroids (https: / / www.qiagen.com / us / products / discovery-and-translational-research / pcr-qpcr-dpcr / dpcr-assays-kits-and-instruments / dpcr-assays / qiacuity-cell-and-gene-therapy-dpcr-assays). HeLa cells were transducted using a multiplicity of infection of 5,000, incubated for 72 hours and the level of red fluorescent protein detected by wide-field microscopy (Zeiss CellDiscoverer) and flow cytometry (Penteon).

[0215] Example 1: MYC contains three independent transcriptional activation domains The inventors identified portions of MYC acting as ADs using a 'sliding window mapping' (SWM) approach that systematically scans the entire MYC N-terminus as a series of fragments of predefined length of 50 amino acids whose endpoints differ by 10 amino acid positions (Fig. 1A). 50aa fragments are commonly considered to be more than capable of encompassing a complete AD (for example, [36, 37]), whereas the 30aa fragments may provide higher resolution concerning AD boundaries, and 70aa fragments may provide further information concerning the influence of neighbouring sequences on AD function

[0038] . The SWM method is objective (selection of boundaries not influenced by personal judgements), high-throughput (generation of large data sets allow detection of unexpected events) and automatable with robotic equipment. Yeast cells (Saccharomyces cerevisiae) naturally lack an endogenous version of MYC and therefore provide a neutral background in terms of potentially interfering biological functions caused by any of these fragments. On the other hand, some aspects of MYC function may require post-translational modifications, or the presence of coactivator- and chromatin functions specific to human cells. The inventors, therefore, carried out AD-mapping studies in both systems for comparison. For the detection of ADs, the MYC fragments were fused C-terminally to the GAL4 DNA-binding domain. Any AD activity displayed by the MYC fragments is subsequently detected in vivo using a yeast strain containing several GAL4 reporter genes, and independently in human cells by co-transfection with a reporter plasmid expressing GAL4-driven secreted Lucia luciferase (Fig. IB). The application of SWM across MYC1-352 reveals the presence of three distinct regions that act as independently active ADs in both yeast and human cells. For each AD detected, one particular construct displays the highest activity and followed by additional constructs shifted by ten or 20 amino acids that display detectable but lower activity (Fig. 1C, D). MYC11-60, MYC81-130, and MYC151-200 display three distinct AD activities and are separable from each other by constructs covering intermediate regions that display no detectable activity at all. The three ADs identified in this screen are subsequently referred to as AD-1, AD-2, and AD-3, respectively. All these positions correspond to region with a local acidic isoelectric point (

[0039] ) and are predictable by the AD prediction program AdPred

[0040] . Furthermore, in agreement with the broad expression spectrum of MYC in all human cells, the inventors found no evidence for any observable degree of cell-type specificity so far: all ADs were found to be active in a variety of human cancer cell lines and primary human fibroblast cells.

[0216] AD-1 (present in MYC11-60; Fig. 1C, D) spans the MB-0 motif that has previously been identified as a MYC AD in human cells [25, 26]. Extensive mutagenesis studies in both yeast and human cells revealed that, in addition to a "central" cluster of bulky hydrophobic amino acids (Y22 / F23 / Y24), the "left" (Y12 / L14 / Y16) and "right" (F31 / Y32) clusters make major contributors to AD-1 activity. Alanine substitutions in any one of these three clusters (either in isolation or in various combinations) do not diminish the activity of AD-1 significantly unless they occur in at least two, and preferably all three clusters; only alanine substitutions in all eight bulky hydrophobic positions (Y12 / L14 / Y16 / Y22 / F23 / Y24 / F31 / Y32) result in a near-complete loss of AD activity. Intriguingly, replacing individual or multiple positions of the central cluster residues with the least tolerated basic residues K or R

[0039] had little measurable effects, although according to the current understanding of activation domain function they are expected to disrupt activity to a considerable extent [36, 37, 40]. Also, substitutions in the central region with the o-helix breaking residue P do not have the predicted negative effects but actually enhanced AD-1 activity above the wildtype level). A previous study claimed that the region surrounding the adjacent MB-1 motif also acts as an AD

[0026] . Based on the results shown (Fig. 3B) and more than a dozen deletion constructs in that area (ROJW, data not shown), the inventors' results have not revealed any evidence to support such a notion. Longer fragments, such as MYC1-80, containing both AD-1 and MB-1, do not show an enhanced activation potential compared to MYC1-50. More significantly, they lack any detectable activity if the six alanine substitutions knocking out AD-1 (alanine substitutions of Y12 / L14 / Y16 / Y22 / F23 / Y24 / F31 / Y32) are introduced into such longer constructs, thus proving that any activation potential in these fragments is solely due to the presence of a functional AD-1.

[0217] AD-2 (present in MYC91-140; Fig. 1C, D) represents a strong activation domain centered around MYC98-108 [25, 27]. A similar site, MYC100-106 acts as a recruitment site for GCN5

[0033] , thus confirming at least two previous lines of evidence for AD activity in this region. Similar to AD-1, a combination of eight alanine substitutions (in F93, L105, L106, F115, F124, M134, W135, and F138) is required to knock out AD-2 function completely. Three of these residues (M134, W135, and F138) are located in the region conventionally referred to as MB-2, the evolutionarily most highly conserved motif in MYC. Since five other residues responsible for AD-2 activity are located N-terminally, it is evident that MB-2 is only part of AD-2. This is noteworthy because numerous studies in the past addressing MYC functions and interactions only deleted the MB-2 motif and therefore did not inactivate AD-2 transcriptional functions comprehensively.

[0218] AD-3 (present in MYC151-200; Fig. 1C, D) maps to a region that has not been up to now associated directly with AD function. It coincides with region containing MB-3 which has been shown to bind WDR5, a WD40-repeat-containing subunit of the MLL / SET histone methyltransferases involved in creating transcriptionally active chromatin

[0041] . Local protein-protein interactions between MYC and WDR5 have been implicated in contributing to promote binding of MYC to very active promoters

[0042] . Deletion of MB-3 reduced the activated transcription of only one (CDC2) of the four target genes tested. The proapoptotic activity of MYC is also enhanced in MB-3 deletion mutants, which may account for the observation that this mutation reduces the penetrance of induced tumours in animal models

[0017] . In both yeast- and human cell-based assays, the transactivation strength of AD-3 is less than the ones observed in AD-1 and AD-2, but still very distinct. A minimum of six alanine substitutions (in L176, Y177, L178, C188, F195, and Y197) is required to knock out AD-3 activity completely. A recent report covalent drug screen identified C171 as target for a covalent drug

[0043] . While this location is slightly N-terminal to the active residues identified, it is possible that the biological effects observed could be due, at least in part, to impacts on AD-3 function. The discovery of AD-3 may also have implications for the function of the naturally occurring MYC-S variant, where the polypeptide initiates at an internal methionine (MYC-M101) and therefore lacks the N-terminal 100 amino acids found in full-length MYC. The expression of MYC-S influences biological activities (stimulation of cell proliferation and apoptosis in immortalized cell lines

[0044] ). The weak transactivation and transrepression properties on endogenous cellular target genes

[0045] could be the result of AD-2 and / or AD-3 activity.

[0219] In summary, all three ADs are discrete motifs that work equally well in yeast and human cells, depend on a series of large hydrophobic residues for their activity and thus fit the general expectation of acidic type activation domain. Unusually though, the residues required for AD function are spread over at least 20 (in the case of AD-2 or AD-3) and as much as 45 positions in AD-2. Such a dispersed and redundant organization pattern make the ADs resilient to inactivation by mutations and offer additional means for regulating / fine-tuning their activities. For example, while mapping the locations and boundaries of AD-2 and AD-3, the inventors encountered an unexpected phenomenon: the activities of these two ADs appeared to be highly dependent on the surrounding sequence. Fragments from the 50aa sliding window screens show readily detectable activity of all three ADs (Fig. 1A), but data from a 70aa screen show an initial increase in AD-2 activity as the sliding window moves into the motif, but this activity is suddenly lost in fragments such as MYC91'160. The larger MYC91'160( 70aa) fragment, that encompasses all the sequence information present in the two smaller fragments, displays no detectable activity in transactivation assays. The same applies to AD-3, where a 70 aa fragment of MYC contains all residues shown to be required for AD-3 activity but fails to transactivate effectively.

[0220] The inventors hypothesize that these experimental observations suggest that AD-2 and AD-3 are potentially subject to conformational (allosteric) access control imposed by neighbouring sequence elements.

[0221] 2: Granular Intrinsic Disorder in the N-terminal Domain of MYC:

[0222] Computational and Experimental Insights

[0223] In the case of MYC, only the secondary structure propensity of a small part of the N-terminus has been determined experimentally so far (MYC1-88;

[0046] ). Therefore, to gain further insights into the regions responsible for conformational access control of AD-2 and AD-3, the inventors focused on further structural characterization of the MYC N-terminal region. The inventors modelled the three-dimensional structure of MYC by computational molecular dynamics simulation and using used local compaction plot (LCP) analysis (Weinzierl ROJ. " Molecular Dynamics Simulations of Human FOXO3 Reveal Intrinsically Disordered Regions Spread Spatially by Intramolecular Electrostatic Repulsion. Biomolecules". 2021 Jun 8;11(6):856. doi:

[0224] 10.3390 / biomll060856. PMID: 34201262; PMCID: PMC8228108.). Briefly, the inventors carried out extensive free-modelling strategy based on accelerated molecular dynamics (aMD,

[0047] ) to simulate the intrinsically disordered part the N-terminal 352 amino acids (MYC1-352) using the Amber ffl4SB force field

[0048] and the TIP4P-D water model optimized for simulating IDPs

[0049] . A starting structure for MYC1'352, an unfolded polypeptide chain lacking any structural definition, was folded initially by implicit MD in the presence of the Amberl4SBonlysc forcefield

[0050] for 250 nanoseconds. The resulting structure was subsequently used as the starting coordinates for twelve independent aMD simulations lasting a full microsecond each. The enhanced sampling conditions reflect time that is likely to be two to three magnitudes longer than the actual simulation time

[0047] ; the motions observed in the simulations are therefore predicted to occur within the hundreds of microseconds to millisecond time range. The aMD simulations confirm the extensively disordered structure of MYC1'352. Analysis of the secondary structures formed during the aMD simulations reveal a heterogeneous distribution of secondary structure elements fluctuating rapidly on a temporal and spatial scale. Extended regions appear fleetingly in several positions, but helices (a-, TT, and 3io -conformations), bends and turns are the dominant secondary structure theme of MYC1-352 (Fig. 2B). The central region (approximately spanning MYC100'230) is predicted to be especially enriched in o-helices that remain stable during around 70% percent of the simulation periods. Although the variability in secondary structure elements suggests a high degree of structural variability within MYC1-352, local compaction plot (LCP) analysis

[0012] reveals evidence for higher order structures, regularities and switching between alternative semi-stable conformations (Fig. 2B). In particular, the simulations suggest that two central semistable o-helices are flanked N- and C-terminally by sequences that are conformationally unusually confined, which we refer to descriptively as " Compaction Region 1" and " Compaction Region 2" (CR-1 and CR-2, respectively; Fig. 2B, C).

[0225] In the LCP, this confinement becomes evident from the shorter (10 - 25 A) distances (in windows covering the distances spaced 40 residues apart) consistently closely superimposed path on the plot (Fig. 2B). Despite the local conformational confinement, the CRs remain only in a metastable, but still disordered structure.

[0226] Representative examples show a tightly coiled, polypeptide strand formed by residues present immediately N-terminal to the MB-2 motif (Fig. 2C). Such local interactions have also been observed in other IDPs and involve residues in the interior that are depleted of solvent interactions and can thus engage in a network of coupled and decoupled portions of the polypeptide chain

[0051] . Overall, the MD results predict that, despite the overall intrinsically disordered nature of MYC1'352, two tightly arranged regions are predicted, intermixed with extended regions that are only capable of forming short and unstable secondary structure elements and can thus be classified as extensively disordered. The mixture of partially folded / compacted and extended conformations is referred to as "granular intrinsic disorder". Various experimental results presented below show that the granular intrinsic disorder of MYC1'352accounts -at least in part - for the difficulties and inconsistencies encountered in the past in defining the precise positions of ADs [25, 26]. The inventors further hypothesize that granular intrinsic disorder is likely to have biological consequences in terms of controlling conformational access to various functional regions in the intrinsically disordered portion of MYC.

[0227] To see if at least some of these internal compactions can be detected experimentally under in vivo conditions, the inventors carried out yeast two-hybrid studies (Y2H;

[0228]

[0052] ) with a range of MYC fragments of different sizes (Fig. 3; ROJW & BT, data not shown). Y2H experiments are conventionally applied to identify protein-protein interactions between two different proteins, but in this instance, we employed the assay as a tool to detect interactions between two polypeptide portions derived from the same protein. Although Y2H assays have been predominantly used in studies involving structured interaction partners, the inventors hypothesized that at least some of the only partially structured internal interactions predicted to occur in the CR-1 and CR-2 regions of MYC may also be observable in such a system. Note that this assay provides evidence for an intermolecular interaction ("trans") between two copies of MYC rather than the intramolecular ("cis") interactions observed in the MD simulations. Fusions of various fragments spanning CR-1 show indeed detectable binding to each other, with the strongest interaction involving two copies of MYC91-160 interacting in a homomeric manner. This assay failed, however, to pick up interactions involving CR-2, suggesting CR-2 binding may be less consistent or weaker than CR-1. To prove that the interaction between two copies of MYC91-160is not solely due to some rather non-specific property based on unusual hydrophobicity or charge composition of this sequence, the inventors prepared a systematic alanine scan of MYC91-160and tested the mutants for their ability to interact in the Y2H assay. The results show that, unexpectedly, the trans interaction property of MYC91-160is abolished by single alanine substitutions in a select number of distinct residues (including D109, Q113, E122, K126, N127, 1129, 1130, D132, M134, F138, L149, and Y152; Fig. 3C, D).

[0229] This result proves that the self-interaction depends strongly on very specific moieties present in key positions of the primary amino acid sequence, thus ruling out interaction models based on generic structural properties (such as excessive local enrichments in charge or hydrophobicity). Protein-protein interactions involving folded proteins are typically dependent on a large interface containing one or more hotspots of critically important residues that need to be mutated simultaneously to abolish an interaction completely

[0053] . The mutagenesis result presented here is therefore unusual in the sense that single alanine substitutions cause an "all-or-nothing" phenotype. In key positions, such as D132, an alanine substitution causes a complete loss of function so that even prolonged incubation of such colonies on selection medium shows no evidence of observable growth. To learn more about the structural requirements for the D132 position, the inventors carried out a complete substitution series with all 19 alternative amino acids. The result shows that only a substitution by another acidic amino acid (D132-E) supports the self-interaction property fully, while Q, N, H and R substitutions provide detectable, but only very partial functionality (Fig.

[0230] 3E). At least five of the residues (1129, 1130, D132, M134 and F138) identified in this assay are part of the MB-2 motif (Fig. 3A, F) thus identifying a potentially novel reason for the high degree of evolutionary conservation in this region. W135, the most highly conserved residue position in MB2, was not, however, identified as playing a significant role in this assay. A model, initially created by AlphaFold 3

[0054] , and subsequently refined by 1 ps of aMD simulation, suggests that the MB-2 motif folds into a semi-stable alpha-helical structure. The inventors propose that D132 acts as a hydrophilic / charged barrier between a cluster of isoleucines (1128 / 1129 / 1130) and the essentially invariant W135 residue. This may explain the mutational sensitivity of this position and why some hydrophilic (N, Q) and even oppositely charged residues (H, R) also display a minimal functionality when placed in this position.

[0231] The inventors have also tested fragments of the CR1 domain and surprisingly found that the self-interaction properties of the CR1 domain were not limited to the full-length CR1 domain (MYC91-160). For example, MYC96-160 (a truncation of 5aa at the N-terminus) and MYC91-160 (delta 101-110) (CR1 domain having a lOaa deletion of residues 101-110 of the MYC protein) bind with comparable affinity to MYC91-160 (full-length CR1 domain) in the Y2H protein interaction assay (Figure 3G).

[0232] In summary, especially for CR-1, the computational predictions and experimental evidence for strongly self-interacting regions in the N-terminal domain of MYC are in exquisite agreement. Rather than displaying an essentially uniform order of intrinsic disorder, the inventors detected evidence of a mixture of loosely and compacted conformations ("granular" disorder). The compacted conformations depend on the identity of key residues within and adjacent to the MB-2 motif.

[0233] Example 3: Functional availability of the MYC ADs

[0234] Studies aimed at identifying different MYC functions have hitherto employed deletions of various motifs (for example, [2, 55]), but considering the computational and experimental evidence for distinct local structures (Fig. 2, 3), such deletions may, even in an IDP, have inadvertently caused larger local and global conformational effects than intended, especially as these deletions also change the local charge density on IDPs that influences conformational states [14-16]. Knowledge of the precise location of the three ADs, in conjunction with specific information regarding the positions of key hydrophobic residues responsible for each of these activities, allowed the inventors to create combinatorial versions of MYC that lack on or more of the AD-functionalities using a minimum number of site-directed mutations (Fig. 4A). Intriguingly, assaying the ADs in full-length MYC detects mainly constructs that contain a functional AD-1 (Fig. 4B). In the 50 aa scanning assays AD-2 is the strongest AD (Fig. 3), but its inactivation in MYC AD-2KO and MYC AD-2KO AD-3KO has no detectable influence on the total activation activity of MYC under these conditions. The same conclusion applies to AD-3. In congruence with observations from the computational simulations, the locations of AD-2 and AD-3 correspond to regions that are predicted to be highly compacted (CR-1 and CR-2, respectively; Fig. IB, C), as well as capable of strong and highly specific internal self-interactions (Fig.

[0235] 2). The inventors, therefore, hypothesize that functional access to the activities of AD-2 and AD-3 are controlled by local and / or global conformations of MYC. These conformationally-restricted aspects are not detectable when the ADs are assayed in smaller fragments lacking the surrounding sequences to take up such higher order structures (50 aa scan; Fig. 3A). On the other hand, the hypothesis predicts that the AD-2 and AD-3 activities would alter if measured within the context of a larger fragment size. This is indeed the case. As already shown in a previous figure (Fig. 2), and shown in more detail (Fig. 4D), the activities of AD-2 and AD-3 are highly context-sensitive. A systematic 70aa scan of MYC shows that both AD-2 and AD-3 are highly susceptible to their larger context, whereas AD-1 remains active in the larger fragments (Fig. 4D). Thus, both the experimental and computational results presented earlier are fully capable of explaining the unusual failure to detect AD-2 and AD-3 activities in full-length MYC in these assays. The inventors further hypothesize that AD-2 and AD-3 do display their transcriptional AD activities in vivo, but only as a consequence of interactions with other proteins - and / or post-translational modification events - that alter the self-interaction properties of CR-1 and CR-2 (most likely by competing for the MYC-internal binding sites; see Discussion for further details).

[0236] Example 4: An anti-MYC directed bioPROTAC employing the MYC CR-1 warhead PROteolysis TArgeting Chimeras (PROTACs), based on bifunctional molecules to induce targeted degradation of an intracellular protein of interest, represents a significant advance in drug discovery and development. MYC-induced metabolic reprogramming of cancer cells make them especially dependent on continuously sustained high levels of MYC expression for survival (" MYC addiction"). Interfering with continuous MYC expression, even for only brief periods, triggers apoptotic cell death pathways [56, 57]. BioPROTACs, containing a protein-recognition "warhead" fused to an E3 ligase, have been shown to be effective in targeting intracellular proteins for specific proteosomal degradation (Fig. 5A, B;

[0058] ). Successful Y2H assay results (Fig.3) typically require Kd values in the micromolar range (or lower) between the protein interaction partners for successful detection, proving that CR-1 is capable of interacting with high affinity with itself. This insight immediately suggests that CR-1 could act as a MYC-specific warhead in a bioPROTAC to target intracellular degradation of MYC (Fig. 5B).

[0237] A previous study showed that the nuclear-located SPOP E3 ligase may be especially effective for degrading nuclear proteins. Furthermore, a small deletion in the SPOP protein (SPOPmut) abolishes the "3-box" motif responsible for binding to Cullin, thus creating an effective negative control for a bioPROTAC construct that still retains the ability to bind to the target protein via its warhead, but subsequently is unable to initiate the ubiquitination pathway

[0058] . The inventors, therefore, created CR-1 fusions to either SPOP or SPOPmut to test whether the trans-binding activity of the domain could also serve as an anti-MYC warhead (Fig. 5A). The CR-1 bioPROTACs were transfected into HEK293T-MYCKO cells in the presence of another plasmid producing increasing amounts of HA-tagged full-length MYC. Western blotting demonstrated the successful expression of the CR-1 BioPROTAC fusion proteins (containing either SPOP or SPOPmut) and showed highly effective degradation of MYC in the functional SPOP-fusion constructs (Fig.5C). Furthermore, results from experiments employing SPOPmut, as well as from cells grown in the presence of the proteosomal inhibitor MG132 (Fig. 5D), are entirely consistent with the expectation that the MYC degradation observed by the functional bioPROTAC is the result of specific targeting of intracellular MYC towards the canonical proteosomal degradation pathway. Apart from quantitating the protein levels of MYC by Western blotting, the inventors also asked whether they could assay the effect of MYC degradation through an independent functional assay. Using a fusion protein containing the intrinsically disordered region of MYC (MYC1-350) fused to the GAL4 DNA-binding domain to drive expression of Lucia-luciferase from a GAL4- responsive reporter construct in HEK293T-MYCKO cells showed that this was indeed the case. Expressing increasing amounts of the MYC1-350-GAL4 in the presence of either CR-l-SPOP or CR-l-SPOPmut shows that the bioPROTAC-dependent effect is both titratable and saturable. The inventors observed a maximum inhibition of the transcriptional activity of MYC of around 80% which gradually diminished when the amount of overexpressed MYC was titrated upward (Fig. 5E). This particular assay also enabled the inventors to gain some preliminary insights into the specificity of the CR-1 bioPROTAC. When the inventors fused MYC AD-1 (MYC11-60; Fig. ID) to GAL4, they observed strong stimulation of reporter gene expression, but no inhibition in the presence of the CR-l-bioPROTAC. This is entirely expected because, in contrast to MYC1-350-GAL4, MYC11-60-GAL4 is not capable of binding to CR-1 and therefore should not be degraded. The inventors observed a similar lack of inhibition with two additional unrelated transcriptional activators (SOX17351-414- GAL4 and VP16-GAL4) proving that the CR-l-bioPROTAC displays specificity towards MYC (Fig. 5F).

[0238] The HEK293T-MYCKO cells used in this assay are intrinsically able to grow in the absence of MYC and, as expected, therefore were not affected in overall viability and proliferation rate by the expression of the CR-1 bioPROTAC.

[0239] 5: Efficient mRNA in primary humanised liver cancer model or B-cell

[0240]

[0241] The inventors then set out to test the efficacy of the bioPROTAC in an in vivo animal cancer model (Figure 6).

[0242] Figure 6A shows testing the anti-MYC bioPROTAC in a liver cancer model. The first step is the engraftment of fluorescently labelled human cells in immune-compromised mouse liver. Next, lipid nanoparticles encapsulating mRNA encoding the anti-MYC bioPROTAC are delivered to the animal cancer model. Tissues were imaged and harvested to determine the anti-tumoral effects of the bioPROTAC.

[0243] Figure 6B shows testing the anti-MYC bioPROTAC in a blood cancer model. Briefly, the cancer model is based on a genetically engineered mouse strain which develops MYC- dependent lymphoma in the absence of any treatment. The efficacy of the anti-MYC bioPROTAC to treat blood cancers is then tested using anti-MYC bioPROTAC delivered as recombinant protein, or mRNA encoding the anti-MYC bioPROTAC delivered by lipid nanoparticles or viral vectors.

[0244] Example 6: alternative designs of a bioPROTAC employing the MYC CR-1 warhead To assay bioPROTAC activity, the inventors developed a MYC-Producer / MYC-Sensor assay (Figure 7A). In this cell transfection -based assay, MYC is expressed from a plasmid (the " MYC Producer") at high and constitutive levels in a human cell knock-out line that is unable to produce its MYC (HEK293T MYCKO). The inventors cotransfected, into the same cells, another plasmid (the " MYC-Sensor") that contains four consensus MYC target sites placed upstream of a minimal promoter driving the expression of a luciferase reported gene. The MYC-Sensor allows intracellular MYC to bind and to stimulate the expression of luciferase. An assay of luciferase activity thus reflects directly the level of the intracellular MYC concentration. The co-transfection of MYC-Producer and MYC-Sensor gives rise to a high luciferase activity due to the produced MYC binding to the target sites on the sensor. The inventors can use this system to assay the effects of our bioPROTACs by measuring the amount of MYC present after co-transfection of a mixture containing MYC-Producer, MYC-Sensor, and bioPROTAC producer.

[0245] Next, the inventors designed and tested different bioPROTACs which comprise the anti-MYC CR1 warhead (light blue) fused via a linker domain (grey) to various types of E2-, or E3-ligases, or the LC3 protein from the autophagesomal pathway (Figure 7B). As discussed above (Example 4), the inventors have also designed a negative control (CRl-SPOPmut) that lacks a small portion of the E3-ligase which inactivates the E3- ligase function.

[0246] Figure 7C shows the results of the MYC-Producer / MYC-Sensor assay for each of the different bioPROTACs. The results show that fusion of CR1 with SPOP gives the highest level of MYC degradation when compared to other E3- (VHL, Socs2) or E2-ligases (E2B, E2D1). In particular, the inventors have surprisingly found that the luciferase levels are reduced to less than 5% after treatment with the CR1-SPOP bioPROTAC. As expected, the inventors found that CRl-SPOPmut did not degrade intracellular MYC, thus proving that a fusion of the anti-MYC CR1 to a functional E3-ligase is crucial. These results clearly demonstrate that the CR1-SPOP bioPROTAC can effectively manipulate the concentration of MYC in cancer cells by reducing intracellular MYC concentrations by more than 95%. The inventors also found that fusing CR1 to LC3, a protein involved in autophageosomal degradation pathways, showed promising evidence of MYC-degradation, although not as effective as SPOP.

[0247] 7: bioPROTAC-mediated of MYC has an h on cells that rely a on MYC cancers and human cultured cancer

[0248]

[0249] cell Hi To study the physiological effects of bioPROTAC-induced degradation of endogenous MYC levels, the inventors employed the Adenovirus Associated Virus (AAV) system as a gene delivery tool. AAV vectors typically achieve 70-90% efficiency of infecting ("transducing") cells in cell culture, so this is a highly effective way of introducing genetic material into cells. The inventors produced constructs that express both the bioPROTAC and a fluorescent red protein from the same transcript (Figure 8A). Such a set-up allowed the inventors to identify precisely which cells have been transduced by their construct due to their red fluorescence so that they could study the physiological effect of expressing the bioPROTAC. The inventors chose the HeLa cancer cell line which overexpresses MYC and depends strongly on high levels of MYC for maintaining high proliferation rates. The results in Figure 8 show that cells transduced with the CR1-SP0P construct show very little fluorescence after 72 hours due to the deleterious effect of the anti-MYC bioPROTAC on cell viability. The inventors found that 94% of HeLa cells infected with the functional bioPROTAC-producing AAV virions die within 48 hours.

[0250] On the other hand, when cells were transduced with the CRl-SPOPmut construct (where SPOP is inactive due to a small deletion preventing its function), the inventors detected high levels of red fluorescent protein expression in all cells and found that less than 3% of cells infected with the mutant bioPROTAC die under the same conditions (this low rate of cell death also occurs in uninfected cells).

[0251] The difference between SPOP and SPOPmut in terms of MYC-degrading ability is also clearly shown in the luciferase reporter gene assays discussed in Example 6. Together, this data clearly shows the efficacy of MYC degradation by the CR-1 bioPROTAC and the therapeutic potential for killing MYC-dependent cancer cells.

[0252] Discussion

[0253] MYC mostly acts as a transcriptional activator in cells - its repression functions depend on association with MIZ1 rather than due to an integral repression domain [42, 59, 60]. An in-depth understanding of its ADs is therefore critical for assessing all major functional aspects of MYC in normal and cancer cells. Using a variety of mutagenesis strategies, the inventors mapped three distinct ADs (AD-1, AD-2, and AD-3) and identified the residues that define them as such. Previous work on mapping ADs in MYC dates back several decades, but the results presented in the research literature are inconsistent and contradictory. This is partially due to the assumption that ADs are linked to the MYC-boxes. While AD-1 and AD-3 do colocalize to MB-0 and MB-3a, respectively, other MYC-boxes, such MB-1 and MB-2, do not display AD activity directly (despite previously published claims;

[0026] ). Another problem afflicting previous studies is that AD-2 and AD-3 activities are typically not functionally detectable in larger constructs, so early attempts to map ADs in MYC using only a small number of fragments of variable size yielded inconsistent results. In general, eukaryotic ADs are intrinsically disordered (albeit with a significant o-helical propensity), which allows them to adopt a folded structure upon their recruitment to coactivators

[0061] . The inventors, therefore, hypothesize that in the context of fragments like MYC91-160, AD-2 is either stabilized in an unusually highly defined, but inactive, conformation or that AD-2 becomes "buried" by the surrounding sequences and thus no longer available for functional interactions with other components of the transcriptional machinery. The inventors further hypothesize that this provides a biologically relevant conformational switch that exposes or hides AD-2 as a consequence of MYC binding to other transcription factors, possibly involving the nearby MB-2 motif. A similar principle appears to apply to AD-3 that is - similar to AD-2 - only detectable in shorter constructs but becomes transcriptionally inactive within the context of additional surrounding sequence (AD-3 is adjacent / overlapping with the MB-3a motif). On the other hand, AD-1 is clearly not subject to such constraints and its presence / activity can be detected easily in larger fragments. Due to its exposed location near the N-terminus of MYC, AD-1 can extend freely without interference from neighbouring sequences

[0062] .

[0254] While many ADs in other transcription factors are generally considered to be freely accessible at all times, one of the first principles emerging from both computational and experimental work reported here is the observation that the intrinsically disordered portion of MYC does not form a completely random structure but contains, albeit in the fluid and flexible manner characteristic of IDPs, additional higher order structures. These Compaction Regions (CR-1 and CR-2) influence at least certain aspects of MYC's biological activity. The influences of CR-1 and CR-2 on adjacent transcriptional activation domains (AD-2 and AD-3, respectively) are clearly detectable in both yeast- and human cell-based assays because their activities are strongly masked in larger fragments containing the CRs. Especially for CR-1, we were able to conduct a more detailed dissection of this domain using systematic alanine-substitutions. While the inhibition of AD-2 activity through CR-1 appears to be mainly due to an in cis (intramolecular) interaction, similar to the models shown in Fig. 2C), Y2H-based protein interaction assays and the CR-l-bioPROTAC experiments clearly demonstrate that CR-1 is also capable of trans interactions between different MYC molecules. These trans interactions are sensitive to mutations. Although several of the point mutants appear to have no detectable effects, the inventors also detected positions where single alanine substitutions weakened or abolished the interaction in a highly effective ("all-or-nothing") manner. Highly specific protein-protein interactions involving IDPs as mutual binding partners have been described in several other systems but the exquisite sensitivity of such interactions to point mutants has not been documented in such a comprehensive manner before. From a small number of previous reports, it is also evident that MYC may not be an isolated example of conformational effects influencing the activities of gene-specific transcription factors: The FOS N-terminal region contains an inhibitor motif (IM 1) which, if mutated, enhances the ability of FOS to activate target promoters. Mutagenesis of two arginine residues within IM1 unmasks the AD present in FOS1-142

[0063] . Conclusions

[0255] Overall, this emphasizes the higher-order structure of the MYC intrinsically disordered domain that has a substantial - and frequently underestimated - influence on the outcome of functional assays. It is likely that the interpretation of other MYC data -such as mapping regions required for oncogenic transformation - will have to be revisited in light of the accessibility of such sequences within the context of MYC higher-order structure.

[0256] The inventors' knowledge and understanding of the positions and functional residues in the three ADs in MYC will benefit research on cancer biology, drug discovery and somatic cell reprogramming towards pluripotent stem cells. Furthermore, the existence of CR-1 and CR-2 have implications for studies aimed at binding drugs to MYC because such regions are more likely to form pockets for the specific binding of small molecules.

[0257] References

[0258] 1. Nair, S. K. and S. K. Burley, X-ray structures of Myc-Max and Mad-Max recognizing DNA. Molecular bases of regulation by proto-oncogenic transcription factors. Cell, 2003. 112(2): p. 193-205.

[0259] 2. Kalkat, M., et al., MYC Protein Interactome Profiling Reveals Functionally Distinct Regions that Cooperate to Drive Tumorigenesis. Mol Cell, 2018. 72(5): p. 836-848 e7.

[0260] 3. Lourenco, C., et al., MYC protein interactors in gene transcription and cancer. Nat Rev Cancer, 2021.

[0261] 4. Das, S. K., B. A. Lewis, and D. Levens, MYC: a complex problem. Trends Cell Biol, 2023. 33(3): p. 235-246.

[0262] 5. Dang, C. V., MYC on the Path to Cancer. Cell, 2012. 149(1): p. 22-35.

[0263] 6. Eilers, M. and R. N. Eisenman, Myc’s broad reach. Genes & Development, 2008.

[0264] 22(20): p. 2755-2766.

[0265] 7. Li, Y., S. C. Casey, and D. W. Felsher, Inactivation of MYC reverses tumorigenesis. J Intern Med, 2014. 276(1): p. 52-60.

[0266] 8. Pelengaris, S., M. Khan, and G. I. Evan, Suppression of Myc-induced apoptosis in beta cells exposes multiple oncogenic properties of Myc and triggers carcinogenic progression. Cell, 2002. 109(3): p. 321-34. 9. Larsson, L. G. and M. A. Henriksson, The Yin and Yang functions of the Myc oncoprotein in cancer development and as targets for therapy. Exp Cell Res, 2010. 316(8): p. 1429-37.

[0267] 10. Fletcher, S. and E. V. Prochownik, Small-molecule inhibitors of the Myc oncoprotein. Biochim Biophys Acta, 2015. 1849(5): p. 525-43.

[0268] 11. Dang, C. V., et al., Drugging the 'undruggable' cancer targets. Nat Rev Cancer, 2017. 17(8): p. 502-508.

[0269] 12. Weinzierl, R. O. J., Molecular Dynamics Simulations of Human FOXO3 Reveal Intrinsically Disordered Regions Spread Spatially by Intramolecular Electrostatic Repulsion. Biomolecules, 2021. 11(6).

[0270] 13. Muhar, M., et al., SLAM-seq defines direct gene-regulatory functions of the BRD4-MYC axis. Science, 2018. 360(6390): p. 800-805.

[0271] 14. Cowling, V. H. and M. D. Cole, Mechanism of transcriptional activation by the Myc oncoproteins. Seminars in Cancer Biology, 2006. 16(4): p. 242-252.

[0272] 15. Herkert, B. and M. Eilers, Transcriptional repression: the dark side of myc. Genes Cancer, 2010. 1(6): p. 580-6.

[0273] 16. Li, L. H., et al., c-Myc represses transcription in vivo by a novel mechanism dependent on the initiator element and Myc box II. EMBO J, 1994. 13(17): p. 4070-9.

[0274] 17. Herbst, A., et al., A conserved element in Myc that negatively regulates its proapoptotic activity. Embo Reports, 2005. 6(2): p. 177-183.

[0275] 18. Nie, Z., et al., c-Myc is a universal amplifier of expressed genes in lymphocytes and embryonic stem cells. Cell, 2012. 151(1): p. 68-79.

[0276] 19. Lin, C. Y., et al., Transcriptional amplification in tumor cells with elevated c-Myc. Cell, 2012. 151(1): p. 56-67.

[0277] 20. Lewis, L. M., et al., Replication Study: Transcriptional amplification in tumor cells with elevated c-Myc. Elife, 2018. 7.

[0278] 21. Nie, Z. Q., et al., Dissecting transcriptional amplification by MYC. Elife, 2020. 9.

[0279] 22. Campbell, KJ. and RJ. White, MYC Regulation of Cell Growth through Control of Transcription by RNA Polymerases I and III. Cold Spring Harbor Perspectives in Medicine, 2014. 4(5).

[0280] 23. Tansey, W. P., Mammalian MYC proteins and cancer. New Journal of Science, 2014. 2014: p. Artilce ID 757534.

[0281] 24. Salghetti, S. E., et al., Regulation of transcriptional activation domain function by ubiquitin. Science, 2001. 293(5535): p. 1651-3.

[0282] 25. Kato, G. J., et al., An amino-terminal c-myc domain required for neoplastic transformation activates transcription. Mol Cell Biol, 1990. 10(11): p. 5914-20.

[0283] 26. Zhang, Q., et al., MB0 and MBI Are Independent and Distinct Transactivation Domains in MYC that Are Essential for Transformation. Genes (Basel), 2017. 8(5). 27. Piskacek, M., et al., The 9aaTAD Activation Domains in the Yamanaka Transcription Factors Oct4, Sox2, Myc, and Klf4. Stem Cell Rev Rep, 2021. 17(5): p.

[0284] 1934-1936.

[0285] 28. Nikiforov, M. A., et al., TRRAP-dependent and TRRAP-independent transcriptional activation by Myc family oncoproteins. Molecular and cellular biology, 2002. 22(14): p. 5054-63.

[0286] 29. Rahl, P. B., et al., c-Myc regulates transcriptional pause release. Cell, 2010. 141(3): p. 432-45.

[0287] 30. Gargano, B., et al., P-TEFb is a crucial co-factor for Myc transactivation. Cell cycle (Georgetown, Tex ), 2007. 6(16): p. 2031-7.

[0288] 31. Eberhardy, S. R. and P. J. Farnham, Myc recruits P-TEFb to mediate the final step in the transcriptional activation of the cad promoter. The Journal of biological chemistry, 2002. 277(42): p. 40156-62.

[0289] 32. McMahon, S. B., et al., The novel ATM-related protein TRRAP is an essential cofactor for the c-Myc and E2F oncoproteins. Cell, 1998. 94(3): p. 363-74.

[0290] 33. Zhang, N., et al., MYC interacts with the human STAGA coactivator complex via multivalent contacts with the GCN5 and TRRAP subunits. Biochimica et biophysica acta, 2014. 1839(5): p. 395-405.

[0291] 34. Feris, E. J., J. W. Hinds, and M. D. Cole, Formation of a structurally-stable conformation by the intrinsically disordered MYC: TRRAP complex. PLoS One, 2019. 14(12): p. e0225784.

[0292] 35. Kessler, J. D., et al., A SUMOylation-Dependent Transcriptional Subprogram Is Required for Myc-Driven Tumorigenesis. Science, 2012. 335(6066): p. 348-353.

[0293] 36. Staller, M. V., et al., A High-Throughput Mutational Scan of an Intrinsically Disordered Acidic Transcriptional Activation Domain. Cell Syst, 2018. 6(4): p. 444-455 e6.

[0294] 37. Sanborn, A. L., et al., Simple biochemical features underlie transcriptional activation domain diversity and dynamic, fuzzy binding to Mediator. Elife, 2021. 10.

[0295] 38. Knight, A. and M. Piskacek, Cryptic inhibitory regions nearby activation domains. Biochimie, 2022. 200: p. 19-26.

[0296] 39. Ravarani, C. N., et al., High-throughput discovery of functional disordered regions: investigation of transactivation domains. Mol Syst Biol, 2018. 14(5): p. e8190.

[0297] 40. Erijman, A., et al., A High-Throughput Screen for Transcription Activation Domains Reveals Their Sequence Features and Permits Prediction by Deep Learning. Mol Cell, 2020. 78(5): p. 890-902 e6.

[0298] 41. Thomas, L. R., et al., Interaction with WDR5 promotes target gene recognition and tumorigenesis by MYC. Mol Cell, 2015. 58(3): p. 440-52. 42. Lorenzin, F., et al., Different promoter affinities account for specificity in MYC-dependent gene regulation. Elife, 2016. 5.

[0299] 43. Boike, L., et al., Discovery of a Functional Covalent Ligand Targeting an Intrinsically Disordered Cysteine within MYC. Cell Chemical Biology, 2021. 28(1): p. 4-+.

[0300] 44. Xiao, Q., et al., Transactivation-defective c-MycS retains the ability to regulate proliferation and apoptosis. Genes Dev, 1998. 12(24): p. 3803-8.

[0301] 45. Hirst, S. K. and C. Grandori, Differential activity of conditional MYC and its variant MYC-S in human mortal fibroblasts. Oncogene, 2000. 19(45): p. 5189-97. 46. Andresen, C., et al., Transient structure and dynamics in the disordered c-Myc transactivation domain affect Bini binding. Nucleic Acids Res, 2012. 40(13): p. 6353-66.

[0302] 47. Pierce, L. C., et al., Routine Access to Millisecond Time Scale Events with Accelerated Molecular Dynamics. J Chem Theory Comput, 2012. 8(9): p. 2997-3002.

[0303] 48. Maier, J. A., et al., ff14SB: Improving the Accuracy of Protein Side Chain and Backbone Parameters from ff99SB. Journal of Chemical Theory and Computation, 2015. 11(8): p. 3696-3713.

[0304] 49. Piana, S., et al., Water dispersion interactions strongly influence simulated structural properties of disordered protein states. J Phys Chem B, 2015. 119(16): p.

[0305] 5113-23.

[0306] 50. Nguyen, H., et al., Folding Simulations for Proteins with Diverse Topologies Are Accessible in Days with a Physics-Based Force Field and Implicit Solvent. J Am Chem Soc, 2014.

[0307] 51. Borgia, A., et al., Extreme disorder in an ultrahigh-affinity protein complex. Nature, 2018. 555(7694): p. 61-66.

[0308] 52. Fields, S. and O. K. Song, A Novel Genetic System to Detect Protein Protein Interactions. Nature, 1989. 340(6230): p. 245-246.

[0309] 53. Cunningham, B. C., et al., Receptor and antibody epitopes in human growth hormone identified by homolog-scanning mutagenesis. Science, 1989. 243(4896): p.

[0310] 1330-6.

[0311] 54. Abramson, J., et al., Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 2024. 630(8016): p. 493-500.

[0312] 55. Akifuji, C., et al., MYCL promotes iPSC-like colony formation via MYC Box 0 and 2 domains. Sci Rep, 2021. 11(1): p. 24254.

[0313] 56. Madden, S. K., et al., Taking the Myc out of cancer: toward therapeutic strategies to directly inhibit c-Myc. Mol Cancer, 2021. 20(1): p. 3.

[0314] 57. Sodir, N. M., et al., Reversible Myc hypomorphism identifies a key Myc-dependency in early cancer evolution. Nat Commun, 2022. 13(1): p. 6782. 58. Lim, S., et al., bioPROTACs as versatile modulators of intracellular therapeutic targets including proliferating cell nuclear antigen (PCNA). Proceedings of the National Academy of Sciences of the United States of America, 2020. 117(11): p. 5791-5800.

[0315] 59. Peukert, K., et al., An alternative pathway for gene regulation by Myc. EMBO J, 1997. 16(18): p. 5672-86.

[0316] 60. Walz, S., et al., Activation and repression by oncogenic MYC shape tumourspecific gene expression profiles. Nature, 2014. 511(7510): p. 483-7.

[0317] 61. Scholes, N. S. and R. O. Weinzierl, Molecular Dynamics of " Fuzzy" Transcriptional Activator-Coactivator Interactions. PLoS Comput Biol, 2016. 12(5): p. e1004935.

[0318] 62. Sullivan, S. S. and R. O. J. Weinzierl, Optimization of Molecular Dynamics Simulations of c-MYC(1-88)-An Intrinsically Disordered System. Life (Basel), 2020. 10(7).

[0319] 63. Klionsky, D. J., et al., Guidelines for the use and interpretation of assays for monitoring autophagy (4th edition)(1). Autophagy, 2021. 17(1): p. 1-382.

[0320] 64. Tudini, E., et al., Caution: Plasmid DNA topology affects luciferase assay reproducibility and outcomes. Biotechniques, 2019. 67(3): p. 94-96.

Claims

Claims1. An isolated polypeptide, derivative or analogue thereof, comprising a Compaction Region 1 (CR1) domain of the oncogenic transcription factor c- MYC (MYC), or a fragment or variant thereof.

2. The polypeptide according to claim 1, wherein the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially set out in SEQ ID NO: 2, or a fragment or variant thereof.

3. The polypeptide according to either claim 1 or 2, wherein the polypeptide, derivative or analogue thereof consists of, or comprises, a fragment of the CR1 domain of a native MYC protein, optionally wherein the fragment of the CR1 domain is a fragment of the CR1 domain having a length, based on the number of amino acid residues, that is at least 70%, 80% or 90% of the CR1 domain of a native MYC protein.

4. The polypeptide according to any one of the preceding claims, wherein the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially set out in SEQ ID NO: 2, wherein at least the last 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid residues from the N- terminus are absent, deleted or removed.

5. The polypeptide according to any one of the preceding claims, wherein the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially set out in SEQ ID No: 7, or a fragment or variant thereof.

6. The polypeptide according to any one of the preceding claims, wherein the polypeptide, derivative or analogue thereof consists of, or comprises, an amino acid sequence as substantially set out in SEQ ID NO: 2, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid residues from internal parts of the peptide sequence are absent, deleted or removed.

7. The polypeptide according to any one of the preceding claims, wherein the polypeptide, derivative or analogue thereof consists of, or comprises, anamino acid sequence as substantially set out in SEQ ID No: 8, or a fragment or variant thereof.

8. The polypeptide according to any one of the preceding claims, wherein the polypeptide, derivative or analogue thereof comprising the CR1 domain of a native MYC protein comprises a conservative substitution in one or more residues selected from D109, Q113, E122, K126, N127, 1129, 1130, D132, M134, F138, L149 and / or Y152 of a native MYC protein.

9. The polypeptide according to claim 8, wherein the polypeptide, derivative or analogue thereof comprising the CRT domain of a native MYC protein comprises a conservative substitution in D132 residue of a native MYC protein, optionally wherein the D132 residue is substituted with E, Q, N, H and R, typically wherein the D132 residue is substituted with E.

10. The polypeptide according to any one of the preceding claims, wherein the polypeptide, derivative or analogue thereof comprises substitutions in one or more residues selected from C133, S136 and S139 of a native MYC protein.

11. The polypeptide according to claim 10, wherein the polypeptide, derivative or analogue thereof comprises:(i) C133 residue substituted with A;(ii) S136 residue substituted with A; and / or(iii) S139 residue substituted with A.

12. A conjugate comprising the polypeptide, derivative or analogue thereof, according to any one of claims 1 to 11, and a payload molecule.

13. The conjugate according to claim 12, wherein the payload molecule is an enzyme, optionally a ligase, typically a ubiquitin ligase.

14. The conjugate according to claim 13, wherein the payload molecule is an E2- and / or E3-ubiquitin ligase.

15. The conjugate according to claim 14, wherein the payload molecule is an E3 ubiquitin ligase.

16. The conjugate according to claim 15, wherein the E3 ubiquitin ligase is selected from gTrCP, FBW7, SKP2, VHL, SPOP, CRBN, DDB2, SOCS2, ASB1, and / or CHIP.

17. The conjugate according to claim 12, wherein the payload molecule is a protein involved in autophagy, optionally wherein the protein is an autophagy-related protein.

18. The conjugate according to claim 17, wherein the autophagy-related protein is a member of the autophagy-related protein 8 (Atg8) protein family, optionally wherein the autophagy-related protein is microtubule-associated protein 1 light chain 3 (LC3).

19. The conjugate according to any one of claims 12-18, wherein the conjugate further comprises a linker.

20. The conjugate according to claim 19, wherein the linker is (GST)n, (GTGSG)n, and / or (GSG)n, wherein n is an integer of at least one.

21. The conjugate according to any one of claims 12-16 or 19-20, wherein the conjugate comprises the structure substantially as represented as SEQ ID NO: 10 or SEQ ID NO: 11, or a fragment or variant thereof.

22. The conjugate according to any one of claims 17-20, wherein the conjugate comprises the structure substantially as represented as SEQ ID NO: 23.

23. A nucleic acid sequence encoding the polypeptide, derivative or analogue thereof according to any one of claims 1-11, or the conjugate according to any one of claims 12-22.

24. A nucleic acid sequence according to claim 23, wherein the nucleic acid sequence is DNA or RNA.

25. A nucleic acid sequence according to claim 24, wherein the nucleic acid sequence is RNA, messenger RNA (mRNA), or synthetic mRNA.

26. A nucleic acid sequence according to claim 24, wherein the nucleic acid sequence comprises a nucleotide sequence substantially as set out in any one of SEQ ID No: 13-21 or 24-27, or a fragment or variant thereof.

27. An expression cassette comprising the nucleic acid sequence according to any one of claims 23-26.

28. A recombinant vector comprising the expression cassette according to claim 27.

29. A lipid nanoparticle comprising the nucleic acid sequence according to any one of claims 23-26, the expression cassette according to claim 27, or the recombinant vector according to claim 28.

30. A polypeptide according to any one of claims 1-11, a conjugate according to any one of claims 12-22, a nucleic acid sequence according any one of claims 23-26, an expression cassette according to claim 27, a vector according to claim 28, or a lipid nanoparticle according to claim 29 for use in therapy or diagnosis.

31. A polypeptide according to any one of claims 1-11, a conjugate according to any one of claims 12-22, a nucleic acid sequence according any one of claims 23-26, an expression cassette according to claim 27, a vector according to claim 28, or a lipid nanoparticle according to claim 29 for use in treating, ameliorating or preventing cancer, optionally liver cancer.

32. A pharmaceutical composition comprising the polypeptide according to any one of claims 1-11, the conjugate according to any one of claims 12-22, the nucleic acid sequence according any one of claims 23-26, the expression cassette according to claim 27, the vector according to claim 28, or the lipid nanoparticle according to claim 29 and a pharmaceutically acceptable vehicle.

33. A process for making the pharmaceutical composition according to claim 32, the process comprising contacting a therapeutically effective amount of the polypeptide according to any one of claims 1-11, the conjugate according to any one of claims 12-22, the nucleic acid sequence according any one of claims 23-26, the expression cassette according to claim 27, the vector according to claim 28, or the lipid nanoparticle according to claim 29 and a pharmaceutically acceptable vehicle.