Selective degradation of proteins
By expressing E3 ubiquitin ligase and fusion protein in host cells, a library of compounds that can bridge E3 ligase and target proteins was screened, which solved the problem of difficult to accurately degrade dysfunctional proteins in the prior art, and improved the production efficiency and activity of polyN-methylated cyclic peptides.
Patent Information
- Application Number
- CN202080051408.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2020-05-15
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2040-05-15
AI Technical Summary
The prior art is difficult to screen out proteins that can accurately and selectively degrade abnormal functional proteins, especially the production methods of polyN-methylated cyclic peptides, and the yield of compounds of natural origin is low, making it difficult to optimize their activity.
By expressing E3 ubiquitin ligase and fusion protein in host cells, using the interaction between the DNA binding moiety and the gene activation moiety, a library of compounds that can bridge E3 ligase and target proteins, including macrocyclic peptides and DNA-encoded peptides, to achieve selective target protein degradation.
Accurate and selective degradation of specific proteins in host cells, improves the productivity and activity of polyN-methylated cyclic peptides, and provides a diverse variety of compound variants for the treatment of various diseases.
Smart Images

Figure CN114127298B_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 848,509, filed May 15, 2019, and U.S. Provisional Application No. 62 / 854,273, filed May 29, 2019, the entire contents of each of which are incorporated herein by reference. Background Art
[0003] Degrading proteins in a precise manner may be key to controlling cellular function. Many pathological conditions are characterized by the aberrant function of cellular pathways, either due to premature protein expression or the expression of dysfunctional variants. Therefore, the ability to specifically and precisely degrade such proteins or the accumulation of dysfunctional variants could be beneficial in treating a variety of diseases. New technologies are being developed to discover and develop novel molecules that mediate protein degradation.
[0004] However, the options for screening molecules that achieve functional degradation in an efficient manner are limited. Therefore, there is a need to develop methods and compositions that achieve selective target protein degradation in a precise and selective manner. The method in the present invention describes a screening platform that can generate new substrates for specific E3 ligases by using a library of compounds that can bridge the interaction between the E3 ligase and the target protein, thereby leading to its degradation. This technology is applicable to various drug moieties, as well as DNA-encoded peptide and macrocyclic compound libraries, which are likely to be close drug candidates due to their bivalent nature. This platform describes a selection method in which only molecules that can produce functional target degradation are present.
[0005] Macrocyclic peptide natural products have been identified and isolated from a variety of species, including bacteria, fungi, plants, algae, mollusks, and mammals. They are recognized sources of a wide range of bioactive molecules. Several cyclic peptides of marine origin have received Food and Drug Administration (FDA) approval, such as ziconotide, a cyclic peptide isolated from the toxin of the cone snail species Conus magus. Ziconotide is an analgesic used for severe and chronic pain that acts by selectively blocking N-type calcium channels that control neurotransmission at many neural synapses. Macrocyclic peptide compounds have also shown considerable promise in a wide range of other therapeutic areas and have led to several clinically approved therapies for cancer, immunomodulation (e.g., cyclosporin A), and fungal infections (e.g., echinocandins).
[0006] The utility and application areas of these compounds are often limited by low yields from their natural sources, challenges in organic synthesis, and the inability to obtain a large number of variants to optimize activity. For this reason, biotechnological or semisynthetic methods starting from natural raw materials are often used for drug manufacturing. A particularly exciting group of macrocyclic peptides are multi-backbone N-methylated cyclic peptides (cyclic peptides N-methylated at multiple positions on the peptide backbone). These compounds have attractive pharmacological properties, such as increased bioavailability due to increased permeability through the intestinal epithelium and increased in vivo half-life due to increased stability to proteases. The archetypal representative of this peptide family is the immunomodulator cyclosporin A. Cyclosporin A is an 11-mer cyclic peptide originally isolated from the ascomycete Tolypocladium inflatum, which is synthesized by a highly complex non-ribosomal peptide synthetase (NRPS) specifically called cyclosporin synthase. Cyclosporin backbone methylation occurs during peptide elongation via a built-in methyltransferase domain within cyclosporin synthase. Over the past few decades, many groups have attempted to reengineer or evolve the NRPS machinery to produce altered versions or diversified derivatives of their natural products (e.g., different amino acids, different size cycles, different N-methylation patterns, etc.), but these efforts have proven futile and challenging.
[0007] Currently established methods for producing these types of poly-N-methylated cyclic peptides involve fermentation of large cultures of the corresponding microorganisms that naturally produce these compounds, followed by elaborate fractionation and purification methods. Some alternatives have been established for a few compounds that rely on total chemical synthesis or hybrid enzymatic and semisynthetic strategies.
[0008] Ribosomal synthesized and post-translationally modified peptides (RiPPs) are a diverse class of ribosomally derived natural products consisting of more than 22 subclasses produced by a variety of organisms, including bacteria, eukaryotes, and archaea. RiPPs are typically produced as all-L preproteins encoded by a gene that is transcribed by a conventional RNA polymerase II structure (in eukaryotes) and then translated by the ribosome. The active macrocycle is encoded within a cassette flanked by N-terminal and C-terminal signal recognition motifs. Post-translation, the all-L preprotein is processed by a group of modifying enzymes that introduce a variety of modifications (e.g., side-chain acylation, isomerization from some or all positions from L-amino acids to D-amino acids, side-chain hydroxylation, backbone N-methylation, end-to-end cyclization, disulfide bond formation, aminocarboxyethylthiotryptophan bridge formation, etc.), which release the cassette peptide from the preprotein and convert it into the final natural product. The N-terminal and C-terminal signal recognition motifs serve as docking sites for the processing enzymes and guide the order and kinetics of catalysis.
[0009] One of the recurring features of RiPP-processing enzymes is that many of them are virtually completely unaware of the sequence of the active peptide encoded within the cassette, thereby providing a high tolerance to substitutions within the cassette. Studies of the amanita toxin / hypodermatin / MSDin family of RiPPs from the poisonous mushroom toxic mushroom have confirmed the notion that the corresponding prolyl oligopeptidase / macrocyclase PopB is promiscuous towards a variety of naturally occurring active peptide cassette sequence variants, as well as many synthetically derived variants. These findings offer the possibility of a simple strategy to generate a wide variety of derivatives of RiPP-based natural products by simply altering the DNA of the cassette-encoding sequence within the proprotein-encoding gene. Summary of the Invention
[0010] Disclosed herein is a method for identifying one or more molecules that trigger degradation of a first test protein in a host cell. The method may include expressing in a host cell (i) an E3 ubiquitin ligase or a functional fragment thereof; (ii) a first fusion protein comprising a first DNA binding portion, a first test protein, and a first gene activation portion. The host cell may comprise a promoter sequence for controlling expression of a death agent, wherein the first DNA binding portion specifically binds to the promoter sequence. The molecule may be delivered to the host cell. In the absence of the molecule, expression of the death agent may be activated. In the presence of the molecule, the first test protein may be degraded by the E3 ubiquitin ligase.
[0011] In some embodiments, the method further comprises expressing a second fusion protein comprising a second DNA binding moiety, a second test protein, and a second gene activation moiety in a host cell, wherein the host cell may further comprise one or more positive selection reporter genes driven by one or more promoters having a sequence specific for the second DNA binding moiety.
[0012] In some embodiments, the method further comprises a plurality of positive selection reporter genes located within the host cell, wherein each positive selection reporter gene in the plurality of positive selection reporter genes is operably linked to a promoter sequence specific for the second DNA binding moiety. In some embodiments, the one or more positive selection reporter genes are encoded in a plasmid located within the host cell.
[0013] In some embodiments, the molecule can be part of a library of molecules. In some embodiments, the molecule from the library can be delivered exogenously. In some embodiments, the molecule can be produced intracellularly. In some embodiments, the molecule can be produced intracellularly from a DNA-encoded library.
[0014] In some embodiments, a method for identifying a molecule that selectively mediates degradation of a specific test protein in a host cell while retaining a second test protein is described. The method can include: expressing in the host cell a first fusion protein comprising a test protein having an activation domain and a DNA binding portion; a second test protein having an activation domain and a DNA binding portion; expressing in the host cell a third protein comprising an E3 ligase and desired ubiquitination machinery components; and delivering the molecule from the library to the host cell, wherein a gene sequence for expressing a negative selection death agent is located within the host cell and is operably linked to a promoter DNA sequence specific for the DNA binding portion of the first fusion protein, wherein a positive selection reporter gene is located within the host cell and is operably linked to a promoter DNA sequence specific for the DNA binding portion of the second fusion protein, and wherein, in the absence of the molecule, expression of the first test protein results in activation of expression of the death agent by the gene activation portion, while expression of the second test protein results in activation of expression of the positive selection reporter gene by the gene activation portion.
[0015] In some embodiments, molecules from the library are delivered exogenously. In some embodiments, the molecules are produced intracellularly. In some embodiments, the molecules are produced intracellularly from a DNA-encoded library. In some embodiments, the host cell comprises more than one sequence for expressing a positive control reporter gene, which is activated by a promoter DNA sequence specific for the DNA binding portion. In some embodiments, the host cell comprises more than one sequence for expressing a death agent, which is activated by a promoter DNA sequence specific for the DNA binding portion. In some embodiments, the host cell comprises an integrated DNA encoding a first fusion protein, an integrated DNA encoding a second fusion protein, an integrated DNA encoding a third fusion protein; a plasmid DNA encoding a death agent; and a plasmid DNA encoding a positive selection reporter gene.
[0016] In some embodiments, the first test protein is a variant of KRAS. In some embodiments, the second test protein is KRAS. In some embodiments, the first test protein is an androgen receptor splice variant ARV (ARV3, ARV7 or ARV9). In some embodiments, the second test protein is a wild-type androgen receptor. In some embodiments, the first test protein is a variant of IDH. In some embodiments, the second test protein is wild-type IDH. In some embodiments, the first test protein is Myc. In some embodiments, the first test protein is CCNE. In some embodiments, the first test protein is estrogen receptor (ER). In some embodiments, the first test protein is IKZF1 or IKZF2. In some embodiments, the first test protein is PD-1 or PDL-1. In some embodiments, the first test protein is CTLA-4. In some embodiments, the first test protein is Tau. In some embodiments, the first test protein is Act1 / CIKS (connected to IκB kinase and stress-activated protein kinase). In some embodiments, the first test protein is an Ets transcription factor variant (ETV1, ETV2, ETV3, ETV4, or ETV5). In some embodiments, the DNA binding portion is derived from LexA, cI, Gli-1, YY1, glucocorticoid receptor, TetR, or Ume6. In some embodiments, the gene activation portion is derived from VP16, GAL4, NF-κB, B42, BP64, VP64, or p65.
[0017] In some embodiments, the killing agent is an overexpression product of a genetic element selected from DNA or RNA. In some embodiments, the genetic element is a growth inhibitory (GIN) sequence, such as GIN11. In some embodiments, the killing agent is a ribosomally encoded xenobiotic agent, a ribosomally encoded poison, a ribosomally encoded endogenous or exogenous gene that causes severe growth defects when slightly overexpressed, a ribosomally encoded recombinase that excises a gene essential for viability, a restriction factor involved in the synthesis of a toxic secondary metabolite, or any combination thereof. In some embodiments, the lethal agent is cholera toxin, SpvB toxin, CARDS toxin, SpyA toxin, HopUl, Chelt toxin, Certhrax toxin, EFV toxin, ExoT, CdtB, diphtheria toxin, ExoU / VipB, HopPtoE, HopPtoF, HopPtoG, VopF, YopJ, AvrPtoB, SdbA, SidG, VpdA, Lpg0969, Lpgl978, YopE, SptP, SopE2, SopB / SigD, SipA, YpkA, YopM, amatoxin, phalloidin, killer toxin KP1, killer toxin KP6, killer toxin Kl, killer toxin K28 (KHR), killer toxin K28 (KHS), anthrax lethal factor endopeptidase, Shiga toxin, saporin toxin, ricin, or any combination thereof.
[0018] In some embodiments, the host cell is a fungus or a bacterium. In some embodiments, the fungus is Aspergillus. In some embodiments, the fungus is Pichia pastoris. In some embodiments, the fungus is Komagataella phaffii. In some embodiments, the fungus is Ustilago maydis. In some embodiments, the fungus is Saccharomyces cerevisiae.
[0019] In some embodiments, the drug is a small molecule. In some embodiments, the small molecule is a peptide mimetic. In some embodiments, the molecule is a peptide or protein. In some embodiments, the peptide or protein is derived from a naturally occurring protein product. In certain embodiments, the peptide or protein is a synthetic protein product. In some embodiments, the peptide or protein is a product of a recombinant gene. In some embodiments, the molecule is a peptide or protein expressed by a test DNA molecule inserted into a host cell, wherein the test DNA molecule comprises a DNA sequence encoding a polypeptide, forming a library. In some embodiments, the library comprises polypeptides having a length of 60 or fewer amino acids. In some embodiments, the DNA sequence encodes the 3'UTR of an mRNA. In some embodiments, the 3'UTR is the 3'UTR of sORF1. In some embodiments, the polypeptide comprises a common N-terminal sequence of methionine-valine-asparagine. In some embodiments, the polypeptides in the library are processed into cyclic peptides or bicyclic peptides in the host cell.
[0020] In certain embodiments, disclosed herein are plasmid vectors. In some embodiments, the plasmid vector comprises a DNA sequence encoding a first polypeptide inserted in a framework with a Gal4-DNA binding domain ("DBD") and a VP16 activation domain (AD), a DNA sequence encoding a second polypeptide inserted in a framework with an Ume6-DNA binding domain ("DBD") and a VP16 activation domain (AD), and a DNA sequence encoding a third polypeptide. In certain embodiments, a host cell comprises a plasmid vector.
[0021] In certain embodiments, disclosed herein are libraries of plasmid vectors, each comprising: a DNA sequence encoding a different peptide sequence operably linked to a first convertible promoter; a DNA sequence encoding a lethal agent under the control of a second convertible promoter; and a DNA sequence encoding a positive selection reporter gene under the control of a third convertible promoter. In some embodiments, the different peptide sequences encode a common N-terminal stabilizing sequence. In some embodiments, the DNA sequence encodes an mRNA sequence comprising a 3' untranslated region (UTR). In some embodiments, the different peptide sequences are 60 amino acids or fewer in length. In some embodiments, the different peptide sequences are randomized. In some embodiments, the different peptide sequences are pre-enriched for target binding. In some embodiments, the libraries are host cell libraries, each comprising a library of plasmid vectors.
[0022] In certain embodiments, disclosed herein are libraries of plasmid vectors, each comprising: a DNA sequence encoding a peptide N-methyltransferase operably linked to a first convertible promoter; and a prolyl oligopeptidase operably linked to a second convertible promoter. In some embodiments, the different peptide sequences are 18 amino acids or less in length. In some embodiments, the different peptide sequences are randomized. In some embodiments, the different peptide sequences are pre-enriched for target binding. In some embodiments, the libraries are host cell libraries, each comprising a library of plasmid vectors.
[0023] Described herein are host cells configured to accelerate the degradation of specific proteins. The host cell can express an E3 ubiquitin ligase or a functional fragment thereof; a first fusion protein comprising a first test protein, a first DNA binding portion, and a first gene activation portion; a death agent, wherein expression of the death agent is controlled by a promoter DNA sequence specific for the first DNA binding portion; and a polypeptide of 60 or fewer amino acids, wherein the polypeptide modulates the interaction between the first fusion protein and the E3 ubiquitin ligase in a manner that results in accelerated degradation of the first fusion protein.
[0024] In some embodiments, the host cell further comprises: a second fusion protein comprising a second DNA binding portion, a second test protein, and a second gene activation portion; and a positive selection reporter gene, wherein expression of the positive reporter gene is controlled by a second promoter DNA sequence specific for the second DNA binding portion.
[0025] In some embodiments, the polypeptide encodes an N-terminal sequence for peptide stabilization. In some embodiments, the polypeptide is a macrocycle. In some embodiments, the polypeptide is an N-methylated macrocycle. In some embodiments, the polypeptide is encoded by an mRNA, wherein the mRNA comprises a 3' UTR. In some embodiments, the mRNA is encoded by a DNA molecule, wherein the DNA molecule is exogenously delivered into a host cell. In some embodiments, libraries of synthetic compounds can be tested.
[0026] In some embodiments, the host cell is a eukaryote or a prokaryote. In some embodiments, the host cell is an animal, a plant, a fungus, or a bacterium. In some embodiments, the host cell is a haploid yeast cell. In some embodiments, the host cell is a diploid yeast cell. In some embodiments, the diploid yeast cell is produced by mating a first host cell comprising a DNA sequence encoding a first chimeric gene, a second chimeric gene, and a third chimeric gene with a second host cell comprising a DNA sequence encoding a death agent, a positive selection reporter gene, and an mRNA comprising a nucleotide sequence encoding a polypeptide. In some embodiments, the fungus is Aspergillus. In some embodiments, the fungus is Pichia pastoris. In some embodiments, the fungus is Ustilago mays.
[0027] Disclosed herein is a kit for accelerated degradation of a selective target protein. The kit may include a first plasmid vector encoding a first fusion protein comprising a first test protein that can be inserted in frame between a first DNA binding moiety and an activation domain; a second fusion protein that can be inserted in frame between a second DNA binding moiety and a second activation domain; and the aforementioned plasmid vector library.
[0028] In some embodiments, the kit further comprises a second plasmid vector configured to express the E3 ligase in the host cell.In some embodiments, the first vector may encode an E3 ubiquitin ligase.
[0029] Disclosed herein is a method for identifying one or more molecules that trigger degradation of a first test protein. The method may include expressing in a plurality of host cells: (i) an E3 ubiquitin ligase or a functional fragment thereof; and (ii) a first fusion protein comprising a first DNA binding portion, a first test protein, and a first gene activation portion. The plurality of host cells may each contain a promoter sequence for controlling expression of a death agent, and the first DNA binding portion may specifically bind to the promoter sequence such that in the absence of a molecule that recruits the E3 ubiquitin ligase to the first fusion protein in a manner that causes ubiquitination and premature degradation of the first fusion protein, expression of the death agent is activated. The method may include delivering a different molecule to each of the plurality of host cells and identifying the molecule that triggers degradation of the first test protein based on the survival of the cell to which the molecule is delivered.
[0030] In certain embodiments, disclosed herein is a method for identifying a molecule that selectively promotes an interaction between a first test protein and a second test protein, leading to their degradation, the method comprising: expressing in a host cell a first fusion protein comprising the first test protein and a DNA binding portion and a gene activation portion; expressing in the host cell a second fusion protein comprising the second test protein, a DNA binding portion, and a gene activation portion; expressing in the host cell a third protein comprising an E3 ubiquitin ligase; and delivering a molecule from a library to the host cell such that the molecule forms a bridging interaction between the first test protein and the E3 ubiquitin ligase, leading to their selective degradation; wherein a gene sequence for expressing a death agent is located within the host cell and is operably linked to a promoter DNA sequence specific for the DNA binding portion of the first fusion protein; wherein a positive selection reporter gene is located within the host cell and is operably linked to a promoter DNA sequence specific for the DNA binding portion of the second fusion protein. The first test protein can form a functional transcription factor that activates expression of the death agent; and the second test protein can form a functional transcription factor that activates expression of the positive selection reporter gene.
[0031] In some embodiments, the host cell comprises more than one sequence for expressing a killing agent that is activated by a promoter DNA sequence specific for the DNA binding moiety. In some embodiments, the host cell comprises more than one sequence for expressing a positive control reporter gene that is activated by a promoter DNA sequence specific for the DNA binding moiety.
[0032] In some embodiments, the host cell comprises integrated DNA encoding a first fusion protein, integrated DNA encoding a second fusion protein, integrated DNA encoding a third fusion protein; plasmid DNA encoding a killing agent; and plasmid DNA encoding a positive selection reporter gene.
[0033] In some embodiments, the DNA binding portion is derived from LexA, cI, Gli-1, YY1, glucocorticoid receptor, TetR or Ume6. In some embodiments, the gene activation portion is derived from VP16, Gal4, NF-κB, B42, BP64, VP64 or p65. In some embodiments, the death agent is a genetic element, wherein overexpression of the genetic material results in growth inhibition of the host cell. In some embodiments, the death agent is an overexpression product of DNA. In some embodiments, the death agent is an overexpression product of RNA. In some embodiments, the gene sequence for expressing the death agent is a growth inhibition (GIN) sequence, such as GIN11. In some embodiments, the death agent is a ribosomally encoded xenobiotic agent, a ribosomally encoded poison, a ribosomally encoded endogenous or exogenous gene that causes severe growth defects when slightly overexpressed, a ribosomally encoded recombinase that excises a gene essential for viability, a restriction factor involved in the synthesis of toxic secondary metabolites, or any combination thereof. In some embodiments, the lethal agent is cholera toxin, SpvB toxin, CARDS toxin, SpyA toxin, HopUl, Chelt toxin, Certhrax toxin, EFV toxin, ExoT, CdtB, diphtheria toxin, ExoU / VipB, HopPtoE, HopPtoF, HopPtoG, VopF, YopJ, AvrPtoB, SdbA, SidG, VpdA, Lpg0969, Lpgl978, YopE, SptP, SopE2, SopB / SigD, SipA, YpkA, YopM, amatoxin, phalloidin, killer toxin KP1, killer toxin KP6, killer toxin Kl, killer toxin K28 (KHR), killer toxin K28 (KHS), anthrax lethal factor endopeptidase, Shiga toxin, saporin toxin, ricin, or any combination thereof.
[0034] In some embodiments, the first test protein is a variant of KRAS. In some embodiments, the second test protein is KRAS. In some embodiments, the first test protein is an androgen receptor splice variant ARV (ARV3, ARV7 or ARV9). In some embodiments, the second test protein is a wild-type androgen receptor. In some embodiments, the first test protein is a variant of IDH. In some embodiments, the second test protein is wild-type IDH. In some embodiments, the first test protein is Myc. In some embodiments, the first test protein is CCNE. In some embodiments, the first test protein is estrogen receptor (ER). In some embodiments, the first test protein is IKZF1 or IKZF2. In some embodiments, the first test protein is PD-1 or PDL-1. In some embodiments, the first test protein is CTLA-4. In some embodiments, the first test protein is Tau. In some embodiments, the first test protein is Act1 / CIKS (connected to IκB kinase and stress-activated protein kinase). In some embodiments, the first test protein is an Ets transcription factor variant (ETV1, ETV2, ETV3, ETV4, or ETV5).
[0035] In some embodiments, the drug is a small molecule. In some embodiments, the small molecule is a peptide mimetic. In some embodiments, the molecule is a peptide or protein. In some embodiments, the peptide or protein is derived from a naturally occurring protein product. In some embodiments, the peptide or protein is a synthetic protein product. In some embodiments, the peptide or protein is a product of a recombinant gene. In some embodiments, the peptide or protein is the expression product of a test DNA molecule inserted into a host cell, wherein the test DNA molecule comprises a DNA sequence encoding a polypeptide, forming a library. In some embodiments, the library comprises sixty or fewer amino acids.
[0036] In some embodiments, the peptide or protein is a product of post-translational modification. In some embodiments, the post-translational modification comprises cleavage. In some embodiments, the post-translational modification comprises cyclization. In some embodiments, the post-translational modification comprises dicyclization. In some embodiments, the cyclization comprises reaction with a prolyl endopeptidase. In some cases, the prolyl endopeptidase can be one selected from SEQ ID NOs: 42-58 or a functional fragment thereof. In some cases, the prolyl endopeptidase can be one having at least 80%, 85%, 90%, 92%, 95%, 97%, or 99% sequence identity to one of SEQ ID NOs: 42-58.
[0037] In some embodiments, the cyclization comprises reacting with a beta-lactamase. In some cases, the beta-lactamase can be one selected from SEQ ID NOs: 119-120 or a functional fragment thereof. In some cases, the beta-lactamase can be one having at least 80%, 85%, 90%, 92%, 95%, 97%, or 99% sequence identity to one of SEQ ID NOs: 119-120.
[0038] In some embodiments, the bicyclization comprises reacting with a hydroxylase and a dehydratase. In some cases, the hydroxylase can comprise SEQ ID NO: 123 or a functional fragment thereof. In some cases, the hydroxylase can be one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to SEQ ID NO: 123. In some cases, the dehydratase can be one selected from SEQ ID NOs: 124-127 or a functional fragment thereof. In some cases, the dehydratase can be one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to one of SEQ ID NOs: 124-127.
[0039] In some embodiments, the bicyclization is formed by an aminocarboxyethylthiotryptophan bridge. In some embodiments, the post-translational modification comprises methylation. In some embodiments, the methylation comprises reaction with an N-methyltransferase. In some cases, the N-methyltransferase is selected from one of SEQ ID NOs: 61-116 or a functional fragment thereof. In some cases, the N-methyltransferase can be one having at least 80%, 85%, 90%, 92%, 95%, 97%, or 99% sequence identity to one of SEQ ID NOs: 61-116. In some embodiments, the post-translational modification comprises halogenation.
[0040] In some embodiments, the post-translational modification comprises glycosylation. In some embodiments, the post-translational modification comprises acylation. In some embodiments, the post-translational modification comprises phosphorylation. In some embodiments, the post-translational modification comprises acetylation.
[0041] In some embodiments, the test DNA molecule comprises a gene sequence that expresses a modifying enzyme.
[0042] In some embodiments, the test DNA molecule comprises a gene sequence that expresses an N-terminal sequence of methionine-valine-asparagine. In some embodiments, the test DNA molecule comprises a gene sequence that encodes a 3'UTR. In some embodiments, the 3'UTR is the 3'UTR of sORF1.
[0043] In some embodiments, the host cell is a eukaryote or a prokaryote. In some embodiments, the host cell is an animal, a plant, a fungus, or a bacterium. In some embodiments, the fungus is Aspergillus. In some embodiments, the fungus is Pichia pastoris. In some embodiments, the fungus is Komagataella phaffii. In some embodiments, the fungus is Ustilago maydis.
[0044] Disclosed herein are compositions and methods comprising genes and peptides related to cyclic and backbone-methylated macrocyclic peptides and macrocyclic peptide production in cells. In particular, the present invention relates to the use of genes and proteins from species of the genus Gymnopus that encode peptides particularly related to naked peptides, in addition to proteins involved in the processing of such cyclic peptides. In a preferred embodiment, the present invention also relates to methods for preparing small peptides and small cyclic peptides (including peptides such as naked peptides) by heterologous expression in eukaryotic, prokaryotic, or cell-free systems.
[0045] The method further describes the use of heterologous enzymes to generate macrocycle libraries with potential N-terminal methylation events within host cells to enable screening for functional molecules. In some cases, the functional molecules can modulate the interaction between two proteins to disrupt or bridge protein-protein interactions. Also described is the selection of modified macrocycles capable of bridging the interaction between an E3 ubiquitin ligase and a protein, leading to functional degradation of the protein.
[0046] In some embodiments, methods for producing cyclic peptides are described herein. The methods for producing cyclic peptides can include recombinantly expressing a prolyl oligopeptidase; and contacting the prolyl oligopeptidase with a linear peptide such that the linear peptide is converted into a cyclic peptide; wherein the active site of the prolyl oligopeptidase does not have a tryptophan residue at a position corresponding to amino acid position 603 of SEQ ID NO: 55, and / or does not have an asparagine residue at a position corresponding to amino acid position 563 of SEQ ID NO: 55.
[0047] In some embodiments, the prolyl oligopeptidase has a leucine residue at amino acid position 603 corresponding to amino acid position 603 of SEQ ID NO: 55, and / or a serine residue at amino acid position 563 corresponding to amino acid position 563 of SEQ ID NO: 55. In some embodiments, the prolyl oligopeptidase comprises a sequence corresponding to any one of SEQ ID NOs: 42-58. In some embodiments, the prolyl oligopeptidase is one having at least 80%, 85%, 90%, 92%, 95%, 97%, or 99% sequence identity to one of SEQ ID NOs: 42-58.
[0048] In some embodiments, contacting the prolyl oligopeptidase with a linear peptide occurs intracellularly. In some embodiments, contacting the prolyl oligopeptidase with a linear peptide does not occur intracellularly. In some embodiments, the linear peptide is recombinantly expressed. In some embodiments, the cyclic peptide comprises 18 or more amino acids.
[0049] In some embodiments, described herein are methods for identifying a cyclic peptide that disrupts an interaction between a first test protein and a second test protein. The method can include: (a) expressing in a host cell: (i) a first fusion protein comprising a first DNA binding portion, a first test protein, and a first gene activation portion; and (ii) the second test protein; and (b) delivering the peptide to the host cell; wherein, in the absence of the cyclic peptide, expression of the killing agent is activated.
[0050] In some embodiments, described herein are methods for identifying a cyclic peptide that bridges an interaction between a first test protein and a second test protein. The method can include: (a) expressing in a host cell: (i) a first fusion protein comprising a first DNA binding portion, a first test protein, and a first gene activation portion; and (ii) the second test protein; and (b) delivering the peptide to the host cell; wherein, in the absence of the cyclic peptide, expression of the killing agent is activated.
[0051] In some embodiments, the present invention provides a method for methylating a peptide. The method may include: recombinantly expressing an N-methyltransferase; and contacting the N-methyltransferase with the peptide such that multiple nitrogen atoms in the peptide backbone are methylated; wherein the N-methyltransferase comprises a sequence of one of SEQ ID NOs: 61-116.
[0052] In some embodiments, the peptide contacted is a cyclic peptide. In some embodiments, the peptide or cyclic peptide comprises 18 or more amino acids. In some embodiments, contacting the N-methyltransferase with the peptide occurs intracellularly. In some embodiments, contacting the N-methyltransferase with the peptide does not occur intracellularly. In some embodiments, the peptide is recombinantly expressed.
[0053] In some embodiments, described herein are methods for identifying a cyclic peptide that disrupts an interaction between a first test protein and a second test protein. The method can include: (a) expressing in a host cell: (i) a first fusion protein comprising a first DNA binding portion, a first test protein, and a first gene activation portion; and (ii) the second test protein; and (b) delivering the peptide to the host cell; wherein, in the absence of the cyclic peptide, expression of the killing agent is activated.
[0054] In some embodiments, described herein are methods for identifying a cyclic peptide that bridges an interaction between a first test protein and a second test protein. The method can include: (a) expressing in a host cell: (i) a first fusion protein comprising a first DNA binding portion, a first test protein, and a first gene activation portion; and (ii) the second test protein; and (b) delivering the peptide to the host cell; wherein, in the absence of the cyclic peptide, expression of the killing agent is activated.
[0055] Incorporation by reference
[0056] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The features of the present disclosure are particularly set forth in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by referring to the following detailed description and the accompanying drawings, which set forth illustrative embodiments in which the principles of the present disclosure are utilized, and in which:
[0058] Figure 1 Diagram shows a platform for identifying compounds that specifically mediate the interaction of a protein (bait) with an E3 ubiquitin ligase, leading to its degradation and consequently the loss of expression of positive markers required for cell growth.
[0059] Figure 2 Illustration of a platform for identifying compounds that mediate the interaction of proteins with E3 ubiquitin ligases, leading to their degradation.
[0060] Figure 3 Illustration of a platform for identifying compounds that specifically mediate the interaction of a protein (bait 2) with an E3 ubiquitin ligase, leading to its degradation, while maintaining active expression of another protein (bait 1).
[0061] Figure 4 A platform for identifying compounds that specifically mediate the expression of proteins (bait MUT ) interacts with an E3 ubiquitin ligase, leading to its degradation while maintaining another protein (the bait WT ) active expression.
[0062] Figure 5A The diagram illustrates the conservation of the tryptophan residue present in the active site of related prolyl oligopeptidases that are unable to macrocyclize larger peptides. The sequence reads illustrate the loss of the tryptophan residue in the active site of a prolyl oligopeptidase isolated from Gymnopus fusipes that is able to macrocyclize larger peptides.
[0063] Figure 5B The structure of a plant prolyl oligopeptidase homolog is shown, with the arrow pointing to the tryptophan residue in the active site of the enzyme at position 603.
[0064] Figure 5C Illustration of the conservation of the asparagine residue present in the active site of related prolyl oligopeptidases that are unable to macrocyclize larger peptides. Sequence reads illustrate the loss of the asparagine residue in the active site of a prolyl oligopeptidase isolated from Thielavia fuscipiens that is able to macrocyclize larger peptides.
[0065] Figure 5D The structure of a plant prolyl oligopeptidase homolog is shown, with the arrow pointing to the asparagine residue in the active site of the enzyme at position 563.
[0066] Figure 6 Diagram shows a vector platform for generating a library of macrocyclic bridgers using a methyltransferase and a prolyl oligopeptidase or lactamase.
[0067] Figure 7 Schematic representation of a vector platform for generating a library of macrocyclic bridgers using methyltransferases and processing enzymes.
[0068] Figure 8 Illustrated using Figure 1 The mechanism described in the assay is a negative readout for target protein degradation.
[0069] Figure 9 Illustrated using Figure 1 The mechanism described in the assay is a positive readout for target protein degradation.
[0070] Figure 10A -C illustrates a positive readout assay using high throughput screening of bridging agents such as NAA, PAA, and 2,4-DCPA to degrade target proteins.
[0071] Figure 11A -C illustrates a positive readout assay using high-throughput screening of bridging agents such as coronatine to degrade target proteins. DETAILED DESCRIPTION
[0072] The present disclosure provides a system that can use a unified eukaryotic or prokaryotic one-hybrid system, wherein the bait expression plasmid is used in both biological environments. In addition, a wide range of leucine zipper fusion proteins of known affinity can be generated to compare the efficiency of interaction detection using such a system. The yeast system can produce quantitative readouts across a dynamic range. In addition, the modified expression vectors disclosed herein can be used to express proteins of interest in both eukaryotic and prokaryotic organisms.
[0073] The present disclosure also provides a system for delivering molecules across cell membranes. The cell membrane is a major challenge in drug discovery, especially for biologics such as peptides, proteins, and nucleic acids. A potential strategy to subvert the membrane barrier and deliver biologics into cells is to attach them to "cell-penetrating peptides" (CPPs). Despite three decades of investigation, the fundamental basis of CPP activity remains elusive. CPPs that enter cells via endocytosis typically exit endocytic vesicles to reach the cytoplasm. Unfortunately, the endosomal membrane has been shown to be a significant barrier to the cytoplasmic delivery of these CPPs, so a negligible portion of the peptide typically escapes into the cell interior. Therefore, what is needed are new scaffolds and structures that impart peptides with highly proficient intrinsic cell penetration capabilities for various cell types. Several naturally occurring polyketides and peptides exhibit significant cell permeability (e.g., cyclosporin and amanitin). These peptides are characterized by specific modifications (e.g., N-methylation and cyclization or dicyclization of the backbone), which may play a key role in their cell membrane permeability. The compositions and methods disclosed herein describe methods and approaches that can be generally utilized to generate compositions that may have high therapeutic value and may be capable of degrading proteins with high selectivity.
[0074] definition
[0075] As used herein, a "reporter gene" refers to a gene whose expression can be measured. Such genes include, for example, LacZ, β-glucuronidase (GUS), amino acid biosynthesis genes, yeast LEU2, HIS3, LYS2, or URA3 genes, nucleic acid biosynthesis genes, mammalian chloramphenicol acetyltransferase (CAT) gene, green fluorescent protein (GFP), or any surface antigen gene for which there is a specific antibody. Reporter genes can result in both positive and negative selection.
[0076] "Allele" refers to the DNA sequence of a gene, including naturally occurring or pathogenic variants of the gene. Expression of different alleles may result in different protein variants.
[0077] A "promoter" is a DNA sequence located near the start of transcription at the 5' end of an operably linked transcribed sequence. A promoter may contain one or more regulatory elements or modules that interact in regulating the transcription of the operably linked gene. Promoters may be switchable or constitutive. Switchable promoters allow for reversible induction or repression of an operably linked target gene upon administration of an agent. Examples of switchable promoters include, but are not limited to, the LexA operon and the alcohol dehydrogenase I (alcA) gene promoter. Examples of constitutive promoters include the human β-actin gene promoter.
[0078] "Operably linked" describes two macromolecular elements arranged so that modulation of the activity of the first element induces an effect on the second element. In this way, modulation of the activity of a promoter element can be used to alter or regulate the expression of an operably linked coding sequence. For example, transcription of a coding sequence operably linked to a promoter element can be induced by factors that activate promoter activity; transcription of a coding sequence operably linked to a promoter element can be inhibited by factors that inhibit promoter activity. Thus, a promoter region is operably linked to a protein coding sequence if transcription of such coding sequence activity is affected by promoter activity.
[0079] As used throughout this document, "in frame" refers to the correct positioning of the desired nucleotide sequence within a DNA segment or coding sequence that is operably linked to a promoter sequence, thereby allowing transcription and / or translation.
[0080] "Fusion construct" refers to a recombinant gene encoding a fusion protein.
[0081] A "fusion protein is" a hybrid protein, i.e., a protein that has been constructed to contain domains from at least two different proteins. The fusion proteins described herein can be hybrid proteins that have (1) a transcriptional regulatory domain from a transcriptional regulatory protein or a DNA binding domain from a DNA binding protein and (2) a heterologous protein whose interaction state is to be determined. The protein from which the transcriptional regulatory domain is derived may be different from the protein from which the DNA binding domain is derived. In other words, the two domains may be heterologous to each other.
[0082] The transcriptional regulatory domain of the bait fusion protein can activate or repress the transcription of the target gene, depending on the biological activity of the domain. The bait protein of the present disclosure can also be part of a fusion protein, in which the protein of interest is operably linked to a DNA binding portion and a transcriptional activation domain.
[0083] A "bridging interaction" refers to an interaction between a first protein and a second protein that occurs only when one or both of the first and second proteins interact with a molecule, such as a peptide or small molecule from a library. In some cases, the bridging interaction between the first and second proteins is direct, while in other cases, the bridging interaction between the first and second proteins is indirect. In some cases, the interaction results in the activity of one protein acting on the second protein, such as ubiquitination and subsequent degradation.
[0084] "Expression" is the process by which the information encoded within a gene is revealed. If the gene encodes a protein, expression includes transcription of the DNA into mRNA, processing of the mRNA (if necessary) into a mature mRNA product, and translation of the mature mRNA into protein.
[0085] As used herein, "cloning vehicle" is any entity capable of delivering a nucleic acid sequence into a host cell for cloning purposes. Examples of cloning vehicles include plasmids or phage genomes. Plasmids that can replicate autonomously in a host cell are particularly desirable. Alternatively, nucleic acid molecules that can be inserted (integrated) into the host cell chromosomal DNA are useful, especially molecules that are inserted into the host cell chromosomal DNA in a stable manner (i.e., in a manner that allows such molecules to be inherited by daughter cells).
[0086] Cloning vehicles typically feature one or a small number of endonuclease recognition sites at which such DNA sequences can be cleaved in a deterministic manner without losing the essential biological function of the vehicle, and DNA can be spliced into it to result in its replication and cloning.
[0087] The cloning vector may further comprise a marker suitable for identifying cells transformed with the cloning vector. For example, the marker gene may be a gene that confers resistance to a particular antibiotic on the host cell.
[0088] The word "vector" is used interchangeably with "cloning vehicle."
[0089] As used herein, an "expression vehicle" is a vehicle or vector similar to a cloning vehicle that is specifically designed to provide an environment that allows expression of a cloned gene after transformation into a host. One way to provide such an environment is to include transcriptional and translational regulatory sequences on such an expression vehicle that can be operably linked to the cloned gene. Another way to provide such an environment is to provide one or more cloning sites on such a vehicle into which the desired cloned gene and the desired expression regulatory elements can be cloned.
[0090] In the expression vector, the gene to be cloned is usually operably linked to certain control sequences, such as a promoter sequence. Expression control sequences will vary depending on whether the vector is designed to express the operably linked gene in a prokaryotic or eukaryotic host and may additionally include transcriptional elements, such as enhancer elements, termination sequences, tissue-specific elements, or translation initiation and termination sites.
[0091] "Host" refers to any organism that serves as a recipient for cloning or expression vehicles. A host can be a bacterial cell, a yeast cell, or a cultured animal cell, such as a mammalian or insect cell. A yeast host can be Saccharomyces cerevisiae.
[0092] As described herein, "host cell" can be a bacterial, fungal or mammalian cell or from an insect or plant. Examples of bacterial host cells are Escherichia coli (E. coli) and Bacillus subtilis (B. subtilis). Examples of fungal cells are Saccharomyces cerevisiae and Schizosaccharomyces pombe (S. pombe). Non-limiting examples of mammalian cells are immortalized mammalian cell lines, such as HEK293, A549, HeLa or CHO cells, or isolated primary tissue cells of patients that have been genetically immortalized (e.g., by transfection with hTERT). Non-limiting examples of plants are Nicotiana tabacum or Physcomitrella patens. Non-limiting examples of insect cells are sf9 (Spodopterafrugiperda) cells.
[0093] A "DNA binding domain (DBD)" or "DNA binding portion" is a portion that directs a particular polypeptide to bind to a specific DNA sequence (i.e., a "protein binding site"). These proteins can be homodimers or monomers that bind DNA in a sequence-specific manner. Exemplary DNA binding domains disclosed herein include LexA, cI, the glucocorticoid receptor binding domain, and the Ume6 domain.
[0094] A "gene activation moiety" or "activation domain" ("AD") is a moiety that is capable of inducing (albeit weakly in many cases) the expression of a gene to which it is bound (an example is an activation domain from a transcription factor). As used herein, "weak" means below the level of activation achieved by the GAL4 activation region II, preferably at or below the level of activation achieved by the B42 activation domain. The level of activation can be measured using any downstream reporter gene system, and the expression level stimulated by the GAL4 region II-polypeptide can be compared in parallel assays to the expression level stimulated by the polypeptide to be tested.
[0095] The term "sequence identity" as used herein in the context of amino acid sequences is defined as the percentage of the amino acid residue in the candidate sequence that is identical with the amino acid residue in the selected sequence after the alignment sequences are aligned and a gap is introduced, if necessary, to achieve maximum percentage sequence identity, and without considering any conservative substitutions as a part for sequence identity. The comparison for determining amino acid sequence identity percentage can be achieved in a variety of ways within the scope of the art, for example, using publicly available computer software, such as BLAST, BLAST-2, ALIGN, ALIGN-2 or Megalign (DNASTAR) software. Those skilled in the art can determine the appropriate parameters for measuring the comparison, including any algorithm required for maximum comparison on the full length of the compared sequence.
[0096] Screening for functional degraders of target proteins
[0097] Selective protein degradation is a unique approach to drug discovery. The ability to selectively degrade aberrant proteins or isoforms thereof provides a controlled approach to selectively targeting certain pathologies, such as cancer. Compounds that achieve selective degradation by bridging to E3 ubiquitin ligases are catalytic in nature and do not require stoichiometric levels, making them highly promising drug compounds. Screening for compounds that selectively link a protein of interest to an E3 ubiquitin ligase does not always guarantee functional degradation of the target protein in question. Screening for compounds that can functionally degrade a target protein by forming a transient ternary complex between itself, the target, and the E3 ligase is difficult to perform. Current screens rely on identifying compounds specific for the target and chemically linking them to another moiety that binds to the specific E3 ligase, generating macromolecules that are limited to targets with known small molecule binders.
[0098] The methods and systems of the present disclosure may involve intracellular selection of peptide-based selective degraders. In other words, the various systems described herein can be used to screen for molecules that selectively cause degradation of a target protein by generating a functional interaction between the target protein and an E3 ubiquitin ligase, or directly cause degradation by the proteasome. Model organisms, such as Saccharomyces cerevisiae, can be used for co-expression of a target of interest with a specific E3 ubiquitin ligase and a test DNA molecule comprising a DNA sequence encoding a library of random peptides. This can allow the use of a selection mechanism (e.g., a stringent viability readout selection mechanism) to select unbiased peptides that cause functional degradation of the target of interest. The method may involve a permutation of the yeast one-hybrid system, which may rely on the degradation of a transcription factor that requires interaction between a test protein fused to a DNA binding domain (DBD) and a transcriptional activation domain (AD) by the proteasome or by a specific E3 ubiquitin ligase via a peptide-mediated interaction (see Figure 1 and 2 ).
[0099] The methods and systems of the present disclosure can use reconstitution of transcription factors mediated by a test protein fused to an AD (e.g., VP16, NF-κB AD, VP64 AD, BP64 AD, B42 acidic activation domain (B42AD) or p65 transactivation domain (p65AD) and a DBD, such as LexA, cI, Gli-1, YY1, glucocorticoid receptor binding domain or Ume6 domain). Similarly, the test protein can comprise an AD and bind to DNA through another binding partner.
[0100] The methods and systems of the present disclosure can also be used with two different proteins or two variants of a protein fused to different DBDs and ADs. The system can identify compounds that bridge one of the proteins to the E3 ligase, leading to its degradation, while retaining the active version of the other test protein. For example, one component of a complex can be degraded without affecting the integrity of the rest of the complex (see Figure 3 This system can also be used to identify selective inhibitors that degrade a specific isoform without affecting the other variant (see Figure 4 ).
[0101] The expression of protein of interest can guide RNA polymerase to specific genomic loci, and allows the expression of genetic elements.Genetic elements can be, for example, genes encoding the protein that enables organisms to grow on selection culture medium.Selection culture medium can be specific to, for example, ADE2, URA3, TRP1, KANR or NATR, and will lack essential components (Ade, Ura, Trp) or include drugs (G418, NAT).Can detect when protein no longer exists mark (for example, when protein is degraded by external composition) can be called counter selection marker, for example URA3 gene, and may be poor or leakage (being easy to be covered by the selection of the mutant that escapes selection).This leakage of selection marker may cause high false positive rate.
[0102] The methods and systems of the present disclosure can combine strong negative selection markers with the intracellular stability of the production of short peptides or macrocycles to screen for mediators of bridging interactions between target proteins and E3 ubiquitin ligases. An inducible one-hybrid approach can be employed that can drive the expression of any one or combination of several cytotoxicity reporter genes (death agents) and positive selection markers. The methods of the present disclosure involving the inducible expression of combinations of cytotoxicity reporter genes in a one-hybrid system can allow for a multiplicative effect in reducing the false positive rate of a one-hybrid assay, since all cytotoxicity reporter genes must be "leaky" at the same time to allow induced cell survival.
[0103] In certain embodiments, disclosed herein is a method for identifying a molecule that can selectively bridge the interaction between a first test protein and an E3 ubiquitin ligase to mediate functional degradation of the test protein in a host cell. A second test protein can be used as a positive control, for example, while the molecule mediates degradation of the first test protein, it may not affect expression of the second test protein. The method can include expressing in a host cell a first fusion protein comprising the first test protein and a DNA binding portion and a gene activation portion; an E3 ubiquitin ligase or fragment thereof, or in some cases, an E3 ubiquitin ligase and related structures thereof; and delivering the molecule from a library to the host cell. The host cell can contain a promoter sequence for controlling expression of a death agent. The promoter can be specific for the DNA binding portion of the first fusion protein, such that in the absence of the molecule, expression of the death agent is activated. When the molecule is present, the first test protein may be degraded by the E3 ubiquitin ligase.
[0104] Figure 1 An example of a method for identifying compounds that bridge protein-protein interactions between a target bait protein and an E3 ubiquitin ligase is shown, wherein the bridging between the two proteins results in functional ubiquitination of the target bait protein and its subsequent degradation. DBD refers to a specific DNA binding domain. AD refers to an activation domain. E3 refers to the E3 ubiquitin ligase of interest. The dotted arrows indicate functional ubiquitination of the bait protein, leading to its degradation and preventing it from activating a death agent. Circles refer to peptides, macrocycles, or small molecules. These may be from a library. In this example, cell viability can be measured with or without a bridging agent. The upper panel illustrates a basic scenario in which the bait is driving expression of a positive marker required for cell growth. The lower panel shows a scenario in which the bridging agent is able to functionally bridge the bait protein to the specific E3 ligase of interest. In this scenario, the bait degrades and the cells are unable to grow because they cannot express the positive marker required for growth.
[0105] Figure 2 An example is shown in which a target bait is operatively linked to a negative selection marker that prevents growth in the presence of the target. Three scenarios are shown; the top panel illustrates a basic scenario in which the bait is driving the expression of a death agent, leading to cell death. The bottom panel illustrates a scenario in which a compound is able to bridge the interaction between the bait protein and the E3 ligase, but does not result in functional ubiquitination and subsequent degradation, leading to expression of the death agent and cell death. The bottom panel shows a scenario in which a compound is able to bridge between the bait protein and the E3 ligase, leading to its degradation and loss of transcription of the death agent, thereby leading to cell survival. In some embodiments, the peptide library can be replaced with an exogenous library comprising compounds other than peptides, such as small molecules. In some embodiments, the small molecule is a peptide mimetic.
[0106] Figure 3 and 4 A platform for identifying compounds that degrade bait target proteins in a variant-specific manner is shown. Figure 3 A similar assay is described in which bait 1 and bait 2 are related (but different) proteins (eg, protein variants). Figure 4 A similar assay is described in which the bait WT refers to the WT allele of the protein, while the bait MUT Refers to the pathological allele targeted for degradation. Figure 2 In Figure 2, DBD refers to the promoter-specific DNA binding domain. AD refers to the activation domain. E3 refers to the E3 ubiquitin ligase of interest. Dashed arrows indicate functional ubiquitination of the bait protein, leading to its degradation and preventing it from activating the death agent. Circles refer to peptides, macrocycles, or small molecules. Three scenarios are shown; the top panel illustrates the basic scenario in which each bait is driving the expression of a positive marker or death agent, leading to cell death. The bottom panel illustrates the scenario in which the compound is able to bridge the interaction between the bait protein and the E3 ligase, but does not result in functional ubiquitination and subsequent degradation, leading to cell death. The bottom panel shows the scenario in which the compound is able to bridge between the bait protein and the E3 ligase, leading to its degradation and loss of death agent transcription, leading to cell survival. In all cases, selection against nonspecific protein degradation is avoided by requiring the functional presence of a nonspecific protein to drive the positive selection marker. Selective peptide-mediated degradation is measured by survival due to (1) the absence of expression of the death agent and (2) expression of the positive selection reporter gene (which provides evidence of selectivity).
[0107] In some embodiments, screening for identifying peptides or small molecules that can mediate target protein degradation can involve testing peptides or small molecules against a host cell population, wherein different cells in the population express different E3 ligases. The host cells can then be transformed with candidate peptides / small molecules from a library or otherwise subjected to candidate peptides / small molecules from a library. In this case, each host cell can contain the same target protein and / or death agent. Surviving cells can be sequenced to identify E3 ligases that successfully interact with the peptide / small molecule. In another example, each well of the assay can contain a plurality of different host cells, wherein different host cells express different E3 ligases. Peptides / small molecules from the library can then be transformed or otherwise introduced into each well to identify peptides / small molecules that successfully interact with the target protein and result in cell survival.
[0108] Examples of targets for degradation are oncogenic proteins, such as K-Ras oncoalleles, cyclin D family, cyclin E family, c-MYC, EGFR, HER2, PDGFR, VEGF, and β-catenin, or oncogenic variants, such as IDH1 (R132H, R132S, R132C, R132G, and R132L) or IDH2 (R140Q, R172K).
[0109] Examples of E3 ubiquitin ligases that can be used on the system can be selected from multi-subunit E3 ligases including, but not limited to, the Culin family (CRL1, CRL2, CRL3, CRL4, CRL5, and CRL7) and single-subunit E3 ligases of the RING, RING-Between-RING (RBR), and HECT families, the HECT family consisting of, but not limited to, Cereblon, Skp2, MDM2, FBXW7, DCAF1, DCAF15, VHL, AFF4, AMFR, ANAPC11, ANKIB1, AREL1, ARIH1, ARIH2, BARD1, BFAR, BIRC2, B IRC3, BIRC7, BIRC8, BMI1, BRAP, BRCA1, CBL, CBLB, CBLC, CBLL1, CCDC36, CCNB1IP1, CGRRF1, CHFR, CNOT4, CUL9, CYHR1, DCST1, DTX1, DTX2, DTX3, DTX3L ,DTX4,DZIP3,E4F1,FANCL,G2E3,HACE1,HECTD1,HECTD2,HECTD3,HECTD4,HECW1,HECW2,HERC1,HERC2,HERC3,HERC4,HERC5,HERC6,HLTF,HUWE1,IRF2 BP1, IRF2BP2, IRF2BPL, Itch, KCMF1, KMT2C, KMT2D, LNX1, LNX2, LONRF1, LONRF2, LONRF3, LRSAM1, LTN1, MAEA, MAP3K1, MARCH1, MARCH10, MARCH11, MAR CH2, MARCH3, MARCH4, MARCH5, MARCH6, MARCH7, MARCH8, MARCH9, Mdm2, MDM4, MECOM, MEX3A, MEX3B, MEX3C, MEX3D, MGRN1, MIB1, MIB2, MID1, MID2, MKRN1, MKRN2, MKRN3, MKRN4P, MNAT1, MSL2, MUL1, MYCBP2, MYLIP, NEDD4, NEDD4L, NEURL1, NEURL1B, NEURL3, NFX1, NFXL1, NHLRC1, NOSIP, NSMCE1, PARK2, PCGF1 , PCGF2, PCGF3, PCGF5, PCGF6, PDZRN3, PDZRN4, PELI1, PELI2, PELI3, PEX10, PEX12, PEX2, PHF7, PHRF1, PJA1, PJA2, PLAG1, PLAGL1, PML, PPIL2, PRPF19,RAD18、RAG1、RAPSN、RBBP6、RBCK1、RBX1、RC3H1、RC3H2、RCHY1、RFFL、RFPL1、RFPL2、RFPL3、RFPL4A、RFPL4AL1、RFPL4B、RFWD2、RFWD3、RING1、RLF、RLIM、 RMND5A、RMND5B、RNF10、RNF103、RNF11、RNF111、RNF112、RNF113A、RNF113B、RNF114、RNF115、RNF121、RNF122、RNF123、RNF125、RNF126、RNF128、RNF13 RNF130、RNF133、RNF135、RNF138、RNF139、RNF14、RNF141、RNF144A、RNF144B、RNF145、RNF146、RNF148、RNF149、RNF150、RNF151、RNF152、RNF157、RNF16 5、RNF166、RNF167、RNF168、RNF169、RNF17、RNF170、RNF175、RNF180、RNF181、RNF182、RNF183、RNF185、RNF186、RNF187、RNF19A、RNF19B、RNF20、RNF20 NF207、RNF208、RNF212、RNF212B、RNF213、RNF214、RNF215、RNF216、RNF217、RNF219、RNF220、RNF222、RNF223、RNF224、RNF225、RNF24、RNF25、RNF26 F31、RNF32、RNF34、RNF38、RNF39、RNF4、RNF40、RNF41、RNF43、RNF44、RNF5、RNF6、RNF7、RNF8、RNFT1、RNFT2、RSPRY1、SCAF11、SH3RF1、SH3RF2、SH3RF3 HPRH、SIAH1、SIAH2、SIAH3、SMURF1、SMURF2、STUB1、SYVN1、TMEM129、Topors、TRAF2、TRAF3、TRAF4、TRAF5、TRAF6、TRAF7、TRAIP、TRIM10、TRIM11、TRIM1 3、TRIM15、TRIM17、TRIM2、TRIM21、TRIM22、TRIM23、TRIM24、TRIM25、TRIM26、TRIM27、TRIM28、TRIM3、TRIM31、TRIM32、TRIM33、TRIM34、TRIM35、TRIM36、TRIM37, TRIM38, TRIM39, TRIM4, TRIM40, TRIM41, TRIM42, TRIM43, TRIM43B, TRIM45, TRIM46, TRIM47, TRIM48, TRIM49, TRIM49B, TRIM49C, TRIM49D1, TRIM5, TRIM50, TRIM51, TRIM52 , TRIM54, TRIM55, TRIM56, TRIM58, TRIM59, TRIM6, TRIM60, TRIM61, TRIM62, TRIM63, TRIM64, TRIM64B, TRIM64C, TRIM65, TRIM67, TRIM68, TRIM69, TRIM7, TRIM71, TRIM72, TRIM73, TR IM74, TRIM75P, TRIM77, TRIM8, TRIM9, TRIML1, TRIML2, TRIP12, TTC3, UBE3A, UBE3B, UBE3C, UBE3D, UBE4A, UBE4B, UBOX5, UBR1, UBR2, UBR3, UBR4, UBR5, UBR7, UHRF1, UHRF2, UNK, UNKL , VPS11, VPS18, VPS41, VPS8, WDR59, WDSUB1, WWP1, WWP2, XIAP, ZBTB12, ZFP91, ZFPL1, ZNF2 80A, ZNF341, ZNF511, ZNF521, ZNF598, ZNF645, ZNRF1, ZNRF2, ZNRF3, ZNRF4, Zswim2 and ZXDC. ,
[0110] Protein expression in host cells
[0111] One or more plasmid constructs can be used to express different proteins in the host cell. The number of plasmids used may depend on the host cell, the presence of the integrated construct in the host cell, and other conditions.
[0112] In some cases, the method for identifying a molecule that triggers degradation of a first test protein can use a protein, such as an E3 ubiquitin ligase, a molecule from a molecular library, and a first fusion protein comprising a first test protein, a first DNA binding portion, and a gene activation domain. The method can also use a promoter that drives expression of a death agent, such as a promoter sequence that is specific for the first DNA binding portion. In addition to this approach, in some cases, the method can also utilize a second fusion protein comprising a second DNA binding domain and a gene activation portion, and a promoter that drives expression of a positive or negative marker (e.g., a promoter sequence) that is specific for the second DNA binding portion.
[0113] The above protein and nucleic acid sequences can be provided to the host cell in the form of a plasmid. In some cases, nucleic acid sequences of proteins and nucleic acids including promoters and death agents / positive and negative markers can be integrated into the host cell. In some cases, the molecules from the molecular library are small molecules / compounds and do not need to be encoded on a plasmid.
[0114] For example, in one example, a first fusion protein can be provided in a plasmid (plasmid 1), an E3 ubiquitin ligase can be provided in a separate plasmid (plasmid 2), and a DNA encoding molecule from a library can be provided in a separate plasmid (plasmid 3). All three plasmids can be transfected into a plurality of host cells. Where a second fusion protein is also used, the second fusion protein can be provided in plasmid 1 or in a separate plasmid (plasmid 4). Expression constructs for plasmids can also be combined in one or two plasmids to reduce the number of plasmids to be transfected. In addition, constructs containing promoters driving death agents or positive / negative selection markers can also be provided in the plasmids, which can otherwise be integrated into the host cells.
[0115] In another embodiment, the first fusion protein can be genetically integrated into the host cell, and plasmid 2 and plasmid 3 containing the E3 ubiquitin ligase and the molecule from the molecular library are transfected into the host cell. In this example, the second fusion protein can also be integrated into the host cell, or in some cases, provided in plasmid form.
[0116] In another embodiment, both the first fusion protein and the E3 ubiquitin ligase are integrated into the host cell, and the molecules from the molecular library are transfected into the host cell in the form of a plasmid. As described above, the second fusion protein can be integrated into the host cell or can be provided in the form of a plasmid. Constructs containing promoters driving death agents or positive / negative selection markers can also be provided in plasmids or they can be integrated into the host cell.
[0117] In another instance, the first fusion protein can be transfected into a plasmid for use / integration into the host cell, but in this case the use of an endogenous E3 ubiquitin ligase may obviate the need for integration or transfection of a plasmid containing the E3 ligase.
[0118] In some cases, the nucleic acid sequence for one or more fusion proteins, the E3 ubiquitin ligase, the promoter driving the death agent (and the promoter driving the positive / negative selection marker, if used) can all be integrated into the host cell. In this case, only a single plasmid containing a molecule from the molecular library can be transfected into the host cell.
[0119] In some embodiments, one or more host cells disclosed herein comprise a plasmid vector. For example, the plasmid may comprise two restriction sites that are capable of integrating two proteins constituting the bait and E3 ligase of interest. The bait protein of interest may be related to oncogenes (e.g., cyclin E family, cyclin D family, c-MYC, EGFR, HER2, K-Ras, PDGFR, Raf kinases, and VEGF). The bait protein of interest may be related to effectors of inflammatory responses (e.g., IL-17RA, IL-17RB, IL-17RC, IL17-RD, IL17-RE, Act1 (CIKS), and IL-23R).
[0120] The plasmid can be configured to express the two proteins that constitute the bait and E3 ligase of interest, as well as an additional factor, such as a variant of one of the bait proteins. The variant used for targeting can be KRAS (G12D, G12V, G12C, G12S, G13D, Q61K or Q61L, etc.), and the control variant is WT KRAS. The additional factor can also be another protein that binds to the bait protein, or another target of the E3 ligase.
[0121] In some embodiments, a host cell disclosed herein comprises a plasmid in which a DNA sequence encoding a first polypeptide is inserted in frame with Gal4-DBD and in frame with VP64-AD, and a DNA sequence encoding a second polypeptide comprising an E3 ubiquitin ligase.
[0122] In some embodiments, the first test protein is a variant of KRAS and the E3 ubiquitin ligase is VHL.
[0123] The plasmid can encode an activation domain or another gene activation portion and a fusion of a DBD with each protein driven by a strong promoter and terminator (e.g., ADH1) or by an inducible promoter (e.g., GAL1). Other exemplary activation domains include those of VP16 and B42AD. In some embodiments, the DNA binding portion is derived from LexA, TetR, LacI, Gli-1, YY1, glucocorticoid receptor, or Ume6 domain, and the gene activation portion is derived from Gal4, B42, or VP64, Gal4, NF-κB AD, Dof1, BP64, B42, or p65. Each protein fusion can be labeled for subsequent biochemical experiments using, for example, FLAG, HA, MYC, or His tags. The plasmid can also include bacterial selection and reproduction markers (i.e., ori and AmpR), and yeast replication and selection markers (i.e., TRP1 and CEN or 2um). The plasmid may contain a variety of bait proteins fused to different DBDs and ADs. Plasmids can also integrate into the genome at a specific locus.
[0124] In certain embodiments, disclosed herein is a library of plasmid vectors, each plasmid vector comprising: a DNA sequence encoding a different peptide sequence operably linked to a first switchable promoter; a DNA sequence encoding a killing agent under the control of a second switchable promoter; and a DNA sequence encoding a positive selection reporter gene under the control of a third switchable promoter.
[0125] Expression of selection markers
[0126] Positive selection markers
[0127] Efficient expression of the test protein can direct RNA polymerase to a specific genomic site and allow expression of a protein that enables the organism to grow on the selective medium. The selective medium can be selective for, for example, ADE2, URA3, TRP1, KAN R or NAT R Is specific and may lack essential components (Ade, Ura, Trp) or may include drugs (G418, NAT). The plasmid may encode one or more positive selection markers that enable the organism to grow on a selective medium.
[0128] Negative selection markers
[0129] An inducible single hybrid approach can be employed that can drive the expression of any one or combination of several cytotoxicity reporter genes (death agents) and positive selection markers. The disclosed method of inducible expression of a combination of cytotoxicity reporter genes in a single hybrid system can allow for a multiplier effect in reducing the false positive rate of a single hybrid assay because all cytotoxicity reporter genes must be "leaky" at the same time to allow induced cell survival. The cytotoxicity reporter gene can comprise or contain a domain of a variety of polypeptides, such as those shown in Table 1.
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151] In some embodiments, the killing agent is an overexpression product of a genetic element selected from DNA or RNA. In some embodiments, the genetic element is a growth inhibitory (GIN) sequence, such as GIN11.
[0152] In some embodiments, the killing agent is a ribosomally encoded xenobiotic agent, a ribosomally encoded poison, a ribosomally encoded endogenous or exogenous gene that causes severe growth defects when mildly overexpressed, a ribosomally encoded recombinase that excises a gene essential for viability, a restriction factor involved in the synthesis of a toxic secondary metabolite, or any combination thereof. In some embodiments, the ribosomally encoded killing agent is cholera toxin, SpvB toxin, CARDS toxin, SpyA toxin, HopUl, Chelt toxin, Certhrax toxin, EFV toxin, ExoT, CdtB, diphtheria toxin, ExoU / VipB, HopPtoE, HopPtoF, HopPtoG, VopF, YopJ, AvrPtoB, SdbA, SidG, VpdA, Lpg0969, Lpgl978, YopE, SptP, SopE2, SopB / SigD, SipA, YpkA, YopM, amatoxin, phalloidin, killer toxin KP1, killer toxin KP6, killer toxin Kl, killer toxin K28 (KHR), killer toxin K28 (KHS), anthrax lethal factor endopeptidase, Shiga toxin, saporin toxin, ricin, or any combination thereof. The cytotoxicity reporter gene or killing agent can be a protein having a sequence selected from SEQ ID NOs: 1-41. The cytotoxicity reporter gene can be a variant of a naturally found cytotoxicity reporter gene. Such variants can have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with SEQ ID NOs: 1-41.
[0153] In addition to one or more positive selection markers, the plasmid can also include one or more negative selection markers that are controlled by different DNA binding sequences to achieve binary selection. The plasmid can encode one or more negative selection markers in Table 1 driven by a promoter that depends on the DBD-DNA binding sequence (DBS) present in the bait protein integration plasmid, for example, a LexAop sequence (DBS), which can be constrained by LexA (DBD). In some embodiments, to ensure the inhibition of the "lethal agent", the plasmid can include a silencing construct, such as a TetR'-Tup11 fusion driven by a strong promoter (e.g., ADH1) to bind to the DBD and silence transcription in the presence of doxycycline. The plasmid can contain bacterial selection and reproduction markers (i.e., ori and AmpR), as well as yeast replication and selection markers (i.e., LEU2 and CEN or 2-micron).
[0154] In certain embodiments, disclosed herein are libraries of plasmid vectors, each comprising a DNA sequence encoding a different peptide sequence operably linked to a first switchable promoter; a DNA sequence encoding a killing agent under the control of a second switchable promoter; and a DNA sequence encoding a positive selection reporter gene under the control of a third switchable promoter. Plasmids comprising promoters driving different E3 ubiquitin ligases may also be included in the vector library.
[0155] Addition or expression of regulators
[0156] Molecules from the library can be screened by using positive and / or negative selection markers in host cells that can selectively bridge the bait protein of interest and a specific E3 ubiquitin ligase, leading to degradation of the bait protein.
[0157] In some embodiments, the drug is a small molecule. In some embodiments, the small molecule is a peptide mimetic. For example, by deleting genes encoding drug efflux pumps, such as PDR5, host cells can be made permeable to small molecules. Genes encoding transcription factors (such as PDR1 and PDR3) induce expression of efflux pumps, including but not limited to the 12 genes described in the 12-gene ΔOHSR (Chinen, 2011). By interfering with ergosterol synthesis and deposition in the plasma membrane, for example, by deleting ERG2, ERG3, and / or ERG6 or driving their expression under a regulatable promoter, host cells can be further permeable to small molecules.
[0158] In other embodiments, the molecule is a peptide, macrocycle, or protein. In some embodiments, the peptide or protein is derived from a naturally occurring protein product. In another embodiment, the peptide or protein is a synthetic protein product. In other embodiments, the peptide or protein is a product of a recombinant gene.
[0159] In some embodiments, the molecule is introduced exogenously into a host cell. In other embodiments, the molecule is the expression product of a test DNA inserted into the host cell, wherein the test DNA comprises a DNA sequence encoding a polypeptide. A DNA encoding library can be formed by delivering multiple test DNA molecules into the host cell. In some embodiments, the peptide sequences of the polypeptides in the library are random. In some embodiments, different peptide sequences are pre-enriched to bind to the target.
[0160] In order to screen for peptides that selectively promote the degradation of proteins of interest, peptides from random peptide libraries can be applied to host cells or expressed from within host cells. Plasmids can further be used to express random peptide libraries (such as random NNK 60-mer sequences). Plasmids can include restriction sites for integrating random peptide libraries driven by strong promoters (e.g., ADH1 promoter) or inducible promoters (e.g., GAL1 promoter).
[0161] In some embodiments, the randomized peptide library is about 60-mer. In some embodiments, the randomized peptide library is about 5-mer to 20-mer. In some embodiments, the randomized peptide library is less than 15-mer.
[0162] The library can also be started with a fixed sequence, for example, methionine-valine-asparagine (MVN) for N-terminal stabilization and / or another combination of high half-life N-terminal residues (see, e.g., Varshavsky. Proc. Natl. Acad. Sci. USA. 93: 12142-12149 (1996)) to maximize the half-life of the peptide and terminated with a 3'UTR of a short protein (e.g., sORF1). The peptide can also be tagged with a protein tag (e.g., Myc). In some embodiments, the N-terminal residue of the peptide comprises Met, Gly, Ala, Ser, Thr, Val, or Pro, or any combination thereof, to minimize proteolysis.
[0163] A plurality of different short peptide sequences can be randomly generated by any method (e.g., NNK or NNN nucleotide randomization). A plurality of different short peptide sequences can also be pre-selected by previous experimental selection for binding to the target, or pre-selected from existing data sets in the scientific literature where rationally designed peptide libraries have been reported.
[0164] In some embodiments, the library comprises polypeptides having a length of about 60 amino acids or less. In another embodiment, the library comprises polypeptides having a length of about 30 or less amino acids. In another embodiment, the library comprises polypeptides having a length of about 20 or less amino acids.
[0165] Promote peptide modification
[0166] The peptide that causes the selective degradation of target can also be the product of post-translational modification.Post-translational modification can comprise any one or combination in cutting, cyclization, dicyclization, methylation, halogenation, glycosylation, acylation, phosphorylation and acetylation.In some embodiments, methylation comprises and N-methyltransferase reaction.In some embodiments, post-translational modification is completed by naturally occurring enzyme.In some embodiments, post-translational modification is completed by synthetase.In some embodiments, synthetase is chimeric.
[0167] The peptides may be ribosomally synthesized and post-translationally modified peptides (RiPPs) whereby the core peptide is flanked by a propeptide sequence including a leader peptide and recognition sequences that signal for the recruitment of maturation, cleavage and / or modification enzymes such as excision enzymes or cyclases, including, for example, lanthipeptide maturases (LanB, LanC, LanM, LanP) from Lactococcus lactis, patellamide biosynthesis factors (PatD, PatG) from cyanobacteria, butelase 1 from Clitoria ternatea, and POPB from Galerina marginata, Lentinula edodes, Omphalotacae olearis, Dendrothele bispora, or Amanita bisporigera, or other species. In some embodiments, the cyclizing or bicyclizing enzyme is a synthetic chimera.
[0168] In one example, the variable peptide library region is embedded in the primary sequence of a modifying enzyme (e.g., a homolog of omphalotin N-methyltransferase from Dendrothele bispora, Marasmius fiardii, Lentinula edula, Fomitiporia mediterranea, Omphalotus olearius, etc.) and contains random residues, some of which may be post-translationally modified by additional modifications such as hydroxylation, halogenation, glycosylation, acylation, phosphorylation, methylation, acetylation. Under the action of a prolyl endopeptidase belonging to the PopB family and an N-methyltransferase belonging to the omphalotin methyltransferase family, this diverse variable region is excised and modified to form an N- to C-cyclized, optionally N-methylated macrocycle. An exemplary list of prolyl endopeptidases is shown in Table 2. The prolyl endopeptidase can be a protein having a sequence selected from SEQ ID NOs: 42-58. The prolyl endopeptidase may be encoded by a nucleotide sequence selected from SEQ ID NO: 59 or 60. The prolyl endopeptidase may be a variant of a naturally occurring prolyl endopeptidase. Such variants may have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 42-58. An exemplary list of N-methyltransferases is shown in Table 3. The methyltransferase may be a protein having a sequence selected from SEQ ID NO: 61-116. The methyltransferase may be encoded by a nucleotide sequence selected from SEQ ID NO: 117 or 118. The prolyl endopeptidase may be a variant of a naturally occurring methyltransferase. Such variants may have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 61-116.
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214]
[0215]
[0216]
[0217]
[0218]
[0219]
[0220]
[0221]
[0222]
[0223]
[0224]
[0225] Gymnopeptide A (GymA) and Gymnopeptide B (GymB) are two related, poly-N-methylated, cyclic octadecapeptides isolated from the spindle-stalked mushroom G. fusipes (also known as Collybia fusipes). GymA and GymB differ at one position (serine in GymA versus threonine in GymB). Several aggressive adherent cancer cell lines (e.g., HeLa, A431, T47D, MCF7, MDA-MB-231) exhibit hypersensitivity to both GymA and GymB, with IC50 values in the low nanomolar range.
[0226] Surprisingly, it was found that NRPS was not utilized to synthesize these peptide macrocyclic compounds, but a gene was encoded by the genome of the fungus Fusobacterium fusobacteria, which contained 18 amino acid nucleic acid sequences encoding GymB. 18 amino acid sequences were located at the C-terminal end of an open reading frame encoding a hypothetical S-adenosylmethionine (SAM) dependent methyltransferase. Hereinafter, the gene encoding the GymB peptide sequence box is referred to as naked peptide precursor gene GymMAB.
[0227] The GymMAB gene is present in a cluster that also includes another open reading frame encoding a prolyl oligopeptidase (GymP), which cleaves and cyclizes the methylated naked peptide cassette. These enzymes share weak similarities with the prolyl oligopeptidase PopB protein from Capsules and Amanita species and the cyclododecapeptide-generating enzyme from O. olearis, and form a unique family of RiPPs / RiPP-processing enzymes with distinct structural and functional features that allow them to accommodate the relatively large size of the 18-mer macrocycle.
[0228] In addition, careful examination of several closely related Gymnopus species such as Gymnopus earle, Gymnopus dryophilus, Gymnopus ocior, Gymnopus acervatus, Gymnopus luxurians, Gymnopus androsaceus (also known as Marasmius androsaceus or Setulipes androsaceus), Micromphale foetidum, Micromphale perforans, and Marasmius fiardii failed to detect any genes encoding orthologs or other genes related to the above enzymes identified in Rhodocollybia maculata and Rhodocollybia butyracea. On the other hand, the biosynthetic gene clusters for the enzymes involved in the production of cyclododecapeptides are present in many closely related species, such as Omphalotous olivascens and species of Lentinula, including Lentinula edura, Lentinula aciculospora, Lentinula raphanica, Lentinula edura, Lentinula boryana, and Lentinula edura. Therefore, the identified genetic clusters appear to have been horizontally transferred.
[0229] Enzymes such as methyltransferases and prolyl oligopeptidases isolated from species such as Fusobacterium fusobacteria can be used to produce methylated macrocycles. The methylated macrocycle can be screened using the methods described herein. The enzyme can be integrated into a host cell and used to generate a DNA-encoded RiPP library. The enzyme can also be used to mass-produce a specific macrocycle of interest in a heterologous prokaryotic or eukaryotic expression system. The use of the enzyme in a heterologous expression system may include, but is not limited to, the reverse Y2H system described in PCT / US2018 / 061292 (published as WO2019 / 099678) and U.S. application No. 15 / 683,586 (published as US20170368132A1), the entire contents of which are incorporated herein by reference (particularly with respect to reverse hybridization and the related yeast systems disclosed therein).
[0230] Macrocycles produced using the methods described herein can be used as pharmaceuticals. Such pharmaceuticals can be used to treat various diseases or conditions. Macrocycles produced using the methods described herein can be used to modulate a protein-protein interaction between a first protein and a second protein. Macrocycles produced using the methods described herein can be used to disrupt a protein-protein interaction between a first protein and a second protein.
[0231] In certain embodiments, disclosed herein is a method for detecting or degrading a target protein mediated by a molecule that links a first target or test protein to a second target protein in a host cell, the method comprising: expressing a first fusion protein comprising a first test protein and a second protein in a host cell; delivering the first molecule to the host cell; modifying the first molecule in the host cell by a modifying enzyme such as a prolyl oligopeptidase and / or a methyltransferase; and allowing the first molecule to bridge the interaction between the first test protein and the second protein, wherein the first molecule is a product of an encoded DNA sequence, wherein the first molecule comprises a random polypeptide library and one or more modifying enzymes, wherein the one or more modifying enzymes modify the random polypeptide library.
[0232] Prolyl oligopeptidase as described herein can be an enzyme that can macrocyclize relatively larger peptides.Prolyl oligopeptidase as described herein can be a kind of that can macrocyclize and comprise at least 5 amino acids, at least 7 amino acids, at least 10 amino acids, at least 15 amino acids, at least 18 amino acids, at least 20 amino acids or at least 25 amino acid whose peptides.Prolyl oligopeptidase as described herein can be a kind of that can macrocyclize and comprise at most 7 amino acids, at most 10 amino acids, at most 15 amino acids, at most 18 amino acids, at most 20 amino acids or at most 25 amino acid whose peptides.
[0233] The tryptophan at position 603 seems to be highly conserved in the related prolyl oligopeptidases that cannot relatively large-scale cyclization peptides. Similarly, the asparagine at position 563 adjacent to the active site serine at position 562 is also conserved in these identical prolyl oligopeptidases. As used herein, "position 603" and "position 563" refer respectively to the position of the active site tryptophan adjacent to the active site serine in the prolyl oligopeptidases of SEQ ID NO:55 and the position of the asparagine, and the corresponding amino acid in other prolyl oligopeptidases. In other words, different from the position 603 or position 563 of the prolyl oligopeptidase of SEQ ID NO:55, it may not necessarily be the 603rd or 563rd amino acid in the protein, but when the prolyl oligopeptidase is compared with it, it is the position compared with position 603 or 563 of SEQ ID NO:55, regardless of the distance between the amino acid and the N-terminal of the protein. Without being bound by theory, these highly conserved tryptophan and asparagine residues are mutated to other amino acids, such as leucine and serine, respectively, which may be the key to making its structure flexible to adapt to peptides (such as larger 18-polymer naked peptides). In addition, replacing tryptophan with another residue (such as leucine) at position 603 may play an important role in expanding the cleavage site recognition specificity of oligopeptidase, from pointing to small secondary amine residues such as proline or sarcosine (N-methyl-glycine) to being able to cut at secondary amine sites (such as N-methyl-valine, N-methyl-isoleucine or N-methyl-leucine) with larger side chains. Consistent with this premise, although the N-terminal cleavage site of naked peptide A / B precursor protein is at the proline residue, the C-terminal cleavage site is at the methyl valine residue. Prolyl oligopeptidase belongs to the serine protease family. The mechanism of action of serine peptidase relates to acyl enzyme intermediate. The formation and decomposition of acyl enzyme are all carried out by the formation of negatively charged tetrahedral intermediate, and this intermediate is stabilised by the oxygen anion binding site that provides two hydrogen bonds for oxygen anion. In the prolyl oligopeptidase, one of hydrogen bonds is formed between the main chain amide groups of oxyanion and asparagine 563, which is directly adjacent to the catalytic serine, serine 562. The second hydrogen bond belongs to this type of serine peptidase, and is provided by the hydroxyl of tyrosine 481 (position 481 of SEQ ID NO:55). In the chymotrypsin type member of the serine protease family of enzyme, hydrogen bonds are provided by the main chain amide groups of the main chain amide groups of the catalytic serine residue and the glycine residue at-2 position of the catalytic serine. The highly conservative asparagine at position 563 is replaced by serine so that the active site serine and glycine hydrogen bond donor positions of serine 563 residues and glycine 561 residues (position 561 of SEQ ID NO:55) are identical with the chymotrypsin type protease.This substitution likely plays an important role in enabling the enzyme to switch between using two different active site serines in each of the two cleavage events. For example, serine 562 is likely the active site residue involved in N-terminal proline-directed cleavage, with two hydrogen bonds to the oxyanion contributed by the backbone amide of the serine residue at position 563 and the hydroxyl group of the tyrosine at position 481, while serine 563 is the active site residue involved in second N-methyl-valine-directed cleavage, with two hydrogen bonds to the oxyanion contributed by the backbone amide of the serine at position 563 and the glycine at position 561, and vice versa. The combination of this novel, wider catalytic pocket due to the replacement of tryptophan 603 with leucine and the switchable active site serine due to the replacement of asparagine 563 with serine makes this new oligopeptidase particularly well-suited to recognize a variety of secondary amine residues with bulky side chains at the cleavage site and contains a larger macrocycle than any previously characterized family member. Figure 5A An alignment of various related enzyme species around this residue is shown. The sequence read from the Fusostigmine enzyme is also shown. Figure 5B The protein structure of the plant Pop homolog is shown, highlighting the location of W603 within the active site (adapted from www.pnas.org / cgi / doi / 10.1073 / pnas.1620499114). The arrow in the lower panel points to the tryptophan residue. Figure 5C An alignment of various related enzyme species around the serine residue is shown. Sequence reads from the Fusostigmine enzyme are also shown. Figure 5D The protein structure of the plant Pop homolog is shown, highlighting the position of N563 within the active site (adapted from www.pnas.org / cgi / doi / 10.1073 / pnas.1620499114). The arrow in the lower panel points to the asparagine residue.
[0234] In some cases, the tryptophan residues (it corresponds to the conservative tryptophan at 603, position of SEQ ID NO:55) in the avtive site of prolyl oligopeptidase can be replaced with different amino acid residues.For example, in some cases, the tryptophan residues in the avtive site of prolyl oligopeptidase can be replaced by leucine residues.In some cases, the prolyl oligopeptidase used herein does not comprise tryptophan residues at 603, position in the avtive site of enzyme, and wherein position 603 corresponds to the avtive site of SEQ ID NO:55.
[0235] In some cases, the asparagine residues (it corresponds to the conservative asparagine at 563, position of SEQ ID NO:55) in the avtive site of prolyl oligopeptidase can be replaced with different amino acid residues.For example, in some cases, the asparagine residues in the avtive site of prolyl oligopeptidase can be replaced by serine residues.In some cases, the prolyl oligopeptidase used herein does not comprise asparagine residues at 563, position in the avtive site of enzyme, and wherein position 563 corresponds to the avtive site of SEQ ID NO:55.
[0236] In other embodiments, cyclization includes reacting with beta-lactamase. Under the action of N-methyltransferase and beta-lactamase family members, the variable region is excised and end-to-end cyclized. Table 4 shows an exemplary list of beta-lactamase and amino acid sequence of the cyclic peptide processed. Beta-lactamase can be a protein with a sequence selected from SEQ ID NO:119-120. Beta-lactamase can be a variant (for example, non-natural variant) of a naturally occurring beta-lactamase. Such variants can have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with SEQ ID NO:119-120. In some embodiments, some side chains of random residues are subsequently isomerized from L-configuration to D-configuration or modified with other modifications such as hydroxylation, halogenation, glycosylation, acylation, phosphorylation, methylation and acetylation.
[0237]
[0238]
[0239]
[0240]
[0241] In some embodiments, cyclization includes reaction with prolyl endopeptidase, N-methyltransferase and hydroxylase. In some embodiments, dicyclization includes further modification of the anchor residue shown in the cyclized peptide to form an internal aminocarboxyethyl thiotryptophan bridge. The first step may involve hydroxylating the 2-position of the indole ring of the tryptophan residue by a hydroxylase belonging to the cytochrome P450 family of oxygenases. Examples of such hydroxylases are shown in Table 5. Hydroxylases can be proteins with a sequence selected from SEQ ID NO: 123. Hydroxylases can be variants (e.g., non-natural variants) of naturally found hydroxylases. Such variants can have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with SEQ ID NO: 123.
[0242]
[0243]
[0244] Step 2 may involve the formation of an aminocarboxyethylthiotryptophan bridge between the 2'-hydroxy position on the tryptophan and the thiol group of the cysteine residue. This condensation reaction is catalyzed by a novel family of dehydratases. Examples of dehydratases are shown in Table 6. The dehydratases may be proteins having a sequence selected from SEQ ID NOs: 124-127. The dehydratases may be variants (e.g., non-natural variants) of naturally occurring dehydratases. Such variants may have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NOs: 124-127.
[0245]
[0246]
[0247]
[0248]
[0249]
[0250]
[0251]
[0252] Step 3 describes the S-oxygenation of aminocarboxyethylthiotryptophan thiol by a flavin monooxygenase that converts the aminocarboxyethylthiotryptophan thiol to the sulfinyl form. Examples of such monooxygenases are shown in Table 7. Step 4 describes potential future modification steps, such as hydroxylation of peptide side chains, such as hydroxylation of position 6 on the indole ring of the tryptophan residue that forms aminocarboxyethylthiotryptophan by a P450 family monooxygenase. The monooxygenase can be a protein having a sequence selected from SEQ ID NO: 128. The monooxygenase can be a variant (e.g., a non-natural variant) of a naturally occurring monooxygenase. Such variants can have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 128.
[0253]
[0254] The sequences flanking the encoded random peptide library can be modified by using N-terminal and C-terminal flanks from MSDIN family genes (toxin preproprotein sequences) identified in the genomes of Amanita bisporus and Amanita phalloides.
[0255] Furthermore, enzymes can be targeted to specific cellular compartments to increase peptide synthesis efficiency and improve yields for peptide production purposes.
[0256] In certain embodiments, disclosed herein is a method for detecting degradation of a target protein mediated by a molecule that links the target or test protein to an E3 ubiquitin ligase in a host cell, comprising: expressing a first fusion protein comprising a first test protein and an E3 ubiquitin ligase in a host cell; delivering the first molecule to the host cell; modifying the first molecule by a modifying enzyme in the host cell; and allowing the first molecule to bridge the interaction between the first test protein and the E3 ubiquitin ligase, wherein the first molecule is a product of an encoded DNA sequence, wherein the first molecule comprises a random polypeptide library and one or more modifying enzymes, wherein the one or more modifying enzymes modify the random polypeptide library.
[0257] host cells
[0258] In some embodiments, the host cell is a eukaryote or a prokaryote. In some embodiments, the host cell is from an animal, a plant, a fungus or a bacterium. In some embodiments, the fungus is Aspergillus or Pichia pastoris. In some embodiments, the host cell is a haploid yeast cell. In other embodiments, the host cell is a diploid yeast cell. In some embodiments, the diploid yeast cell is produced by mating a first host cell comprising a DNA sequence encoding a first chimeric gene, a second chimeric gene, and a third chimeric gene with a second host cell comprising a DNA sequence encoding a death agent, a positive selection reporter gene, and an mRNA comprising a nucleotide sequence encoding a polypeptide. In some embodiments, the plant is Nicotiana tabacum or Physcomitrella patens. In some embodiments, the host cell is an sf9 (Spodoptera frugiperda) insect cell.
[0259] In certain embodiments, disclosed herein is a host cell configured to express a first fusion protein comprising a first test protein, a first DNA binding portion, and a first gene activation portion; an E3 ubiquitin ligase; a death agent, wherein expression of the death agent is controlled by a promoter DNA sequence specific for the DNA binding portion, and a polypeptide of 60 amino acids or less, wherein the polypeptide modulates the interaction between the first test protein and the E3 ubiquitin ligase to result in accelerated degradation of the first test protein.
[0260] In some embodiments, the host cell may further comprise a second fusion protein comprising a second DNA binding portion, a second test protein, and a second gene activation portion; and a positive selection reporter gene, wherein expression of the positive reporter gene is controlled by a second promoter DNA sequence specific for the second DNA binding portion.
[0261] The host cell may have a mutational background that enables the uptake of the small molecule. In some cases, the host cell has a mutational background that enables increased transformation efficiency.
[0262] In certain embodiments, disclosed herein are host cells comprising a plasmid vector, wherein a DNA sequence encoding a first polypeptide is inserted in a frame with Gal4-DBD and VP64-AD, and a second polypeptide is inserted in a frame with LexA-DBD and VP64-AD, and wherein the DNA sequence encodes an E3 ubiquitin ligase.
[0263] In certain embodiments, disclosed herein are kits comprising the plasmid; and transfectable host cells compatible with the plasmid, or any combination thereof. In some embodiments, the host cells provided have been transfected with components of the plasmid. In some embodiments, the kit includes selectable reagents for use with host cells transfected with the plasmid. In some embodiments, a variant library of any plasmid is provided, wherein more than one pair of bait proteins or E3 ubiquitin ligases are provided. Such libraries can be used, for example, to screen for reagents with selective protein targeting. In some embodiments, a variant library of a polypeptide plasmid is provided, wherein a plurality of different short test polypeptide sequences for screening are provided. A plurality of different short peptide sequences can be randomly generated by any method (e.g., NNK or NNN nucleotide randomization). A plurality of different short peptide sequences can also be preselected for binding to the target by previous experimental selection, or preselected from existing data sets in scientific literature where rationally designed peptide libraries have been reported.
[0264] Host cells can be made further permeable to small molecules, for example, by deleting genes encoding drug efflux pumps such as PDR5. Genes encoding transcription factors such as PDR1 and PDR3 induce expression of efflux pumps, including but not limited to the 12 genes described in the 12-gene ΔOHSR (Chinen, 2011). Host cells can be further permeable to small molecules by interfering with ergosterol synthesis and deposition in the plasma membrane, for example, by deleting ERG2, ERG3, and / or ERG6 or driving their expression under a regulatable promoter.
[0265] Host cells can also carry mutations to achieve more efficient vector transformation and / or more efficient small molecule uptake.
[0266] The mentioned plasmids can be used in various arrangements. In some embodiments, integration of the plasmid into the genome of the host cell is followed by transformation of the library with randomly encoded peptides using, for example, NNK or NNN codons.
[0267] In some embodiments, in order to screen to identify peptides that can mediate degradation of a target protein, host cells are propagated in a selective medium to ensure the presence of the desired plasmid and expression of non-target proteins (e.g., on a medium lacking a positive selection marker for yeast, or in a medium containing antibiotics for human or bacterial cells). The host cells can then be transformed with a peptide library plasmid and immediately transferred to a selective medium to ensure that all components are present (i.e., on a medium lacking plasmid selection markers for yeast or antibiotics for bacterial or mammalian cells), and to induce expression of the target protein that activates expression of any inducible component, such as a death agent (e.g., with Gal, doxycycline, etc.).
[0268] In other embodiments, plasmids are used as a "plug-and-play platform" utilizing a yeast mating type system, wherein one or more (or two or more) plasmids (or the genetic elements therein) are introduced into the same cell by cell fusion or cell fusion followed by meiosis rather than transfection. This cell fusion involves two different yeast host cells with different genetic elements. In this embodiment, yeast host cell 1 is one of MATa or MATα and includes the integration of target protein and E3 ubiquitin ligase plasmids. In this embodiment, yeast host cell 1 strain can be propagated on positive selection medium to ensure the presence of protein. In this embodiment, yeast host cell 2 can be of the opposite mating type. This strain carries (or has integrated) a random peptide library and a "death agent" (e.g., a cytotoxic reporter gene) plasmid. Yeast host cell 2 can be generated by a large-scale, efficient transformation protocol, which ensures highly diverse library variations in cell culture. Aliquots of the library batch can then be frozen to maintain consistency. In this embodiment, the strains are mated in batches to produce a diploid strain carrying all markers, target protein, E3 ligase, positive selection, "death agent" and peptides. This batch culture can then be propagated on solid medium that selects for all system components (ie, medium lacking both positive selection markers) and induces expression of any inducible components (ie, using Gal).
[0269] Colonies that survive limiting dilution experiments on host cells carrying the target protein and E3 ligase, along with the library / cytotoxicity construct (introduced into the cells by transfection or mating), can constitute colonies with the specific target protein that has been degraded by the peptide and no longer triggers the death cascade triggered by the encoded "death agent" (e.g., a cytotoxicity reporter gene), while maintaining expression of the bait variant protein that drives the positive selection marker. The peptide sequence can be obtained by DNA sequencing the peptide coding region of the plasmid in each surviving colony.
[0270] To ensure that survival is due to degradation of the target and not random chance or erroneous gene expression, an inducible promoter can be used to inactivate the E3 ligase or peptide production and confirm specificity. In some embodiments, cell survival is observed only on medium containing galactose, in which all components are expressed; when peptide expression is lost, no survival is observed on medium without galactose.
[0271] Plasmids can also be isolated and re-transformed into fresh host cells to confirm specificity. Biochemical fractionation of live host cells containing the target, E3 ligase, peptide, positive selection, and "killing agent" followed by pull-down experiments can be performed to confirm the interaction between the peptide sequence and the target protein or E3 ligase using a coding tag as part of the fusion construct (e.g., Myc-tag, HA-tag, His-tag). This also facilitates SAR to determine the binding interface.
[0272] The peptides used in screening assays can be derived from complex libraries that involve post-translational modification enzymes. Modified peptides can be analyzed by methods such as mass spectrometry and can also be sequenced to identify (ID) the primary sequence. The intrinsic membrane permeability of peptides can also be tested by reapplying the peptide exogenously (from lysate) to host cells and observing the inactivation or activation of reporter genes.
[0273] Once a sufficient number of viable host cell colonies have been sequenced, highly conserved sequence patterns will emerge and can be easily identified using multiple sequence alignments. Any such patterns can be used to "anchor" residues in the library peptide insert sequence and arrange variable residues to generate diversity and achieve tighter binding. In some embodiments, this can also be accomplished using algorithms developed for pattern recognition and library design. After convergence, the patterns of disrupted peptides determined by sequencing can be used to define peptide disruption sequences. Convergence is defined as no new sequences being retrieved in the last iteration relative to the second-to-last iteration.
[0274] In some embodiments, peptide libraries can be generated and / or used for screening as described herein. The peptides in the generated library can be peptides with drug-like properties. The peptides used in the screening assay can be derived from processes involving enzymes that modify peptides after translation. For example, in order to generate a peptide library with an N-methylated backbone or a macrocyclic structure, methyltransferases (e.g., those described in Table 3) and prolyl oligopeptidases (e.g., those described in Table 2) can be used to generate the library, such as Figure 6As shown. In some cases, the diversified core peptide sequence can be initially flanked by a homologous recognition site (RS) for the corresponding prolyl oligopeptidase or lactamase, and subsequently released from the linear product to form a cyclic peptide. The core peptide sequence can be further post-translationally modified using enzymes, such as those described in Tables 5-7.
[0275] In an alternative approach, to generate libraries of drug-like N-methylated and / or macrocyclic peptides (e.g., for designing systems for identifying "bridging" peptides or peptides that inhibit protein-protein interactions), methyltransferases (e.g., such as those described in Table 3 ) can be used, wherein a protease cleavage site (e.g., TEV protease) is inserted upstream of the diversified core peptide sequence, as described in Table 3 . Figure 7 As shown. After the methylation of the core sequence, protease induction can release the core peptide flanked by recognition sites (N-terminal or C-terminal sites) for other processing enzymes that can realize cyclization, such as butterfly bean viscose or sorting enzyme. Alternatively, an intein donor site can be included to induce cyclization. In some cases, the cyclization process may introduce a minimal "scar" sequence in the final macrocycle. Using the enzyme described in Table 5-7, the core peptide sequence can be further post-translationally modified.
[0276] Example
[0277] Example 1: Methods for identifying molecules that lead to selective protein degradation.
[0278] This is an embodiment of a system that uses two variants of a protein fused to different DBDs to identify promoters of degradation of specific variants. The integration plasmid is used to integrate into the Saccharomyces cerevisiae protein that constitutes the protein of interest and the E3 ligase. The plasmid encodes a fusion of AD (VP64) and DBD (Gal4) with KRas (G12D), as well as another fusion construct of AD (VP64) and DBD (LexA) with KRas, and the E3 ubiquitin ligase Cereblon (CRL4-CRBN). The protein fusion sequence is tagged with FLAG, MYC or HA. The plasmid also includes yeast replication and selection markers (TRP1 and CEN). The plasmid also has a site for integration into the genome at a specific locus.
[0279] Saccharomyces cerevisiae is co-transformed with a selection and library plasmid expressing a random peptide library, NNK 20-mer sequences. The selection plasmid is driven by the strong ADH1 promoter. The selection and library plasmid also contains a sequence encoding a HIS tag.
[0280] The selection and library plasmids also contain a LexAop sequence that, when bound to a functional transcription factor formed by the Gal4-KRas(G12D)-VP64 fusion protein, induces expression of a "death agent" (a cytotoxic reporter gene). The selection and library plasmids also contain the positive selection marker ADE2, which is controlled by the LexA-KRas-VP64 fusion protein and results in expression of the positive selection marker upon expression of the fusion protein. The plasmids also contain yeast replication and selection markers (TRP1 and CEN).
[0281] Screening is performed by mating a batch of strains to produce a diploid strain that carries all markers, target protein, E3 ligase, positive selection, killing agent, and peptide. This batch culture is then propagated on solid medium, which allows selection of all system components (medium lacking two nutrient components) and induction of expression of any inducible components with Gal.
[0282] Surviving colonies consisted of cells with degraded KRas(G12D) that no longer triggered the death cascade induced by the encoded death agent, whose degradation was facilitated by a peptide bridged to Cereblon. The same cells also expressed WT Kras, which was not targeted and was driving positive selection for survival.
[0283] By DNA sequencing of the peptide coding region of the selection and library plasmids in each surviving colony, the peptide sequence capable of selectively degrading KRas(G12D) was obtained.
[0284] To confirm specificity, an inducible marker is used to inactivate E3 ligase production and confirm specificity. The plasmid is then isolated and re-transformed into a fresh parental strain to confirm specificity.
[0285] Live strains containing the target, E3 ligase, peptide, selection marker, and killing agent are biochemically fractionated, followed by pull-down experiments to confirm the interaction between the peptide sequence and either protein using the encoded tag.
[0286] An alternative embodiment can be accomplished by converting LexA with Gal4. In another alternative embodiment, the fusion protein in any construct is driven by the inducible promoter GAL1 instead of the ADH1 promoter. In another embodiment, the yeast selection marker 2um is included in the target and E3 ligase integration plasmids and the selection and library plasmids instead of CEN. Similarly, the yeast selection marker LEU2 can be used alternatively in another embodiment. In yet another embodiment, the N-terminus of the peptide translated from the selection and library plasmid can be alternatively glycine, alanine, serine, threonine, valine or proline. In other embodiments, the genetic reporter gene in the confirmation plasmid is HIS3 or URA3 instead of ADE2. Any mating type of the haploid state of Saccharomyces cerevisiae can be used as the background strain in the alternative embodiment. In other embodiments, the peptide library can be expressed from a scaffold capable of post-translational modification. In other embodiments, the background strain also expresses enzymes for peptide cyclization and methylation, such as lanthipeptide compound maturases (LanB, LanC, LanM, LanP) from Lactococcus lactis, patamide biosynthesis factors (PatD, PatG) from cyanobacteria, butterfly pea myxase 1 from butterfly pea, and GmPOPB from Capsulophora striata or other species.
[0287] Example 2: Negative readout of target protein degradation
[0288] In this embodiment, the target bait is operatively linked to a positive selection marker that enables growth in the absence of an essential nutrient (e.g., Figure 1 Schematic diagram shown). In this case, cells are plated on a degradation readout medium and viability is measured in the presence or absence of a bridging agent. If the bridging agent is able to functionally bridge the bait protein to the specific E3 ligase of interest, the bait is degraded and the cells are unable to grow on the degradation readout medium because they do not express the positive marker required for growth. As previously described (Chinen, 2011), Saccharomyces cerevisiae cells are modified to minimize drug efflux by deleting the drug efflux pump and the transcription factor encoding its expression. The E3 ligase used in this example is TIR1. The target is a plant protein (auxin-responsive protein - AXR) that is ubiquitinated and degraded in response to indoleacetic acid fused to the TetR DBD and AD. The DBD upstream of the positive selection marker, in this case ADE2, is bound. Cells are plated at an OD of 1.0 in a 5-fold dilution series, of which 2 μl is spotted on selection plates without adenine, with or without 25 μM of the bridging agent naphthaleneacetic acid (NAA). The plates were incubated at 30°C and imaged after 2 days. Figure 8 middle.
[0289] Example 3: Positive readout of target protein degradation
[0290] In this example, the target decoy is operatively linked to a "death agent" negative selection marker that arrests cell growth when expressed, also as Figure 2 The survival of cells expressing different E3 ligases is assessed in the presence of exogenously provided bridging agents. In cases where the bridging agent is able to functionally bridge the bait protein to the specific E3 ligase of interest in a manner that results in ubiquitination of the bait / DBD fusion protein, the bait is degraded and the cells grow because they do not express the "death agent" marker.
[0291] As previously described (Chinen, 2011), cells were modified to minimize drug efflux by deleting the drug efflux pump and the transcription factors encoding its expression. The E3 ligase used in this example was TIR1 and the control E3 ligase used was CRBN. The target was a plant protein that is ubiquitinated and degraded in response to indoleacetic acid fused to the TetR DBD and AD. The cells also contained a death agent regulated by the DBD-containing fusion protein. The cells were seeded at OD-1.0 in a 5-fold dilution series, of which 2ul was spotted on a selection plate containing 25uM of the bridging agent NAA. The plates were incubated at 30°C and imaged after 2 days. The results are shown in Figure 9 middle.
[0292] In another example, survival assays are performed on cells expressing a heterologous E3 ligase (in this case, TIR1) to identify bridging agents that can bridge TIR1 to the bait and cause degradation, thereby enabling cell growth ( Figure 10A The target is a plant protein (AXR) that is ubiquitinated and degraded in response to indoleacetic acid fused to the TetR DBD and AD. Cells containing the E3 ligase and the death agent are mated with cells containing the target. The mated cells are then cultured at a starting OD of 600 = 0.05 was aliquoted into 96-well plates. In this example, each compound was used at a concentration of 10 uM in four replicate wells and the starting OD 600 =0.05 plated cells (border wells on each side served as negative controls). Plates were incubated at 30°C with shaking and OD values were calculated in a Cytation 1 instrument. 600 Growth was continuously monitored. From the conditions showing positive growth, the bridging agent could be identified. Functional bridging agents were NAA (pores C, D / 6, 7), 2,4-DCPA (pores E, F / 8, 9) and PAA (pores G, H / 8, 9) (structure as Figure 10B As shown). Figure 10C The final time point shown is 48 hours.
[0293] In another example, survival assays are performed on cells expressing a heterologous E3 ligase (in this case COI1b) to identify bridging agents that can bridge COI1b to the bait and cause degradation, thereby enabling cell growth (e.g., Figure 11A As previously described (Chinen, 2011), Saccharomyces cerevisiae cells were modified to minimize drug efflux by deleting the drug efflux pump and the transcription factor encoding it. The target used was a plant protein (AXR) that is ubiquitinated and degraded in response to jasmonate-isoleucine fused to the TetR DBD and AD. Cells containing the E3 ligase and the death agent were mated with cells containing the target. The mated cells were then cultured at a starting OD of 600 = 0.05 was aliquoted into 96-well plates. Each compound was tested in duplicate wells at concentrations below 100 μM. Wells showing cell growth contained various concentrations of coronatine (structured as Figure 11B The plate was incubated at 30°C with shaking and the OD values were calculated in a Cytation 1 instrument. 600 Monitor growth continuously. Figure 11C The final time point shown is 84 hours.
Claims
1. A method for identifying a molecule that triggers degradation of a first test protein in a host cell, the method comprising: (a) expressing in said host cell: (i) E3 ubiquitin ligase; and (ii) a first fusion protein comprising a first DNA binding moiety, the first test protein, and a first gene activation moiety; wherein (i) the host cell comprises a promoter sequence for controlling expression of the killing agent, and (ii) the first DNA binding moiety specifically binds to the promoter sequence; and (b) delivering the molecule to the host cell; wherein, in the absence of the molecule, expression of the death agent is activated, and wherein, in the presence of the molecule, the first test protein is degraded by the E3 ubiquitin ligase.
2. The method of claim 1 , further comprising expressing in the host cell a second fusion protein comprising a second DNA binding portion, a second test protein, and a second gene activation portion, wherein the host cell further comprises one or more positive selection reporter genes driven by one or more promoters having a sequence specific for the second DNA binding portion.
3. The method of claim 2, wherein a plurality of positive selection reporter genes are located within the host cell, wherein each positive selection reporter gene in the plurality of positive selection reporter genes is operably linked to a promoter sequence specific for the second DNA binding moiety.
4. The method of claim 2, wherein the one or more positive selection reporter genes are encoded in a plasmid located within the host cell.
5. The method according to any one of claims 1 to 4, wherein the molecule is from a molecular library.
6. The method of any one of claims 1-4, wherein the molecule is delivered exogenously.
7. The method according to any one of claims 1 to 4, wherein the host cell comprises sequences of more than one gene for expressing a death agent that is activated by a promoter DNA sequence specific for the first DNA binding moiety.
8. The method of any one of claims 1-4, wherein the host cell comprises integrated DNA encoding the fusion protein, integrated DNA encoding the E3 ubiquitin ligase, and plasmid DNA encoding the killing agent.
9. The method of any one of claims 1-4, wherein the first test protein is a KRAS variant.
10. The method of claim 9, wherein the second test protein is KRAS.
11. The method of any one of claims 1-4, wherein the E3 ubiquitin ligase comprises Cereblon.
12. The method of any one of claims 1-4, wherein the first DNA binding moiety is derived from LexA, cI, Gli-1, YY1, glucocorticoid receptor, TetR, or Ume6.
13. The method of any one of claims 1-4, wherein the first gene activation moiety is derived from VP16, GAL4, NF-κB, B42, BP64, VP64, or p65.
14. The method according to any one of claims 1 to 4, wherein the killing agent is an overexpression product of a genetic element selected from DNA or RNA.
15. The method of claim 14, wherein the genetic element is a growth inhibition (GIN) sequence.
16. The method of any one of claims 1-4, wherein the killing agent is a ribosomally encoded xenobiotic agent, a ribosomally encoded poison, a ribosomally encoded endogenous or exogenous gene that causes severe growth defects when mildly overexpressed, a ribosomally encoded recombinase that excises a gene essential for viability, a restriction factor involved in the synthesis of a toxic secondary metabolite, or any combination thereof.
17. The method of claim 16, wherein the lethal agent is cholera toxin, SpvB toxin, CARDS toxin, SpyA toxin, HopUl, Chelt toxin, Certhrax toxin, EFV toxin, ExoT, CdtB, diphtheria toxin, ExoU / VipB, HopPtoE, HopPtoF, HopPtoG, VopF, YopJ, AvrPtoB, SdbA, SidG, VpdA, Lpg0969, Lpgl978, YopE, SptP, SopE2, SopB / SigD, SipA, YpkA, YopM, amatoxin, phalloidin, killer toxin KP1, killer toxin KP6, killer toxin Kl, killer toxin K28 (KHR), killer toxin K28 (KHS), anthrax lethal factor endopeptidase, Shiga toxin, saporin toxin, ricin, or any combination thereof.
18. The method of any one of claims 1-4, wherein the host cell is a eukaryote or a prokaryote.
19. The method according to any one of claims 1 to 4, wherein the host cell is from an animal, a plant, a fungus or a bacterium.
20. The method of claim 19, wherein the host cell is derived from a fungus.
21. The method of claim 20, wherein the fungus is an Aspergillus species.
22. The method of claim 19, wherein the host cell is a microorganism.
23. The method of claim 22, wherein the microorganism is Pichia pastoris.
24. The method of claim 22, wherein the microorganism is Saccharomyces cerevisiae.
25. The method of claim 19, wherein the host cell is derived from bacteria.
26. The method of any one of claims 1-4, wherein the molecule is a small molecule.
27. The method of claim 26, wherein the small molecule is a peptidomimetic.
28. The method of any one of claims 1-4, wherein the molecule is a peptide or a protein.
29. The method of claim 28, wherein the peptide or protein is derived from a naturally occurring protein product.
30. The method of claim 28, wherein the peptide or protein is a synthetic protein product.
31. The method of claim 30, wherein the peptide or protein is a product of a recombinant gene.
32. The method of any one of claims 1-4, wherein the molecule is a polypeptide expressed by a test DNA molecule inserted into the host cell, wherein the test DNA molecule encodes the polypeptide.
33. The method of claim 32, wherein the test DNA molecule is from a library of DNA sequences encoding different polypeptides.
34. The method of claim 33, wherein each DNA sequence in the DNA sequence library is located in a separate vector.
35. The method of claim 33, wherein the library encodes polypeptides of 60 or fewer amino acids in length. The method according to claim 33 , wherein the DNA sequence of the DNA sequence library encodes the 3′UTR of mRNA. The method of claim 36 , wherein the 3′UTR is the 3′UTR of sORF1.
38. The method of claim 33, wherein the polypeptide comprises an N-terminal sequence of methionine-valine-asparagine.
39. The method of claim 32, wherein the polypeptide is processed into a cyclic peptide or a bicyclic peptide in the host cell.
40. The method of claim 32, wherein the polypeptide is a product of a post-translational modification.
41. The method of claim 40, wherein the post-translational modification comprises cleavage.
42. The method of claim 40, wherein the post-translational modification comprises cyclization.
43. The method of claim 40, wherein the post-translational modification comprises bicyclization.
44. The method of claim 42, wherein the cyclization comprises reaction with a prolyl endopeptidase.
45. The method of claim 44, wherein the prolyl endopeptidase is one selected from SEQ ID NOs: 42-58 or a functional fragment thereof.
46. The method of claim 44, wherein the prolyl endopeptidase is one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to one of SEQ ID NOs: 42-58.
47. The method of claim 42, wherein the cyclization comprises reaction with a beta-lactamase.
48. The method according to claim 47, wherein the lactamase is one selected from SEQ ID NO: 119-120 or a functional fragment thereof.
49. The method of claim 47, wherein the lactamase is one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to one of SEQ ID NOs: 119-120.
50. The method of claim 42, wherein the cyclization comprises reaction with a hydroxylase and a dehydratase.
51. The method of claim 50, wherein the hydroxylase comprises SEQ ID NO: 123 or a functional fragment thereof.
52. The method of claim 50, wherein the hydroxylase is one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to SEQ ID NO:
123.
53. The method of claim 50, wherein the dehydratase is one selected from SEQ ID NOs: 124-127 or a functional fragment thereof.
54. The method of claim 50, wherein the dehydratase is one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to one of SEQ ID NOs: 124-127.
55. The method of claim 42, wherein the cyclization is formed by an aminocarboxyethylthiotryptophan bridge.
56. The method of claim 40, wherein the post-translational modification comprises methylation.
57. The method of claim 56, wherein the methylation comprises reaction with an N-methyltransferase.
58. The method according to claim 57, wherein the N-methyltransferase is one selected from SEQ ID NOs: 61-116 or a functional fragment thereof.
59. The method of claim 57, wherein the N-methyltransferase is one having at least 80%, 85%, 90%, 92%, 95%, 97% or 99% sequence identity to one of SEQ ID NOs: 61-116.
60. The method of claim 40, wherein the post-translational modification comprises halogenation.
61. The method of claim 40, wherein the post-translational modification comprises glycosylation.
62. The method of claim 40, wherein the post-translational modification comprises acylation.
63. The method of claim 40, wherein the post-translational modification comprises phosphorylation.
64. The method of claim 40, wherein the post-translational modification comprises acetylation.
65. The method of claim 32, wherein the test DNA molecule comprises a gene sequence that expresses a modifying enzyme.
66. A host cell configured to express: E3 ubiquitin ligase; a first fusion protein comprising a first test protein, a first DNA binding portion, and a first gene activation portion; a killing agent, wherein expression of the killing agent is controlled by a promoter DNA sequence specific for the first DNA binding moiety; and A polypeptide of 60 amino acids or less, wherein the polypeptide modulates the interaction between the first fusion protein and the E3 ubiquitin ligase in a manner that results in accelerated degradation of the first fusion protein.
67. The host cell of claim 66, wherein the host cell further comprises: a second fusion protein comprising a second DNA binding portion, a second test protein, and a second gene activation portion; and A positive selection reporter gene, wherein expression of the positive reporter gene is controlled by a second promoter DNA sequence specific for the second DNA binding moiety.
68. The host cell of claim 66 or 67, wherein the polypeptide encodes an N-terminal sequence for peptide stabilization.
69. The host cell of claim 66 or 67, wherein the polypeptide is encoded by an mRNA, wherein the mRNA comprises a 3'UTR.
70. The host cell of claim 69, wherein the mRNA is the encoded product of a DNA molecule, wherein the DNA molecule is exogenously delivered into the host cell.
71. The host cell of claim 66 or 67, wherein the host cell is a eukaryote or a prokaryote.
72. The host cell of claim 66 or 67, wherein the host cell is from a plant, an animal, a fungus, or a bacterium.
73. The host cell of claim 72, wherein the host cell is from a fungus.
74. The host cell of claim 73, wherein the host cell is a haploid yeast cell.
75. The host cell of claim 73, wherein the host cell is a diploid yeast cell.
76. The host cell of claim 75, wherein the diploid yeast cell is produced by mating a first host cell comprising a DNA sequence encoding the first fusion protein and the E3 ubiquitin ligase with a second host cell comprising a DNA sequence encoding the death agent and an mRNA comprising a nucleotide sequence encoding a polypeptide of 60 or fewer amino acids.
77. The host cell of claim 66 or 67, wherein the host cell has a mutation background that increases small molecule uptake.
78. The host cell of claim 66 or 67, wherein the host cell has a mutation background that increases transformation efficiency.
79. The host cell of claim 73, wherein the fungus is an Aspergillus species.
80. The host cell of claim 73, wherein the fungus is Pichia pastoris.
81. A method for identifying a molecule that triggers degradation of a first test protein, the method comprising: (a) Expression in multiple host cells: (i) E3 ubiquitin ligase; and (ii) a first fusion protein comprising a first DNA binding moiety, the first test protein, and a first gene activation moiety; wherein (i) the plurality of host cells each comprises a promoter sequence for controlling expression of a death agent, and (ii) the first DNA binding moiety specifically binds to the promoter sequence such that expression of the death agent is activated in the absence of a molecule that recruits the E3 ubiquitin ligase to the first fusion protein in a manner that results in ubiquitination and premature degradation of the first fusion protein; (b) delivering a different molecule to each of the plurality of host cells; (c) identifying a molecule that triggers degradation of the first test protein based on the survival of the cell to which the molecule is delivered.
82. The method of claim 81 , further comprising expressing in the plurality of host cells a second fusion protein comprising a second DNA binding portion, a second test protein, and a second gene activation portion in the host cells, wherein the host cells further comprise one or more positive selection reporter genes driven by one or more promoters having a sequence specific for the second DNA binding portion.
83. The method of claim 81 or 82, wherein the different molecules are from a molecular library.
84. The method of claim 81 or 82, wherein the different molecule is delivered exogenously.
85. The method of claim 81 or 82, wherein the first test protein is a KRAS variant.
86. The method of claim 85, wherein the second test protein is KRAS.
87. The method of claim 81 or 82, wherein the E3 ubiquitin ligase comprises Cereblon.
88. The method of claim 81 or 82, wherein the first DNA binding moiety is derived from LexA, cI, Gli-1, YY1, glucocorticoid receptor, TetR, or Ume6.
89. The method of claim 81 or 82, wherein the first gene activation moiety is derived from VP16, GAL4, NF-κB, B42, BP64, VP64, or p65.
90. The method of claim 81 or 82, wherein the killing agent is an overexpressed product of a genetic element selected from DNA or RNA.
91. The method of claim 90, wherein the genetic element is a growth inhibition (GIN) sequence.
92. The method of claim 81 or 82, wherein the killing agent is a ribosomally encoded xenobiotic agent, a ribosomally encoded poison, a ribosomally encoded endogenous or exogenous gene that causes severe growth defects when mildly overexpressed, a ribosomally encoded recombinase that excises a gene essential for viability, a restriction factor involved in the synthesis of a toxic secondary metabolite, or any combination thereof.
93. The method of claim 92, wherein the lethal agent is cholera toxin, SpvB toxin, CARDS toxin, SpyA toxin, HopUl, Chelt toxin, Certhrax toxin, EFV toxin, ExoT, CdtB, diphtheria toxin, ExoU / VipB, HopPtoE, HopPtoF, HopPtoG, VopF, YopJ, AvrPtoB, SdbA, SidG, VpdA, Lpg0969, Lpgl978, YopE, SptP, SopE2, SopB / SigD, SipA, YpkA, YopM, amatoxin, phalloidin, killer toxin KPl, killer toxin KP6, killer toxin Kl, killer toxin K28 (KHR), killer toxin K28 (KHS), anthrax lethal factor endopeptidase, Shiga toxin, saporin toxin, ricin, or any combination thereof.
94. The method of claim 81 or 82, wherein the plurality of host cells are eukaryotic or prokaryotic.
95. The method of claim 81 or 82, wherein the host cell is from an animal, a plant, a fungus, or a bacterium.
96. The method of claim 95, wherein the fungus is Aspergillus or Pichia pastoris.
97. The method of claim 81 or 82, wherein the molecule is a small molecule.
98. The method of claim 97, wherein the small molecule is a peptide mimetic.
99. The method of claim 81 or 82, wherein the molecule is a peptide or protein.
100. The method of claim 99, wherein the peptide or protein is derived from a naturally occurring protein product, a synthetic protein, or a recombinant protein.
101. The method of claim 81 or 82, wherein the molecule is a polypeptide expressed by a test DNA molecule inserted into the host cell, wherein the test DNA molecule encodes the polypeptide.
102. The method of claim 101, wherein the test DNA molecule is from a library of DNA sequences encoding different polypeptides.
103. The method of claim 102, wherein the library encodes polypeptides of 60 or fewer amino acids in length.
104. The method of claim 103, wherein the polypeptide is a product of a post-translational modification.
Citation Information
Patent Citations
Protein interfaces
US10188691B2
Protein interfaces
US20170368132A1
Selective modulation of protein-protein interactions
WO2019099678A1