Multivalent protein and screening method
By developing multivalent protein scaffolds, the problems of high production cost and high risk of immune response of existing protein therapeutics have been solved, and efficient and low-cost screening and design of protein therapeutics have been achieved.
Patent Information
- Application Number
- CN202380081869.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-16
- Filing Date
- 2023-09-28
- Publication Date
- 2025-09-16
AI Technical Summary
Existing protein therapeutics have high production costs, complex production processes and low efficiency, while antibody therapeutics diffuse slowly in the body and have a high risk of immune response, making them difficult to use widely.
Develop multivalent protein scaffolds, including oligomeric cores of multiple subunit monomers and orthogonal binding sites, prepared through genetic engineering or chemical coupling methods to achieve efficient binding and screening of peptide targets.
It enables the rapid, scalable, and adaptable identification and design of protein therapeutics, reducing production costs, improving treatment efficiency, and reducing the risk of immune response.
Smart Images

Figure CN120659804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to multivalent protein scaffolds and their use as modular systems for phenotypic screening of combinations of target molecules and as therapeutic agents. The invention also relates to multidomain polypeptide constructs comprising multiple binding domains and a structural domain. The invention further relates to methods for identifying new therapeutic agents using the protein scaffolds, and to therapeutic agents that can be identified in this manner. Background Art
[0002] There is a continuous need to identify new therapeutic modalities for a variety of pathological conditions.
[0003] Protein-based therapies offer an attractive approach to address many common diseases. They have demonstrated high rates of clinical success, and many protein therapeutics have been approved for clinical use by regulatory agencies around the world.
[0004] Protein therapeutics have a variety of modes of action—for example, replacing defective or abnormal proteins; enhancing existing pathways; providing new functions or activities with therapeutic utility; interfering with molecules or microorganisms; and delivering other compounds or proteins such as radionuclides, cytotoxic drugs, or effector proteins. Therapeutic proteins can be classified based on their physical and structural properties and can be divided into, for example, antibody-based drugs, Fc fusion proteins, anticoagulants, blood factors, bone morphogenetic proteins, genetically engineered protein scaffolds, enzymes, growth factors, hormones, interferons, interleukins, thrombolytics, etc. Therapeutic proteins can also be classified based on their molecular mechanism of activity. For example, monoclonal antibodies generally act by non-covalently binding to their target. Enzymes can affect covalent bonds within the target. Other proteins, such as serum albumin, can exert their activity without specific interactions.
[0005] Therapeutic antibodies are a class of protein therapeutics that have been successfully used in the clinic. Antibodies, also known as immunoglobulins (Ig), have been studied and evaluated as potential treatments for a variety of disease conditions. For example, monoclonal antibody therapies have been used to treat diseases including rheumatoid arthritis, multiple sclerosis, psoriasis, and various forms of cancer. Marketed antibody therapeutics include muromomab, abciximab, rituximab, daclizumab, basiliximab, palivizumab, infliximab, trastuzumab, etanercept, gemtuzumab, alemtuzumab, ibritomomab, and rituximab. b), adalimumab, alefacept, omalizumab, tositumomab, efalizumab, cetuximab, bevacizumab, natalizumab, ranibizumab, panitumumab, eculizumab, and certolizumab.
[0006] Antibodies typically contain four polypeptide chains that form an Fc region and two antigen-binding (Fab) regions. Each Fab region contains a variable region (Fv) that forms a paratope and contacts the antigen. The variable regions of naturally occurring antibodies are generally symmetrically bound. However, this limitation means that the variable region of a given antibody can usually only target a single type of receptor (or other target).
[0007] To address this issue, much research interest has shifted to bispecific antibodies. Bispecific antibodies differ from traditional monospecific antibodies in that their two Fab sites can each bind to a different antigen. Bispecific antibodies are generally categorized as Ig-like or non-Ig-like, the latter consisting of chemically linked Fab regions.
[0008] Currently, people are actively studying the clinical application of bispecific antibodies. Two examples of bispecific antibodies on the market are: Blinatumomab, trade name Blincyto, which contains both CD3 sites targeting T cells and CD19 sites targeting B cells, and can be used to treat Philadelphia chromosome-negative relapsed or refractory acute lymphoblastic leukemia; and Emicizumab, trade name Hemlibra, which simultaneously targets both coagulation factors IXa and X and is used to treat hemophilia A. Bispecific antibodies are often used to bind to multiple types of cells simultaneously, for example, by recruiting cytotoxic immune cells while binding to tumor cell receptors.
[0009] Although some bispecific antibodies have good prospects, problems still exist. The production cost of antibody therapeutics is high, especially due to their size and complex post-translational modification chemistry, including complex glycosylation patterns. Antibody production requires extremely large-scale mammalian cell culture and then a large number of purification steps, which leads to extremely high production costs and limits the widespread use of such drugs. In addition, antibodies are known to have limited application in cancer treatment due to poor tumor targeting (for example, studies have shown that the proportion of antibodies that interact with tumors after administration to mouse xenograft models is generally less than 20%). The Fc part of antibodies (such as IgG antibodies) can interact with various receptors expressed on the surface of multiple cells, thereby increasing their retention in the blood circulation. In addition, antibodies diffuse slowly in the body due to their large size. IgG-like antibodies can be immunogenic, leading to harmful immune responses downstream due to Fc receptor activation. Specifically, for bispecific antibodies, the existing "knobs into holes" approach for Ig-like antibodies is difficult to modify for screening a large number of antigen-binding domains due to insufficient modularity. Furthermore, this bispecific antibody approach is practically limited to screening of Fv / Fab regions and cannot be applied to investigate the therapeutic potential of domains from proteins other than immunoglobulins. While tandem fusion approaches for non-Ig-like antibodies are more easily engineered, they are difficult to scale up.
[0010] Blanco-Toribio et al. (Monoclonal Antibodies (MAbs), January 1, 2013, Vol. 5, No. 1, pp. 70-79) describe the generation and characterization of monospecific and bispecific hexavalent trimers. These molecules, termed "trimers," utilize a modified version of the N-terminal trimerization region of the non-collagen 1 (NC1) domain of human collagen XVIII flanked by two flexible linkers as a trimerization scaffold. By simultaneously fusing single-chain variable fragments (scFvs) with identical or different specificities to the N- and C-termini of the trimerization scaffold domain, the authors generated monospecific and bispecific hexavalent molecules that were efficiently secreted as soluble proteins from transfected mammalian cells. Bispecific trimers with anti-laminin and anti-CD3 properties at the N- and C-termini, respectively, were found to also form trimers in solution. A drawback of this approach is that it is cumbersome to manufacture due to the transfection involved.
[0011] WO-A-2020 / 0188346 describes a bispecific antigen-binding protein in which two antigen-binding domains ("ABD") are covalently bound to a fusion protein composed of two or more domains that form an isopeptide bond containing the antigen-binding protein. Such domains forming the isopeptide bond are typically catcher domains such as SpyCatcher (SC), and the resulting bispecific protein has the form ABD-SC-SC-ABD.
[0012] Brune et al. (Bioconjugate Chem., 2017, Vol. 28, No. 5, pp. 1544–1551) describe a "plug-and-play" synthetic assembly using dual-antigen immunoorthogonally reactive proteins. The authors genetically engineered a multimeric coiled-coil structure, IMX313, and two orthogonally reactive shedding proteins to create a dual-addressable synthetic nanoparticle. This construct, named SpyCatcher-IMX-SnoopCatcher (SnoopCatcher), provides a modular platform for multimerizing SpyCatcher with antigen and SnoopCatcher with antigen on opposite sides of the particle simply by mixing.
[0013] Therefore, a new model is needed to identify and design protein therapeutics that can overcome all or part of the above problems in a rapid, scalable and adaptable manner, especially a new platform technology that is comparable to the traditional antibody platform. Summary of the Invention
[0014] The present inventors have recognized the above problems. Currently, non-antibody assembly platforms are recognized as customizable, reproducible, scalable, and adaptable scaffolds that can be used to screen, identify, and develop new therapeutic agents. The methods described in this application can evaluate the potential therapeutic benefits of a variety of different protein geometries, potencies, and / or functions.
[0015] The inventors' proposed approach aims, in part, to develop protein constructs with advantageous properties. Such constructs can be produced recombinantly by expression as fusion proteins, or their component domains can be linked by other methods known in the art, such as chemical coupling. The inventors have discovered that the advantage of simultaneously modified polypeptides can be achieved by modifying both the N-terminus and the C-terminus. Such modifications typically involve the addition of polypeptide domains capable of separately binding to target molecules, such as antigen-binding regions or regions forming isopeptide bonds. The N-terminus and C-terminus can each bind to a different target molecule, creating a so-called bispecific binding construct. The resulting protein constructs are capable of binding to target molecules at both the modified amino terminus and the modified C-terminus. The inventors have specifically engineered protein constructs with the same general orientation for both the N-terminus and the C-terminus. The resulting constructs bind to binding partners at each terminus when the binding partners at each terminus have the same general spatial orientation (e.g., when bound to a solid surface such as a plate or bead, or when located on a cell surface). This approach can be referred to as imparting a "cis" orientation to the modified N-terminus and C-terminus of a single polypeptide chain. Typically, bispecific constructs are in cis orientation. Two or more such protein constructs can be combined to form an oligomeric protein.
[0016] These protein constructs can be used to create combinatorial systems that can be used to screen for useful combinations from combinations of effector moieties such as binding regions (e.g., antigen binding regions). Furthermore, after identifying a useful combination, the construct can be modified to remove portions required for combinatorial screening (or, for example, replace them with linkers), thereby obtaining a simpler protein construct with the identified advantageous binding region combination. Such constructs can be used as therapeutic, diagnostic, or analytical agents, particularly in monomeric form or in the form of oligomers composed of more than one construct.
[0017] Accordingly, the present application provides a multivalent protein scaffold comprising:
[0018] - an oligomeric core comprising a plurality of subunit monomers; and
[0019] - at least one first binding site orthogonal to at least one second binding site,
[0020] The first binding site and the second binding site are located on the same surface of the support.
[0021] Also provided is a multivalent protein scaffold comprising:
[0022] - an oligomeric core comprising multiple subunit monomers;
[0023] - at least one first binding site orthogonal to at least one second binding site,
[0024] wherein the first binding site and the second binding site are located on the same side of the scaffold; and
[0025] The first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target.
[0026] Also provided is a multivalent protein scaffold comprising:
[0027] - an oligomeric core comprising multiple subunit monomers;
[0028] - at least one first binding site orthogonal to at least one second binding site,
[0029] wherein the first binding site and the second binding site are located on the same side of the scaffold; and
[0030] The oligomer core does not include the antibody Fc region.
[0031] Preferably, the oligomer core comprises at least three subunit monomers. More preferably, the oligomer core comprises 3 to 6 subunit monomers.
[0032] Preferably, in one embodiment, the subunit monomers are non-covalently linked together. Preferably, in another embodiment, the subunit monomers are covalently linked together. Preferably, when the subunit monomers are covalently linked together, the subunit monomers are genetically fused together. In some embodiments, the subunit monomers are expressed as a single polypeptide chain derived from a recombinant nucleic acid.
[0033] In one embodiment, the oligomer core is preferably a homo-oligomer core. In such embodiments, each monomer in the oligomer core preferably includes at least one first binding site and at least one second binding site, and wherein the at least one first binding site is orthogonal to the at least one second binding site. Preferably, in one aspect, each monomer includes a first binding site connected to the first end of the monomer and a second binding site connected to the second end of the monomer. Preferably, the first end and the second end of each monomer are located on the same face of the monomer. Preferably, in another aspect, each monomer includes a first binding site connected to the first end of the monomer and a second binding site connected to the first binding site.
[0034] In another embodiment, the oligomer core is preferably a hetero-oligomer core. Preferably, in such embodiments, the core comprises: at least one first subunit monomer comprising a first binding site; and at least one second subunit monomer comprising a second binding site, and wherein the first binding site is orthogonal to the second binding site.
[0035] Preferably, in the present invention, the protein scaffold provided herein is generally as follows: each subunit monomer comprises less than 300 amino acids, preferably less than 200 amino acids, more preferably less than 150 amino acids. Preferably, the oligomeric core has a molecular weight of less than about 150 kDa, preferably less than about 100 kDa, more preferably less than about 70 kDa.
[0036] In some embodiments, the oligomer core does not include an Fc region of an antibody. In some embodiments, the oligomer core does not include a CH2 domain. In some embodiments, the oligomer core does not include a CH3 domain. In some embodiments, the oligomer core does not include a CH2 domain or a CH3 domain.
[0037] Preferably, the oligomer core and / or scaffold does not produce an immune response when administered to a human subject. This application will be described in further detail.
[0038] In some embodiments, the oligomer core and / or scaffold or domain does not produce a deleterious immune response when administered to a human subject. For example, no active B cell or T cell response is generated against the domain, and / or the domain does not specifically bind to an immunoglobulin receptor or activate antibody-dependent cell-mediated cytotoxicity (ADCC).
[0039] Preferably, the oligomeric core comprises a soluble multimeric structural unit of a multimeric protein. Preferably, the multimeric protein comprises a collagen NC (non-collagenous) domain (such as the NC1 domain), CutA1, the C1q head domain, TNF, p53, fibrinogen, C4, Bacillus subtillus AbrB, or a homologue or paralogue thereof.
[0040] Preferably, the multimeric protein comprises collagen VIII NC1 (non-collagenous) domain, collagen X NC1 (non-collagenous) domain, C1q head domain, CutA1 protein, macrophage migration inhibitory factor (MIF) or macrophage migration inhibitory factor 2 (MIF-2), tumor necrosis factor (TNF), TNF family proteins including TL1A, CD40L or OX40L, or homologs or paralogs thereof.
[0041] Preferably, the multimeric structural unit comprises a polypeptide having at least 30% or at least 50% amino acid identity with SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:29, SEQ ID NO:60, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:42, SEQ ID NO:31, SEQ ID NO:58, SEQ ID NO:78, SEQ ID NO:80 or SEQ ID NO:19.
[0042] In particular, the inventors have identified CutA1, typically human CutA1, as a beneficial structural element. In some embodiments, the CutA1 protein is genetically engineered to contain one or more substitutions or deletions relative to wild-type CutA1. CutA1 domains, optionally as genetically engineered, have found particular utility as scaffolds. Thus, in some embodiments, the CutA1 domain is provided as a fusion protein linked to a different polypeptide domain via a peptide bond at the N-terminus and / or C-terminus of CutA1.
[0043] In certain embodiments, the multimerization structural element is human CutA1 (SEQ ID NO: 19). In certain embodiments, the multimerization structural element is a truncated form of SEQ ID NO: 19 (human CutA1), as described elsewhere herein. In certain embodiments, the multimerization structural element is a modified form of SEQ ID NO: 19 (human CutA1), wherein at least one cysteine residue in SEQ ID NO: 19 is substituted with a different amino acid residue, as described elsewhere herein. In certain embodiments, the multimerization structural element is a modified and truncated form of SEQ ID NO: 19 (human CutA1), wherein at least one N- and / or C-terminal amino acid residue is removed, and wherein at least one cysteine residue in SEQ ID NO: 19 is substituted with a different amino acid residue, as described elsewhere herein.
[0044] In certain embodiments, the multimerization structural element is a cytokine. In certain embodiments, the multimerization structural element belongs to the TNF superfamily. In certain embodiments, the multimerization structural element is TNF (SEQ ID NO: 80) or OX40L (SEQ ID NO: 78) or CD40L (SEQ ID NO: 58) or TL1A (SEQ ID NO: 31) or is derived from TNF or OX40L or CD40L or TL1A (e.g., is a truncated and / or modified form thereof), as described elsewhere herein. Thus, the multimerization structural element may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity with the amino acid sequence of human CutA1 (SEQ ID NO: 19), OX40L (SEQ ID NO: 78), CD40L (SEQ ID NO: 58), TL1A (SEQ ID NO: 31) or TNF (SEQ ID NO: 80).
[0045] In certain embodiments, the multimerization structural element is a modified and truncated form of SEQ ID NO:78, SEQ ID NO:58, SEQ ID NO:80, or SEQ ID NO:31, wherein at least one N-terminal and / or C-terminal amino acid residue is removed, and wherein at least one cysteine residue is substituted with a different amino acid residue, as described elsewhere herein.
[0046] CutA1 is structurally similar to the non-human family of PII proteins (Bagautdinov et al., Acta Crystallogr Sect F Struct Biol Cryst Commun. 2008 May 1; 64(Pt 5):351-7). In some embodiments, the structural element comprises or consists of a protein derived from the PII protein family or a truncated or modified form thereof. PII proteins include, for example, the signal transduction protein PII GlnB from Synechococcus elongatus PCC 7942 (SEQ ID NO: 81) (Xu et al., Acta Crystallogr D Biol Crystallogr. 2003; 59: 2183-2190), GlnK from Escherichia coli (SEQ ID NO: 82) (Xu et al., J Mol Biol. 1998; 282: 149-165), the PII-like domain of YqfO from Bacillus cereus (SEQ ID NO: 83) (Godsey et al., Protein Sci. 2007 Jul; 16: 1285-1293), or the hypothetical protein SA13888 from Staphlococcus aureus (SEQ ID NO: 84) (Saikatendu et al., BMC Structure Biol. 2006; 6:27), or a truncated or modified form thereof as described by Lüddecke et al., Sci Rep.
[0047] Preferably, the first binding site and / or the second binding site comprises a protein domain. Generally, the first binding site comprises a first protein domain and the second binding site comprises a second protein domain. Preferably, the first binding site and / or the second binding site are genetically fused to the subunit monomer to which they are attached to form a single polypeptide chain.
[0048] Preferably, the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target. Preferably, the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target. More preferably, the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target. Preferably, the first protein domain is capable of forming an isopeptide bond with the first polypeptide target, and the second protein domain is capable of forming an isopeptide bond with the second binding target.
[0049] Preferably, each of the first and second binding sites comprises a different split-ligand binding protein domain. More preferably, one of the first and second binding sites comprises a shed Streptococcus pyogenes fibronectin binding protein domain, and the other of the first and second binding sites comprises a shed Streptococcus pneumoniae adhesion protein domain.
[0050] Preferably, each of the first binding site and the second binding site independently has at least 50% amino acid identity with any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18. In some embodiments, each of the first binding site and the second binding site independently has at least 60%, at least 70%, at least 80%, or at least 90% amino acid identity with any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18.
[0051] The present application also provides a protein complex comprising the protein scaffold as described herein, wherein the first binding site binds to a first polypeptide target linked to a first effector portion, and the second binding site binds to a second polypeptide target linked to a second effector portion.
[0052] Preferably, in the complex, each of the first binding site / polypeptide target pair and the second binding site / polypeptide target pair is independently selected from the following combinations: (i) any one of SEQ ID NO:4, 6 or 8 and any one of SEQ ID NO:5, 7 or 9; (ii) SEQ ID NO:12 and SEQ ID NO:13 or 15; (iii) SEQ ID NO:5 and SEQ ID NO:11; (iv) SEQ ID NO:15 and SEQ ID NO:16; (v) SEQ ID NO:17 and SEQ ID NO:18; or (vi) SEQ ID NO:23 and SEQ ID NO:16.
[0053] Also provided is a screening platform comprising a library, wherein the library comprises a plurality of protein complex populations as described herein, wherein each of the protein complex populations comprises a different combination of a first effector moiety, a second effector moiety and / or an oligomer core.
[0054] Also provided is a method for identifying a therapeutic drug or drug analog, the method comprising:
[0055] Providing a protein complex as described herein;
[0056] contacting the protein complex with a biological system; and
[0057] Measuring whether a protein complex induces a desired change in the characteristic function of a biological system,
[0058] And optionally further comprising: selecting a protein complex that causes a desired change in a property of the biological system.
[0059] The method may further comprise synthesizing a therapeutic drug or drug candidate comprising an oligomeric core of the scaffold of the protein complex of the identified therapeutic drug analog linked to a first effector moiety and a second effector moiety of the protein complex.
[0060] Also provided is a therapeutic drug candidate obtainable according to the method described herein.
[0061] Also provided is a therapeutic drug obtainable according to the method described in the present application.
[0062] Also provided is a therapeutic agent or therapeutic agent candidate comprising or consisting of one or more constructs or polypeptides as described herein.
[0063] The present application also provides a therapeutic drug or drug candidate comprising an oligomer core comprising a plurality of subunit monomers linked to one or more first effector moieties and one or more second effector moieties, wherein the one or more first effector moieties and the one or more second effector moieties are located on the same face of the oligomer core, and wherein: (1) the one or more first effector moieties comprise two or more first effector moieties, and the one or more second effector moieties comprise two or more second effector moieties; and / or the oligomer core does not comprise an antibody or antibody fragment. Preferably, the oligomer core of the therapeutic drug counterpart is an oligomer core as described in further detail herein.
[0064] Preferably, in the therapeutic drug candidates of the present invention, the oligomeric core comprises a plurality of subunit monomers, and: (i) each subunit monomer comprises collagen NC1 domain, CutA1, C1q domain, TNF, p53, fibrinogen, C4, Bacillus subtilis AbrB or a homolog or paralog thereof; and / or (ii) each subunit monomer comprises a multimeric structural unit comprising a polypeptide having at least 50% amino acid identity with SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3 or SEQ ID NO: 19.
[0065] Preferably, in the therapeutic agent or drug candidate of the present invention, the oligomeric core comprises a plurality of subunit monomers, and: (i) each subunit monomer comprises collagen VIII NC1 (non-collagenous) domain, collagen X NC1 (non-collagenous) domain, C1q head domain, CutA1 protein, macrophage migration inhibitory factor (MIF) or macrophage migration inhibitory factor 2 (MIF-2), tumor necrosis factor (TNF), TNF family proteins including TL1A or CD40L or OX40L, or homologs or paralogs thereof; and / or (ii) each subunit monomer comprises a multimeric structural unit comprising SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 29, SEQ ID NO: 60, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 42, SEQ ID NO: 31, SEQ ID NO: 58, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85 A polypeptide having at least 50% amino acid identity to SEQ ID NO: 80 or SEQ ID NO: 19.
[0066] In some aspects, the present invention provides a polypeptide comprising a first binding domain at the N-terminus and / or a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain. The polypeptide is generally a single genetically engineered polypeptide chain expressed as a fusion protein derived from a recombinant nucleic acid. Generally, the first binding domain and the second binding domain are capable of binding to their targets when the target molecule is expressed in a single cell or fixed on a plate or a single bead. In this application, this approach is sometimes referred to as causing the first binding domain and the second binding domain to have a "cis" orientation.
[0067] Thus, in some aspects, the first binding domain is at the N-terminus of the polypeptide and the second binding domain is at the C-terminus, wherein the first binding domain and the second binding domain are separated by a domain. However, in some aspects, the first binding domain is connected to the second binding domain, and the second binding domain is connected to the domain. Typically, the second binding domain is connected to the N-terminus or the C-terminus of the domain. As described elsewhere herein, the first binding domain and the second binding domain are typically different, but in some embodiments may be the same. A linker may be included between any two domains, or between all domains of a construct, or between any selected domains in a construct, such as a linker peptide of 2 to 30 amino acid residues, typically 5 to 25 amino acid residues.
[0068] In some embodiments, the first binding domain and the second binding domain are different antigen binding domains. Accordingly, the construct is a bispecific construct.
[0069] In some embodiments, the first binding domain and / or the second binding domain is a protein or peptide that can specifically bind to a biomolecule, such as a signaling molecule that can specifically interact with a binding partner such as a protein or peptide ligand or receptor (e.g., a cytokine or cell surface receptor).
[0070] In other embodiments, the first and second binding domains are catcher domains (i.e., split-ligand binding protein domains), each of which is capable of forming an isopeptide bond with a cognate peptide. Such cognate peptides are often referred to as tag peptides. For example, as is known in the art and described below, SpyTag forms an isopeptide bond with the SpyCatcher domain. Generally, the cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain. However, in some embodiments, the cognate peptides of the first and second binding domains may be the same. A polypeptide having a catcher at each end and a tag-containing molecule (e.g., a protein) can be provided, for example, as part of a kit. In some embodiments, each tag peptide is covalently bound to its paired catcher domain, optionally wherein one or both cognate peptides are linked to the first catcher domain and / or the second catcher domain via an isopeptide bond. In some embodiments, one or both cognate peptide tags are provided as a fusion polypeptide with an effector moiety (generally an antigen binding domain). Thus, by linking a catcher to its cognate peptide tag, the effector moiety (e.g., an antigen binding domain) can be linked to its binding domain.
[0071] In embodiments where the homologous peptides are identical for both the first and second binding domains, preferential, selective or sequential binding or coupling can be achieved by temporal or sequential control, spatial control or spatiotemporal control. This can be, for example, by activating or inactivating one binding domain and / or by competitive binding. Without limitation, activation can be achieved, for example, by inhibiting enzymatic cleavage of the peptide (e.g., as described in Driscoll et al., Sep 2023, bioRxiv2023.08.31.555700; doi: https: / / doi.org / 10.1101 / 2023.08.31.555700 , which are incorporated herein by reference in their entirety), or by light activation, for example, by light-triggered covalent bond formation in a photocaged binding domain formed by site-specific incorporation of a non-natural residue such as coumarin-lysine at the reaction site (e.g., see SpyCatcher003, Rahikainen et al., 2023 “Visible light-induced specific protein reaction delineates early stages of cell adhesion”, bioRxiv2023.07.21.549850; doi: https: / / doi.org / 10.1101 / 2023.07.21.549850(described herein, which is incorporated herein by reference in its entirety.) Another photoactivation method using photocaged glutamate analogs is described in Yang et al., Angewendte Chemie, Vol. 62, No. 40, October 2, 2023 (published online August 16, 2023), “Photoactivatable Protein with Genetically Encoded Photocaged Glutamic Acid.” Alternatively, control of reaction site availability can be achieved by removing any inhibitory domains, for example, by steric hindrance, competitive binding, or blocking reactive or catalytic residues.
[0072] In some embodiments, two antigen binding domains, each containing an identical isopeptide bond-forming tag, can be precisely and selectively coupled to a construct of the invention comprising two identical isopeptide bond-forming catcher domains (to which the tag is bound), one of which is unreactive until its reactivity is revealed using light activation. Typically, one of the catcher domains is unreactive because it is bound by a non-natural photoreactive residue in the reactive isopeptide bond-forming site, such as a coumarin-lysine (7-hydroxycoumarolysine, "HCK") amino acid or a photocaged glutamate analog. Once the first tagged peptide is attached to the first catcher by forming an isopeptide bond between the first catcher and the tag, the photocaged residue can be released from the second catcher by applying appropriate light (e.g., 405 nm light). The second catcher is then available for binding, and a second tagged peptide can be added.
[0073] In some embodiments, two antigen binding domains, each containing an identical isopeptide bond-forming tag, can be precisely and selectively coupled to a construct of the present invention comprising two identical isopeptide bond-forming catcher domains (to which the tag is bound), one of which is non-reactive until its reactivity is revealed using a site-specific protease. Typically, one of the catcher domains is non-reactive due to the presence of a non-reactive tag mutant sequence fused to one of the catcher domains via a flexible linker containing a protease cleavage site. Once the first tagged peptide is attached to the first catcher by forming an isopeptide bond between the first catcher and the tag, the addition of a protease can release the non-reactive tag mutant from the second catcher. The second catcher is then available for binding, and the second tagged peptide can be added.
[0074] In certain embodiments, the non-reactive tag mutant is SpyTag003 D117A (SpyTag003DA) (Keeble et al., 2019 Proc. Natl. Acad. Sci. USA 116, 26523–26533, which is incorporated herein by reference in its entirety), which is fused to the C-terminus of a second SpyCatcher003 via a flexible linker containing a tobacco etch virus (TEV) protease cleavage site. After cleavage at the TEV site, the SpyTag003DA peptide will be free to dissociate, exposing the reactive Lys of SpyCatcher003 to enable it to react with the second provided SpyTag-linked conjugate. Cleavage can be performed with any suitable protease, such as SuperTEV protease.
[0075] In some embodiments, the polypeptide comprises a first binding domain at its N-terminus and a second binding domain at its C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain, the first binding domain and the second binding domain are catcher domains, each catcher domain is capable of forming an isopeptide bond with a cognate peptide, wherein the first catcher domain is linked to its cognate peptide tag via an isopeptide bond, and wherein the second catcher domain is linked to its cognate peptide tag via an isopeptide bond. In some embodiments, each peptide tag is linked to an antigen binding domain.
[0076] The polypeptides described in the four paragraphs above can be used as monomers, or combined to form oligomers.
[0077] In some aspects, oligomers are provided that include two or more polypeptides as described above and elsewhere herein.
[0078] In some aspects, the polypeptides or oligomers described in the above paragraphs include the features described elsewhere in this application. For example, the domains of the polypeptide constructs are generally "subunit monomers" described in various sections of this application, and thus the descriptions and definitions of subunit monomers apply equally to such domains. Similarly, the first binding domain and the second binding domain are generally the first binding site and the second binding site as described elsewhere in this application, and thus the descriptions and definitions of the first binding site and the second binding site apply equally to the first binding domain and the second binding domain of the polypeptide construct.
[0079] In some specific aspects, the domain of the polypeptide construct (or the subunit monomer of the oligomer core) is a CutA1 protein, typically a human CutA1 protein. The data provided below, in particular Figure 23, demonstrating various fusions of HsCutA1 ("Homo sapiens CutA1") with effector proteins, and the ability of these direct fusions to exert biological effects. Thus, in certain aspects, the present invention provides polypeptides comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a human CutA1 domain. In certain aspects, CutA1, typically human CutA1, is genetically engineered to remove one or more cysteine residues from the native sequence (herein as SEQ ID NO: 19). Generally, this is a case where one or more cysteine residues are replaced by one or more non-cysteine residues. The removal of one or more cysteine residues advantageously allows for targeted cysteine coupling at non-native sites or in fusion proteins. Without wishing to be bound by theory, unpaired cysteines typically interfere with stability and downstream applications, so removing unpaired cysteines may be advantageous. In some embodiments, it is advantageous to remove one or more unpaired cysteines and reintroduce one or more cysteine residues only at optimized targeting positions.
[0080] In some embodiments, one or more cysteine residues in CutA1 are substituted with one or more alanine residues. In some embodiments, one or more cysteine residues in CutA1 are substituted with one or more valine residues. In some embodiments, one or more cysteine residues in CutA1 are substituted with one or more serine residues. In some embodiments, the substitution comprises or consists of two cysteines being substituted with two alanines, which is referred to herein as a "CACA" substitution. In some embodiments, the substitution comprises or consists of one cysteine being substituted with valine and one cysteine being substituted with serine, which is referred to herein as a "CVCS" substitution. In some embodiments, the cysteine residues at positions 75 and 96 of wild-type human CutA1 (e.g., SEQ ID NO: 19) are substituted with different residues. In some embodiments, human CutA1 is genetically modified to replace two cysteine residues, wherein the cysteine substitution comprises or consists of (i) C75A, C96A or (ii) C75V, C96S. Thus, in some embodiments, the domain of the polypeptide construct (or the subunit monomer of the oligomer core) is a human CutA1 protein that has been genetically engineered to have two cysteine residues substituted, wherein the cysteine substitutions comprise or consist of (i) C75A, C96A or (ii) C75V, C96S. The generation and biological effects of polypeptide constructs comprising such cysteine substitutions in human CutA1 domains are described herein. Figure 25 ).
[0081] In some embodiments, it is advantageous to remove one or more unpaired cysteines and reintroduce one or more cysteine residues only at the optimized target position. The reintroduced cysteine can be used as, for example, a conjugation site for a drug or dye, for example, in the production of a labeled antibody (or antibody type molecule) or an antibody drug conjugate ("ADC") type molecule. Example 13 illustrates that non-cysteine residues are substituted for cysteine residues, and it is observed that CutA1 with newly introduced Cys residues has a higher conjugation efficiency when a dye or drug is conjugated to CutA1 compared to the wild type and negative control (CC041, i.e., CutA1 CACA genetically engineered to remove natural cysteine). Newly introduced one or more cysteines can be substituted into the sequence at a favorable position. As shown in Example 13, exemplary positions in human CutA1 include one or more of V64, E78, K79, K82, E83, K91, Q102, K110, E114, F136, S139, F158, and Q166. In some embodiments, human CutA1 (e.g., SEQ ID NO: 19) is substituted with cysteine at 2, 3, 4, 5, or 6 of residues E78, K82, Q102, E114, F136, and Q166.
[0082] In some embodiments, CutA1 is coupled to an effector molecule, which is any molecule that provides the desired effect. Effector molecules are sometimes referred to as "payloads" in the art, particularly when effector molecules are coupled to antibodies to form antibody-drug conjugates or ADCs. Therefore, such payloads or effector molecules are well known in the art. Generally, effector molecules will comprise or consist of a label, a dye molecule, or a drug, i.e., a substance that has a therapeutic or preventive physiological effect on the human or animal body when ingested by the human or animal body. The drug is typically a pharmaceutical drug.
[0083] In some embodiments, the effector molecule is a nucleic acid, polynucleotide or oligonucleotide. The nucleic acid, polynucleotide or oligonucleotide can be DNA, RNA, XNA, LNA or a mixture thereof. Generally, it is DNA or RNA. The nucleic acid, polynucleotide or oligonucleotide can be single-stranded or double-stranded. In some embodiments, the oligonucleotide comprises 3 to 50 nucleotides (or nucleotide pairs when double-stranded), for example 5 to 30 nucleotides (or pairs), or consists of them. Antibody oligonucleotide conjugates are generally referred to as a subset of ADCs in the art.
[0084] In some embodiments, CutA1 is conjugated to a drug to form a CutA1 drug conjugate. The drug can be of any type, as will be apparent to one skilled in the art. Drugs that can be conjugated to CutA1 include, but are not limited to, analgesics, antibiotics, anticancer drugs, anticoagulants, antidepressants, antidiabetics, antiepileptics, antipsychotics, antispasmodics, antivirals, cardiovascular drugs, depressants, sedatives, and stimulants.
[0085] In some embodiments, two or more different drugs are conjugated to a single CutA1 molecule. Antibody-drug conjugates with dual payloads are described in Yamazaki et al., Nature Communications, Vol. 12, Article No. 3528 (2021), where both MMAE and MMAF are conjugated and are described for use in combating breast tumor heterogeneity and drug resistance.
[0086] In some embodiments, the drug is an anticancer drug, such as a cytotoxic drug. In some embodiments, the anticancer drug is a tubulin inhibitor, such as a maytansinoid such as mertansine (also known as DM1) or auristatin, or a paclitaxel derivative. Examples of drugs that inhibit tubulin polymerization are auristatins, such as tissue factor-directed monomethyl auristatin E (MMAE) and monomethyl auristatin F (MMAF), compounds derived from dolastatin 10 (e.g., TZT-1027 described by Kobayashi et al., Jpn J Cancer Res. 1997 March; 88(3): 316-27), and tubulysin, such as tubulysin A. In some embodiments, the anticancer drug is a DNA damaging agent that causes cell death by cleaving DNA and / or causing DNA alkylation. Examples of DNA damaging molecules are duocarmycin, calicheamicin and pyrrolobenzodiazepines Class (pyrrolobenzodiazepine). In some embodiments, the drug is a topoisomerase I inhibitor, which binds to the complex between topoisomerase I and DNA to stabilize it, thereby preventing DNA from reconnecting, causing DNA damage. Examples of drugs that inhibit topoisomerase I are DXd and SN-38. In some embodiments, the drug is an RNA polymerase II inhibitor, such as α-amanitin. In some embodiments, the drug is an immunomodulator, such as a TLR agonist or a STING agonist (as desired, for example, Fu et al., Signal Transduction and Targeted Therapy, Volume 7, Article Number: Table A of 93 (2022)).
[0087] In some embodiments, the drug is a small molecule drug. Generally, the molecular weight of a small molecule drug is 1000 Da or less, more typically 750 Da or less, or 500 Da or less. This molecular weight is the molecular weight of the drug molecule itself and excludes any linker.
[0088] In some embodiments, the drug is a cytotoxic drug, typically a small molecule cytotoxic drug, such as DXd, mertansine, monomethyl auristatin E (MMAE), or monomethyl auristatin F (MMAF). Cytotoxic small molecule drugs are often referred to as chemotherapy or chemotherapeutic agents, particularly in the context of cancer treatment.
[0089] The drug can be conjugated to CutA1 by any known method. Typically, the drug is conjugated to CutA1 at a cysteine residue, more typically a cysteine residue not present in the native sequence. When conjugated to cysteine, the molecule to be conjugated may contain a maleimide. Cysteine residues readily react with maleimide to form succinimidyl thioether conjugates.
[0090] In some embodiments, the drug has a molecular weight greater than 1000 Da. In some embodiments, the drug is or comprises an oligopeptide or polypeptide comprising 2 or more, e.g., 3, 4, 5, 6, 7, 8, 9, 10 or more, or 20 or more, or 50 or more, or 100 or more amino acid residues covalently linked by peptide bonds. In some embodiments, the drug is an immunotoxin, e.g., Pseudomonas exotoxin A (PE), described by Wolf & Beile (Int J Med Microbiol. 2009 Mar; 299(3): 161-76. doi: 10.1016 / j.ijmm.2008.08.003) as an anticancer agent.
[0091] In some embodiments, CutA1 is coupled to a dye molecule. In some embodiments, the dye molecule is a small molecule dye. Typically, the molecular weight of the small molecule dye is 1000 Da or less, more typically 750 Da or less, or 500 Da or less. In some embodiments, the dye is fluorescent, such as fluorescein. Coupling can be performed by any known method. Typically, the dye is coupled to CutA1 at a cysteine residue, more typically a cysteine residue not present in the native sequence. When coupled to cysteine, the molecule to be coupled can contain maleimide. Cysteine residues readily react with maleimide to form succinimidyl thioether conjugates.
[0092] In some embodiments, the effector molecule is a nanoparticle, such as a gold nanoparticle, a silica nanoparticle, a lipid nanoparticle, or a lipid-polydopamine hybrid nanoparticle ("LPN"), as described in Yang et al., ActaPharmaceutica Sinica B, Vol. 10, No. 11, 2020.11, pp. 2212-2226. Typically, the nanoparticle is coupled to one or more cysteine residues on CutA1 by site-specific coupling to the maleimide group of the lipid nanoparticle or polydopamine (PDA) hybrid nanoparticle.
[0093] In certain aspects, CutA1, typically human CutA1 (e.g., SEQ ID NO: 19), is genetically engineered to delete one or more residues from either or both ends of the native sequence to form a truncated CutA1 domain that is incorporated into the constructs of the invention. In some embodiments, one or more residues are deleted from the N-terminus of CutA1, typically human CutA1. Typically, 5 to 70 residues, such as 10 to 59 residues, are deleted from the N-terminus of CutA1, typically human CutA1. In some embodiments, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 59, 60, 61, 62, 63, 64, 65, or 66 residues are deleted from the N-terminus of CutA1. In some embodiments, the truncated CutA1 begins at residue 30 (i.e., 29 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 33 (i.e., 32 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 44 (i.e., 43 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 60 (i.e., 59 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 67 (i.e., 66 N-terminal residues are missing), for example Figure 23"HsCutA1" shown in b 67-171 CACA" direct fusion construct.
[0094] In some embodiments, one or more residues are deleted from the C-terminus of CutA1, typically human CutA1. A C-terminal deletion can replace or supplement an N-terminal deletion. Typically, 5 to 20 residues are deleted from the C-terminus of CutA1, typically human CutA1, for example, 6 to 12 residues are deleted. In some embodiments, about 8 residues are deleted. Deleting 8 residues in human CutA1 results in a truncated protein with the C-terminus located at Figure 24 and residue 171 (valine) in the wild-type sequence set forth in SEQ ID NO: 19. In some embodiments, the C-terminus of the truncated protein is located at Figure 24 and residue 168 (threonine) in the wild-type sequence set forth in SEQ ID NO: 19. In some embodiments, the C-terminus of the truncated protein is located at Figure 24 and residue 169, 170, 172, 173, 174, 175 or 176 of the wild-type sequence set forth in SEQ ID NO:19.
[0095] In some embodiments, the truncated CutA1 begins at any one of residues 30 to 67 of SEQ ID NO: 19. In some embodiments, the truncated CutA1 begins at any one of residues 44 to 67 of SEQ ID NO: 19. In certain embodiments, the truncated CutA1 consists of residues 44-179 (SEQ ID NO: 29), residues 61-168 of SEQ ID NO: 19, or residues 60-171 of SEQ ID NO: 19, e.g. Figure 24 In some embodiments, the truncated CutA1 begins at any one of residues 44 to 67 of SEQ ID NO: 19 and ends at any one of residues 168 to 179. In other embodiments, the truncated CutA1 begins at any one of residues 30 to 65 of SEQ ID NO: 19 and ends at any one of residues 165 to 179. In some embodiments, the truncated CutA1 begins at any one of residues 44 to 67 of SEQ ID NO: 19 and ends at any one of residues 171 to 179.
[0096] Truncation of CutA1 can improve the accuracy of constructing fusion constructs. Figure 24 The generation, yield and biological effects of polypeptide constructs comprising this truncated human CutA1 domain are shown.
[0097] In certain aspects, CutA1 is human CutA1 that has been genetically engineered to remove one or more cysteine residues from the native sequence and is truncated at the N-terminus and / or C-terminus. In some embodiments, the truncated CutA1 consists of residues 44-179 or residues 60-171 of SEQ ID NO: 19, such as Figure 24 In other embodiments, the truncated CutA1 begins at any one of residues 30 to 65 of SEQ ID NO: 19 and ends at any one of residues 165 to 179 and also has a cysteine substitution comprising (i) C75A, C96A or (ii) C75V, C96S or consisting thereof.
[0098] In certain aspects, CutA1 is human CutA1 that has been genetically engineered to remove one or more cysteine residues from the native sequence, is truncated at the N-terminus and / or C-terminus, and at least one non-cysteine residue is substituted with a cysteine residue. In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19), has a cysteine residue substituted in the native CutA1 sequence that includes or consists of (i) C75A, C96A, or (ii) C75V, C96S, and has at least one cysteine at a residue position that is not a cysteine residue in the native sequence.
[0099] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19) with cysteine substitutions (substitutions-out) including or consisting of C75A, C96A; and cysteine residue substitutions (substituted-in) at one, two, or three of residues K82, E114, and F136.
[0100] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19) with cysteine substitutions (substitutions-out) including or consisting of C75A, C96A, and cysteine residue substitutions (substituted-in) at one, two, or three of residues V64, E78, K79, E83, K91, Q102, K110, S139, F158, and Q166. In some embodiments, the variant CutA1 has cysteine residue substitutions at 2, 3, 4, 5, 6, 7, 8, 9, or 10 of residues V64, E78, K79, E83, K91, Q102, K110, S139, F158, and Q166.
[0101] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19), with a cysteine substitution comprising or consisting of C75A, C96A, and a cysteine residue substitution at one or more of residues E78, Q102, and Q166. In some embodiments, the variant CutA1 has a cysteine residue substitution at two or three of residues E78, Q102, and Q166.
[0102] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19) with a cysteine substitution comprising or consisting of C75A, C96A, and a cysteine residue substitution at one or more of residues E78, K82, Q102, E114, F136, and Q166. In some embodiments, the variant CutA1 has a cysteine residue substitution at 2, 3, 4, 5, or 6 of residues E78, K82, Q102, E114, F136, and Q166.
[0103] In some embodiments, the variant CutA1 comprises a sequence as defined above, or consists of a sequence as defined above, or consists of any of the specific sequences described in the Examples, and comprises 1 to 10 amino acid substitutions at residues not specified as cysteine residue substitutions (substituted-in) or cysteine residue substitutions (substitutions-out). For example, in some embodiments, a variant human CutA1 sequence can be provided that contains 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions at positions other than 95, 64, 76, 78, 79, 82, 83, 91, 95, 102, 110, 114, 136, 139, 158, and 166. These 1 to 10 substitutions can be any of the standard 20 amino acids, although generally they will not replace a non-cysteine residue with a cysteine. Generally, these 1 to 10 substitutions will be conservative substitutions. In some embodiments, one, two, three or more of these 1 to 10 substitutions can introduce non-natural amino acids (UAA), also referred to as non-proteinogenic amino acids. Non-natural amino acids can be used to achieve coupling of target molecules by biorthogonal chemistry, for example, by click chemistry. Non-natural amino acids are known in the art and include D-amino acids, homoamino acids, N-methyl amino acids, hydroxyproline (Hyp), β-alanine, citrulline (Cit), ornithine (Orn), norleucine (Nle), 3-nitrotyrosine, nitroarginine and pyroglutamic acid (Pyr). For example, propargyl lysine is a non-natural amino acid that, when incorporated into proteins, can be utilized to connect commercially available fluorescent azide dyes by copper-catalyzed alkyne-azide cycloaddition click reactions (also referred to as click reactions).Other UAAs suitable for site-specific modification of polypeptide sequences include 1: 3-(6-acetylnaphth-2-ylamino)-2-aminopropionic acid (AnAP), 2: (S)-1-carboxy-3-(7-hydroxy-2-oxo-2H-chromen-4-yl)propan-1-ammonium (CouAA), 3: 3-(5-(dimethylamino)naphthalene-1-sulfonamide)propionic acid (dansylalanine), 4: Nε-azidobenzyloxycarbonyl lysine (PABK), 5: propargyl-L-lysine (PrK), 6: Nε-(1-methylcycloprop-2-enecarboxamido)lysine (CpK), 7: Nε-acryloyllysine (AcrK), 8: Nε-(cyclooct-2-yn-1-yloxy)carbonyl)L-lysine (CoK), 9: bicyclo[6.1.0]non-4-yn-9-ylmethanol lysine (BCNK), 10: trans 11: trans-cyclooctyl-4-ene lysine (4′-TCOK), 12: dioxo-TCO lysine (DOTCOK), 13: 3-(2-cyclobuten-1-yl) propionic acid (CbK), 14: Nε-5-norbornene-2-oxycarbonyl-L-lysine (NBOK), 15: cyclooctyne lysine (SCOK), 16: 5- Norbornene-2-ol tyrosine (NOR), 17: cyclooct-2-ynol tyrosine (COY), 18: (E)-2-(cyclooct-4-en-1-yloxy)ethanol tyrosine (DS1 / 2), 19: azidohomoalanine (AHA), 20: homopropargylglycine (HPG), 21: azidonorleucine (ANL), 22: Nε-2-nitroethoxycarbonyl-L-lysine (NEAK).
[0104] Without wishing to be bound by theory, trimerization or multimerization of binding proteins may be a useful property for improving or achieving biological effects, and new components for trimerization or multimerization have been extensively studied (Cuesta, AM et al., Trends in Biotechnology, 2010). In addition to its suitability for fusing effector proteins or binding domains to the N- and C-termini of CutA1, CutA1 provides a new component for stable multimerization of binding proteins. The data provided below, in particular Figure 22 and Figure 23 , demonstrating that CutA1 can generally be a suitable domain for coupling or fusion to a binding protein or effector protein, for example, to promote trimerization to exert a biological effect after coupling to a binding protein via a catcher domain or directly fusion to a binding protein. Therefore, the present invention also provides polypeptides comprising a binding domain or effector protein fused to the N-terminus or C-terminus of CutA1. In some embodiments, CutA1, typically human CutA1, is modified as described elsewhere herein.
[0105] In some aspects, the domain (or subunit monomer of the oligomer core) of the polypeptide construct is a TNF family protein, including a TNF protein, a TL1A protein, an OX40L protein, or a CD40L protein, typically a human protein. In some aspects, the TNF, TL1A, OX40L, or CD40L protein, typically a human protein, is genetically engineered to remove one or more cysteine residues from the native sequence. Typically, this is where one or more cysteine residues are replaced with one or more non-cysteine residues. Removal of one or more cysteine residues advantageously allows for targeted cysteine coupling at non-native sites or in fusion proteins. Without wishing to be bound by theory, unpaired cysteines often interfere with stability and downstream applications, and therefore, removal of unpaired cysteines may be advantageous. In some aspects, the TNF, TL1A, OX40L, or CD40L protein, typically a human protein, is genetically engineered to delete one or more residues from either or both ends of the native sequence to form a truncated TNF, TL1A, OX40L, or CD40L domain that is incorporated into the constructs of the invention. In some embodiments, one or more residues are deleted from the N-terminus of the native sequence, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more residues. In some embodiments, one or more residues are deleted from the C-terminus of the native sequence, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more residues. In some embodiments, one or more residues are deleted from the N-terminus of the native sequence, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more residues, and one or more residues are deleted from the C-terminus of the native sequence, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more residues.
[0106] In some aspects, as shown in the Examples, the modular assembly provided by the present invention can be combined with simple post-assembly purification (clean-up). This is particularly useful for manufacturing drug candidates for downstream analysis. In some embodiments, these methods can be automated. In some embodiments, purification uses beads, such as paramagnetic beads, which can be combined with a suitable catcher to quench uncoupled tagged protein from the assembly of the tagged binder and the catcher core (see, e.g., Figure 37 This makes purification of the assembly (by removing unbound tagged conjugates) cost-effective and rapid, and easily scalable.
[0107] Purification methods typically involve contacting the post-coupling reaction mixture with a purification binding domain that can bind to unbound binder. The purified binding domain that binds to the previously unbound binder can then be removed from the reaction mixture.
[0108] These purification methods are particularly well-suited to plate-based formats that can be automated using liquid handling robots. For example, the GST-catcher can be loaded onto glutathione-coupled beads (e.g., Figure 35 ), and unbound excess glutathione was removed by washing with PBS (e.g., Figure 36 ). GST catcher-bound beads can then be added to each coupling reaction, typically in excess relative to the corresponding uncoupled tagged conjugate (e.g., a 2-fold excess of GST catcher). The beads can then be removed by appropriate means, such as by magnetic capture particles if the beads are magnetic, leaving a supernatant containing only the fully assembled molecules (e.g., Figure 38 This approach is advantageous over covalent coupling of the catcher to the beads because it can be performed on demand at various scales. It also does not require complex chemistry or reducing conditions. Furthermore, the design of the catcher construct, the ratio of the catcher protein, or the resin can be easily varied as needed.
[0109] This purification technique is a general improvement over existing techniques and may also be applied to other modular assembly methods, notably Driscoll et al., September 2023, bioRxiv 2023.08.31.555700; doi: https: / / doi.org / 10.1101 / 2023.08.31.555700 The method described.
[0110] This purification technique can be adapted for methods such as sortase-based assembly, as described in Andres et al., Mol Cancer Ther (2020) 19(4): 1080–1088 “High-Throughput Generation of Bispecific Binding Proteins by Sortase A–Mediated Coupling for Direct Functional Screening in Cell Culture”.
[0111] The purification step typically uses a fusion protein that is capable of binding excess binder in the reaction mixture after assembly of the catcher and the tagged binder. These purification fusion proteins typically comprise a protein useful in the purification method just described above, fused to at least one binding domain as described herein, typically a catcher domain capable of binding excess binder in the reaction mixture. Proteins useful in the purification step typically form associations based on protein domains that can be used in the purification process. This association can be covalent (e.g., HaloTag, a modified haloalkane dehalogenase designed for covalent binding to a synthetic ligand comprising a chloroalkane linker attached to a variety of useful molecules, such as fluorescent dyes, affinity handles, or solid surfaces) or non-covalent (e.g., maltose binding protein [MBP]). As described above, the construct is typically a glutathione-S-transferase "GST" sequence capable of binding to GSH on beads. In some embodiments, the purification fusion protein is attached to a solid support such as beads. In some embodiments, the beads can be magnetic or paramagnetic.
[0112] GST / GSH is a particularly suitable system because paramagnetic beads with high protein capacity are commercially available at relatively low cost.
[0113] The purified fusion protein may comprise a single type of binding domain that is capable of purifying an excess of a conjugate. The single binding domain may be provided as a single copy (e.g., GST-SpyCatcher), or optionally provided as multiple copies (e.g., in the form of GST-SpyCatcher-SpyCatcher). When there are two or more conjugates to be purified after assembly in excess, two or more purified fusion proteins may be used in combination, each with a different single type of binding domain. GST-SpyCatcher and GST-DogCatcher are examples of purified fusion proteins. In some exemplary embodiments, for example, the first purified fusion protein is GST-SpyCatcher and the second purified fusion protein is GST-DogCatcher. Specific sequences are provided in the examples of GST-SpyCatcher003 and GST-DogCatcher. In other exemplary embodiments, the purified fusion protein comprises GST-SpyCatcher002 or GST-SpyCatcher-003. In other exemplary embodiments, the purified fusion protein comprises GST-SpyCatcher002 and GST-SpyCatcher-003, used in combination. Other purified fusion proteins may include inactivated variants of SpyCatcher, such as SpyCatcher002 KA, SpyDock, or SpySwitch (e.g., as described in Khairil Anuar et al., Nature Communications, Vol. 10, Article No. 1734 (2019), in particular Figure 7 ).
[0114] MBP-SpyCatcher and MBP-DogCatcher are other examples of purified fusion proteins. In some exemplary embodiments, for example, the first purified fusion protein is MBP-SpyCatcher and the second purified fusion protein is MBP-DogCatcher. In other exemplary embodiments, the purified fusion protein comprises MBP-SpyCatcher002 or MBP-SpyCatcher-003. In other exemplary embodiments, the purified fusion protein comprises MBP-SpyCatcher002 and MBP-SpyCatcher-003 used in combination.
[0115] HaloTag-SpyCatcher and HaloTag-DogCatcher are other examples of purified fusion proteins. In some exemplary embodiments, for example, the first purified fusion protein is HaloTag-SpyCatcher and the second purified fusion protein is HaloTag-DogCatcher. In other exemplary embodiments, the purified fusion protein comprises HaloTag-SpyCatcher002 or HaloTag-SpyCatcher-003. In other exemplary embodiments, the purified fusion protein comprises HaloTag-SpyCatcher002 and HaloTag-SpyCatcher-003 used in combination.
[0116] In some embodiments, two or more binding domains are connected to a single protein that can be used for the above-mentioned purification method. Multiple binding domains, typically multiple catcher domains, can be connected to any suitable position on the protein (e.g., GST, MBP or HaloTag) for purification. Multiple binding (e.g., catcher) domains can be arranged in series at one end of the purified protein, or at different ends. As in other embodiments described herein, a linker peptide (e.g., 2-30 amino acid residues) can be appropriately included between the separate components of the fusion. Some exemplary forms in which each catcher is different include catcher 1-GST-catcher 2, catcher 1-catcher 2-GST, GST-catcher 1-catcher 2, catcher 1-MBP-catcher 2, catcher 1-catcher 2-MBP, MBP-catcher 1-catcher 2, catcher 1-HaloTag-catcher 2, catcher 1-catcher 2-HaloTag or HaloTag-catcher 1-catcher 2.
[0117] In some embodiments, the present invention provides polypeptides comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a domain that can be used for the purification method just described above, and wherein the first binding domain and the second binding domain are the same or different. In these embodiments, the domains are generally capable of forming an association based on a protein domain that can be used for the purification process. This binding can be covalent (e.g., HaloTag) or non-covalent (e.g., maltose binding protein [MBP]). As described above, this construct is typically a glutathione-S-transferase "GST" sequence that can bind to GSH on the beads.
[0118] In some aspects, the present disclosure provides TNF receptor superfamily (TNFRSF) agonists, e.g., TRAIL receptor agonists, e.g., DR5 agonists, comprising CutA1 or a variant thereof as described herein. TNFRSF agonists are known in the art. Typically, the agonism is provided by one or more binding regions (e.g., Fab, ScFv, Nanobody), e.g., linked to CutA1 at the N-terminus and / or C-terminus. Agonist antibodies for TRAL receptors such as DR5 are known in the art, such as IGM8444 and TAS266 (see, for example, (i) Wang et al., ASCO 2020 Annual Meeting Abstract: “IGM-8444 as a potent agonistic Death Receptor 5 (DR5) IgM antibody: Induction of tumor cytotoxicity, combination with chemotherapy and in vitro safety profile.”); (ii) Huet et al., Cancer Res (2012) 72(8_Suppl): 3853. “Abstract 3853: TAS266, a novel tetrameric nanobody agonist targeting death receptor 5 (DR5), exhibits superior anti-tumor efficacy compared to conventional DR5 targeting approaches.” agonist targeting death receptor 5 (DR5), elicits superior antitumor efficacy than conventional DR5-targeted approaches”); and (iii) WO-A-2017 / 011837).
[0119] DR5 binders themselves may not be agonists. In general, without wishing to be bound by theory, anti-DR5 antibodies (including antigen-binding fragments and genetically engineered forms thereof, including those according to the present invention) need to have a specific valency, such as divalent, trivalent, or tetravalent, to achieve agonistic effects on DR5. The valency of the binder can be increased by binding to a multimeric scaffold (e.g., CutA1). At least the data in Examples 11 and 14 indicate that at least trivalency provides effective agonistic effects with the constructs of the present invention. This suggests that CutA1 is particularly advantageous for targeting TNFRSF, likely due to its geometry (including its trimerization properties) and small size similar to the natural ligand of the TNFRSF receptor. Figure 11 Ligands of the TNFSF family are exemplified for reference. Thus, in one embodiment, a trivalent CutA1-DR5 agonist is provided. In some embodiments, a hexavalent CutA1-DR5 agonist is provided.
[0120] The positions of the C- and N-termini of the two trimers can vary widely and one may be more preferred than the other. For example, for DR5, a CutA1 C-terminal fusion may sometimes be preferred (see, Figure 22 , Figure 23 C, Figure 32 ).
[0121] In some embodiments, multimeric CutA1-DR5 agonists are provided. In some embodiments, trimers of CutA1-DR5 agonist constructs are provided. In some embodiments, the multimers can be dimers, or can be higher order, such as tetramers, pentamers, hexamers, heptamers, or even higher order.
[0122] Typically, a CutA1-DR5 agonist construct includes a second binding domain for a different target, i.e., a bispecific construct. The second binding domain is typically provided at a different position on CutA1, most typically at an opposite end. However, as described elsewhere herein, it may also be provided in tandem with a DR5 agonist at the same position on CutA1. Thus, suitable construct formats include (described in the usual N to C orientation): binding domain 2—CutA1—DR5 binder; DR5 binder—CutA1—binding domain 2; DR5 binder—binding domain—CutA1; CutA1—DR5 binder—binding domain 2; binding domain 2—DR5 binder—CutA1; or CutA1—binding domain 2—DR5 binder. Each of these constructs is provided as described above, for example, in a multivalent (e.g., trivalent, hexavalent) and / or multimeric (e.g., trimer or hexameric) form.
[0123] TAS266 is a tetrameric nanobody agonist targeting death receptor 5 (DR5). As discussed in the examples, the inventors hypothesized that the severe liver toxicity that led to the clinical failure of TAS266 was due to excessive aggregation of DR5 dependent on anti-drug antibodies and subsequent excessive induction of apoptosis in hepatocytes (Inhibrx, WO2017-A-011837). These antibodies are present in commercially available intravenous immunoglobulin (IVIG) solutions (Inhibrx, WO2017-A-011837). Compared to TAS266, the CutA1-based form (exemplified by L11-HsCutA1-L7) does not increase IVIG-dependent toxicity in the HepG2 hepatocyte cell line (see, e.g., Figure 41 Without wishing to be bound by theory, the inventors observed that reducing the valency of DR5-targeting binders from four to three and simultaneously introducing bispecificity allowed CutA1-based constructs, exemplified by L11-HsCutA1-L7, to cause less severe side effects and show enhanced tumor specificity.
[0124] In some embodiments, a construct comprising CutA1 or a variant thereof and a DR5 binding domain is provided. Generally, the construct also comprises a second binding domain for another target. In some embodiments, the DR5 binding domain competes with the TAS266 tetramer construct for binding to DR5. In some embodiments, the DR5 binding domain comprises a "VHH" binding domain from TAS266 (e.g., SEQ ID NO: 106). In some embodiments, the construct comprises CutA1 or a variant thereof as described herein and a "VHH" binding domain from TAS266 (e.g., as shown in SEQ ID NO: 102) and optionally a second binding domain (e.g., scFv, VHH, or Fab) that binds to a different target protein. A linker sequence may be included between the anti-DR5 VHH and CutA1, and / or between CutA1 and the second binding domain. Generally, the two binding domains are located at opposite ends of CutA1, but they may be located sequentially at one end of CutA1.
[0125] SEQ ID NO: 106 provides an example of a DR5 binder that can be used to construct a multivalent DR5 agonist, such as tetravalent TAS266 or trimeric CutA1-DR5. In some embodiments, the DR5 binder competes with SEQ ID NO: 106 for DR5 binding. This competition can be determined, for example, in a surface plasmon resonance assay, such as a Biacore assay. In some embodiments, the DR5 agonist is at least 70% identical, at least 80% identical, at least 90% identical, such as at least 95% identical, or at least 99% identical to SEQ ID NO: 106. Generally, any variation is limited to framework residues, and variation in CDR residues is not permitted.
[0126] An example of an agonist hsCutA1-DR5 binder is provided in SEQ ID NO: 102. Sequences having at least 70% identity, at least 80% identity, at least 90% identity, such as at least 95% identity, or at least 99% identity to SEQ ID NO: 102 can be used. Generally, any variation is limited to framework residues, while variation in CDR residues is not permitted.
[0127] In other aspects, the present invention provides single-chain CutA1, which is a further genetically engineered variant of CutA1, as shown in Example 14. This is a genetically engineered variant of CutA1 comprising a plurality of CutA1 molecules (full-length or truncated, and optionally substituted or otherwise genetically engineered, as broadly described herein), wherein each "monomer" of CutA1 is covalently linked, optionally via a linker.
[0128] In one embodiment of this aspect, a single polypeptide chain comprises a plurality of CutA1 sequences, optionally wherein a linker polypeptide of 3 to 30 amino acids in length is located between some or all of the CutA1 sequences, optionally wherein at least some and optionally all of the CutA1 sequences are as defined elsewhere herein. In other embodiments, a single polypeptide chain comprises three human CutA1 sequences, each separated by a linker polypeptide, optionally comprising an isopeptide bond-forming domain at the N-terminus and / or C-terminus.
[0129] Typically, a single-chain CutA1 is expressed as a fusion protein comprising two or more CutA1 molecules, such as three or more CutA1 molecules, for example a fusion protein of 4, 5, 6, 7, 8 or more CutA1 molecules. Three CutA1 molecules are a common situation. Some or all of the CutA1 molecules in the chain are typically separated by a polypeptide linker, typically wherein the polypeptide linker is 3 to 30 amino acid residues in length, more typically 10 to 20 amino acid residues in length, for example 12, 13, 14, 15, 16 or 17 amino acids in length. Typically, the non-limiting linker is 15 amino acids in length, for example (GGGGS)3.
[0130] For example, in a trimeric single-chain CutA1 construct, each of the three CutA1 components can be connected by a linker to produce a construct. The construct can also include a target binding moiety, typically (although not necessarily) located at the end. An exemplary single-chain CutA1 construct is: SpyCatcher-Linker-CutA1-Linker-CutA1-Linker-CutA1-DogCatcher. This single-chain construct may be particularly useful for providing a 1+1 bivalent control with the same geometry as the 3+3 molecule. This single-chain CutA1 has a variety of potential uses, including but not limited to: as an in vitro or in vivo control scaffold compared to 3+3 CutA1; as a scaffold to evaluate 1+1 in a fixed geometry; and in vivo applications when combined with other CutA1 technologies such as ADCs (e.g., for the preparation of high drug-to-antibody ratio "DAR" ADCs).
[0131] The present invention enables highly adaptable screening of a large number of effector moieties. The effector moiety can be any protein domain, not limited to Fab / Fv regions or other antigen binding domains. The present invention can also be used to study the effects of molecules achieved through higher titer interactions, which cannot be achieved through traditional bispecific antibodies or other means. Accordingly, the present invention provides a system for high-throughput screening of bispecific molecule combinations that can achieve effects that can only be achieved through higher titer interactions. The present invention also provides novel candidate therapeutic agents that can be identified according to the methods provided herein. Such candidate therapeutic agents have the advantages of multifunctionality and high titer.
[0132] Limited applications of multivalent protein scaffolds have been described in the art. One such attempt is described in Brunet et al. (Bioconjugate Chemistry), Vol. 28, No. 5, pp. 1544-1551. This work involved the use of heptameric assemblies in which antigens were attached to opposite sides of the IMX313 heptameric core as potential malaria vaccines. However, the relative presentation of the attached antigens makes these constructs unsuitable for binding to multiple cellular receptors, limiting their application in cell-bound treatments for conditions such as cancer and autoimmune diseases.
[0133] Therefore, there is a need for new and / or improved methods for phenotypic screening of combinations of effector moieties (also referred to herein as ligands) and for developing novel therapeutic agents. Prior art methods have been unable to improve the potency of the functional combinations studied and / or have been unable to identify synergistic benefits when antigens are presented on the same side of a protein scaffold. Furthermore, there is a need for new and / or improved therapeutic agents, including the therapeutic agents of the present application that can be designed and identified according to the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0134] Figure 1 Schematic diagram of a multivalent protein scaffold as described herein, wherein the first binding site and the second binding site are located on the same face of the oligomer core, and thus on the same face of the multivalent protein scaffold.
[0135] Figure 2 Schematic diagram of a multivalent protein scaffold as described herein, wherein the first binding site and the second binding site are located on opposite sides of the oligomer core, and thus on opposite sides of the multivalent protein scaffold.
[0136] Figure 3 Schematic diagram of a multivalent protein scaffold as described herein, wherein the plurality of first binding sites and the plurality of second binding sites are located on the same face of the oligomer core, and thus on the same face of the multivalent protein scaffold for binding to a surface.
[0137] Figure 4 Schematic diagram of a multivalent protein scaffold as described herein, wherein the plurality of first binding sites and the plurality of second binding sites are located on opposite sides of the oligomer core, and thus on opposite sides of the multivalent protein scaffold, and therefore cannot engage with a surface simultaneously.
[0138] Figure 5 Schematic diagram of a multivalent protein scaffold as described herein, wherein the first binding site and the second binding site are located in a bipositive orientation on the same face of the oligomer core, and thus on the same face of the multivalent protein scaffold.
[0139] Figure 6 is a schematic diagram of a multivalent protein scaffold as described herein, wherein the first binding site and the second binding site are located in an edge-on orientation on the same face of the oligomer core, and thus on the same face of the multivalent protein scaffold.
[0140] Figure 7 Schematic diagram of a multivalent protein scaffold as described herein, wherein the first binding site and the second binding site are located on the same face of the oligomer core, and thus on the same face of the multivalent protein scaffold, in an orientation intermediate between the double positive orientation and the side positive orientation.
[0141] Figure 8 Schematic diagram of the angle (X) formed between the first binding site and the second binding site of a subunit monomer attached to the oligomeric core of a multivalent protein scaffold as described herein.
[0142] Figure 9 Schematic diagram of one embodiment of the present invention wherein a fusion of a first binding site and a second binding site is linked to an oligomeric core as described herein to generate a multivalent protein scaffold. Various geometries and stoichiometric ratios can be generated using the methods disclosed herein.
[0143] Figure 10 The cis-two-dimensional representation of the fusion site is shown. a) For a given target plane, a distance d is drawn from the target plane. c The longest cross section of the core protein is determined by means of orthogonal lines. The parallel plane passing through the midpoint of the cross section is shown in the figure. Among them, for all coupling sites, when the distance of the shortest path from the coupling site to the target plane that does not intersect the protein surface is less than the distance d c A certain percentage (such as less than d c 50% of d ), protein coupling sites (stars, circles) are considered to be preferentially cis. b) In an exemplary protein, all binding sites are cis. c) In an exemplary protein, since all secondary binding sites (circles) have just less than 50% d c The minimum path length for this threshold is such that, according to a), all binding sites are considered cis with respect to each other. d) Although, as in c), all binding sites have the same projected distance (through the protein) to the target plane, all second binding sites (circles) are not accessible from the same side of the scaffold and / or are not located on the same side of the scaffold due to the protein's geometry hindering the shortest path to them. e-f) Binding sites are too far apart to be considered cis with respect to any target plane.
[0144] Figure 11 This is a schematic diagram of the protein structure described in this application. The protein structure for a given PDB ID is visualized in animated form, with each chain tinted differently. The N- and C-termini of the individual monomers are labeled and oriented toward the binding surface. The symmetries of the protein structure, such as C3 cyclic symmetry, C4 cyclic symmetry, and D2 dihedral symmetry, are indicated in brackets. *: Protein symmetries estimated from the NMR structure. It is a heteromer with C1 symmetry, but its domains are homomers and arranged in a manner similar to C3 symmetry. By selecting a multimeric protein core with an appropriate monomer configuration, multiple binding sites can be extended from each monomer to form a single binding surface. In addition to recombinant fusion, such proteins can then be used to achieve modular assembly, for example, by recombinant fusion with SpyCatcher and SnoopCatcher or DogCatcher at the N-terminus and C-terminus (for example, the N-terminus of the monomer of the oligomeric core protein is recombinantly fused to SpyCatcher and the C-terminus is recombinantly fused to SnoopCatcher), thereby allowing peptides or proteins modified in an appropriate manner to quickly acquire multiple valencies or other properties. Examples of proteins include: (1) the highly thermostable human copper-binding protein HsCutA1, which has a C3 geometry, with the N-terminus and C-terminus of each monomer close to each other and extending into the same plane of the assembled trimer (PDB ID 2ZFH); (2) the highly thermostable homolog PhCutA1 from Pyrococcus horikoshii (PDB ID 4NYO), which is highly similar to HsCutA1 in structure; (3) the NC1 domain of collagen X (PDB ID: 1GR3); (4) the NC1 domain of collagen VIII (PDB ID: 1O91); (5) macrophage migration inhibitory factor 2 (PDB ID: 7MSE); (6) tumor necrosis factor (PDB ID: 1TNF); and (7) the TNF-like protein TL1A (PDB ID: 2RE9).
[0145] Figure 12Shown is a cis-oriented multimeric protein complex that is readily expressed and readily prepared by standard protein purification methods. a) Ni-NTA purification of H6-SpC-PhCutA1-SnC (SEQ ID NO: 21). H6-SpC-PhCutA1-SnC is readily expressed in E. coli BL21(DE3) and retains the intact trimeric structure characteristic of the highly stable PhCutA1 even after boiling in SDS-PAGE loading buffer. The first and second washes used 10 column volumes of equilibration buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 10 mM imidazole). The third and fourth washes used 10 column volumes of wash buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 30 mM imidazole). All elution steps used 2 column volumes of elution buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 200 mM imidazole). Samples were analyzed on 12% SDS-PAGE gels and stained with Coomassie. P – lysate pellet; CL – clarified lysate; FT – flow-through; W – wash; E – elution. b-c) Size exclusion chromatography results for H6-SpC-PhCutA1-SnC using a HiLoad 16 / 600 Superdex 200 pg column. b) UV A280 absorbance chromatogram highlighting the H6-SpC-PhCutA1-SnC peak. c) SDS-PAGE of 2 mL fractions of H6-SpC-PhCutA1-SnC obtained from the highlighted region in the chromatogram.
[0146] Figure 13 Shown is the transition of a circularly symmetric trimeric core protein from a dihedral hexamer to a cis-oriented presentation. a) A homohexameric antiparallel coiled-coil structure with the N- and C-termini on opposite sides of the protein assembly (PDB ID: 5W0J). b) Heteromeric assemblies (PDB ID: 5VTE) can be obtained by inducing point mutations in homomeric assemblies (e.g., by introducing or modifying salt bridges to "lock" the assembly into one orientation). c) Homomeric assemblies adapted for presentation in the cis-oriented presentation can be obtained by ligating the ends of heteromeric assemblies (compared to HIV GP41 (PDB ID: 1I5Y)). The direction of the transition from black to white (structure, schematic) and the direction of the arrow (schematic) indicate the orientation from the N- to C-terminus. The structures were visualized using PyMOL.
[0147] Figure 14Shown is the highly stable trimeric protein SpC-PhCutA1-SnC. a) SpC-PhCutA1-SnC samples were heated in PBS with 0%, 0.5%, or 1% SDS at 97°C for 2 hours before the addition of SDS loading dye. Samples were analyzed by 12% SDS-PAGE and Coomassie staining. Lane 1 shows a control sample that was not heated and did not contain SDS. As the SDS concentration increased, the trimer partially monomerized, confirming that the trimer was not covalently cross-linked. b) SpC-PhCutA1-SnC after extended storage. Protein aliquots were stored at 4°C, room temperature (21°C), or 37°C for 7 days. Samples were then prepared with SDS loading buffer and analyzed by SDS-PAGE. Compared to storage at 4°C, the protein showed little degradation when stored at 21°C to 37°C. Optionally, the protease inhibitor PMSF can be added, which has a similar effect.
[0148] Figure 15 Shown are the preparation of SpyTag- and SnoopTag-bearing ligand components for modular assembly with platform proteins. a) SnT-L1 was purified from a 200 mL BL21(DE3) culture by Ni-NTA chromatography. Each wash step used 5 mL of Ni-NTA wash buffer. Each elution step used 2 mL of Ni-NTA elution buffer. Samples were analyzed on a 12% SDS-PAGE gel and Coomassie-stained. Pel. – lysate pellet; FT – flow-through; W – wash; E – elution; L – molecular weight marker. b) Purification of L2-SpT was performed in the same manner as in a). c) After Ni-NTA purification, SnT-L1 was analyzed by size-exclusion chromatography. The UV A280 absorbance chromatogram of SnT-L1 was obtained using an AKTA Pure 25 equipped with a HiLoad Superdex 16 / 600 75 pg column. Inset: SDS-PAGE results of 2 mL fractions of SnT-L1 obtained from the highlighted region in the chromatogram. d) Size exclusion chromatography analysis of L2-SpT performed in the same manner as in c).
[0149] Figure 16Shown is the effect of SpC-PhCutA1-SnC on promoting the stable trimerization of SpyTag- and SnoopTag-bearing proteins. a) H6-SpC-PhCutA1-SnC was conjugated with excess SnT-L1 or L2-SpT at a 1:2:2 molar ratio for the indicated times. The samples were then supplemented with SDS loading buffer and denatured by boiling at 95°C for 5 minutes. The samples were analyzed on 8% and 16% SDS-PAGE gels and Coomassie-stained. The amount of SpC-PhCutA1-SnC conjugated to SnT-L1 or L2-SpT was observed over time, as measured by the consumption of the ligand moiety. b) SpC-PC-SnC was conjugated to SnT-L1 and L2-SpT in the same manner as in a). c) Coupling of SpC-PhCutA1-SnC with SnT-L1 and L2-SpT. SpC-PhCutA1-SnC was incubated with excess SnT-L1 and / or L2-SpT at a molar ratio of 1:2:2 at 25°C for 64 hours. Samples were analyzed on 16% SDS-PAGE gels and Coomassie-stained. Notably, SpC-PhCutA1-SnC was fully coupled to SnT-L1 / L2-SpT and retained the high thermal stability characteristic of PhCutA1. Coupling of SpC-PC-SnC with SnT-L1 and L2-SpT was performed in the same manner as in c). The coupling of SpC-PC-SnC with SnT-L1 / L2-SpT was continued until the reaction was complete.
[0150] Figure 17Shown are high-molecular-weight scaffolds that allow post-assembly purification by dialysis. a) SpC-PhCutA1-SnC after Ni-NTA purification. b) The SpC-PhCutA1-SnC:SnT-L1:L2-SpT assembly was dialyzed using a 96-well dialysis plate. a) SpC-PhCutA1-SnC:SnT-L1:L2-SpT assembly was purified by Ni-NTA chromatography, and the elution fractions were pooled and concentrated. b) Protein coupling of SpC-PhCutA1-SnC, SnT-L1, and L2-SpT was performed at 25°C for 2 hours. Dialysis was performed using a 96-well high-throughput dialysis plate equipped with a 100 kDa molecular weight cutoff (MWCO) membrane at a 1:1 sample to dialysis buffer ratio. Dialysis was performed at room temperature with orbital shaking. The PBS dialysate was changed every 30 minutes for the first 90 minutes. Dialysis for 24 hours demonstrated the removal of protein impurities (a-b) and unconjugated ligand (b), and the achievement of equilibrium (sample to dialysis buffer ratio of 1:1). Samples were boiled in reducing SDS loading dye, analyzed by 12% SDS-PAGE, and Coomassie-stained. c) Alternatively, the SpC-PhCutA1-SnC:SnT-L1:L2-SpT conjugate was purified by dialysis in a 12-well plate format using a sample to dialysis buffer ratio of 1:30. SpC-PhCutA1-SnC, SnT-L1, and L2-SpT were subjected to a 1:1 conjugation at 25°C for 24 hours. Dialysis was performed for 16 hours at room temperature using a 12-well high-throughput dialysis plate equipped with a 100 kDa MWCO cellulose membrane without stirring. A 100 μL sample volume and a 3 mL PBS dialysate volume, both containing 1× PMSF, were used. Samples and dialysate were collected at the following time points: 2 hours, 4 hours, 8 hours, and 16 hours. Samples were analyzed on 14% SDS-PAGE gels and Coomassie-stained. S = dialyzed sample; D = dialysate (PBS).
[0151] Figure 18Shown are various variations in core components (PhCutA1 to MIF2m or HsCutA1), protein components for conjugation (SpC / SnC to SpC3 / DgC), and variable linker lengths (GGGGSGGGGSGGGGS for MIF2m and GGGGS for HsCutA), highlighting the potential for rapid prototyping. a-b) Samples obtained by Ni-NTA purification of H6-SpC3-HsCutA1-DgC or H6-SpC3-MIF2m-DgC. TL – total lysate; P – lysate pellet; CL – clarified lysate; FT – flow-through; W – wash; E – elution. Samples were analyzed on SDS-PAGE gels and Coomassie-stained. c) Both SpC3-MIF2m-DgC and SpC3-HsCutA1-DgC can be rapidly conjugated to DogTag- or SpyTag-bearing proteins. The platform proteins were incubated at a molar ratio of 1:1.5:1.5 at 25°C for 16 hours. d) Incubation of H6-SpC3-HsCutA1-DgC with 0.1% glutaraldehyde demonstrated cross-linking of the trimeric proteins in solution. 10 μM H6-SpC3-HsCutA1-DgC was cross-linked with 0.1% glutaraldehyde for 0-20 minutes. The reaction was incubated at 37°C and terminated by the addition of 100 mM Tris (pH 8.8). Trimeric cross-links formed rapidly, and with extended incubation times, monomeric H6-SpC3-HsCutA1-DgC was consumed, and a cross-link with the molecular weight expected for cross-linking between two trimers was formed. e) SpC3-HsCutA1-DgC was reacted with L1-SpT and L3-DgT at a molar ratio of 1:2:2 at 25°C for 16 hours. Subsequently, each sample of pure SpC3-HsCutA1-DgC, L1-SpT:SpC3-HsCutA1-DgC, SpC3-HsCutA1-DgC:L3-DgT, and L1-SpT:SpC3-HsCutA1-DgC:L3-DgT was subjected to size exclusion chromatography using a Superose 6Increase 5 / 150GL column. The analysis results showed that the hydrodynamic radius of each protein scaffold or protein assembly increased. The peak fractions taken from the peak region indicated by the shaded portion of each chromatogram were injected into an SDS-PAGE gel to show only the results of the scaffold and assembled proteins after the excess ligand was removed. These samples were used as Figure 20 c's input.
[0152] Figure 19Shown is how Alphafold version 2.0 was used to predict the cis-orientation of fusion proteins. The top-scoring model for each SpC3-scaffold-DgC assembly was visualized using PyMOL. Single chains highlight SpyCatcher3 (medium gray), the scaffold (dark gray), and the DogCatcher (light gray). Other chains are animated as white, transparent surfaces. In all simulations, a GSGS linker was placed between the catcher and the scaffold. In the case of collagen XV NC1, because the GSGS linker caused the structure prediction calculations to crash, the prediction process was repeated using a (GGGGS)2 linker (its structure is shown here). Monomer sequences used in Alphafold predictions: TL1A (SEQ ID NO:32); Col XV NC1 (SEQ ID NO:33); MIF2m (SEQ ID NO:34); HsCutA1 (SEQ ID NO:54); PhCutA1 (SEQ ID NO:55); Col X NC1 (SEQ ID NO:56); TNF (SEQ ID NO:57).
[0153] Figure 20The scaffolds are suitable for in vitro assays. a) H6-SpC-PhCutA1-SnC was conjugated to two different ligands, SnT-L1 and L2-SpT. Serum-starved NCI-N87 cells were cultured for 7 days in the presence of relevant growth factors and the dual-conjugated assemblies (H6-SpC-PhCutA1-SnC:SnT-L1:L2-SpT), single-conjugated assemblies (H6-SpC-PhCutA1-SnC:SnT-L1, H6-SpC-PhCutA1-SnC:L2-SpT), and simple ligand controls (SnT-L1, L2-SpT, SnT-L1+L2-SpT). Cell viability was measured using the MTT assay. Control antibodies against L1, L2, and L1+L2 at 10 nM were used as controls (data not shown), and the results obtained were similar to those of the single- and double-conjugated assembly samples. Error bars represent the error from three technical replicates (n=3). b) The fully assembled scaffold H6-SpC-PhCutA1-SnC:SnT-L1:L2-SpT inhibits Akt and Erk1 / 2 activation. Results from 1-hour treatment of NCI-N87 cells with the scaffold alone (S; H6-SpC-PhCutA1-SnC), the ligands alone (L1; SnT-L1 and L2; L2-SpT), the single-conjugated assemblies (SxL1, SxL2), and the double-conjugated assembly (SxL1xL2) showed that the fully assembled H6-SpC-PhCutA1-SnC:SnT-L1:L2-SpT inhibited activation of the downstream Akt / ERK signaling pathway. c) H6-SpC-HsCutA1-SnC was conjugated to two different ligands, SnT-L1 and L3-SpT. Serum-starved NCI-N87 cells were treated for 2 days with the dual-conjugated assemblies (H6-SpC-HsCutA1-SnC:SnT-L1:L3-SpT), single-conjugated assemblies (H6-SpC-HsCutA1-SnC:SnT-L1, H6-SpC-HsCutA1-SnC:L3-SpT), and ligand-only controls (SnT-L1, L3-SpT, SnT-L1+L3-SpT). Cell viability was measured using the MTT assay. Error bars represent the error from three technical replicates (n=3).
[0154] Figure 21Shown are the purifications of L1-PhCutA1-L2, a direct fusion multi-domain polypeptide, by both Ni-NTA and size exclusion chromatography. a) L1-PhCutA1-L2 was readily expressed in E. coli BL21(DE3) and purified by Ni-NTA chromatography using HisPur resin (ThermoFisher). The first and second washes used 10 column volumes of equilibration buffer (50 mM Tris, pH = 7.8; 300 mM NaCl; 10 mM imidazole). The third and fourth washes used 2 column volumes of wash buffer (50 mM Tris, pH = 7.8; 300 mM NaCl; 30 mM imidazole). All elution steps used 2 column volumes of elution buffer (50 mM Tris, pH = 7.8; 300 mM NaCl; 200 mM imidazole). Samples were analyzed on a 12% SDS-PAGE gel and Coomassie-stained. P – lysate pellet; CL – clarified lysate; FT – flow-through; W – wash; E – elution. b) Size-exclusion chromatography of L1-PhCutA1-L2 using a HiLoad 16 / 600 Superdex 200 pg column. b) UV A280 absorbance chromatogram highlighting the L1-PhCutA1-L2 peak. Inset: SDS-PAGE of 2 mL fractions of L1-PhCutA1-L2 obtained from the highlighted region in the chromatogram.
[0155] Figure 22 Shown are the proteins composed of various ligands (L1, L2, L4, L5, L6, L7) and SpC3-HsCutA1 44-179 Cell survival of two cancer cell lines treated with an assembly consisting of SpyTag-DgC (SEQ ID NO: 24). Cell viability of Colo 205 and NCI-N87 cells is presented in the heat map. Cells were treated with a single assembly (either SpyTag-ligand or DogTag-ligand present only on the scaffold) and a fully coupled assembly. Colo 205 and NCI-N87 were treated with doses ranging from 0.16 to 100 nM and 4 to 100 nM, respectively. The data shown represent cell survival at the highest dose. The figure below shows the relative cell survival ratio.
[0156] Figure 23Shown are direct fusions of HsCutA1 (without the modular protein-binding catcher). a) Structural models comparing the lengths of SpyCatcher003 (SpC3) and SpyTag003 (SpT3), as well as DogCatcher (DgC) and DogTag (DgT), are also compared with alternative linkers, such as (G4S)n, where n = 1, 2, 3, (EAAAK)n, where n = 1, 2, 3, and the ribosomal L9 linker, to mediate the connection between the polypeptide (protein binder) and the multivalent scaffold, replacing the modular binding site (SpyCatcher003 and DogCatcher). When the protein binder is fused to the N- or C-terminus of the tag, the length between the corresponding tag and catcher is shown. b) SpC3-(G4S)-HsCutA1-(G4S)-DgC (SEQ ID NO: 24) assembled with the indicated protein pairs (e.g., L5 / L4, where L5 carries SpyTag and L4 carries DogTag). Compared to the cytotoxicity of direct fusions, the modular display platform of direct fusions was replaced as follows: L5 / L4 to L5-GGGGS-[C75A, C96A]HsCutA1 67-171 -GGGGS-L4, L1 / L2 to L1-(GGGGS)3-[C75A,C96A]HsCutA1 67-171 -(GGGGS)3-L2, L2 / L4 to L2-(GGGGS)4-HsCutA1 44-179 -(GGGGS)4-L4. Cell survival of NCI-N87 cells was determined after 7 days of treatment with 20 nM of the assembly and direct fusion. c) SpC3-[C75V,C96S]HsCutA1 44-179 -DgC (SEQ ID NO: 73) assembled with L7 at SpC3 or DgC to form trivalent monospecific molecules at the N-terminus or C-terminus, and compared with their direct fusion equivalents. The direct fusion equivalent is L7-(GGGGS)5-[C75V,C96S]HsCutA1 44-179 and [C75V,C96S]HsCutA1 44-179 -(GGGGS)8-L7. The fraction of surviving Colo205 cells was determined 48 hours after the start of treatment at the indicated doses. The same assay was performed for d) and e). d) SpC3-[C75V,C96S]HsCutA1 44-179 -DgC (SEQ ID NO: 73) and SpC3-[C75V, C96S]HsCutA1 60-171-DgC (SEQ ID NO: 76) assembled with L7 at SpC3 or DgC to form a hexavalent monospecific molecule, and compared with their direct fusion equivalents. The direct fusion equivalent is L7-(GGGGS)5-[C75V,C96S]HsCutA1 44-179 -(GGGGS)8-L7 and L7-(GGGGS)5-[C75V,C96S]HsCutA1 60-171 e) Levels of apoptosis induced by direct fusions of several different cores fused to L7 at the N- and C-termini in the form of L7-(GGGGS)5-[core]-(GGGGS)8-L7, where [core] is [C75V,C96S]HsCutA1. 44-179 (SEQ ID NO:68), [C75V,C96S]HsCutA1 60-171 (SEQ ID NO: 70), OX40L (SEQ ID NO: 78), CD40L (SEQ ID NO: 79), and full-length soluble TNF (SEQ ID NO: 80).
[0157] Figure 24 Shown are terminal truncations of HsCutA1. a) Schematic representation of the coding sequence of differentially truncated HsCutA1 fused to SpyCatcher003 (SpC3) and DogCatcher (DgC) via example linkers L1 and L2. The terminal truncations are designated 44-179 (SEQ ID NO: 29) and 60-171 (SEQ ID NO: 66), respectively. Sequences were compared in a multiple sequence alignment of WT HsCutA1 (33-179) and HsCutA1 truncated to 44-179 and 60-171. b) Animation of the truncation sites on HsCutA1. Residues 44-59 are shown at the N-terminus and residues 172-179 are shown at the C-terminus. Animation of the modeling by PDB ID: 2ZFH. c) HsCutA1 44-179 HsCutA1 60-171 Comparison of expression of truncation proteins. E. coli cell lysates were compared to show any changes in protein expression caused by truncation. d) SpC3-(GGGGS)3-[C75V,C96S]HsCutA1 44-179 -(GGGGS)3-DgC (SEQ ID NO:73) and SpC3-(GGGGS)3-[C75V,C96S]HsCutA1 60-171-(GGGGS)3-DgC (SEQ ID NO: 76) variably assembles with L7 at SpC3, DgC, or both, forming trivalent or hexavalent cytotoxic monospecific molecules. The fraction of surviving Colo 205 cells was determined 48 hours after initiation of treatment at the indicated doses.
[0158] Figure 25 Shown are the removal of cysteine from HsCutA1. a) Animation of the naturally occurring cysteine sites in HsCutA1 for removal. Animation modeled by PDB ID: 2ZFH. b) Table of the frequency of amino acids at sites similar to C75 and C96 of HsCutA1 after a multiple sequence alignment (MSA) of 406 homologs. Amino acids with a frequency of zero are not shown. c) HsCutA1 with native WT cysteines (C75, C96), cysteines mutated to alanine (C75A, C96A), and the mutations shown by MSA (C75V, C96S). 44-179 HsCutA1 60-171 Comparison of expression of E. coli cell lysates was performed to show any changes in protein expression caused by mutagenesis. d) SpC3-GGGGS-HsCutA1 44-179 -GGGGS-DgC(SEQ ID NO:24),SpC3-(GGGGS)3--[C75V,C96S]HsCutA1 44-179 -(GGGGS)3-DgC(SEQ ID NO:73), SpC3-(GGGGS)3-HsCutA1 60-171 -(GGGGS)3-DgC(SEQ ID NO:74), SpC3-(GGGGS)3--[C75A,C96A]HsCutA1 60-171 -(GGGGS)3-DgC (SEQ ID NO:75), and SpC3-(GGGGS)3--[C75V,C96S]HsCutA1 60-171 -(GGGGS)3-DgC (SEQ ID NO: 76) assembles with L7 at both SpC3 and DgC, forming a monospecific molecule that induces cytotoxicity. The fraction of surviving Colo 205 cells was determined 48 hours after initiation of treatment at the indicated doses.
[0159] Figure 26Shown are typical thiol-maleimide reactions with monomeric proteins and various linker-payload structures. A) The thiol group of the monomeric protein reacts with the maleimide moiety of the payload-linker compound to produce a sulfosuccinimide payload-linker protein adduct. B) Examples of linker-payload compounds conjugated to the catcher core protein: 1-Fluorescein-5-maleimide; 2-Deruxtecan; 3-Maleimidocaproyl-Val-Cit-PAB-DM1; 4-Maleimidocaproyl-Val-Cit-PAB-MMAF.
[0160] Figure 27 Shown are cartoon images of the CutA1 structure (modeled by PDB ID: 2ZFH) and the locations of Cys mutations on a model of a fluorescein-5-sulfosuccinimide-CutA1 conjugate. A) Spatial location of the Cys mutations on one monomer of the CutA1 structure. Cys thiols are shown as spheres. B) Model of the fluorescein-5-sulfosuccinimide conjugate on the K82C mutant of CutA1 (SEQ ID NO: 86). All images were generated in PyMOL.
[0161] Figure 28 Shown are deconvoluted ESI mass spectra of all Cys mutant CutA1 catcher core variants (SEQ ID NOs:95, 93, 96, 92, 94, and 97), wild-type (SEQ ID NO:24), and cysteine-free (SEQ ID NO:91) controls coupled to fluorescein-5-maleimide, as well as unconjugated CC042 and fluorescein-5-sulfosuccinimide-CC042 conjugates. A) Dye / protein ratios for all mutants were calculated using the method for binding samples after PD-10 purification. B) 12% SDS-PAGE gel of unconjugated and conjugated CutA1 catcher core mutants. Little to no conjugation was observed for the CC041 negative control, while all other mutants were conjugated to fluorescein-5-maleimide. CC denotes the catcher core protein gel band. C) Deconvoluted ESI mass spectrum of the unconjugated catcher core: the major peak is within one mass unit of the expected molecular weight of the CC042 monomer. D) Mass spectrum of fluorescein-5-sulfosuccinimide-CC042 conjugate: the main peak is approximately 18 Da larger than the expected molecular weight of the monomer, probably due to hydrolysis of the sulfosuccinimide ring after conjugation.
[0162] Figure 29Shown is SDS-PAGE (4-20% gradient) of fluorescein-5-maleimide coupling and catcher / tag assembly of CC042 (SEQ ID NO: 93) and CC041 cysteine-free negative control (SEQ ID NO: 91). Only CC042 was effectively coupled to fluorescein-5-maleimide. For CC041, a low level of background signal indicating coupling was observed, indicating that maleimide coupling was primarily specific for Cys residues. The coupled CC042 showed a significant migration shift in the gel when reacted with a DogTag binder (L7), a SpyTag binder (L8), or both binders simultaneously. For the CC041 control, a similar band distribution was observed, indicating that maleimide coupling did not impair assembly efficiency.
[0163] Figure 30 Shown are specific membrane binding and internalization of RTK conjugates, single assemblies, and complete assemblies. HsCutA1 refers to CC7, while HsCutA1* refers to CC042 conjugated to fluorescein-5-maleimide. A) Complete assembly of His-tagged conjugates L1 and L2 with His-tagged CC7 (SEQ ID NO:76) demonstrates membrane binding and internalization of all compounds when stained with an anti-His antibody. (Note the differential ratio of the His-tag.) B) CutA1 staining of the complete assembly of CC7 with conjugates L1 and L2 confirms membrane binding and internalization. C) Fluorescein-5-maleimide (F-5-M)-conjugated CC042 (SEQ ID NO:93) alone, as a single assembly, and as a complete assembly demonstrates RTK-specific internalization of all assemblies, with no off-target binding by HsCutA1*.
[0164] Figure 31 Shown are the structures of cytotoxic drug-linker compounds and the corresponding deconvoluted ESI mass spectra of CC042 (SEQ ID: 93) conjugated to these compounds. For reference, the mass spectrum of unconjugated CC042 can be found at Figure 28 C. A) Derutecan: The major peak of the conjugate is within one mass unit of the expected molecular weight. B) Maleimidocaproyl-Val-Cit-PAB-DM1: The major peak of the conjugate is within one mass unit of the expected molecular weight. C) Maleimidocaproyl-Val-Cit-PAB-MMAF: The major peak of the conjugate is within one mass unit of the expected molecular weight.
[0165] Figure 32Shown are cell survival of two cancer cell lines after treatment with assemblies consisting of SpyTag ligand L7 and / or DogTag ligand L7 and SpC3-(GGGGS)3--[C75A,C96A]-HsCutA 160-171 -(GGGGS)3-DgC (SEQ ID NO: 75), SpC3-(GGGGS)--scHsCutA1-(GGGGS)-DgC (SEQ ID NO: 99), or SpC3-(GGGGS)3-L9-(GGGGS)3-DgC (SEQ ID NO: 98). Colo 205 and HCT116 cells were treated with doses ranging from 0.0008 to 2 nM. The data shown are normalized to mock-treated controls.
[0166] Figure 33 Shown is cell survival of the cancer cell line Colo 205 after treatment with an assembly consisting of the DogTag ligand L7 and SpyTag ligand L8 or -L10 in scFv and Fab form, and SpC3-GGGGS--[C75A,C96A,K82C]-HsCutA 44-179 -GGGGS-DgC (SEQ ID NO: 93). The hexavalent Nanobody-based L7 assembly was used as a control. Colo 205 cells were treated with doses ranging from 0.0008 to 2 nM. The data shown were normalized to the mock-treated control.
[0167] Figure 34 Shown is a [C75A, C96A]HsCutA-based assay using a conjugate containing L7 and L11. 160-171 [C75A,C96A]HsCutA 160-171 Cell survival of the cancer cell line HCT116 after treatment with a multivalent bispecific fusion of [C75A, C96A]HsCutA1 or a failed clinical candidate tetrameric L7 nanobody fusion. 60-171 Served as a control. HCT116 cells were treated with doses ranging from 0.000256 to 0.8 nM. The data shown were normalized to the mock-treated control.
[0168] Figure 35Shown is a schematic diagram of the coupling of SpyCatcher003 and DogCatcher to magnetic beads for use in the assembly purification pipeline. Glutathione (GSH)-coupled magnetic beads were mixed with GST-SpC3 and GST-DgC at equimolar concentrations. The binding of the GST-catchers to the beads was incubated at 25°C for 1 hour. The beads were then pelleted using a magnetic block and washed three times with PBS, or until a baseline absorbance reading at 280 nm was achieved. The GST-bound beads containing both catchers were retained for use in the assembly purification process.
[0169] Figure 36 Shown is an SDS-PAGE (4-20% gradient) of glutathione-coupled beads prepared with GST-SpC3 and GST-DgC. MagneGST resin was prepared with SDS-loading buffer to check for the presence of contaminating proteins in the stock solution. W1-3: PBS wash.
[0170] Figure 37 Shown is a schematic diagram of another variation of the catcher-based protein assembly and purification. SpC3-core-DgC was combined with SpT-protein and DgT-protein, with the tagged protein in 1.4-fold excess relative to the catcher core. The coupling reaction was incubated at 25°C for 1 hour to allow complete capture of the tagged protein. Figure 35 Preformed paramagnetic glutathione-coupled beads were bound to GST-SpC3 and GST-DgC as described in , so that each GST catcher was in 2-fold excess to its corresponding tagged protein. The sample was transferred to a magnetic block to capture the bead-GST-catcher-tag conjugates, and the supernatant containing the core-catcher-tag conjugates was retained for further analysis and downstream assays.
[0171] Figure 38 Shown is SDS-PAGE (4-20% gradient) of protein assembly and MagneGST purification. S p T :CC7:L2 D g T Protein assembly (L7 S pT:CC7:L2 D g T (Front)) shows two excess conjugates, while purified L7 SpT :CC7:L2 DgT Protein assembly (L7 SpT :CC7:L2 DgT (Back) Only fully assembled molecules are shown. GST catcher + resin demonstrates that excess binder has been captured during the MagneGST purification process. (CC7 corresponds to SEQ ID NO: 76)
[0172] Figure 39 Figure 2. Design and production of single-chain SpC3-HsCutA1-DgC. (A) Schematic diagram of SpC3-HsCutA1-DgC as a linear construct and tertiary arrangement. The predicted structure of the single-chain HsCutA1 core generated by AlphaFold2 is also shown. (B) UV A280 absorbance of single-chain SpC3-HsCutA1-DgC by size exclusion chromatography using HiLoad 16 / 600 Superdex 200pg is shown. 2 mL fractions from the prominent peak were separated by SDS-PAGE, and fractions B6-B11 were retained as purified single-chain SpC3-HsCutA1-DgC.
[0173] Figure 40 Shown are the apoptosis induced by trivalent and hexavalent L7 assemblies, measured as activation of caspase 3 / 7. 44-179 -DgC (SEQ ID NO: 24) coupling produces trivalent and hexavalent assembled molecules. Colo205 cells were treated with the assemblies, and caspase 3 / 7 activity was measured 1 hour (A) and 24 hours (B) after treatment. After 1 hour, only the hexavalent molecule (but not the trivalent or a combination of two trivalent molecules) was able to induce apoptosis. After 24 hours, caspase 3 / 7 activity was observed after treatment with the trivalent assemblies.
[0174] Figure 41 Shown is cell survival of the liver cancer cell line HepG2 treated with the L11-HsCutA1-L7 direct fusion, the failed clinical asset TAS266, and HsCutA1 in the presence or absence of intravenous immunoglobulin (IGIV 0.2-10 mg / ml). Untreated cells grown in the absence of IVIG are considered 100% viable. In contrast to TAS266, L11-HsCutA1-L7 did not induce ADA-associated increased cytotoxicity. DETAILED DESCRIPTION
[0175] The present invention will be described below with reference to specific embodiments and with reference to certain drawings, but the invention is not limited thereto but is limited only by the claims. Any reference signs in the claims should not be construed as limiting the scope of the invention. Of course, it should be understood that any one embodiment of the present invention need not achieve all aspects and advantages. Thus, for example, it should be understood by those skilled in the art that the present invention can be implemented or realized in a manner that achieves or optimizes one or a group of advantages described in this application without realizing other aspects or advantages that may be described or suggested in this application.
[0176] Furthermore, throughout this specification and the appended claims, unless the context clearly indicates otherwise, references to "a scaffold" encompass both "a single scaffold" and "a scaffold" encompass "two or more scaffolds"; references to "an oligomer" encompass "two or more such oligomers"; and so on.
[0177] All publications, patents, and patent applications cited throughout this application are hereby incorporated by reference in their entirety.
[0178] definition
[0179] The following words or definitions are only used to facilitate understanding of the present invention. Unless expressly stated in this application, the meaning of all words used in this application is the same as that understood by those skilled in the art to which the present invention belongs. For definitions and terms in the art, practitioners may refer in particular to: Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th edition (Cold Spring Harbor Press, Plainsview, New York, 2012); and Ausubel et al., Current Protocols in Molecular Biology, 114th Supplement, (John Wiley & Sons, New York, 2016). The scope of the various definitions in this application should not be interpreted in a manner narrower than that understood by those of ordinary skill in the art.
[0180] As used herein, "about" when referring to measurable values such as quantities and durations encompasses values that deviate from the specified value. As long as such deviations are suitable for performing the disclosed methods, "about" encompasses values that deviate from the specified value by ±20% or ±10%, more preferably by ±5%, even more preferably by ±1%, and even more preferably by ±0.1%.
[0181] As used in this disclosure, the term "amino acid" has the broadest meaning it can encompass and is intended to include organic compounds containing amino (NH2) and carboxyl (COOH) functional groups and side chains (e.g., R groups) that are unique to each amino acid. In some embodiments, an amino acid refers to a naturally occurring L-α-amino acid or residue. Commonly used single-letter and three-letter abbreviations for naturally occurring amino acids are used in this application: A = Ala (alanine); C = Cys (cysteine); D = Asp (aspartic acid); E = Glu (glutamic acid); F = Phe (phenylalanine); G = Gly (glycine); H = His (histidine); I = Ile (isoleucine); K = Lys (lysine); L = Leu (leucine); M = Met (methionine); N = Asn (aspartic acid). =Amide); P = Pro (proline); Q = Gln (glutamine); R = Arg (arginine); S = Ser (serine); T = Thr (threonine); V = Val (valine); W = Trp (tryptophan); Y = Tyr (tyrosine) (Lehninger, A.L., 1975, Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes chemically modified amino acids such as D-amino acids, retro-inverso amino acids, amino acid analogs, natural amino acids that are not normally found in proteins, such as norleucine, and chemically synthesized compounds known in the art to have unique properties of amino acids, such as β-amino acids. For example, analogs or mimetics of phenylalanine or proline that can impart the same conformational constraints to peptide compounds as naturally occurring phenylalanine or proline also fall within the definition of an amino acid. In this application, such analogs and mimetics are referred to as "functional equivalents" of the corresponding amino acids. Other examples of amino acids are found in Roberts and Vellaccio, Peptides: Analysis, Synthesis, Biology (Gross and Meiehofer ed., vol. 5, p. 341, Academic Press, Inc., New York, 1983), which is incorporated herein by reference.
[0182] The terms "polypeptide" and "peptide" are used interchangeably in this application to refer to polymers of amino acid residues and variants and synthetic analogs thereof. Thus, these terms apply not only to amino acid polymers in which one or more amino acid residues are synthetic non-natural amino acids (e.g., chemical analogs of the corresponding natural amino acids), but also to natural amino acid polymers occurring in nature. Polypeptides may be subjected to maturation or post-translational modification treatments, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. Peptides can be made by recombinant techniques (e.g., by expressing recombinant or synthetic polynucleotides). Peptides made recombinantly are generally substantially free of culture medium. For example, culture medium accounts for less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of the protein preparation.
[0183] The term "protein" refers to a folded polypeptide with a secondary or tertiary structure. Proteins can consist of a single polypeptide or multiple polypeptides assembled into a multimer. Multimers can be homo-oligomers or hetero-oligomers. Proteins can be naturally occurring proteins (also called wild-type proteins) or modified proteins (also called non-naturally occurring proteins). Non-naturally occurring proteins differ from wild-type proteins by the addition, substitution, or deletion of one or more amino acids.
[0184] "Variants" of proteins include peptides, oligopeptides, polypeptides, proteins and enzymes that have amino acid substitutions, deletions and / or insertions relative to the corresponding unmodified (i.e., wild-type) protein from which they are derived and that have similar biological and functional activities. The term "amino acid identity" as used in this application refers to the degree of sequence identity when an amino acid comparison is performed amino acid by amino acid within a comparison window. Therefore, the "percentage of sequence identity" is calculated as follows: two sequences that are maximally aligned within the comparison window are compared: the number of positions in the two sequences having the same amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, Met) is determined to obtain the number of matched positions; the number of matched positions is divided by the total number of positions in the comparison window (i.e., the window size); the result is multiplied by 100 to obtain the percentage of sequence identity.
[0185] In all aspects and embodiments of the invention, a "variant" generally has at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98% or 99% overall sequence identity to the corresponding wild-type protein amino acid sequence. Sequence identity can also apply to fragments or portions of a full-length polynucleotide or polypeptide. Thus, while a sequence may have only 50% overall sequence identity to a full-length reference sequence, a particular region, domain or subunit may have 80% or 90% sequence identity, or even as high as 99% sequence identity to the reference sequence.
[0186] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is designated as the "normal" form or "wild-type" simply because it is the most common form among all genes. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence modifications (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., altered properties) compared to the wild-type gene or gene product. It is important to note that naturally occurring mutants can be identified and isolated based on the altered properties compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring natural amino acids are well known in the art. For example, methionine (M) can be substituted for arginine (R) by replacing the methionine (ATG) codon at the relevant position in the polynucleotide encoding the mutant monomer with an arginine (CGT) codon. Methods for introducing or substituting unnatural amino acids are also well known in the art. For example, introduction of unnatural amino acids can be achieved by incorporating synthetic aminoacyl-tRNAs into the IVTT system used to express the mutant monomer. Alternatively, non-natural amino acids can be introduced by expressing mutant monomers in E. coli that are deficient in specific amino acids in the presence of synthetic (i.e., non-natural) analogs of such amino acids. In addition, in the case of obtaining mutant monomers by partial peptide synthesis, non-natural amino acids can also be generated by naked ligation. Conservative substitution is used to replace amino acids with other amino acids having similar chemical structures, similar chemical properties, or similar side chain volumes. The introduced amino acid can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the substituted amino acid. Alternatively, conservative substitution can also be used to replace an existing aromatic or aliphatic amino acid with another aromatic or aliphatic amino acid. Conservative modification of amino acids is well known in the art and can be selected based on the properties of the 20 main amino acids listed in Table 1 below. For amino acids with similar polarity, further reference can be made to the amino acid side chain hydrophilicity scale listed in Table 2.
[0187] Table 1: Chemical properties of amino acids
[0188] Ala Aliphatic, hydrophobic, neutral Met Hydrophobic, neutral Cys Polar, hydrophobic, neutral Asn Polar, hydrophilic, neutral Asp Polar, hydrophilic, charged (-) Pro Hydrophobic, neutral Glu Polar, hydrophilic, charged (-) Gln Polar, hydrophilic, neutral Phe Aromatic, hydrophobic, neutral Arg Polar, hydrophilic, charged (+) Gly Aliphatic, neutral Ser Polar, hydrophilic, neutral His Aromatic, polar, hydrophilic, charged (+) Thr Polar, hydrophilic, neutral Ile Aliphatic, hydrophobic, neutral Val Aliphatic, hydrophobic, neutral Lys Polar, hydrophilic, charged (+) Trp Aromatic, hydrophobic, neutral Leu Aliphatic, hydrophobic, neutral Tyr Aromatic, polar, hydrophobic
[0189] Table 2: Hydrophilicity Scale
[0190]
[0191] The mutated or modified protein, monomer or peptide can also be chemically modified at any site in any manner. Chemical modification of the mutated or modified monomer or peptide can be achieved by attaching the molecule to one or more cysteines (cysteine attachment), or attaching the molecule to one or more lysines, or attaching the molecule to one or more non-natural amino acids, or enzymatically modifying the epitope, or modifying the termini. Suitable methods for performing such modifications are well known in the art. Mutants of the modified protein, monomer or peptide can be chemically modified by attaching any molecule. For example, mutants of the modified protein, monomer or peptide can be chemically modified by attaching a dye or a fluorescent group.
[0192] polypeptide constructs
[0193] The present invention relates in part to multi-domain polypeptide constructs that are used to combine two or more domains to form oligomeric proteins and, in some embodiments, can also be used as monomers. Multi-domain polypeptides are generally genetically engineered to combine domains that do not naturally occur together. In some embodiments, three, four, five, or six polypeptide constructs are combined to form an oligomer. In some embodiments, three constructs are combined to form a trimer, such as a homotrimer.
[0194] This application describes an oligomeric core of subunit monomers. Subunit monomers are generally domains of multidomain polypeptide constructs. Such domains can form the core of a multivalent protein scaffold through oligomerization.
[0195] Accordingly, the description of the characteristics of the subunit monomers applies equally to the disclosure and definition of the domains of the respective polypeptide constructs.
[0196] Similarly, depending on the context, the first binding domain and the second binding domain of the polypeptide construct may constitute the first binding site and the second binding site described elsewhere in this application, or may constitute the first effector moiety and the second effector moiety described elsewhere in this application. For example, when a binding domain is a "catch" domain that constitutes an isopeptide bond (or other binding site as described herein), it is a binding site as described elsewhere below. When the binding domain is, for example, an antibody, an antigen-binding fragment, an antibody mimetic, a protein ligand or a peptide ligand, a protein signaling molecule or a peptide signaling molecule (such as a cytokine), a biological receptor, or other molecule described as an effector moiety in this application, the binding domain is an effector moiety as described elsewhere in this application.
[0197] Accordingly, the definitions and descriptions of the first and second binding sites or the first and second effector moieties apply equally to the first and second binding domains of the polypeptide construct, where appropriate.
[0198] In some embodiments, a polypeptide construct includes a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain. The N-terminus refers to the terminal amino acid residue at the amino terminus of the polypeptide. The C-terminus refers to the terminal amino acid residue at the carboxyl terminus of the polypeptide. Generally, when the target molecule is expressed on a single cell or fixed to a plate or a single bead, the first binding domain and the second binding domain are able to bind to their target molecule. This situation is sometimes referred to as "providing a first binding domain and a second binding domain in a cis orientation" in this application. Generally, the first binding domain and the second binding domain are able to bind to their cellular targets on the surface of a single cell and cluster their targets within the cell membrane. Therefore, certain cis-acting agents (i.e., agents that act on a single cell) prefer a cis orientation. For a description of the cis orientation of bispecific antibodies and the opposite "trans" orientation, see Dickopf et al. (Computational and Structural Biotechnology Journal, Vol. 18, 2020, pp. 1221-1227).
[0199] The “cis orientation” used in this application (“cis” for cis comes from Latin) refers geometrically to the spatial arrangement of two components on the same side of a plane, and is similar to “trans” or “trans orientation” (“trans” for trans comes from Latin) which geometrically represents two components across a plane (i.e., on different sides of a plane). These two definitions are similar to cis-trans isomers and have been used to describe the structure of bispecific antibodies before. This geometric definition is different from the “cis action” and “trans action” in terms of biological action, which respectively refer to a single bispecific molecule acting on a single or adjacent cell (cis) and on different cell populations (trans), for example, recruiting effector cells to target cells. Exploring different forms of geometric structures is one direction of current research (for example, see Dengl et al., Nature Communications (Nat Commun.), 2020, 11, 4974). For certain "trans-acting" bispecific antibodies, a "cis orientation" can be advantageous, for example, due to the ability to shorten the distance between molecules (Dijkopf et al., Journal of Computational and Structural Biotechnology, Vol. 18, 2020, pp. 1221-1227). For multivalent cis orientations of cis-acting bispecific antibodies that achieve multivalent binding to targets on individual cells or clustering them into clusters, a cis orientation is particularly advantageous (e.g., compared to bispecific tandem fusions described by Veggiani et al. (Biochemistry, Jan. 19, 2016, Vol. 113, No. 5, pp. 1202-1207) and higher-order monospecific clustering (Khairil Anuar et al., Nature Communications, Vol. 10, Article No. 1734 (2019)).
[0200] The function of the domain is to provide defined structural support for the binding domain. The domain has the advantage of ensuring that the binding domain has the desired orientation for binding to its target (generally, both binding domains are in a cis orientation). Thus, the construct can present a single binding surface.
[0201] In certain embodiments, the attachment site on the domain (oligomer core) for the binding domain enables binding even with a short linker.
[0202] The domain can be any polypeptide domain that includes a specific secondary structure (generally an alpha helix or a beta sheet). In some particularly advantageous embodiments, the N-terminus and the C-terminus of the domain are located in the same spatial region, for example, substantially adjacent to each other or adjacent to each other. In this way, after the two binding domains are connected to the two ends of the domain, the two binding domains can be made substantially adjacent in a three-dimensional conformation. In some embodiments, the N-terminus and the C-terminus are substantially oriented in the same direction. As described in other parts of this application, by making the N-terminus and the C-terminus adjacent in space, the binding domains can be located on the same face of the construct, or on the same face of an oligomer containing multiple constructs. In this way, the construct generally has a single binding surface. The binding region of the construct is generally cis-oriented.
[0203] In some embodiments, the domain presents the binding domain in a suitable orientation (e.g., suitable relative positioning of the N-terminus and C-terminus of a single monomer) through the tertiary structure of the monomer. In some embodiments, the domain presents the binding domain in a suitable orientation (e.g., suitable relative positioning of the N-terminus and / or C-terminus between monomers) through the association of the monomer with the quaternary structure. In some preferred embodiments, the domain presents the binding domain in a suitable orientation (e.g., suitable relative positioning of the N-terminus and / or C-terminus within and between monomers) through the tertiary structure of the monomer in combination with the association of the monomer with the quaternary structure.
[0204] A domain may comprise a single polypeptide chain, or may be composed of two or more separate polypeptide chains that associate to form a single domain (e.g., two antiparallel (NC CN) α-helices or two or more β-strands that associate to form a β-sheet). In some embodiments, after identifying two or more polypeptide chains with suitable properties, they are fused, typically by recombination to form a single polypeptide chain (i.e., a fusion protein), but may also be chemically coupled or bonded to form a single covalent molecule.
[0205] Since the structural domain is different from the two binding domains, when the binding domain is a catcher polypeptide such as SpyCatcher, DogCatcher or SnoopCatcher, the structural domain is not a catcher polypeptide.
[0206] Typically, the domain does not include a CH2 domain. Typically, the domain does not include a CH3 domain. In some embodiments, the domain includes neither a CH2 domain nor a CH3 domain.
[0207] The inventors have identified several exemplary domains, which are described below and in the Examples. In some embodiments, the domain comprises or consists of the collagen X NC1 domain (SEQ ID NO: 2), or comprises a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, such as at least 90% or at least 95% identity thereto. In some embodiments, the domain comprises or consists of the collagen VIII NC1 domain (SEQ ID NO: 3), or comprises a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, such as at least 90% or at least 95% identity thereto. In some embodiments, the domain comprises or consists of the CutA1 polypeptide (e.g., SEQ ID NO: 1 or SEQ ID NO: 19), or comprises a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, such as at least 90% or at least 95% identity thereto.
[0208] In certain embodiments, the structural element comprises or consists of a polypeptide that has at least 50%, such as at least 90%, amino acid identity to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 19, SEQ ID NO: 29, SEQ ID NO: 60, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 42, SEQ ID NO: 31, SEQ ID NO: 78, SEQ ID NO: 80, or SEQ ID NO: 58.
[0209] A suitable domain is the C4 domain of collagen IV (also known as collagen IV NC1 domain) (PDB ID: 1m3d, SEQ ID NO: 49), or a polypeptide having at least 50%, at least 60%, at least 70% or at least 80%, such as at least 90% or at least 95% identity thereto.
[0210] In general, collagen NC1 domains including collagen IV, collagen VIII, and collagen X NC1 domains can be used as the structural domains of the present invention.
[0211] However, not all collagen NC1 domains are suitable as structural domains. In particular, the collagen XV and collagen XVIII NC1 domains do not have the required orientation.
[0212] In certain embodiments, the domain comprises human macrophage migration inhibitory factor (MIF) (PDB ID: 1CA7 or SEQ ID NO: 25, or PDB ID: 6OY8 with a Y99G mutation), or human macrophage migration inhibitory factor 2 (MIF2) (PDB ID: 7MSE or SEQ ID NO: 26, or SEQ ID NO: 27 with S62A and F99A mutations), or a homolog or paralog thereof.
[0213] In certain embodiments, the domain comprises a TNF family protein including TNF (PDB ID: 1TNF, SEQ ID NO:42; full length soluble domain SEQ ID NO:80), TL1A (PDB ID: 2RE9, SEQ ID NO:31), OX40L (SEQ ID NO:78), or CD40L (PDB ID: 3LKJ, SEQ ID NO:58).
[0214] Other domains include, for example, an antiparallel coiled-coil hexamer modified in a suitable manner (PDB ID: 5WOJ, see Example 4, SEQ ID NO: 43); HIV-1 GP41 core (PDB ID: 1I5Y or SEQ ID NO: 44); cytochrome c555 (PDB ID: 5Z25 or SEQ ID NO: 45); the invariant chain (Ii) of the MHC class II-associated chaperone and targeting protein (PDB ID: 1iie or SEQ ID NO: 46), p53 (PDB ID: 1C26 or SEQ ID NO: 47); a fibrinogen-like domain (PDB ID: 4M7F or SEQ ID NO: 48); Bacillus subtilis AbrB (PDB ID: 1YFB or SEQ ID NO: 50); phage lambda head protein D (e.g., PDB ID: 1C5E or PDB ID: 1C5E, or SEQ ID NO: 51); a domain-swapped trimeric variant of HCRBPII (PDB ID: 1C5E or SEQ ID NO: 52); ID: 6VIS or SEQ ID NO: 52); T1L reovirus attachment protein σ1 (PDB ID: 4ODB or A, B, C chains of SEQ ID NO: 53).
[0215] In some specific aspects, the domain of the polypeptide construct (or the subunit monomer of the oligomer core) is a CutA1 protein, typically a human CutA1 protein. The data provided below, in particular Figure 23, demonstrating various fusions of HsCutA1 ("Homo sapiens CutA1") with effector proteins, and the ability of these direct fusions to exert biological effects. Thus, in certain aspects, the present invention provides a polypeptide comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a human CutA1 domain.
[0216] In certain aspects, CutA1, typically human CutA1, is genetically engineered to remove one or more cysteine residues from the native sequence (herein as SEQ ID NO: 19). Typically, this is where one or more cysteine residues are substituted with one or more non-cysteine residues. Removal of one or more cysteine residues advantageously allows for targeted cysteine coupling at non-native sites or in fusion proteins. Without wishing to be bound by theory, unpaired cysteines often interfere with stability and downstream applications, and therefore removal of unpaired cysteines may be advantageous.
[0217] In some embodiments, one or more cysteine residues in CutA1 are substituted with one or more alanine residues. In some embodiments, one or more cysteine residues in CutA1 are substituted with one or more valine residues. In some embodiments, one or more cysteine residues in CutA1 are substituted with one or more serine residues. In some embodiments, the substitution comprises or consists of two cysteines being substituted with two alanines, which is referred to herein as a "CACA" substitution. In some embodiments, the substitution comprises or consists of one cysteine being substituted with valine and one cysteine being substituted with serine, which is referred to herein as a "CVCS" substitution. In some embodiments, the cysteine residues at positions 75 and 96 of wild-type human CutA1 (e.g., SEQ ID NO: 19) are substituted with different residues. In some embodiments, human CutA1 is genetically engineered to replace two cysteine residues, wherein the cysteine substitution comprises or consists of (i) C75A, C96A or (ii) C75V, C96S. Thus, in some embodiments, the domain of the polypeptide construct (or the subunit monomer of the oligomer core) is a human CutA1 protein that has been genetically engineered to have two cysteine residues substituted, wherein the cysteine substitutions comprise or consist of (i) C75A, C96A or (ii) C75V, C96S. The generation, yield, and biological effects of polypeptide constructs comprising such cysteine-substituted human CutA1 domains are described below ( Figure 25 ).
[0218] In certain aspects, CutA1, typically human CutA1 (e.g., SEQ ID NO: 19), is genetically engineered to delete one or more residues from either or both ends of the native sequence to form a truncated CutA1 domain that is incorporated into the constructs of the invention. In some embodiments, one or more residues are deleted from the N-terminus of CutA1, typically human CutA1. Typically, 5 to 70 residues, such as 10 to 59 residues, are deleted from the N-terminus of CutA1, typically human CutA1. In some embodiments, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 59, 60, 61, 62, 63, 64, 65, or 66 residues are deleted from the N-terminus of CutA1. In some embodiments, the truncated CutA1 begins at residue 33 (i.e., 32 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 44 (i.e., 43 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 60 (i.e., 59 N-terminal residues are missing). In some embodiments, the truncated CutA1 begins at residue 67 (i.e., 66 N-terminal residues are missing), for example Figure 23 "HsCutA1" shown in b 67-171 CACA" direct fusion construct.
[0219] In some embodiments, one or more residues are deleted from the C-terminus of CutA1, typically human CutA1. The C-terminal deletion may replace or supplement the N-terminal deletion. Typically, 5 to 20 residues are deleted from the C-terminus of CutA1, typically human CutA1, for example, 6 to 12 residues are deleted. In some embodiments, about 8 residues are deleted, which leaves a truncated protein in human CutA1 with its C-terminus located at Figure 24 and residue 171 (valine) in the wild-type sequence set forth in SEQ ID NO: 19. In some embodiments, the C-terminus of the truncated protein is located at Figure 24 and residue 168 (threonine) in the wild-type sequence set forth in SEQ ID NO: 19. In some embodiments, the C-terminus of the truncated protein is located at Figure 24 The wild-type sequence shown and residue 169, 170, 172, 173, 174, 175 or 176 in SEQ ID NO: 19.
[0220] In certain embodiments, the truncated CutA1 consists of residues 44-179 or residues 60 to 171 of SEQ ID NO: 19, e.g. Figure 24In some embodiments, the truncated CutA1 begins at any one of residues 44 to 67 of SEQ ID NO: 19 and ends at any one of residues 168 to 179. In other embodiments, the truncated CutA1 begins at any one of residues 30 to 65 of SEQ ID NO: 19 and ends at any one of residues 165 to 179. In some embodiments, the truncated CutA1 begins at any one of residues 44 to 60 of SEQ ID NO: 19 and ends at any one of residues 171 to 179.
[0221] Truncation of CutA1 can improve the accuracy of constructing fusion constructs. Figure 24 The generation, yield, and biological effects of polypeptide constructs comprising this truncated human CutA1 domain are shown.
[0222] In certain aspects, CutA1 is human CutA1 that has been genetically engineered to remove one or more cysteine residues from the native sequence and is truncated at the N-terminus and / or C-terminus. In some embodiments, the truncated CutA1 consists of residues 44-179 or residues 60-171 of SEQ ID NO: 19, such as Figure 24 In some embodiments, the truncated CutA1 starts at any one of residues 30 to 65 of SEQ ID NO: 19 and ends at any one of residues 165 to 179 and also has a cysteine substitution comprising (i) C75A, C96A or (ii) C75V, C96S or consisting thereof. In other embodiments, the truncated CutA1 starts at any one of residues 30 to 65 of SEQ ID NO: 19 and ends at any one of residues 165 to 179 and also has a cysteine substitution comprising (i) C75A, C96A or (ii) C75V, C96S or consisting thereof.
[0223] The polypeptide construct of the present invention further comprises a first binding domain and a second binding domain in addition to the above-mentioned structural domains.
[0224] In certain embodiments, the binding domain is capable of forming an isopeptide bond with a cognate peptide (e.g., various catcher domains well known in the art). Constructs containing such domains that form isopeptide bonds are particularly suitable for screening different pairs of effector molecules, such as antigen binding proteins. As described elsewhere in this application, many effector molecule combinations can be linked to constructs containing binding domains that form isopeptide bonds via isopeptide-forming peptide tags. Therefore, constructs containing binding domains that can form isopeptide bonds with cognate peptides are particularly useful as drug discovery platforms.
[0225] The aspects of the present invention relating to isopeptide bond formation are generally described using the example of a larger molecule (domain) (generally referred to as the "catcher") being attached to the domain and a smaller polypeptide or peptide (generally referred to as the "tag") forming the target portion of the binding region (e.g., antigen binding domain). However, all aspects and embodiments can also be performed in the reverse manner - the larger molecule (e.g., the catcher) forming the target portion of the binding region (e.g., antigen binding domain) and the smaller peptide tag forming the binding domain attached to the domain.
[0226] In some embodiments, the first binding domain and the second binding domain in the polypeptide construct are effector molecules such as antigen binding domains. In such embodiments, the construct is particularly suitable for use as a diagnostic agent, analytical agent, or therapeutic agent.
[0227] In some embodiments, a pair of candidate or effective antigen binding regions (e.g., where the construct contains a binding domain that forms an isopeptide bond) is first identified using the drug discovery platform of the present invention, and then the construct is expressed without the binding domain that forms an isopeptide bond and with the identified antigen binding domain (or other effector moiety) combination, wherein the antigen binding domain is directly linked to the structural domain, without a catcher domain on the structural domain, and without a peptide tag on the antigen binding domain. For the avoidance of doubt, it is expressly noted that such direct fusion constructs may still include a linker region between the terminal residue of the structural domain and the terminal residue of each effector moiety (e.g., antigen binding region), as described in detail elsewhere in this application.
[0228] Accordingly, one aspect of the present invention provides a system for large-scale high-throughput screening of multiple pairs of potentially useful effector molecules using paired tagged effector proteins. Useful combinations identified can be used as they are provided in screening constructs or converted to a simpler form (e.g., as candidate therapeutic agents) by directly fusing effector molecules (e.g., antigen-binding regions) to the same domains used in the drug discovery platform. This provides a technology for identifying and developing bispecific and multispecific therapeutics in a simple, rapid, and reliable manner.
[0229] Antigen binding domains are typical domains that can be used and applied to the present invention. In certain aspects of the present invention, the antigen binding domain comprises a peptide tag such as SpyTag or SnoopTag, which can form an isopeptide bond and can be bound to a construct comprising a paired catcher domain via an isopeptide bond, for example, to create a platform for combinatorial screening or modular screening. In other aspects of the present invention, the construct of the present invention comprises a first antigen binding domain at the N-terminus and a second antigen binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain. Optionally, there is a linker sequence between one or each antigen binding domain and the structural domain. As described below, the length of a suitable peptide linker for connecting the binding domain (binding site) and the structural domain (monomer subunit) is generally 1 to 150, 1 to 100, 1 to 50, 1 to 25, 1 to 20, 1 to 15 or 1 to 10 amino acids. Examples of linker sequences are GSGS, GGGGS, GGGGSGGGGS or GGGGSGGGGSGGGGS. Other suitable linkers include linkers that form α-helical secondary structures, such as EAAAK and multiples thereof, such as (EAAAK)2, (EAAAK)5 or (EAAAK)7, or an α-helical linker derived from the ribosomal L9 protein, with the sequence PANLKALEAQKQKEQRQAAEELANAKKLKEQLEK (Kuhlman et al., 1997 J Mol Biol 270, 5, 640-647) or a derivative (e.g., a derivative comprising 1, 2, 3, 4, 5 or up to 10 substitutions, insertions or deletions) or multiples thereof, such as 2, 3, 4, 5 or up to 10 repeats of the L9 linker. Other suitable linkers include linkers that confer variable function to the fusion protein, such as those discussed in Chen et al. 2013 (Adv Drug Deliv Rev. 2013, 65, 10, 1357-1369), including, but not limited to, PAPAP, CC, AP, VSQTSKLTRAETVFPDV, PLGLWA, RVLAEA, EDVVCCSMSY, GGIEGRGS, TRHRQPRGWE, AGNRVRRSVG, RRRRR, RRRRRRRRR, GFLG or LE, or derivatives (e.g., derivatives comprising 1, 2, 3, 4, 5 or up to 10 substitutions, insertions or deletions) or multiples thereof, for example, 2, 3, 4, 5 or up to 10 repeats of the PAPAP linker. The catcher domain has structural similarity to an antibody domain having an immunoglobulin fold (Kang et al., Science. 318, 5856, 1625-1628). Replacing the capture domain with a human antibody domain provides an alternative route to generate humanized direct fusion constructs.Thus, other suitable linkers include antibody fragment domains, such as CH1, CH2 or CH3 from various antibody isotypes or from various species, and modifications thereof, preferably of human origin. Examples include a CH1 domain from a human IgG1 isotype with the sequence ASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVL QSSGLYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKKV, a CH2 domain from a human IgG1 isotype with the sequence PCPAPELLGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVE VHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPIEKTISK AK, or a CH3 domain from a human IgG1 isotype with the sequence GQPREPQVYTLPPSRDELTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPP VLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK.
[0230] The antigen binding domain is generally an antigen binding fragment of an antibody. The antigen binding fragment of an antibody is not a complete antibody of full length and typically lacks at least the CH2 and / or CH3 domains. Such antigen binding fragments are well known and include Fab, F(ab')2, Fv, or single-chain Fv fragments (scFv). The antigen binding fragment generally includes the CDRs required for antigen binding (generally six CDRs) and the framework residues required to maintain the correct CDR structure. In some embodiments, the antigen binding domain includes a heavy (H) chain variable domain sequence (VH) and a light (L) chain variable domain sequence (VL).
[0231] In some embodiments, the antigen binding region can be a single domain antibody. Single domain antibodies (sdAbs) may include antibodies whose complementary determining regions are part of a single domain polypeptide, such as, but not limited to, heavy chain antibodies, antibodies naturally free of light chains, derivative single domain antibodies of traditional four-chain antibodies, genetically engineered antibodies, and non-antibody derived single domain scaffolds. Single domain antibodies can be any single domain antibody in the art or any future single domain antibody. Single domain antibodies can be derived from species, including but not limited to mice, humans, camels, llamas, fish, sharks, goats, rabbits, and cattle. Single domain antibodies can be natural single domain antibodies existing in nature, i.e., heavy chain antibodies without light chains. Such single domain antibodies are, for example, disclosed in WO-A-94 / 04678. For clarity, such variable domains derived from natural heavy chain antibodies without light chains are sometimes referred to as VHH or nanobodies to distinguish them from the VH of traditional four-chain immunoglobulins. Such VHH molecules can be derived from antibodies raised in species of the Camelidae family, such as camels, llamas, dromedaries, alpacas and guanacos. Other species outside the Camelidae family can also produce heavy chain antibodies that are naturally free of light chains, and such VHHs are also within the scope of the present invention.
[0232] The antigen-binding domain may also comprise or consist of an antibody mimetic such as an affibody or DARPin. As is known in the art, an affibody is a small polypeptide comprising three alpha helices, typically having approximately 58 amino acids and a molecular weight of approximately 6 kDa. As is known in the art, a DARPIN (designed ankyrin repeat protein) is a genetically engineered antibody mimetic protein that generally exhibits high specificity and high affinity binding to a target protein.
[0233] The binding domain may include natural ligands such as cytokines that exist in nature, which can serve as substitutes for the antigen binding domain.
[0234] When two different antigen-binding domains are present on a bispecific molecule, they typically bind to different epitopes. These different epitopes can be different epitopes of the same target molecule or different epitopes of different target molecules. In some embodiments, both epitopes are located on a therapeutic target, where binding of the antigen-binding domain to the therapeutic target results in a change in a biological mechanism (generally a pathogenic mechanism), thereby achieving a therapeutic benefit.
[0235] In some embodiments, each antigen binding domain can exert an agonistic effect on a biological target. In some embodiments, each antigen binding domain can exert an antagonistic effect on a biological target. In some embodiments, one antigen binding domain can exert an agonistic effect on a first target, and another antigen binding domain can exert an antagonistic effect on a second target.
[0236] In some embodiments, the construct may include two binding regions that bind to the same epitope but with different affinities for that epitope. In some embodiments, the construct may include two binding regions that bind to the same epitope, optionally with different affinities. Furthermore, the two binding regions may have different formats, for example, one antigen-binding region may be a scFv and the other antigen-binding region may be a Fab.
[0237] The multi-domain polypeptide constructs of the invention generally have the following form, where the orientation shown is the conventional N-terminus to C-terminus direction:
[0238] (Binding domain 1)-Linker 1-Domain-Linker 2-(Binding domain 2),
[0239] Wherein, linker 1 and linker 2 are optional linker sequences, optionally having 1 to 20 amino acids, such as GSGS. A purification tag, such as a His tag (eg, a 6×His tag), is optionally added to either end of the construct.
[0240] Each binding domain can be identical, but is typically different. The binding domain can typically be a catcher polypeptide or an antigen binding domain. Several constructs are provided below for illustrative purposes.
[0241] SpC-Linker 1-CutA1-Linker 2-SpC,
[0242] SnC-Linker 1-CutA1-Linker 2-SnC,
[0243] SnC-Linker 1-CutA1-Linker 2-SpC,
[0244] SpC-Linker 1-CutA1-Linker 2-SnC,
[0245] SpC3-Linker 1-CutA1-Linker 2-DgC,
[0246] scFv-Linker 1-CutA1-Linker 2-ScFv,
[0247] Fab-Linker 1-CutA1-Linker 2-Fab,
[0248] ScFv-Linker 1-CutA1-Linker 2-Fab,
[0249] Fab-Linker 1-CutA1-Linker 2-ScFv,
[0250] Nanobody 1-Linker 1-CutA1-Linker 2-Nanobody 2,
[0251] Nanobody-Linker 1-CutA1-Linker 2-DgC,
[0252] SpC3-Linker 1-CutA1-Linker 2-Nanobody.
[0253] Wherein, SpC is SpyCatcher, SpC3 is SpyCatcher003, SnC is SnoopCatcher, Fab is the antigen-binding fragment of an antibody, and scFv is a single-chain Fv. In all cases, Linker 1 and Linker 2 are optional. The CutA1 sequence can be human, derived from Pyrococcus Horikoshi, or a homolog from another species, or have at least 30%, at least 50%, at least 70%, or at least 90% identity to a human sequence or a Pyrococcus Horikoshi sequence.
[0254] Other constructs for illustrative purposes are as follows:
[0255] SpC-Linker 1-NC1-Linker 2-SpC,
[0256] SnC-connector 1-NC1-connector 2-SnC,
[0257] SnC-Linker 1-NC1-Linker 2-SpC,
[0258] SpC-Linker 1-NC1-Linker 2-SnC,
[0259] scFv-Linker 1-NC1-Linker 2-ScFv,
[0260] Fab-Linker 1-NC1-Linker 2-Fab,
[0261] ScFv-Linker 1-NC1-Linker 2-Fab,
[0262] Fab-Linker 1-NC1-Linker 2-ScFv,
[0263] Nanobody 1-Linker 1-NC1-Linker 2-Nanobody 2,
[0264] Nanobody-Linker 1-NC1-Linker 2-DgC,
[0265] SpC3-Linker 1-NC1-Linker 2-Nanobody.
[0266] wherein NC1 is the collagen NC1 domain derived from collagen VIII or collagen X. In all cases, linker 1 and linker 2 are optional.
[0267] Other constructs for illustrative purposes include macrophage migration inhibitory factor (MIF) (SEQ ID NO: 25), or macrophage migration inhibitory factor 2 (MIF2) (SEQ ID NO: 26), or the S62A-F99A mutant of MIF2 (MIF2m, SEQ ID NO: 27) as domains. These constructs include:
[0268] SpC-Linker 1-MIF2-Linker 2-SpC
[0269] SnC-Linker 1-MIF2-Linker 2-SnC
[0270] SnC-Linker 1-MIF2-Linker 2-SpC
[0271] SpC-Linker 1-MIF2-Linker 2-SnC
[0272] scFv-Linker 1-MIF2-Linker 2-ScFv
[0273] Fab-Linker 1-MIF2-Linker 2-Fab
[0274] ScFv-Linker 1-MIF2-Linker 2-Fab
[0275] Fab-Linker 1-MIF2-Linker 2-ScFv
[0276] Nanobody 1-Linker 1-MIF2-Linker 2-Nanobody 2
[0277] Nanobody-Linker 1-MIF2-Linker 2-DgC
[0278] SpC3-Linker 1-MIF2-Linker 2-Nanobody
[0279] In addition to the above-mentioned exemplary domains, the above-mentioned constructs can use any suitable domains, in particular any domains described herein. Therefore, the above-mentioned exemplary forms are intended to give forms for general domains and for the domains described herein.
[0280] The orientation of the binding domains in monomeric or oligomeric forms of the multi-domain constructs can be functionally evaluated using a variety of assays. A list of exemplary assays is provided below:
[0281] FRET: FRET analysis can be used to demonstrate that a selected scaffold has a cis orientation compared to a non-cis oriented protein. A scaffold polypeptide with a catcher moiety (e.g., SpC3-HsCutA1-DgC) can be coupled to a pair of fluorescent protein FRETs fused to corresponding tags (e.g., mCherry(6+)-SpT3-H6 and H6-DgT-mCitrine(4-)). After the tagged FRET pair is coupled to the scaffold protein, the luminescence of the acceptor FRET protein can be measured using standard fluorescence readout methods and compared to the sensitized luminescence of the donor FRET protein. Protein scaffolds that preferentially adopt a cis orientation will exhibit higher luminescence from the acceptor, while those that preferentially adopt a trans orientation will exhibit higher luminescence from the donor.
[0282] SPR: SPR experiments can be used to demonstrate that a cis-oriented scaffold conjugated with an appropriate ligand preferentially binds to a target in the same plane as a non-cis-oriented protein. Target proteins specific to the ligands conjugated to the scaffold (e.g., targets for L1 and L2 in SpC-PhCutA1-SnC:SnT-L1:L2-SpT) can be immobilized on the surface of an SPR sensor chip. Both targets can be immobilized on the sensor chip together, or only L1 or L2 can be immobilized separately for control purposes. Subsequently, scaffolds conjugated to L1 and L2 in either a cis- or non-cis-oriented configuration can be loaded onto an SPR sensor chip with the L1 and / or L2 targets, and the binding of the coupled assembly to the immobilized targets can be determined. In the case where two targets are immobilized on the same chip, an assembly with a cis orientation should have readily measurable binding to both L1 and L2 targets, whereas an assembly with a non-cis orientation should not have readily measurable binding to such a chip. Furthermore, in the case where only L1 or L2 targets are immobilized, both assemblies should have readily measurable binding.
[0283] SEC-MALS: SEC-MALS experiments can measure the native oligomeric state of scaffold and assembly proteins in solution. Scaffold and assembly protein preparation methods are described in the Methods section of this article. After preparation, the sample is injected into an FPLC instrument coupled to a MALS instrument and detector to separate the sample by size. The native protein mass is approximated by calculating the light scattering behavior. The oligomeric state of the protein is then determined by dividing the predicted monomer mass calculated by software such as ProtParam by the native protein mass. Scaffold and assembly proteins (e.g., SpC-PhCutA1-SnC and SpC-PhCutA1-SnC:SnT-L1:L2-SpT) should have an oligomeric state value of 3.
[0284] Cross-linking and LC-MS / MS: To measure the binding of a cis-oriented coupled assembly to two targets simultaneously (e.g., targets for L1 and L2), target-expressing cells can be incubated with a biotinylated coupled assembly (biotin-SpC-PhCutA1-SnC:SnT-L1:L2-SpT). The binding between the target and the biotinylated assembly can then be cross-linked using BS3 (disuccinimidyl suberate). Cells are then lysed, and the cross-linked target-assembly complexes are extracted with streptavidin. The complexes are then trypsin-digested and loaded onto LC-MS / MS to determine the binding of L1 and L2 to their respective targets. This approach can also be repeated for a non-cis-oriented assembly to compare the LC-MS / MS data outputs for the two protein assemblies. The cis-oriented assembly should preferentially bind to both targets, L1 and L2, while the non-cis-oriented assembly may preferentially bind to one of the two targets.
[0285] Multivalent protein scaffolds
[0286] One aspect of the present invention relates to a modular system for screening target molecules. This system enables multivalent presentation of target molecules. In one aspect, a multivalent protein scaffold is provided. The multivalent protein scaffold comprises an oligomeric core comprising a plurality of subunit monomers. The multivalent protein scaffold also comprises at least one first binding site orthogonal to at least one second binding site. Suitable binding sites are described in further detail herein.
[0287] The scaffold is used as a platform for other molecules to bind. Different combinations of molecules can be attached to the scaffold in a modular manner. The scaffold allows multivalent binding of molecules. The molecules bound to the scaffold generally have potential therapeutic benefits, and the scaffold after binding to the molecules can be used to study whether the multivalent assembly of different molecules may have the desired effect. For example, different combinations of polypeptides with potential anti-cancer effects can be connected to the scaffold first, and then the resulting assembly can be used for screening analysis to see whether the combination has an effect on cancer cells (such as binding to cancer cells and causing cancer cell death). After identifying the molecules, therapeutic drug candidates can be obtained by modifying the multivalent protein scaffold so that the drug can be directly connected to the identified molecules without a modular system. The multivalent protein scaffold presents the molecules on the same side of the scaffold, so that all molecules have the potential to interact with target cells.
[0288] The scaffolds of the present invention generally include at least two first binding sites and at least two second binding sites. By providing at least two first and at least two second binding sites, the scaffolds can be used to screen for multivalent interactions, which is not always possible when screening with bispecific antibodies, for example. Furthermore, they can serve as a "toolkit" for testing different scaffolds against the same combination.
[0289] Accordingly, the present application provides a multivalent protein scaffold comprising:
[0290] - an oligomeric core comprising multiple subunit monomers; and
[0291] - at least two first binding sites that are orthogonal to at least two second binding sites,
[0292] The first binding site and the second binding site are located on the same surface of the support.
[0293] The scaffolds of the present invention generally comprise binding sites capable of forming covalent bonds with their respective targets. Covalent bonds enable strong, irreversible associations. The complexes produced by the covalent linkage of the scaffolds of the present invention to the binding site targets are not only physically robust but also readily available. Such complexes can be produced in high yields and with high uniformity. Consequently, when such complexes are administered to a biological system, such as a subject described herein, the biological responses produced are reproducible and controllable.
[0294] Accordingly, the present application also provides a multivalent protein scaffold comprising:
[0295] - an oligomeric core containing multiple subunit monomers;
[0296] - at least one first binding site orthogonal to at least one second binding site,
[0297] wherein the first binding site and the second binding site are located on the same side of the scaffold; and wherein the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target.
[0298] As described herein, the scaffold provided by the present application has significant advantages compared to traditional antibodies, including bispecific antibodies. The oligomer core of the scaffold of the present invention generally does not include an antibody Fc region. In some embodiments, the oligomer core does not include a CH2 region. In some embodiments, the oligomer core does not include a CH3 region. In some embodiments, the oligomer core includes neither a CH2 region nor a CH3 region. When using the immunoglobulin domain of an antibody (usually a constant domain in a region such as Fc), the advantages of the scaffold of the present invention described in this application are generally not obtained. For example, bispecific antibodies do not have the degree of modularity of the present invention and therefore may not be used for studies of multivalent interactions.
[0299] Accordingly, the present application also provides a multivalent protein scaffold comprising:
[0300] - an oligomeric core containing multiple subunit monomers;
[0301] - at least one first binding site that is orthogonal to at least one second binding site;
[0302] wherein the first binding site and the second binding site are located on the same face of the scaffold; and wherein the oligomer core does not include an antibody Fc region.
[0303] The present application also provides a multivalent protein scaffold comprising:
[0304] - an oligomeric core containing multiple subunit monomers;
[0305] - at least one first binding site that is orthogonal to at least one second binding site;
[0306] wherein the first binding site and the second binding site are located on the same face of the scaffold; and wherein the oligomer core does not include the CH2 domain of an antibody, or does not include the CH3 domain of an antibody, or does not include either the CH2 domain or the CH3 domain.
[0307] The multivalent protein scaffold comprises an oligomer core and at least one first binding site and at least one second binding site. The multivalent protein scaffold may further comprise other parts such as linkers, insertion domains and / or functional groups as described in further detail in this application.
[0308] Preferably, the diameter of the multivalent protein scaffold is less than about 100 nm, such as less than about 50 nm, such as less than about 25 nm, such as less than about 10 nm. Preferably, the height of the multivalent protein scaffold is less than about 100 nm, such as less than about 50 nm, such as less than about 30 nm, such as less than about 20 nm, such as less than about 10 nm. The size of the multivalent protein scaffold is preferably For example For example For example
[0309] Preferably, the multivalent protein scaffold itself does not substantially induce an immune response in a biological system, cell culture, or subject, such as a human subject. That is, generally, due to the absence of binding sites on the oligomeric core and / or the effector moiety attached to the binding site of the protein scaffold, when the protein scaffold is administered to a biological system, such as a human subject, no immune response is induced (or substantially no immune response is induced, e.g., the immune response induced is no greater than the immune response upon administration of a non-immunogenic protein). For example, upon administration of a protein scaffold (without binding sites on the multivalent protein scaffold and / or the effector moiety attached to the binding site of the multivalent protein scaffold) to a biological system, such as a subject as described herein, no innate or acquired immunity is induced in the biological system. For example, the protein scaffold generally does not result in activation of the complement system, B cells, T cells, natural killer cells, mast cells, basophils, eosinophils, neutrophils, dendritic cells, or macrophages.
[0310] The multivalent protein scaffold preferably does not include an antibody or antibody fragment, but as further described in this application, antibodies and / or antibody fragments can be attached to the scaffold as effector moieties. The multivalent protein scaffold (e.g., without any effector moieties) more preferably does not include an antibody Fc region. The Fc region is the tail region of an antibody that interacts with cell surface receptors called Fc receptors and some proteins of the complement system. In some cases, the multivalent protein scaffold does not include an immunoglobulin constant region. In some embodiments, the multivalent protein scaffold does not include a CH2 domain. In some embodiments, the multivalent protein scaffold does not include a CH3 domain. In some embodiments, the multivalent protein scaffold does not include either a CH2 domain or a CH3 domain.
[0311] Preferably, the multivalent protein scaffold is thermodynamically stable. Preferably, the multivalent protein scaffold is stable at a temperature of about 0 to about 100°C, such as about 4°C to about 90°C, such as about 10°C to about 50°C, such as about 20 to about 38°C, such as about 25 to about 37°C. That is, preferably, when in an aqueous solution at a temperature of about 0 to about 100°C (e.g., about 4°C to about 90°C, such as about 10°C to about 50°C, such as about 20 to about 38°C, such as about 25 to about 37°C), the multivalent protein scaffold does not dissociate into substituted subunit monomers and / or the binding site does not dissociate from the oligomer core. For example, when in an aqueous solution having the above temperature, at least 90%, such as at least 95%, such as at least 99%, such as at least 99.9%, such as at least 99.99% or 99.999% of the multivalent protein scaffold does not dissociate into substituted subunit monomers and / or the binding site does not dissociate from the oligomer core. More preferably, the multivalent protein scaffold is stable at a temperature of about 0 to about 100° C., such as about 4° C. to about 90° C., such as about 10° C. to about 50° C., such as about 20 to about 38° C., such as about 25 to about 37° C. When measured at a temperature of about 0° C. to about 100° C., such as about 4° C. to about 90° C., such as about 10° C. to about 50° C., such as about 20 to about 38° C., such as about 25 to about 37° C., the lifespan of the multivalent protein scaffold is preferably at least 10 minutes, more preferably at least one hour, such as at least one day, such as at least one week, such as at least one month or at least one year.
[0312] Preferably, the interactions between the components of the multivalent protein scaffold are not weak transient interactions. Weak transient complexes are in different oligomeric states that change dynamically in vivo, whereas strong transient complexes change their quaternary state only when triggered, for example, by ligand binding. Weak transient interactions are characterized by a dissociation constant (K D ) are in the micromolar range and have a lifetime of seconds. Strong transient interactions can have longer lifetimes and low K in the nanomolar range due to stabilization by effector molecule binding. D value. Preferably, the components of the multivalent protein scaffold interact with each other at least through a strong transient reaction. More preferably, the components of the multivalent protein scaffold form permanent interactions. Permanent interactions mean that under normal conditions (e.g., 20°C to 40°C, pH = 6-8), the multivalent protein scaffold does not dissociate into its component parts. Multivalent protein scaffolds whose components form permanent interactions generally only dissociate under denaturing conditions that can denature the tertiary structure of the subunit monomers themselves.
[0313] Thus, at a temperature of about 0°C to about 100°C, such as about 4°C to about 90°C, such as about 10°C to about 50°C, such as about 20°C to about 38°C, such as about 25°C to about 37°C, the K of the multivalent protein scaffold components when interacting with each other isD The value is preferably less than 1 μM, such as less than 100 nM, more preferably less than 10 nM.
[0314] Preferably, the multivalent protein scaffold and its components are stable against proteases. For example, when the multivalent protein scaffold and its components are exposed to a protease such as trypsin, the scaffold does not lose its tertiary or quaternary structure. At a temperature of about 10 to about 40°C, such as about 20 to about 38°C, such as about 25 to about 37°C, the multivalent protein scaffold and its components are stable against proteases for at least 1 hour, such as at least 2 hours, such as at least 4 hours, such as at least 8 hours, such as at least 24 hours or more.
[0315] oligomer core
[0316] As described above, the multivalent protein scaffold provided by the present application includes an oligomer core, which includes a plurality of subunit monomers. The subunit monomers of the oligomer core are generally the domains of the polypeptide constructs described in other parts of the present application.
[0317] The number of subunit monomers can be any suitable number. For example, the oligomer core may include about 2 to about 20 subunit monomers, such as about 2 to about 10 subunit monomers, more preferably 3 to 7 subunit monomers, and more preferably 3 to 6 subunit monomers. For example, the oligomer core may include two, three, four, five, six, seven, eight, nine, or 10 subunit monomers. Preferably, the oligomer core includes at least 3 subunit monomers. Most preferably, the oligomer core includes three subunit monomers. Preferably, the oligomer core does not include, or is not composed of, 7 subunit monomers.
[0318] Preferably, the subunit monomers have rotational symmetry after polymerization, such as three-fold rotational symmetry, four-fold rotational symmetry, five-fold rotational symmetry, six-fold rotational symmetry or seven-fold rotational symmetry. The oligomer core may have C2, C3, C4, D2, C5, C6, D3, C7, C8, D4, C9, C10, D5, C11, C12, D6 or T symmetry.
[0319] Some or all of the subunit monomers of the oligomer core may be non-covalently linked together. Some or all of the subunit monomers of the oligomer core may be covalently linked together. The subunit monomers of the oligomer core may be covalently and non-covalently mixed. For example, the oligomer core may include a first monomer and a second monomer that are covalently bound to form a heterodimer. The oligomer core may include at least two such heterodimers that are non-covalently bound together. For example, the oligomer core may include three non-heterodimers that are non-covalently linked together, wherein each heterodimer includes two monomers that are covalently bound together.
[0320] Some or all of the subunit monomers of the oligomer core may be linked by non-covalent interactions. Suitable non-covalent interactions include, but are not limited to, electrostatic interactions such as ionic bonds, hydrogen bonds, and halogen bonds, and van der Waals forces such as orientation forces, π-π stacking interactions, cation-π interactions, anion-π interactions, or polar-π interactions.
[0321] Some or all of the subunit monomers of the oligomer core may be covalently linked together. When the subunit monomers are covalently linked together, the monomers are generally of an amino acid sequence corresponding to the original or native monomer domain.
[0322] Two or more subunit monomers can be covalently bound together via disulfide bonds. Disulfide bonds are typically formed between cysteine residues in polypeptides. Artificial amino acids with free thiol groups can also participate in disulfide bond formation.
[0323] Two or more subunit monomers can be covalently linked by chemical crosslinking. Crosslinkers include homobifunctional crosslinkers, heterobifunctional crosslinkers, and photoreactive crosslinkers. Homobifunctional crosslinkers have identical reactive groups at both ends. Examples of homobifunctional crosslinkers include disuccinimidyl suberate (DSS), disuccinimidyl tartrate (DST), and dithiodisuccinimidyl propionate (DSP). Commonly used thiol-thiol crosslinkers include, for example, BMOE and DTME. Heterobifunctional crosslinkers have two different reactive groups and can be used to link different functional groups. Examples of heterobifunctional crosslinkers include MDS (m-maleimidobenzoic acid-N-hydroxysuccinimide ester), GMBS (N-γ-maleimidobutyryloxysuccinimide ester), EMCS (N-(ε-maleimidocaproyloxy)succinimide ester), and sulfo-EMCS (N-(ε-maleimidocaproyloxy)sulfosuccinimide ester). Photoreactive crosslinkers are heterobifunctional crosslinkers that become reactive only when exposed to ultraviolet or visible light. Two common chemical groups used as photoreactive crosslinkers are aryl azides and bisaziridine compounds. Among these, aryl azides (N-((2-pyridyldithio)ethyl)-4-azidosalicylamide) are widely used. Upon exposure to ultraviolet light between 250 and 350 nm, these agents promote the formation of a nitrene group, which undergoes an addition reaction with the double bond. Furthermore, these cross-linkers can initiate the formation of C-H insertion products or react with nucleophiles. Some commonly used cross-linkers include ANB-NOS (N-5-azido-2-nitrobenzyloxysuccinimide) and thio-SANPAH. Bisaziridinium NHS esters (i.e., azidopentanamide esters) contain a photoactivatable bisaziridinium ring and an N-hydroxysuccinimide (NHS) ester that reacts efficiently with primary amino groups in neutral to alkaline buffers to form stable amide bonds. Compared to phenylazido groups, they exhibit greater photostability and are easily activated by long-wavelength UV light (330–370 nm) to generate carbon guest intermediates that form covalent bonds with any peptide scaffold or amino acid side chain within the distance of a spacer arm.
[0324] More preferably, two or more subunit monomers within the oligomer core can be genetically fused. By encoding the subunit monomers within a single polynucleotide sequence, they can be expressed within a single polypeptide chain, thereby achieving a genetic approach. Accordingly, when the subunit monomers are genetically fused, the oligomer core can comprise a single polypeptide chain.
[0325] The genetically fused subunit monomers can be genetically fused together through a peptide linker. Peptide linkers suitable for connecting subunit monomers are amino acid sequences, and such amino acid sequences include amino acid sequences that can serve as hinge regions between subunit monomers, so that the subunit monomers can fold independently of each other while having flexibility sufficient to enable the subunit monomers to maintain their multimerization capabilities. In general, the length, flexibility and hydrophilicity of the peptide linker are designed so that the subunit monomers can be easily assembled to form an oligomer core. Preferably, the subunit monomers connected by the peptide linker can be assembled to form an oligomer core, wherein, when adjacent subunit monomers are not the same subunit monomers, the interaction between adjacent subunit monomers is substantially the same as the interaction between the same subunit monomers.
[0326] The length of the peptide linker suitable for connecting the monomer subunits of the oligomer core is generally 1 to 100, 1 to 50, 1 to 25, 1 to 20, 1 to 15 or 1 to 10 amino acids. Such linkers can, for example, be composed of one or more of the following amino acids: lysine, serine, arginine, proline, glycine and alanine. Suitable flexible peptide linkers are, for example, amino acid sequence stretches consisting of 2 to 20 (e.g., 4, 6, 8, 10 or 16) serine and / or glycine. Rigid linkers are, for example, amino acid sequence stretches consisting of 2 to 30 (e.g., 4, 6, 8, 16 or 24) proline. Suitable linkers include, but are not limited to, the following linkers: GGGS, PGGS, PGGG, RPPPPP, RPPPP, VGG, RPPG, PPPP, RPPG, PPPPPPP, PPPPPPPPPP, RPPG, GG, GGG, SG, SGSG, SGSGSG, SGSGSGSGSG, SGSGSGSGSG, and SGSGSGSGSGSGSGSG, wherein G represents glycine, P represents proline, R represents arginine, S represents serine, and V represents valine. Other linkers include, for example, GSGS, GGGGS, GGGGSGGGGS, and GGGGSGGGGSGGGGS. Suitable linking groups can be designed using conventional modeling techniques. The flexibility of the linker is generally sufficient to allow its monomers or subunits to assemble into corresponding protein oligomers.
[0327] Preferably, the total molecular weight of the oligomer core is less than about 1000 kDa, such as less than about 500 kDa, for example, less than about 250 kDa. The total molecular weight of the oligomer core is preferably from about 10 kDa to about 1000 kDa, such as from about 10 kDa to about 500 kDa, such as from about 10 kDa to about 250 kDa, such as from about 10 kDa to about 150 kDa. More preferably, the total molecular weight of the oligomer core is from about 20 kDa to about 150 kDa.
[0328] Preferably, the diameter of the oligomer core is less than about 100 nm, such as less than about 50 nm, such as less than about 30 nm, such as less than about 20 nm, such as less than about 10 nm. Preferably, the height of the oligomer core is less than about 100 nm, such as less than about 50 nm, such as less than about 30 nm, such as less than about 20 nm, such as less than about 10 nm. The size of the oligomer core is preferably 1 to 50 nm, such as 2 to 40 nm, such as about 2 nm to about 20 nm, such as about 5 nm to about 10 nm.
[0329] Preferably, the oligomer core is thermodynamically stable. Preferably, the oligomer core is stable at a temperature of about 0°C to about 50°C. That is, preferably, the oligomer core does not spontaneously dissociate into substituted monomers in a solution at a temperature of about 0°C to about 50°C. More preferably, the oligomer core is stable at a temperature of about 10°C to about 40°C, such as about 20°C to about 38°C, for example, about 25°C to about 37°C.
[0330] Preferably, the subunit monomers form an oligomer core through stable interactions. The interactions between the subunit monomers are preferably not weak transient interactions. Weak transient complexes present different oligomeric states that change dynamically in vivo, while strong transient complexes only change their quaternary state, for example, when triggered by ligand binding, and exist in a single major oligomeric state (for example, under standard conditions, at least 90%, for example at least 95%, for example at least 99%, for example at least 99.9%, for example at least 99.99% or 99.999% of the complex can exist in a certain stable oligomeric state). The characteristic of weak transient interactions is that the dissociation constant (K D ) are in the micromolar range and have lifetimes of seconds. Strong transient interactions generally have longer lifetimes and low K in the nanomolar range due to stabilization by effector molecule binding. D value. The subunit monomers more preferably interact with each other at least through a strong transient reaction, and more preferably form a permanent interaction. Permanent interaction means that under normal conditions (for example, in an aqueous solution at a temperature of about 0°C to about 100°C; under these conditions, the oligomer core is usually at least 90%, such as at least 95%, such as at least 99%, such as at least 99.9%, such as at least 99.99% or 99.999% does not dissociate), the oligomer does not dissociate or substantially does not dissociate into the subunit monomers that are its components. The oligomer core generally dissociates only under denaturing conditions that can denature the tertiary structure of the subunit monomer itself. Accordingly, the multimerization K of the subunit monomers DThe value is preferably less than 1 μM, for example less than 100 nM, more preferably less than 10 nM. The life cycle of the oligomer core is generally at least 10 minutes, more preferably at least one hour, for example at least one day, for example at least one week, for example at least one month or at least one year. The life cycle can be measured at any suitable temperature (e.g., about 0°C to about 100°C, for example, about 4°C to about 90°C, for example, about 10°C to about 50°C, for example, about 20°C to about 38°C, for example, about 25°C to about 37°C).
[0331] Preferably, the oligomer core is stable to proteases. For example, the oligomer core may not lose its tertiary or quaternary structure for a limited period of time, such as 4 hours, when exposed to a protease such as trypsin at relatively dilute concentrations.
[0332] Preferably, the oligomer core is a human or humanized oligomer core. The human oligomer core is a polymer region of a human protein. The humanized oligomer core is a non-human protein polymer region that has been modified to be closer to the corresponding human protein polymer region. The amino acid sequence of the humanized oligomer core and the corresponding human protein polymer region may have at least 50% amino acid identity, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98% or at least 99% amino acid identity. The corresponding human protein polymer region is the polymer region of the human protein that has the greatest amino acid sequence identity with the humanized oligomer core. This information can be obtained by searching BLAST (blast.ncbi.nlm.nih.gov) and is limited to human information.
[0333] Preferably, the oligomer core of the multivalent protein scaffold itself does not induce an immune response in a biological system, cell culture, or subject (e.g., a non-human subject or a human subject). That is, generally, since there are no binding sites on the oligomer core and / or the effector portion attached to the oligomer core binding site, when the oligomer core is administered to a biological system, no immune response is induced. For example, after the oligomer core (no binding site on the oligomer core and / or the effector portion attached to the oligomer core binding site) is administered to a biological system (e.g., a subject as described in the present application), the innate or acquired immunity of the biological system is generally not induced. For example, the oligomer core generally does not trigger activation of the complement system, B cells, T cells, natural killer cells, mast cells, basophils, eosinophils, neutrophils, dendritic cells, or macrophages. However, the molecules attached to the oligomer core can be specifically designed to produce an immune response in a human subject.
[0334] Preferably, the oligomeric core of the protein scaffold does not include an antibody or antibody fragment, but as further described in this application, an antibody and / or antibody fragment may be attached to the oligomeric core as an effector moiety. The oligomeric core (e.g., without any effector moiety) more preferably does not include an antibody Fc region. In some cases, the oligomeric core does not include an immunoglobulin constant region.
[0335] In some embodiments, the oligomeric core of the multivalent protein scaffold can be a homo-oligomeric core, that is, the oligomeric core can include only the same monomer. The homo-oligomeric core can include two or more, such as three or more, four or more, five or more, six or more, or seven or more identical subunit monomers.
[0336] In some other embodiments, the oligomeric core of the multivalent protein scaffold can be a hetero-oligomeric core, that is, the oligomeric core can include more than one monomer. Different types of subunit monomers can form the oligomeric core. That is, the subunit monomers can be linked together. The hetero-oligomeric core can include two or more, such as three or more, four or more, five or more, six or more, or seven or more different subunit monomers. For example, when the hetero-oligomeric core includes three monomers, it can include two first monomers and one second monomer; or, it can include one first monomer, one second monomer, and one third monomer. When the hetero-oligomeric core includes four monomers, it can include: two first monomers and two second monomers; or two first monomers, one second monomer, and one third monomer; or one first monomer, one second monomer, one third monomer, and one fourth monomer. Preferably, when the oligomeric core is a hetero-oligomeric core, the oligomeric core includes two subunit monomers. That is, for a hetero-oligomer core comprising species A and species B monomers and having a total of n subunit monomers, the stoichiometric formula of the hetero-oligomer core can be expressed as (A a B b ), wherein a+b=n. When the oligomer core is a hetero-oligomer core, the subunit monomers included therein can be modified so that a first subunit monomer preferentially binds to a second subunit monomer rather than preferentially binding to another first monomer (that is, for a hetero-oligomer core comprising monomers A and B, the hetero-oligomer core has the form of ABABAB... rather than AAABBB...).
[0337] In further embodiments, the oligomeric core may comprise multiple polymeric subunits. For example, two monomers may be fused into one (e.g., by fusion), and the resulting monomer fusion may be assembled to form the oligomeric core. The fused monomers may be the same or different.
[0338] For example, two or more identical monomers can be fused together, and the resulting fusion product can be further assembled with other identical fusion products to form a homo-oligomer core. The oligomer core may include multiple homodimers, where, for example, the homodimer is "AA", the oligomer core may include "AA", "AAAA", "AAAAAA", etc.
[0339] Alternatively, two or more different monomers can be fused together, and the resulting fusion product can be further assembled with other identical fusion products to form an oligomer core. In this application, such oligomer cores are generally referred to as homo-oligomer cores, wherein each fusion product is considered a monomer subunit. Such oligomer cores can include multiple heterodimers, where, for example, if the heterodimer is "AB," the oligomer core can include "AB," "ABAB," "ABABAB," and so on.
[0340] Alternatively, two or more identical monomers can be fused together, and the resulting fusion product can be further assembled with other different fusion products to form a hetero-oligomeric core. The oligomeric core can include multiple homodimers, where, for example, the first homodimer is "AA" and the second homodimer is "BB", and the oligomeric core can include "AABB", etc.
[0341] Alternatively, two or more different monomers can be fused together, and the resulting fusion product can be further assembled with other different fusion products to form a hetero-oligomeric core. The oligomeric core can include multiple heterodimers, where, for example, the first heterodimer is "AB" and the second heterodimer is "CD," and the oligomeric core can include "ABCD," etc.
[0342] The homo-oligomer core comprises a plurality of subunit monomers, wherein each monomer comprises at least one first binding site and at least one second binding site. For example, when the homo-oligomer core comprises three subunit monomers, the oligomer core will comprise at least three first binding sites and at least three second binding sites. When the homo-oligomer core comprises four, five, six or seven subunit monomers, the oligomer core will respectively comprise: at least four first binding sites and at least four second binding sites; or at least five first binding sites and at least five second binding sites; or at least six first binding sites and at least six second binding sites; or at least seven first binding sites and at least seven second binding sites. Therefore, the multivalent protein scaffold preferably comprises at least two first binding sites and at least two second binding sites, that is, wherein all first binding sites are identical to other first binding sites and all second binding sites are identical to other second binding sites. The multivalent protein scaffold more preferably comprises at least three first binding sites and at least three second binding sites. In some embodiments, the multivalent protein scaffold comprises at least four, at least five, at least six, at least seven or at least eight first binding sites and second binding sites.
[0343] The hetero-oligomer core comprises a plurality of subunit monomers, wherein the subunit monomers comprise at least two subunit monomers, wherein the first subunit monomer comprises at least one first binding site, and the second subunit monomer comprises at least one second binding site. For example, when the hetero-oligomer core comprises three subunit monomers, it may comprise three different binding sites; alternatively, it may comprise two first binding sites and one second binding site. When the hetero-oligomer core comprises four subunit monomers, it may comprise four different binding sites; alternatively, it may comprise two first binding sites, one second binding site, and one third binding site; alternatively, it may comprise two first binding sites and two second binding sites.
[0344] The binding sites are described in further detail in this application.
[0345] monomer
[0346] As described herein, the multivalent protein scaffolds provided herein include an oligomeric core comprising a plurality of subunit monomers. The subunit monomers are generally domains of the multidomain polypeptide constructs described elsewhere in this application.
[0347] Each subunit monomer (excluding any binding sites attached thereto, which will be described in further detail herein) preferably comprises less than 300 amino acids, preferably less than 200 amino acids, and more preferably less than 150 amino acids. For example, the molecular weight of each subunit monomer (excluding any binding sites attached thereto) is preferably less than 40 kDa, such as less than 30 kDa, such as less than 20 kDa. Protein scaffolds comprising such monomers as described herein can have a relatively small mass to achieve efficient in vivo diffusion. Such monomers can generally be expressed and correctly folded in bacterial cell expression systems or yeast cell expression systems. Such expression systems can generally achieve yields far higher than those of mammalian cell cultures typically required for antibody production.
[0348] The subunit monomer preferably does not include, or is not composed of, an antibody or antibody fragment. The oligomer core or subunit monomer preferably does not include, or is not composed of, an antibody Fc region. In some embodiments, the subunit monomer does not include, or is not composed of, a CH2 domain. In some embodiments, the subunit monomer does not include, or is not composed of, a CH3 domain. In some embodiments, the subunit monomer does not include, or is not composed of, a CH2 domain or a CH3 domain.
[0349] Preferably, each monomer subunit of the oligomeric core is human or humanized. A human monomer is a monomer of a human oligomeric protein. A humanized monomer is a non-human oligomeric protein monomer that has been modified to more closely resemble a corresponding human protein monomer. Thus, a humanized monomer may have at least 50% amino acid identity with the corresponding human protein amino acid sequence, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% amino acid identity. Generally, a human or humanized protein does not induce a harmful immune response in a patient to whom it is administered.
[0350] Each subunit monomer comprised within the oligomeric core preferably includes a multimerization building block, which is a structural and / or functional portion of the subunit monomer that enables multimerization of the subunit monomer.
[0351] The multimerization structural element can be a protein domain. The multimerization structural element of the oligomer core can be a multimerization domain (or original multimerization domain) of a naturally occurring multimeric protein. A protein domain is an autonomous folding unit of a protein. A multimerization domain is generally a protein domain involved in protein-protein interactions with other protein domains. The multimerization structural unit is preferably a soluble unit that makes the monomer and oligomer core soluble.
[0352] Based on the disclosure of this application, a person skilled in the art should be able to identify multimeric structural units suitable for use in the present invention. For example, a person skilled in the art can identify multimeric proteins. For example, databases such as the NCBI database (www.ncbi.nlm.nih.gov) and the Protein Data Bank (PDB, www.rscb.org) list a variety of multimeric proteins, and multimeric proteins with rotational symmetry axes can be obtained by searching these databases.
[0353] Preferably, the determined multimeric protein is a homo-oligomer, such as a homodimer, a homotrimer, a homotetramer, a homopentamer, a homohexamer, a homoheptamer, etc. The multimeric protein may be a hetero-oligomer, such as a heterodimer, a heterotrimer, a heterotetramer, a heteropentamer, a heterohexamer, a heteroheptamer, etc. The domain of the multimeric protein responsible for multimerization (i.e., the multimerization domain) may be determined through functional and / or structural information.
[0354] The multimerization structural element preferably includes the multimerization interface of the multimerization domain (i.e., the multimerization domain realizes the structural or functional element of the multimerization of each domain). In addition to the multimerization function of the subunit monomer, other aspects of the multimerization domain can be modified. Therefore, the subunit monomer of the oligomer core preferably includes a multimerization structural unit. The subunit monomer and the multimerization domain from which it is derived preferably have at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity. The subunit monomer retains the ability to form a multimer (i.e., the oligomer core).
[0355] Preferably, the subunit monomers of the oligomeric core comprise soluble multimerization structural elements of a multimeric protein. Soluble domains are preferred over intramembrane multimerization domains. Preferably, each subunit monomer of the oligomeric core comprises a soluble multimerization structural element of a soluble multimeric protein.
[0356] The multimeric structural elements can be derived from multimeric proteins with suitable symmetry (such as rotational symmetry or dihedral symmetry, which will be described in further detail in this application), such as collagen (such as collagen NC1 domain), CutA, TNF, p53, fibrinogen, C4, Bacillus subtilis AbrB or their homologs or paralogs.
[0357] Preferably, the subunit monomer may include a monomer or multimerization domain of a protein selected from: collagen X (PDB ID: 1GR3) (such as its NC1 domain); collagen VIII (PDB ID: 1o91) (such as its NC1 domain); C1q head domain (for example, PDB ID: 1PK6 is the globular head of human C1q); CutA protein (copper tolerance protein A), such as from Pyrococcus Horikoshi, human (PDB ID: 2ZFH), Thermus thermophiles (PDB ID: 1V6H), Oryza sativa (PDB ID: 2ZOM) or Shewanella sp. SIB1 (PDB ID: 2ZOM). ID: 3AHP); or a polypeptide having at least 30% or at least 50% amino acid sequence identity with any of the above polypeptides, more preferably a polypeptide having at least 60%, at least 70%, at least 80%, at least 90%, at least 95% or at least 97%, at least 98% or at least 99% amino acid sequence identity with any of the above polypeptides.
[0358] In some embodiments, the subunit monomer may include collagen X (PDB ID: 1GR3 or SEQ ID NO: 2) (such as its NC1 domain), collagen VIII (PDB ID: 1o91 or SEQ ID NO: 3) (such as its NC1 domain), heteromeric C1q head domain (for example, PDB ID: 1PK6 is the globular head of human C1q, see SEQ ID NOs: 36 to 38), CutA protein (copper tolerance protein A) (such as CutA1 protein derived from Pyrococcus Horikoshi (PDB ID: 4YNO or SEQ ID NO: 1), human (PDB ID: 2ZFH or SEQ ID NO: 19), Thermus thermophilus (PDB ID: 1V6H or SEQ ID NO: 39), rice (PDB ID: 2ZOM or SEQ ID NO: 40) or Shewanella SIB1 (PDB ID: 3AHP or SEQ ID NO: 41)), TNF-like protein TL1A (PDB ID: 2RE9 or SEQ ID NO: 50). NO:31), TNF (PDB ID: 1TNF or SEQ ID NO:42; full-length soluble domain SEQ ID NO:80), TNF family protein CD40L (PDB ID: 3LKJ or SEQ ID NO:58), human macrophage migration inhibitory factor (MIF) (PDB ID: 1CA7 or SEQ ID NO:25, or with Y99G mutation PDB ID: 6OY8), human macrophage migration inhibitory factor 2 (MIF2) (PDB ID: 7MSE or SEQ ID NO:26, or with S62A and F99A mutations SEQ ID NO:27) or their homologs or paralogs, or consisting thereof.
[0359] Other multimerization domains include those of the following: antiparallel coiled-coil hexamer (PDB ID: 5WOJ, see Example 4, SEQ ID NO: 43); HIV-1 GP41 core (PDB ID: 1I5Y or SEQ ID NO: 44); cytochrome c555 (PDB ID: 5Z25 or SEQ ID NO: 45); MHC class II-associated chaperone and targeting protein invariant chain (Ii) (PDB ID: 1iie or SEQ ID NO: 46); p53 (PDB ID: 1C26 or SEQ ID NO: 47); fibrinogen-like domain (PDB ID: 4M7F or SEQ ID NO: 48); collagen IV C4 (PDB ID: 1LI1 or SEQ ID NO: 49); Bacillus subtilis AbrB (PDB ID: 1YFB or SEQ ID NO: 51); NO: 50); or a polypeptide having at least 50% amino acid sequence identity with any of the above polypeptides, more preferably a polypeptide having at least 60%, e.g., at least 70%, at least 80%, at least 90%, at least 95% or at least 97%, at least 98% or at least 99% amino acid sequence identity with any of the above polypeptides.
[0360] Other multimerization domains include the multimerization domains of the following: bacteriophage lambda head protein D (such as PDB ID: 1C5E or PDB ID: 1C5E or SEQ ID NO: 51); a domain-swapped trimeric variant of HCRBPII (PDB ID: 6VIS or SEQ ID NO: 52); T1L reovirus attachment protein σ1 (PDB ID: 4ODB or chain A, chain B, chain C of SEQ ID NO: 53); or a polypeptide having at least 50% amino acid sequence identity with any of the above proteins, more preferably a polypeptide having at least 60%, such as at least 70%, at least 80%, at least 90%, at least 95% or at least 97%, at least 98% or at least 99% amino acid sequence identity with any of the above proteins.
[0361] Preferably, in one embodiment, the oligomeric core may comprise monomers derived from the multimeric building blocks of Pyrococcus Horikoshi CutA1. Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 1.
[0362] As described above and elsewhere in this application, CutA1 (eg, Pyrococcus Horikoshi) is a typical domain of a multi-domain polypeptide construct.
[0363] In certain embodiments, the oligomeric core may comprise monomers derived from the multimeric structural elements of human CutA1 (SEQ ID NO: 19). Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 19.
[0364] In certain embodiments, the multimerization building block is a truncated form of SEQ ID NO: 19 (human CutA1), as described elsewhere herein. In certain embodiments, the multimerization building block is a modified form of SEQ ID NO: 19 (human CutA1), wherein at least one cysteine residue in SEQ ID NO: 19 is substituted with a different amino acid residue, as described elsewhere herein.
[0365] Without wishing to be bound by theory, one advantage of CutA1, TNF, OX40L, CDL40L, or TL1A is the reduced multimeric state using a single surface display, even when used as a monospecific assembly (e.g., as a "direct fusion" construct, such as a bispecific therapeutic). For example, the inventors observed that effective inhibition could be achieved by catcher-based display of CutA1 as hexavalent and trivalent assemblies with a trimeric core protein. The reduced multimeric state is expected to reduce molecular weight and / or increase stability.
[0366] In certain embodiments, the oligomeric core may comprise a monomer derived from a multimeric building block of human TNF (SEQ ID NO: 80), OX40L (SEQ ID NO: 78), CD40L (SEQ ID NO: 58), or TL1A (SEQ ID NO: 31). Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 80, 78, 58, or 31.
[0367] In certain embodiments, the multimerization building block is a truncated form of SEQ ID NO: 80, 78, 58, or 31 as described elsewhere herein. An exemplary truncated sequence is SEQ ID NO: 79. In certain embodiments, the multimerization building block is a modified form of SEQ ID NO: 80, 78, 58, or 31, wherein at least one cysteine residue in the sequence is replaced with a different amino acid residue, as described elsewhere herein. In certain embodiments, the multimerization building block is truncated and at least one cysteine residue is substituted relative to SEQ ID NO: 80, 78, 58, or 31.
[0368] Preferably, in another embodiment, the oligomeric core may comprise monomers derived from the multimeric structural unit of collagen X NC1. Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity with the amino acid sequence of SEQ ID NO: 2.
[0369] As described above and elsewhere in this application, collagen X NC1 is a typical domain of a multi-domain polypeptide construct.
[0370] Preferably, in another embodiment, the oligomeric core may comprise monomers derived from the multimeric structural unit of collagen VIIINC 1. Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity with the amino acid sequence of SEQ ID NO: 3.
[0371] As described above and elsewhere in this application, collagen VIIINC1 is a typical domain of a multi-domain polypeptide construct.
[0372] Preferably, in another embodiment, the oligomeric core may comprise monomers derived from the multimeric structural unit of human CutA1. Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity with the amino acid sequence of SEQ ID NO: 19.
[0373] As described above and elsewhere in this application, CutA1 (eg, human) is a typical domain of a multi-domain polypeptide construct. In addition, the inventors have genetically engineered CutA1 to optimize and improve its characteristics and suitability as a domain, as described in the Examples.
[0374] Preferably, in another embodiment, the oligomeric core may comprise monomers derived from the multimeric building blocks of MIF or MIF-2. Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 25, SEQ ID NO: 26 or SEQ ID NO: 27.
[0375] As described above and elsewhere in this application, MIF is a typical domain of a multi-domain polypeptide construct. As described above and elsewhere in this application, MIF-2 is a typical domain of a multi-domain polypeptide construct.
[0376] Preferably, in another embodiment, the oligomeric core may comprise monomers derived from multimeric building blocks of TNF family proteins, including TNF and TNF-like TL1A, OX40L, or CD40L. Accordingly, each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 42, SEQ ID NO: 31, or SEQ ID NO: 58.
[0377] As described above and in other parts of this application, TNF is a typical domain of a multi-domain polypeptide construct. As described above and in other parts of this application, proteins from the TNF superfamily are typical domains of a multi-domain polypeptide construct. The TNF superfamily is a protein superfamily of type II transmembrane proteins that are well-known to contain TNF homology domains and form trimers. This superfamily contains 19 members, including CD40L, OX40L and TNF-α (also referred to herein as "TNF"), which are combined with 29 members of the TNF receptor superfamily. Other members of this superfamily include TNF-b, TNF-g, Fas ligand, CD27 ligand, CD30 ligand, CD137 ligand, CD253, CD354, APO-3L, CD256, CD257, CD258, TL1, TITR ligand and ED1-A1.
[0378] When the oligomer core is a hetero-oligomer core, each monomer within the oligomer core may be derived from the same protein. For example, each monomer within the oligomer core may be derived from one of the aforementioned proteins. Even when the multimerization domains of each monomer subunit are identical, the oligomer core may be a heteromer due to differences in the binding sites attached to the monomer subunits. Even when all monomer subunits are derived from the same protein, the oligomer core may be a heteromer due to differences in the multimerization domains of the monomer subunits.
[0379] It will be understood by those skilled in the art that the monomer subunits of the oligomeric core of the multivalent protein scaffold provided herein can be further varied to provide other functions or beneficial properties.
[0380] For example, a monomer can be a fragment, derivative or variant of a monomer or multimerization structural unit described herein. It will be understood by those skilled in the art that a fragment of an amino acid sequence includes a deletion variant of such a sequence, wherein the number of amino acids deleted is one or more, for example, at least 1, 2, 5, 10, 20, 50 or 100. The deletion position can be the C-terminus or N-terminus of the native sequence, or inside the native sequence. Generally, the deletion of one or more amino acids will not affect the surrounding residues of the adjacent subunit monomer multimerization structural unit.
[0381] Amino acid sequence derivatives include post-translationally modified sequences, including in vivo or in vitro modified sequences. A variety of different protein modification methods are known to those skilled in the art, including methods for introducing new functions into amino acid residues; methods for protecting reactive amino acid residues; and methods for coupling amino acid residues to chemical moieties such as linkers or reactive functional groups on the bottom (surface) to be attached to such amino acid residues.
[0382] Amino acid sequence derivatives also include insertion variants of such sequences, wherein the number of amino acids added or introduced into the native sequence is one or more, for example, at least 1, 2, 5, 10, 20, 50 or 100. The insertion position can be the C-terminus or N-terminus of the native sequence, or within the native sequence. Generally, the insertion of one or more amino acids will not affect the surrounding residues of the adjacent subunit monomer multimerization structural unit.
[0383] Variants of amino acid sequences include sequences in which one or more (e.g., at least 1, 2, 5, 10, 20, 50, or 100) amino acid residues in a native sequence are replaced with one or more non-natural residues. Thus, such variants may include point mutants or may have a greater degree of mutation. For example, a non-natural amino acid sequence may be spliced into a portion of a native sequence by natural chemical ligation to obtain a variant of a native enzyme. Variants of amino acid sequences include both sequences containing natural amino acids and sequences containing non-natural amino acids.
[0384] Variants, derivatives, and functional fragments of the above amino acid sequences generally retain the oligomerization ability of the wild-type sequence. Preferably, the variants, derivatives, and functional fragments of the above sequences have better properties than the wild-type (native) sequence, such as higher stability, lower toxicity, and other functions including binding sites.
[0385] Binding site
[0386] The multivalent protein scaffold comprises at least one first binding site and at least one second binding site. In some embodiments, generally in a modular system that can be used to identify useful effector molecule combinations in the drug discovery process, at least one first binding site is orthogonal to at least one second binding site. That is, the chemical reaction that binds the first binding site to its target (the first target) is orthogonal to the chemical reaction that binds the second binding site to its target (the second target). Thus, the first target binds to the first binding site but not to the second binding site; and the second target binds to the second binding site but not to the first binding site. It can be seen that the meaning of the word "orthogonal" is consistent with its usual meaning in the field of protein-protein interactions, and the first binding interaction (i.e., the interaction of the first binding site with the first ligand) is independent of the second binding interaction (i.e., the interaction of the second binding site with the second ligand).
[0387] The binding sites of the multivalent protein scaffold enable the scaffold to be used as a modular system for binding effector moieties.The first binding site and the second binding site bind to their ligand target on the effector moiety.
[0388] The first binding site and the second binding site can be incorporated into the multivalent protein scaffolds provided herein in any suitable manner. In one embodiment, the first binding site and the second binding site are attached as a tandem fusion to one or each monomer of the oligomer core as described herein to form a multivalent protein scaffold. For comparison purposes, SEQ ID NO: 22 is an example of two binding sites (described herein) provided as a fusion connected by an αH linker.
[0389] The binding site is described in further detail below.
[0390] The interaction between the binding site and its target can be a non-covalent interaction. Preferably, each binding site can form a covalent bond with its corresponding target. The reactive functional groups in the subunit monomer or effector portion can be natural functional groups or functional groups introduced, for example, by genetic manipulation or chemical modification of the monomer. The reactive groups can be derived from non-natural amino acids incorporated into the monomer during synthesis or expression (e.g., cell-free expression achieved by in vitro transcription / translation, etc.).
[0391] The binding sites on the multivalent protein scaffold can bind to their targets via reactive groups. Any suitable reactive group can be used. For example, the reactive group can be an amine reactive group, a carboxyl reactive group, a thiol reactive group, or a carbonyl reactive group. The reactive group can include a cysteine reactive group. The reactive group can include a maleimide, an azide, a thiol, an alkyne, an NHS ester, or a haloacetamide.
[0392] The reactive group may be a group that reacts with an unnatural amino acid, such as 4-azido-L-phenylalanine (Faz) and the amino acids described in Liu C.C. and Schultz P.G. (Annu. Rev. Biochem., 2010, Vol. 79, pp. 413-444). Figure 1 Such groups are particularly useful when the corresponding unnatural amino acid is contained within the binding site and the ligand target.
[0393] The reactive group may be a click chemistry reactive group. The term "click chemistry" was first coined in 2001 by Kolb et al. (Kolb H.C., Finn M.G., and Sharpless K.B.) to describe a class of highly potent, selective, modular components that function reliably in both large and small-scale applications. (Click Chemistry: A Variety of Chemical Functionalities Formed by Several Exemplary Combinations, Angew. Chem. Int. Ed., vol. 40, 2001, pp. 2004-2021). The authors set a strict set of criteria for click chemistry reactions: "Such reactions must be modular; have a wide range; produce high yields; produce only harmless byproducts that can be removed by non-chromatographic methods; and be stereospecific (but not necessarily enantioselective). Required process features include: simple reaction conditions (ideally, such processes should be insensitive to oxygen and water); readily available starting materials and reagents; no solvents, or the use of benign (such as water) or easily removable solvents; and easy product isolation. When purification is required, it must be accomplished by non-chromatographic methods such as crystallization or distillation, and the product must be stable under physiological conditions."
[0394] For example, the first binding site and the second binding site may comprise orthogonal click chemistry reagents. Suitable click chemistry reactions include, but are not limited to, the following reactions:
[0395] (i) A copper-free variant of the 1,3-dipolar cycloaddition reaction, in which an azide reacts with an alkyne under strained conditions (e.g., in the ring of a cyclooctane);
[0396] (ii) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other linker;
[0397] (iii) Staudinger ligation reaction, in which specific reaction with azide can be achieved by replacing the alkyne moiety with an aryl phosphine to produce an amide bond;
[0398] (iv) Azone dipolar cycloaddition reaction;
[0399] (v) norbornene cycloaddition reaction;
[0400] (vi) Oxanorbornadiene cycloaddition reaction;
[0401] (vii) tetrazine ligation reaction;
[0402] (viii) [4+1] cycloaddition reaction;
[0403] (ix) tetrazole photoclick chemistry reaction; and
[0404] (x) Tetracycloheptane ligation reaction.
[0405] The reactive group may be a haloacetamide, such as iodoacetamide, bromoacetamide or chloroacetamide.
[0406] The reactive group may be selected from vinyl, TCO, tetrazine and strained alkyne; DBCO; activated acids such as acid chlorides; and piperazine and reactive amines.
[0407] Host-guest chemistry can also be used to react a binding site with its target. For example, a binding site may include a ligand for binding to a metal complex, while the target may include the metal complex, or vice versa. That is, a binding site may include a metal complex capable of non-covalent interaction via chelation or supramolecular association, while the target may include a site capable of acting as a ligand for stable association with a modifying molecule, or vice versa.
[0408] The reactive group can be any of the reactive groups disclosed in Sakamoto and Hamachi, "Recent Advances in Chemical Modification of Proteins," Anal. Sci., 2019, Vol. 35, pp. 5-27; and McKay and Finn, "Click Chemistry Reactions in Complexation Mixtures: Bioorthogonal Bioconjugation Reactions," Chem. Biol., 2014, Vol. 21, No. 9, pp. 1075-1101. Both of these articles are incorporated herein by reference in their entirety.
[0409] The binding sites of the multivalent protein scaffold preferably include polypeptides such as protein domains. More preferably, the first binding site includes a first protein domain and the second binding site includes a second protein domain.
[0410] When the first binding site includes a first protein domain and the second binding site includes a second protein domain, the first binding site and / or the second binding site are preferably genetically fused to the subunit monomer to which they are connected to form a single polypeptide chain. Generally, the first binding site and / or the second binding site and the subunit monomer to which they are connected are expressed as a single polypeptide chain (e.g., as a fusion protein derived from a recombinant nucleic acid molecule). The benefit of this approach may be that a multivalent protein scaffold can be easily expressed for the binding effector portion without further chemical modification (e.g., for binding to click chemistry reagents). Below, the connection between the protein binding site and the protein to which it is connected (such as the monomer subunit of the oligomeric core of the multivalent protein scaffold) is described.
[0411] The first binding site may include a first protein domain capable of forming a non-covalent bond with the first polypeptide target, and the second binding site may include a second protein domain capable of forming a non-covalent bond with the second polypeptide target. More preferably, the first binding site includes a first protein domain capable of forming a covalent bond with the first polypeptide target, and the second binding site includes a second protein domain capable of forming a covalent bond with the second polypeptide target. The covalent bond formed can be any suitable covalent bond, examples of which are described above. Preferably, the first protein domain is capable of forming an isopeptide bond with the first polypeptide target, and the second protein domain is capable of forming an isopeptide bond with the second binding target. An isopeptide bond is an amide bond that can be formed, for example, between the carboxyl group of one amino acid and the amino group of another amino acid. At least one of such linking groups is generally part of the side chain of one of the above-mentioned amino acids.
[0412] Preferably, the first binding site and the second binding site include different shed protein domains, such as split-type ligand binding protein domains. In the present application, the ligand binding protein domain refers to the domain of the protein that binds to the ligand. Among them, although any suitable protein can be used, proteins stabilized in a natural manner by intrachain covalent bonds such as isopeptide bonds are particularly beneficial. In this case, a portion of the protein containing isopeptide bond donor residues is separated from a portion of the peptide containing isopeptide bond acceptor residues. The two protein fragments can be connected to other polypeptides such as monomers and / or polypeptide targets of the oligomer core as described in the present application by gene fusion. When the two separated fragments are in contact, an isopeptide bond that generally irreversibly connects the two fragments together is produced. Accordingly, the shed protein method is particularly useful in generating binding sites and complementary tags. Since a fragment of a protein preferentially binds to or only to its natural partner (i.e., the complementary part of the protein from which it is derived) compared to any other possible partner, such paired binding sites / tags are generally orthogonal. Such principles can be found, for example, in Reddington and Howarth, Current Opinion in Chemical Biology (Curr. Op. Chem. Biol.), 29, 94-99, 2015; and Keeble et al., Proceedings of the National Academy of Sciences of the United States of America (PNAS), 116, 52, 265-23, 2019).
[0413] Preferably, one of the first and second binding sites comprises a shed Streptococcus pyogenes fibronectin binding protein domain and the other of the first and second binding sites comprises a shed Streptococcus pneumoniae adhesin protein domain.
[0414] Preferably, each of the first protein domain and the first polypeptide target and the second protein domain and the second polypeptide target may include a pair of peptide linkers, such as the peptide linkers disclosed in the following documents: WO 2016 / 193746 A1; WO 2018 / 197854A1; WO 2018 / 189517 A1; Kiber et al., Proceedings of the National Academy of Sciences of the United States of America, Vol. 116, No. 52, 2019, pp. 26523-26533; Fierer et al., Proceedings of the National Academy of Sciences of the United States of America, Vol. 111, No. 13, 2014, pp. E1176-E1181).
[0415] Preferably, each of the first binding site and the second binding site independently has at least 50% amino acid identity with any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18.
[0416] Preferably, one of the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are independently selected from the following combinations: (i) any one of SEQ ID NO:4, 6 or 8 and any one of SEQ ID NO:5, 7 or 9; (ii) SEQ ID NO:12 and SEQ ID NO:13 or 15; (iii) SEQ ID NO:5 and SEQ ID NO:11; (iv) SEQ ID NO:15 and SEQ ID NO:16; (v) SEQ ID NO:17 and SEQ ID NO:18; or (vi) SEQ ID NO:23 and SEQ ID NO:16.
[0417] More preferably, each of the first protein domain and the first polypeptide target and the second protein domain and the second polypeptide target is independently selected from the following pairs:
[0418]
[0419] The protein domain and the targeting domain can have at least 50% amino acid identity (e.g., at least 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99% or 100% amino acid identity) with the above sequences while retaining the ability of the protein domain to specifically bind to the targeting domain.
[0420] In some embodiments, the first binding site is a protein domain, the first target is a tag that binds to the first protein domain; and the second binding site is a protein domain, the second target is a tag that binds to the second protein domain. In some embodiments, the first binding site is a tag, the first target is a protein domain that binds to the first tag; and the second binding site is a tag, the second target is a protein domain that binds to the second tag. In some embodiments, the first binding site is a protein domain, the first target is a tag that binds to the first protein domain; and the second binding site is a tag, the second target is a protein domain that binds to the second tag. Preferably, both the first binding site and the second binding site are protein domains, and the first target and the second target are tags that specifically bind to the first protein domain and the second protein domain, respectively.
[0421] More preferably, each of the first protein domain and the first polypeptide target and the second protein domain and the second polypeptide target is independently selected from the following pairs:
[0422]
[0423] The protein domain and the targeting domain can have at least 50% amino acid identity (e.g., at least 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99% or 100% amino acid identity) with the above sequences while retaining the ability of the protein domain to specifically bind to the targeting domain.
[0424] The above binding groups and targets can be divided into the following subgroups:
[0425] Subgroup A:
[0426] -SpyCatcher(SEQ ID NO:4) / SpyTag(SEQ ID NO:5);
[0427] -SpyCatcher(SEQ ID NO:4) / SpyTag002(SEQ ID NO:7);
[0428] -SpyCatcher(SEQ ID NO:4) / SpyTag003(SEQ ID NO:9);
[0429] -SpyCatcher002(SEQ ID NO:6) / SpyTag(SEQ ID NO:5);
[0430] -SpyCatcher002(SEQ ID NO:6) / SpyTag002(SEQ ID NO:7);
[0431] -SpyCatcher002(SEQ ID NO:6) / SpyTag003(SEQ ID NO:9);
[0432] -SpyCatcher003(SEQ ID NO:8) / SpyTag(SEQ ID NO:5);
[0433] -SpyCatcher003(SEQ ID NO:8) / SpyTag002(SEQ ID NO:7);
[0434] -SpyCatcher003(SEQ ID NO:8) / SpyTag003(SEQ ID NO:9);
[0435] -SpyTag (SEQ ID NO: 5) / K tag (SEQ ID NO: 11) (mediated by SpyLigase: (SEQ ID NO: 10))
[0436] Subgroup B:
[0437] -SnoopCatcher(SEQ ID NO:12) / SnoopTag(SEQ ID NO:13);
[0438] -SnoopCatcher(SEQ ID NO:12) / SnoopTagJr(SEQ ID NO:15);
[0439] -SnoopTagJr (SEQ ID NO: 15) / DogTag (SEQ ID NO: 16) (mediated by SnoopLigase (SEQ ID NO: 14))
[0440] -DogCatcher(SEQ ID NO:23) / DogTag(SEQ ID NO:16)
[0441] Subgroup C:
[0442] - Pilin C (SEQ ID NO: 17) / Isopeptide Tag (SEQ ID NO: 18)
[0443] Preferably, the first binding site / target pair is selected from subgroup A, and the second binding site / target pair is selected from subgroup B and subgroup C; or, the first binding site / target pair is selected from subgroup B, and the second binding site / target pair is selected from subgroup A and subgroup C; or the first binding site / target pair is selected from subgroup C, and the second binding site / target pair is selected from subgroup A and subgroup B.
[0444] Further preferably, the first protein domain / polypeptide target pair and the second protein domain / polypeptide target pair can be selected from the group consisting of: (i) SpyCatcher / SpyTag and SnoopCatcher / SnoopTag; (ii) SpyCatcher002 / SpyTag002 and SnoopCatcher / SnoopTag; (iii) SpyCatcher003 / SpyTag003 and SnoopCatcher / SnoopTag; (iv) SpyCatcher / SpyTag and Pilar Protein (v) SpyCatcher002 / SpyTag002 and Pilin C / Isopeptide Tag; (vi) SpyCatcher003 / SpyTag003 and Pilin C / Isopeptide Tag; (vii) Pilin C / Isopeptide Tag and SnoopCatcher / SnoopTag; (viii) SpyCatcher / SpyTag and SnoopTagJr / DogTag; (ix) SpyCatcher002 / SpyTag002 and SnoopTagJr / DogTag; (x ) SpyTag / K tag and SnoopCatcher / SnoopTag; (xi) SpyTag / K tag and SnoopTagJr / DogTag; (xii) SpyTag / K tag and pilin C / isopeptide tag; (xiii) SnoopTagJr / DogTag and pilin C / isopeptide tag; (xiv) SpyCatcher003 / SpyTag003 and DogCatcher / DogTag; and (xv) SpyCatcher003 / SpyTag002 and DogCatch er / DogTag; (xvi) SpyCatcher003 / SpyTag and DogCatcher / DogTag; (xvii) SpyCatcher003 / SpyTag003 and SnoopCatcher / SnoopTagJ r; (xviii) SpyCatcher003 / SpyTag002 and SnoopCatcher / SnoopTagJr; (xix) SpyCatcher003 / SpyTag and SnoopCatcher / SnoopTagJr.
[0445] It will be understood by those skilled in the art that when SpyLigase / SpyTag / K tag or SnoopLigase / SnoopTagJr / DogTag is used, the connection of the two tags is catalyzed by a "ligase". Accordingly, the first binding site and the first polypeptide target and the second binding site and the second polypeptide target are selected from the two "tags" mentioned above. The ligase can be added exogenously to catalyze the connection of the two "tags", or can be non-covalently or covalently associated with the multivalent protein scaffold (for example, genetically fused to the multivalent protein scaffold). The tags included in the multivalent protein scaffold are interchangeable.
[0446] Other binding site / tag pairs include SdyTag / SdyCatcher (Tan et al., PLOS One, vol. 11, no. 1, e0165074) and the Cpe0147439-563 / Cpe0147565-587 pair derived from the Clostridium perfringens cell surface adhesion protein Cpe0147 (Young et al., Chem Comm., vol. 53, no. 9, p. 1502).
[0447] In this application, when describing the binding between a binding site and its target, "specific binding" refers to the ability of a binding site to bind to its complementary binding site with a greater affinity than when bound to an unrelated control. The unrelated control can be an unrelated control protein. For example, SnoopCatcher specifically binds to SnoopTag with a greater affinity than when bound to an unrelated control protein. The binding is preferably covalent (e.g., forming an isopeptide bond). Preferably, the control protein is bovine serum albumin, and the affinity of the binding site when binding to the complementary binding site is at least 10 times, at least 50 times, at least 100 times, at least 500 times, or at least 1000 times the affinity of the binding site when binding to the control protein. Affinity can be determined by methods known in the art. For example, affinity can be determined by enzyme-linked immunosorbent assay (ELISA), biofilm interferometry, surface plasmon resonance, kinetic methods, or equilibrium / solution methods. One skilled in the art will understand which pairs of binding sites can specifically bind to form protein complexes that can be used in the methods of the present invention.
[0448] The at least one first binding site and the at least one second binding site preferably do not comprise an antibody or antibody fragment. The at least one first binding site and the at least one second binding site more preferably do not comprise an antigen-binding fragment of an antibody, such as a Fab or Fc region.
[0449] For all content in this application involving a "binding site" that binds to a "target", it should be easily understood by those skilled in the art that the various chemical binding groups are interchangeable and reversed, that is, although "binding" is described above as a reaction between the reactive group A of the binding site and the corresponding reactive group B of the target of the binding site, "binding" can also be an equivalent chemical reaction between the reactive group B of the binding site and the reactive group A of the target.
[0450] When the multivalent protein scaffold comprises one or more binding sites that are protein domains (e.g., protein domains that are attached to monomeric subunits of the oligomeric core of the multivalent protein scaffold), the protein domains can be attached to the multivalent protein scaffold (e.g., attached to monomeric subunits of the oligomeric core of the multivalent protein scaffold) by any suitable means.
[0451] The binding site can be connected to the multivalent protein scaffold (e.g., to the monomer subunits of the oligomeric core of the multivalent protein scaffold) via a linker. In one embodiment, the same linker can be used at each end of the subunit monomers of the oligomeric core. In another embodiment, different linkers can be used at each end of the subunit monomers of the oligomeric core.
[0452] The binding site is preferably covalently linked to the oligomer core (or subunit monomer). The covalent bond can be, for example, a peptide bond, a disulfide bond, or a click chemistry bond. More preferably, the covalent bond comprises at least one amino acid (i.e., a peptide linker) and forms part of the same polypeptide chain as the subunit monomer of the binding site.
[0453] The peptide linker used to connect the monomer subunits of the oligomer core of the multivalent protein scaffold to the binding site can be genetically fused to the subunit monomers and / or the binding site. When the linker and the subunit monomers and / or the binding site are expressed as the same construct derived from the same polynucleotide coding sequence, it is called the genetic fusion of the linker. The length, flexibility and hydrophilicity of the peptide linker are generally designed so that each binding site can be located on the same face of the oligomer core or the multivalent protein scaffold. The peptide linker generally can achieve directional binding of the binding site.
[0454] The length of a peptide linker suitable for connecting a binding site to a monomer subunit is generally 1 to 100, 1 to 50, 1 to 25, 1 to 20, 1 to 15 or 1 to 10 amino acids. The linker may, for example, be composed of one or more of the following amino acids: lysine, serine, arginine, proline, glycine and alanine. Suitable flexible peptide linkers are, for example, amino acid sequence stretches consisting of 2 to 20 (e.g., 4, 6, 8, 10 or 16) serine and / or glycine. Rigid linkers are, for example, amino acid sequence stretches consisting of 2 to 30 (e.g., 4, 6, 8, 16 or 24) proline. Suitable linkers include, but are not limited to, the following linkers: GGGS, PGGS, PGGG, RPPPPP, RPPPP, VGG, RPPG, PPPP, RPPG, PPPPPPP, PPPPPPPPPP, RPPG, GG, GGG, SG, SGSG, SGSGSG, SGSGSGSGSG, SGSGSGSGSG, and SGSGSGSGSGSGSGSG, wherein G represents glycine, P represents proline, R represents arginine, S represents serine, and V represents valine. Other linkers include, for example, GSGS, GGGGS, GGGGSGGGGS, and GGGGSGGGGSGGGGS. Suitable linking groups can be designed using conventional modeling techniques. The flexibility of the linker is generally sufficient to allow the binding site and monomer subunits to assemble into corresponding protein oligomers.
[0455] The oligomeric core of the multivalent protein scaffold preferably includes at least one first binding site located at the end of the subunit monomer and at least one second binding site located at the end of the subunit monomer. For example, the first binding site may be located at the first end of the subunit monomer and the second binding site may be located at the second end of the subunit monomer.
[0456] Each end is preferably determined by reference to the end of the subunit monomer of the oligomer core that does not include any joint or binding site. More preferably, the end of the subunit monomer is selected from the N-terminus and / or C-terminus of the subunit monomer. When the binding site (such as a protein domain) constitutes a part of the same polypeptide as the subunit monomer, the N-terminus and C-terminus preferably refer to the amino acids corresponding to the corresponding ends of the monomer that does not contain the binding site. Similarly, when the joint constitutes a part of the same polypeptide as the subunit monomer, the N-terminus and C-terminus preferably refer to the amino acids corresponding to the corresponding ends of the monomer that does not contain the joint.
[0457] Preferably, as described in detail above, the subunit monomer ends to which the binding sites are attached are located on the same face of the oligomeric core or multimeric protein scaffold.
[0458] In some cases, the oligomer core includes only a single end of each subunit monomer located on a given face. In this case, at least one first binding site and at least one second binding site are generally connected to the same end at the same time, thereby being located on the same face of the multivalent protein scaffold. For example, the oligomer core may include a plurality of subunit monomers, wherein each subunit monomer includes a first binding site connected to the first end of the monomer and a second binding site connected to the first binding site. The oligomer core may include a plurality of subunit monomers, wherein at least one first binding site is connected to the first end of the first subunit monomer and at least one second binding site is connected to the second end of the second subunit monomer (such as a hetero-oligomer core). The oligomer core may include a combination of the various binding modes described above.
[0459] Each subunit monomer within the oligomeric core of the multivalent protein scaffold preferably includes two ends located on the same face of the monomer (and therefore located on the same face of the oligomeric core and the multivalent protein scaffold). The two ends are preferably the N-terminus and C-terminus of the monomeric polypeptide. More preferably, each monomer includes a first binding site connected to the first end of the monomer and a second binding site connected to the second end of the monomer. The first binding site and the second binding site can be the N-terminus and the C-terminus, respectively, or the C-terminus and the N-terminus, respectively.
[0460] Sometimes, a monomer may include more than one binding site at each terminus. For example, a subunit monomer may include at each terminus: (i) a first binding site linked to the monomer terminus and a second binding site linked to the first binding site (or vice versa); (ii) a first binding site linked to the monomer terminus and at least one additional first binding site linked to the first binding site; (iii) a second binding site linked to the monomer terminus and at least one additional second binding site linked to the second binding site; (iv) one first binding site or one second binding site linked to the monomer terminus.
[0461] Localization of binding sites on multivalent protein scaffolds
[0462] As described above, in the multivalent protein scaffolds provided herein, at least one first binding site and at least one second binding site are located on the same side of the scaffold. Similarly, in typical embodiments of multi-domain polypeptide constructs, the first binding domain and the second binding domain are located on the same side of the polypeptide construct.
[0463] By being located on the same face of the multivalent protein scaffold (or multi-domain polypeptide construct), the at least one first binding site and the at least one second binding site are arranged in such a manner that the effector moieties that ultimately bind to the multivalent protein scaffold via these binding sites can interact with their corresponding biological targets (e.g., receptors on the cell surface) on the same surface or plane.
[0464] The at least one first binding site and the at least one second binding site are located on the same face of the multivalent protein scaffold (or multi-domain polypeptide construct). Preferably, the at least one first binding site and the at least one second binding site are located on the same face of the oligomer core. Preferably, the at least one first binding site and the at least one second binding site are located on the same face of the subunit monomer to which they are attached. As described herein, a subunit monomer is generally a domain of a multi-domain polypeptide construct.
[0465] The term "at least one first binding site and at least one second binding site are located on the same side of the scaffold" may be combined with Figure 1 and Figure 2 Understand it as follows.
[0466] The multivalent protein scaffold (1) or oligomeric core (10) comprises an imaginary rotational symmetry axis (20) corresponding to the number of monomers in the core. For example, a homotrimeric core comprises a C3 symmetry axis. For example, a homopentameric core comprises a C5 symmetry axis. Similarly, a heterooligomeric core comprises an imaginary rotation axis passing through the center of the oligomeric core and parallel to the interfaces between the subunits, for example: the rotation axis of a heterodimer passes through the oligomeric core and is parallel to the length of the interfaces between the monomers; the rotation axis of a heterotrimer passes through the oligomeric core and is parallel to the length of at least two interfaces between the monomers. A plane (21) can be defined that is perpendicular or approximately perpendicular (e.g., between about 80° and about 100°, such as between about 85° and about 95°, such as between about 88° and about 92°, such as about 90°) to the rotation axis and passing through the center of the oligomeric core. The at least one first binding site (11) and the at least one second binding site (12) are located on the same side of the plane and thus on the same face of the multivalent protein scaffold (1). Figure 1 As shown in the schematic diagram, in the trimer oligomer core, only one first binding site and one second binding site are shown for clarity. Figure 2 The comparative figure shows a situation where the at least one first binding site (11) and the at least one second binding site (12) are located on opposite sides of the plane (21) and thus on opposite sides of the multivalent protein scaffold (1).
[0467] It should be understood by those skilled in the art that when the monomers of the oligomer core are linked, for example, by covalent fusion as described herein, the imaginary symmetry axes are evenly distributed.
[0468] Thus, "the same face of the protein scaffold" can be a solvent-accessible surface of the multivalent protein scaffold on one side of a plane perpendicular to the highest-order rotational symmetry axis of the oligomeric core of the multivalent protein scaffold and passing through the center of the multivalent protein scaffold. Similarly, a face of the oligomeric core can be a solvent-accessible surface of the oligomeric core (which, in this definition, is preferably not associated with a binding site) on one side of a plane perpendicular to the highest-order rotational symmetry axis of the oligomeric core and passing through the center of the oligomeric core.
[0469] Preferably, one side of the multivalent protein scaffold is the solvent accessible portion of the multivalent protein scaffold in contact with a single surface (e.g., a cell surface such as a cell wall, a cell membrane surface, or a protein complex surface). Figure 3 Schematic diagram (for the sake of clarity, the figure shows a plurality of first binding sites and second binding sites (11, 12) attached to an oligomeric core (10) of a multivalent protein scaffold (1), at least one first binding site (11) and at least one second binding site (12) are preferably located on the multivalent protein scaffold (1) in such a way that they can both come into contact with a surface (30). This cannot be achieved in the case where at least one first binding site (11) and at least one second binding site (12) are located on the multivalent protein scaffold (1). Figure 4 This is achieved when the binding sites are located on opposite sides of the multivalent protein scaffold (1) as shown in the schematic diagram, wherein, for example, while the first binding site may be able to contact the surface (30), the second binding site may not be able to contact the surface (30).
[0470] The at least one first binding site and the at least one second binding site are preferably arranged in a double positive orientation, a side positive orientation or any posture therebetween. Figure 5 As shown in the schematic diagram, "double positive orientation" means that the first binding site and the second binding site are both located on the same side of the multivalent protein scaffold, and their connection direction is basically parallel to the rotational symmetry axis of the multivalent protein scaffold; side positive orientation means that, Figure 6 As shown in the schematic diagram, one of the first binding site and the second binding site is located on one side of the multivalent protein scaffold and the connection direction is substantially parallel to the rotational symmetry axis of the multivalent protein scaffold, and the other of the first binding site and the second binding site is located on the same side of the multivalent protein scaffold and the binding direction is substantially perpendicular to the rotational symmetry axis of the multivalent protein scaffold (i.e., substantially parallel to plane (21)). Of course, any posture between these two extremes can also be adopted. For example, Figure 7 As shown in the schematic diagram, the first binding site and / or the second binding site are located on the same face of the multivalent protein scaffold, and the connection direction forms an angle of approximately 45° with the rotational symmetry axis of the multivalent protein scaffold.
[0471] It will be understood by those skilled in the art that the angle between the one or more first binding sites (11) and the axis (20) and the angle between the one or more second binding sites (12) and the axis (20) do not need to be the same. For example, the one or more first binding sites (11) may be located on the "front" of the multivalent protein scaffold, while the one or more second binding sites (12) may be located on the "side" of the multivalent protein scaffold (i.e., the above-mentioned side-positive orientation). Alternatively, the one or more first binding sites (11) may be located on the "side" of the multivalent protein scaffold, while the one or more second binding sites (12) may be located on the "front" of the multivalent protein scaffold (i.e., the above-mentioned side-positive orientation). The one or more second binding sites (12) and the one or more second binding sites (12) may both be located on the "front" of the multivalent protein scaffold (i.e., the above-mentioned double-positive orientation).
[0472] Under the effect that the first binding site and the second binding site are located on the same surface of the multivalent protein scaffold, when the first binding site and the second binding site are simultaneously connected to a given subunit monomer (2), the angle formed between the first binding site and the second binding site and the center of the monomer ( Figure 8 The angle (X) in (I) is typically at most 160°, such as at most 140°, such as at most 120°, such as at most 100° or at most 90°. Typically, the angle formed between the first and second binding sites and the center of the monomer is at least 10°, such as at least 20°, such as at least 30°, such as at least 45° or at least 60°.
[0473] In addition, the visualization of structure can also be achieved by placing the target plane in a three-dimensional coordinate system so that it does not intersect any position of the oligomer core surface determined by protein structure data (NMR, X-ray) or structure prediction methods. For each fusion site, the distance from the target plane shortest path under the premise that the target plane does not intersect with the oligomer core surface (except when the original fusion site intersects) can be determined. Preferably, for a given structure, the target plane position that makes all such shortest paths less than 50%, 45% or 40% of the maximum protein cross-sectional length orthogonal to the target plane can be determined.
[0474] In some embodiments, the maximum shortest path length from the same plane is less than 100 nm, such as less than 50 nm, such as less than 20 nm, such as less than 10 nm, such as less than 5 nm, such as less than 2 nm. In a preferred embodiment, all shortest path lengths from the target plane are within a circular region on the target plane, and the radius of the circular region is less than 50 nm, such as less than 25 nm, such as less than 10 nm, such as less than 5 nm.
[0475] In some embodiments, the cis orientation of a protein fused to a scaffold core can be determined by structure prediction methods. In addition to distance, Alphafold can also take into account the linker geometry and the interaction between the binding domain and the scaffold core protein after fusion. Preferably, the scaffold core is predicted to retain its oligomerization properties even after fusion to the predicted binding site via a linker (preferably a short linker, such as GSGS, such as GGGGS, such as GGGGSGGGGS, such as GGGGSGGGGSGGGGS), and the binding site is predicted to be approximately in a cis geometry.
[0476] Insert Field
[0477] The multivalent protein scaffold comprises an oligomeric core comprising multiple subunit monomers and at least one first binding site orthogonal to at least one second binding site. Preferably, the at least one first binding site and the at least one second binding site are located on the same side of the multivalent protein scaffold. The multivalent protein scaffold may further comprise an insertion domain. The insertion domain is a protein domain. The insertion domain may be located on the same side of the multivalent protein scaffold as the binding site, or on a different side.
[0478] In some cases, the oligomeric core and / or subunit monomer includes at least one free end that is not connected to the binding site, and the multivalent protein scaffold includes an insertion domain located at the free end. The insertion domain may also be located within the loop region of the oligomeric core, and thus may be located within the loop region of the subunit monomer. Preferably, the multimeric protein scaffold includes at least one insertion domain located on the side of the oligomeric core and / or multimeric protein scaffold opposite to the side where the binding site is located (e.g., within 90° of the opposite end of the axis).
[0479] In this application, an insertion domain refers to a polypeptide sequence encoding a protein domain, which is an autonomous folding functional unit of a protein. The insertion domain does not interfere with the above structure and the folding of the oligomer core or binding site.
[0480] The insertion domain preferably has an effector function. The insertion domain may include an antibody, an antibody fragment, or an antigen binding fragment, such as an antigen binding fragment that can bind to CD3 or CD16. For example, the insertion domain can be combined with an immunomodulatory protein such as a cytokine, a chemotherapeutic agent, or a cancer immunotherapy (i.e., a therapy that uses the immune system of the subject to be treated to treat cancer) agent. The insertion domain may constitute a protein that induces cell death when in contact with a biological system. The insertion domain may induce apoptosis, enhance anti-tumor response, or have other beneficial activities. The insertion domain may have complement inhibition or complement stimulation activity.
[0481] For the avoidance of doubt, it is explicitly stated here that an inserted domain is typically inserted within a domain of a multi-domain polypeptide construct as described herein.
[0482] protein complexes
[0483] The present application also provides a protein complex comprising a multivalent protein scaffold, as described in further detail herein, linked to at least one first effector moiety and at least one second effector moiety. Each first effector moiety is linked to a first target that is bound to a first binding site on the multivalent protein scaffold. Each second effector moiety is linked to a second target that is bound to a second binding site on the multivalent protein scaffold.
[0484] The target is preferably a polypeptide target, and more preferably a partner of the above-mentioned paired peptide linker.
[0485] Each effector portion is bound to the multivalent protein scaffold by being connected to the target. The first effector portion and the second effector portion may be the same or different, and are preferably different. Among them, any connection approach described above for the binding site can be adopted. Each effector portion can be directly connected to the target through traditional organic chemical reaction pathways available to those skilled in the art, thereby binding to the multivalent protein scaffold. Suitable chemical methods are shown in textbooks such as Advanced Organic Chemistry (Wiley, 2020) by March.
[0486] Preferably, each effector moiety is covalently linked to the target. More preferably, the target is a polypeptide target and can be genetically fused to the effector moiety. That is, preferably, one or each effector moiety is genetically fused to the polypeptide target by being encoded in the same polynucleotide as the polypeptide target in such a way that it is expressed as the same polypeptide chain as the polypeptide target. The effector moiety can be genetically fused to a first polypeptide target, a cleavage site, and a second polypeptide target, wherein the first polypeptide target is orthogonal to the second polypeptide target. The cleavage site can be a TEV cleavage site. When the first polypeptide target and the second polypeptide target are simultaneously present on the effector moiety, only the polypeptide target at the end is functional (i.e., able to bind to its cognate binding site on the multivalent protein scaffold). The terminal polypeptide target can be separated by the cleavage site so that only a single target is present. After the effector moiety is fully coupled, the terminal polypeptide target can be further specifically used.
[0487] The effector moiety is preferably a protein domain. The protein domain is preferably a soluble protein domain. The protein domain preferably comprises a domain of a secreted protein or an extracellular domain of a transmembrane protein. More preferably, the protein domain comprises an extracellular domain of a cell surface receptor, such as a human cell surface receptor, or a ligand for such a cell surface receptor.
[0488] The effector portion is preferably a portion that exerts a therapeutic effect when in contact with a biological system. The effector portion may, for example, be an immunomodulatory protein such as a cytokine, a chemotherapeutic agent, or a cancer immunotherapy (i.e., a therapy that utilizes the immune system of the subject to be treated for cancer) agent. The effector portion may induce cell death when in contact with a biological system. The effector portion may induce apoptosis, enhance anti-tumor response, or have other beneficial activities. The effector portion may have complement inhibition or complement stimulation activity. The effector portion may cause changes in gene expression, receptor internalization, cytokine release, cell death, or sensitivity to therapeutic agent molecules.
[0489] In one embodiment, the effector moiety can be a synthetic organic or inorganic molecule. A suitable molecule can be a chemotherapeutic agent. A suitable molecule can be a toxic agent, such as one having an EC50 of less than about 100 μM, for example, less than about 10 μM, for example, less than about 1 μM or less than about 100 nM. Wherein, the EC50 is the concentration of the toxic agent required to cause 50% cytotoxicity when evaluated in a suitable cell-based assay. A suitable cell-based assay can be, for example, a sulforhodamine B (SRB) assay.
[0490] Suitable synthetic molecules may be enzyme activators or enzyme inhibitors. Suitable molecules may be inhibitors of one or more of serine / threonine / tyrosine kinases, matrix metalloproteinases (MMPs), heat shock proteins (HSPs), and proteasomes. Suitable molecules can serve as: alkylating agents (e.g., nitrogen mustards, nitrosoureas, tetrazines, aziridines, cisplatin and its derivatives); antimetabolites (e.g., antifolates, fluoropyrimidines, deoxynucleoside analogs and thiopurines); antimicrotubule agents (e.g., vinca alkaloids or taxanes); topoisomerase inhibitors (e.g., topoisomerase I inhibitors such as irinotecan and topotecan, topoisomerase II inhibitors such as etoposide, doxorubicin, mitoxantrone and teniposide, or topoisomerase II inhibitors such as novobiocin, merbarone and aclarubicin); or cytotoxic antibiotics (e.g., anthracyclines and bleomycin). Suitable molecules may have a molecular weight of about 50 g / mol to about 5000 g / mol, such as about 100 g / mol to about 1000 g / mol, for example about 250 g / mol to about 500 g / mol.
[0491] In another embodiment, the effector moiety preferably comprises an antibody or antigen-binding fragment thereof. As used herein, the term "antibody or antigen-binding fragment thereof" in relation to an effector moiety may refer to a complete antibody (i.e., each unit comprising two heavy chains and two light chains interconnected by disulfide bonds) and antigen-binding fragments thereof. Antibodies generally comprise the immunologically active portion of an immunoglobulin (Ig) molecule (i.e., a molecule containing an antigen-binding site that specifically binds (immunoreacts) with an antigen). As used herein, the terms "specifically bind" or "immunoreacts" when describing the interaction of an antibody or fragment thereof with an antigen indicate that the antibody preferentially reacts with one or more antigenic determinants of the target antigen compared to other polypeptides. Each heavy chain is composed of a heavy chain variable region (abbreviated herein as HCVR or VH), i.e., at least one heavy chain constant region. Each light chain is composed of a light chain variable region (abbreviated herein as LCVR or VL) and a light chain constant region. The variable regions of the heavy and light chains contain the binding domain that interacts with the antigen. The VH and VL regions can be further subdivided into hypervariable regions, complementarity determining regions (CDRs), and framework regions (FRs). The complementarity determining regions are interspersed within the framework regions, which are more conserved. Antibodies may include, but are not limited to, polyclonal antibodies, monoclonal antibodies, chimeric antibodies, dAb antibodies (single domain antibodies), single-chain antibodies, Fab, Fab' and F(ab')2 fragments, scFV, and Fab expression libraries. Antibodies may be selected, for example, from the group consisting of: single-chain antibodies; single-chain variable fragments (scFv); variable fragments (Fv); antigen-binding fragments (Fab); recombinant antibodies; monoclonal antibodies; fusion proteins comprising the antigen-binding domains of natural antibodies or nucleic acid aptamers; single-domain antibodies (sdAbs), also known as VHH antibodies; nanobodies (single domain antibodies derived from camelids); single-domain antibody fragments derived from shark IgNARs, known as VNARs; single-chain antibody dimers (diabodies); single-chain antibody trimers (triabodies); anticalins; nucleic acid aptamers (DNA or RNA aptamers); or active ingredients or fragments thereof.
[0492] A "Fab fragment" (also called an antigen-binding fragment or Fab region) comprises the constant domain of the light chain (CL) and the first constant domain of the heavy chain (CH1), as well as the variable domains VL and VH, located on the light and heavy chains, respectively. The variable domains include the complementarity-determining loops (CDRs, also called hypervariable regions) involved in antigen binding. Fab' fragments differ from Fab fragments by the addition of several residues to the carboxyl terminus of the heavy chain CH1 domain, including one or more cysteines from the antibody hinge region.
[0493] A "single-chain Fv" (scFv) comprises the VH and VL domains of an antibody, wherein these domains are present within a single polypeptide chain. In one embodiment, the Fv polypeptide further comprises a polypeptide linker between the VH and VL domains, which enables the scFv to form the required structure for antigen binding. For a review of scFv, see Pluckthun, Pharmacology of Monoclonal Antibodies (Rosenburg and Moore, ed., Vol. 113, Springer-Verlag, New York, pp. 269-315, 1994). Examples of scFv fragments include antibody scFv fragments described in WO93 / 16185, U.S. Patent No. 5,571,894, and U.S. Patent No. 5,587,458.
[0494] The effector moiety can be a Fab region of a therapeutic antibody. For example, the effector moiety can be a Fab region of a monoclonal antibody, such as muromomab, abciximab, rituximab, daclizumab, basiliximab, palivizumab, infliximab, trastuzumab, etanercept, gemtuzumab, alemtuzumab, ibritumomab, tiuximab, clopidogrel, clopidogrel, clopidogrel, trastuzumab, etanercept, gemtuzumab, alemtuzumab, tiuximab ...tiuximab, trastuzumab, etanercept, gemtuzumab, alemtuzumab, ibritomomab, adalimumab, alefacept, omalizumab, tositumomab, efalizumab, cetuximab, bevacizumab, natalizumab, ranibizumab, panitumumab, eculizumab, or certolizumab.
[0495] The effector moiety can target any receptor associated with a pathological condition, such as those described herein. For example, the effector moiety can target any receptor that confers a clinical benefit by binding, such as a hormone receptor.
[0496] In some embodiments, an effector moiety may have a target (eg, a receptor) that was not previously known to be associated with a pathological condition and that, for example, has been found to provide therapeutic benefit by targeting it in certain circumstances.
[0497] It will be appreciated by those skilled in the art that the protein complexes provided herein can be used to simultaneously bind to two targets within a biological system, thereby enabling simultaneous contact. The targets can, for example, originate from the same cell. The protein complexes provided herein can, for example, be used to bind to two different receptors on the surface of the same cell.
[0498] The protein complexes provided herein generally include multiple first binding sites and multiple second binding sites on a multivalent protein scaffold, and can therefore bind to multiple first effector moieties and multiple second effector moieties. This is particularly beneficial because such "high titer" compounds may enhance effector function or achieve unprecedented effector function. It has been previously found that when multiple clones of a single effector moiety are contacted with a biological system, the therapeutic response can be improved (e.g., Brunet et al. (see above); and Carey Anuar et al. (Nature Communications, October 1, 2019, pp. 1-13). However, contacting multiple copies of multiple different effector moieties is a complex technical problem, and the protein complexes provided herein can solve this problem.
[0499] In some embodiments, the effector function may be achieved only when the combination of effector moieties interacts with the biological system to which they are contacted. For example, a first effector moiety (e.g., an effector moiety attached to a first binding site) may exhibit a therapeutic effect only when combined with a second effector moiety (e.g., an effector moiety attached to a second binding site), whereas neither the first effector moiety nor the second effector moiety alone has a therapeutic effect.
[0500] It will be appreciated by those skilled in the art that the platform and method of the present application can screen new effector moiety combinations useful in therapy and identify useful candidate combinations.
[0501] Screening Platform
[0502] The present application also provides a screening platform. The screening platform comprises a library, wherein the library comprises a plurality of populations of protein complexes of the present invention. Each population of protein complexes comprises a different combination of a first effector moiety, a second effector moiety, and / or an oligomer core. The present application also provides such a library.
[0503] Libraries can be used, for example, to screen for new effector moiety combinations. Accordingly, a library can include multiple samples of different protein complexes. Each sample can be a homogenous sample, that is, each sample can contain only one type of protein complex. Each sample can be different from all other samples. That is, each sample contains a protein complex that includes a combination of a first effector moiety and a second effector moiety that is different from the combination of the first effector moiety and the second effector moiety included in the protein complexes of all other samples. A library can, for example, include from about 1 or 2 to about 1,000,000 samples, such as from about 10 to about 100,000 samples, such as from about 50 to about 50,000 samples, such as from about 100 to about 10,000 samples, such as from about 500 to about 1,000 samples. Each sample can include a different type of protein complex, wherein the protein complex in each sample has a different combination of the first effector moiety, the second effector moiety, and the oligomer core than the protein complexes of all other samples.
[0504] In some embodiments, the library can be a "one-dimensional" library. In some embodiments, all samples in the library may have the same or substantially the same oligomer core and first effector moiety, and may differ from each other in the second effector moiety. In other embodiments, all samples in the library may have the same or substantially the same oligomer core and second effector moiety, and may differ from each other in the first effector moiety. In some other embodiments, all samples in the library may have the same or substantially the same first effector moiety and second effector moiety, and may differ from each other in the oligomer core. A polypeptide substantially identical to a given polypeptide (such as an oligomer core or a first polypeptide binding site or a second polypeptide binding site) may, for example, have at least 90% sequence identity, such as at least 95% sequence identity, such as at least 97%, 98%, 99%, 99.9% or 99.99% sequence identity with the given polypeptide. A polypeptide substantially identical to a given polypeptide (such as an oligomer core or a first polypeptide binding site or a second polypeptide binding site) may differ from the given polypeptide, for example, by including one or more sequence additions, deletions, insertions or variations as described herein. A polypeptide substantially identical to a given polypeptide (e.g., an oligomer core or a first polypeptide binding site or a second polypeptide binding site) may differ from the given polypeptide, for example, in that the polypeptide has undergone a post-translational modification, such as a modification of its glycosylation or phosphorylation pattern.
[0505] In some embodiments, the library can be a "two-dimensional" library. In some embodiments, all samples in the library can have the same or substantially the same oligomer core and can differ from each other in the combination of the first effector moiety and the second effector moiety. In other embodiments, all samples in the library can have the same or substantially the same first effector moiety and can differ from each other in the combination of the oligomer core and the second effector moiety. In some other embodiments, all samples in the library can have the same or substantially the same second effector moiety and can differ from each other in the combination of the oligomer core and the first effector moiety.
[0506] In some embodiments, the library can be a "three-dimensional" library, wherein, in some embodiments, all samples within the library can differ from each other in the combination of oligomer core, first effector moiety, and second effector moiety.
[0507] In addition to the library, the screening platform may also include other components. For example, the screening platform may include any or all of the following components:
[0508] - Biological systems that come into contact with samples in the library;
[0509] - a detection system for detecting changes in a biological system caused by contact between the biological system and a sample in a reservoir;
[0510] - Reagents and / or buffer solutions; and
[0511] - Optical, electrical or spectroscopic means for detecting changes reported by a detection system.
[0512] The biological system can be a cell culture, such as a mammalian cell culture, preferably a human cell culture, more preferably an immune cell culture and / or a cancer cell line culture. The biological system can be a biological sample, such as a blood sample, a serum sample, a plasma sample, or a tissue or organ sample. Biological samples include tumor samples, cells, cell lysates, urine, amniotic fluid, and other biological fluids. The biological sample is preferably a mammalian sample. The sample can be a human or non-human sample.
[0513] The detection system can be any suitable detection system. The detection system can be a dye or stain, such as a cell viability stain. Suitable stains can include, for example, trypan blue, fluorescein diacetate (green), propidium iodide, and Hoechst 33258.
[0514] Reagents include components required for cell viability, including components of cell growth medium, and may include therapeutic molecules.
[0515] Buffers include aqueous compositions that may, for example, contain buffer salts. Preferred buffer salts that may be used include: Tris; phosphate; citric acid / Na2HPO4; citric acid / sodium citrate; sodium acetate / acetic acid; Na2HPO4 / NaH2PO4; imidazole (isopyrazole) / HCl; sodium carbonate / sodium bicarbonate; ammonium carbonate / ammonium bicarbonate; MES; Bis-Tris; ADA; ACES; PIPES; MOPSO; Bis-Tris propane; BES; MOPS; TES; HEPES; DIPSO; MOBS; TAPSO; Trizma; HEPPSO; POPSO; TEA; EPPS; Tris (hydroxymethyl) methylglycine (Tricine); Glycylglycine (Gly-Gly); N,N-bis (2-hydroxyethyl) glycine (Bicine); HEPBS; TAPS; AMPD; TABS; AMPSO; CHES; CAPSO; AMP; CAPS; and CABS. Achieving the desired pH value by selecting an appropriate buffer is routine practice for those skilled in the art. Guidance can be found, for example, at http: / / www.sigmaaldrich.com / life-science / core-bioreagents / biological-buffers / learning-center / buffer-reference-center.html. The buffer salt is preferably used in solution at a concentration of 1 mM to 1 M, preferably 10 mM to 100 mM, for example, approximately 50 mM.
[0516] Devices for detecting changes reported by the detection system include: microscopes (optical or electronic); electrical devices such as electrophysiological devices (such as patch clamps); and spectroscopic devices, such as equipment for UV / VIS spectroscopy, NMR spectroscopy, mass spectroscopy, infrared spectroscopy, Raman spectroscopy, circular dichroism spectroscopy, etc.
[0517] method
[0518] Also provided is a method for identifying a therapeutic drug analog, the method comprising:
[0519] Providing a protein complex as described herein;
[0520] contacting the protein complex with a biological system; and
[0521] Measure whether a protein complex causes a desired change in the properties of a biological system.
[0522] Optionally, the method may further comprise selecting a protein complex that causes a desired change in a property of the biological system.
[0523] Also provided is a method of identifying a therapeutic combination of effector molecules (e.g., antigen binding domains), the method comprising:
[0524] Providing a protein complex as described herein;
[0525] contacting the protein complex with a biological system; and
[0526] Measure whether a protein complex causes a desired change in the properties of a biological system.
[0527] The biological system may be a cell culture, such as a mammalian cell culture, preferably a human cell culture, more preferably an immune cell culture and / or a cancer cell line culture.
[0528] The biological system can be a biological sample such as a blood sample, a serum sample, a plasma sample, or a tissue or organ sample. Biological samples include tumor samples, cells, cell lysates, urine, amniotic fluid, and other biological fluids. The biological sample is preferably a mammalian sample. The sample can be a human or non-human sample.
[0529] The change in the property of a biological system can be any change that is relevant to the desired activity of the target therapeutic agent. In some embodiments, the desired change is cell death. This situation can be particularly useful in the development of cancer therapeutics.
[0530] Other changes include changes in effector function. Accordingly, the method may include the step of measuring whether the protein complex elicits an effector function in the biological system.
[0531] Effector function changes can include changes in gene expression, changes in protein modification functions such as phosphorylation, receptor internalization, cytokine release, cell death, or sensitivity to therapeutic molecules. The effector function can be affinity binding to a biological sample, and such binding can be measured by various techniques such as ELISA. Affinity binding to a target biological system can be as described above, so that the effector domain specifically acts on the target biological system, which can be a specific cell type such as a cancer cell.
[0532] Effector function can be assessed by comparison with a control, which can be a protein complex without the effector moiety.
[0533] The control can be a single protein complex with only one type of effector moiety attached to the effector moiety (i.e., only one type of effector moiety is attached to the multivalent protein scaffold). In this case, the method can be used to identify effector moieties that have "synergistic function" or "synergistic biological function." "Synergistic function" or "synergistic biological function" refers to the following effector function or level of effector function: the fusion protein components alone do not have this function, but only when used in a bispecific multivalent protein complex; or the activity is greater or less than when the first and second effector moieties of the protein complex are used alone, that is, the activity is only achieved when the two effector moieties are used together in the complex.
[0534] The method may further comprise the step of determining a molecule of the biological system that is bound by the effector moiety of the protein complex. The method may preferably comprise selecting an effector moiety combination that specifically binds to the same molecule of the biological system as the selected protein complex (e.g. the effector moiety itself).
[0535] The method may further include: synthesizing a therapeutic drug candidate or therapeutic drug or its analogue comprising a selected effector portion combination. The therapeutic drug candidate or therapeutic drug may include an oligomer core and an effector portion of a therapeutic drug analogue, wherein, however, the functional substitution of the binding site and the target is a covalent bond, such as the gene fusion key described in further detail herein. The therapeutic drug or drug candidate may include an oligomer core identical to the therapeutic drug analogue identified with the disclosed method. Alternatively, the therapeutic drug candidate may include an oligomer core different from the therapeutic drug analogue identified with the disclosed method. The therapeutic drug candidate may have an oligomer core selected or designed to confer other therapeutic benefits (such as other effector functions).
[0536] The present application also provides a therapeutic drug candidate obtainable according to the disclosed method.
[0537] therapeutic drug candidates, therapeutic drugs
[0538] Also provided is a candidate therapeutic agent comprising an oligomeric core comprising a plurality of subunit monomers linked to one or more first effector moieties and one or more second effector moieties, wherein the one or more first effector moieties and the one or more second effector moieties are located on the same face of the oligomeric core, and wherein: (1) the one or more first effector moieties comprise two or more first effector moieties, and the one or more second effector moieties comprise two or more second effector moieties; and / or the oligomeric core does not comprise an antibody or antibody fragment.
[0539] Also provided is a therapeutic drug having the same characteristics.
[0540] In general, the oligomer core is an oligomer core as described in further detail herein. In general, the subunit monomers are as described in further detail herein. In general, the first effector moiety and the second effector moiety are as described in further detail herein. The first effector moiety and the second effector moiety can be connected to the subunit monomers of the oligomer core by any suitable means, including any connection method described herein. In some embodiments, the connection between the first effector moiety and the second effector moiety includes the first binding site and the second binding site and the first polypeptide target and the second polypeptide target as described herein. However, in other embodiments, the connection between the first effector moiety and the second effector moiety does not include the first binding site and the second binding site and the first polypeptide target and the second polypeptide target as described herein, but may include simple covalent connections such as gene fusion bonds and / or click chemistry bonds as described herein.
[0541] Specific implementation plan
[0542] In a first preferred aspect, the present application provides:
[0543] - A multivalent protein scaffold comprising an oligomeric core, the oligomeric core comprising a plurality of (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 30% or at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 1; wherein each monomer comprises a first binding site and a second binding site, wherein the first binding site is orthogonal to the second binding site, and wherein each of the first binding site and the second binding site independently has an amino acid sequence identical to SEQ ID NO: 1. NO: 4-9, 11-13, 23 or 15-18 have at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity); and wherein each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first binding site and the second binding site has at least 50% amino acid identity with SEQ ID NO: 4, 6 or 8, and the other has at least 50% amino acid identity with SEQ ID NO: 12. A preferred multivalent protein scaffold of the present invention comprises a monomer consisting of SEQ ID NO: 21 or a fragment thereof (e.g., including residues 14 to 348).
[0544] - A protein complex comprising the multivalent protein scaffold of the first aspect, wherein the first binding site binds to a first polypeptide target linked to a first effector moiety; the second binding site binds to the first polypeptide target linked to a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and wherein each of the first binding site / polypeptide target pair and the second binding site / polypeptide target pair is independently selected from the following combinations: (i) any one of SEQ ID NO: 4, 6 or 8 and any one of SEQ ID NO: 5, 7 or 9; (ii) SEQ ID NO: 12 and SEQ ID NO: 13 or 15; (iii) SEQ ID NO: 5 and SEQ ID NO: 11; (iv) SEQ ID NO: 15 and SEQ ID NO: 16; (v) SEQ ID NO: 17 and SEQ ID NO: 18; or (vi) SEQ ID NO: 23 and SEQ ID NO: 16.
[0545] - A screening platform comprising a library comprising a plurality of sets of protein complexes according to the first aspect, wherein each set comprises a different combination of a first effector moiety and a second effector moiety.
[0546] - A therapeutic drug candidate, comprising an oligomeric core, the oligomeric core comprising a plurality of (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 1; wherein each monomer is directly linked (e.g., genetically fused) to a first effector moiety and a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and, preferably, wherein each monomer is directly linked to the first effector moiety and the second effector moiety to which it is linked via a polypeptide linker.
[0547] In the second preferred aspect, the present application specifically provides:
[0548] - A multivalent protein scaffold comprising an oligomeric core, the oligomeric core comprising a plurality (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 2; wherein each monomer comprises a first binding site and a second binding site, wherein the first binding site is orthogonal to the second binding site, and wherein each of the first binding site and the second binding site independently has at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with any one of SEQ ID NOs: 4-9, 11-13, 23 or 15-18; and wherein each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first binding site and the second binding site has at least 50% amino acid identity with SEQ ID NO: 4, 6 or 8, and the other has at least 50% amino acid identity with SEQ ID NO: 12. A preferred multivalent protein scaffold of the present invention comprises a monomer consisting of SEQ ID NO: 20 or a fragment thereof (e.g., including residues 14 to 380).
[0549] - A protein complex comprising the multivalent protein scaffold of the second aspect, wherein the first binding site binds to a first polypeptide target linked to a first effector moiety; the second binding site binds to the first polypeptide target linked to a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and wherein each of the first binding site / polypeptide target pair and the second binding site / polypeptide target pair is independently selected from the following combinations: (i) any one of SEQ ID NO: 4, 6 or 8 and any one of SEQ ID NO: 5, 7 or 9; (ii) SEQ ID NO: 12 and SEQ ID NO: 13 or 15; (iii) SEQ ID NO: 5 and SEQ ID NO: 11; (iv) SEQ ID NO: 15 and SEQ ID NO: 16; (v) SEQ ID NO: 17 and SEQ ID NO: 18; or (vi) SEQ ID NO: 23 and SEQ ID NO: 16.
[0550] - A screening platform comprising a library comprising a plurality of sets of protein complexes according to the second aspect, wherein each set comprises a different combination of a first effector moiety and a second effector moiety.
[0551] - A therapeutic drug candidate, comprising an oligomeric core, the oligomeric core comprising a plurality of (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 2; wherein each monomer is directly linked (e.g., genetically fused) to a first effector moiety and a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and, preferably, wherein each monomer is directly linked to the first effector moiety and the second effector moiety to which it is linked via a polypeptide linker.
[0552] In the third preferred aspect, the present application specifically provides:
[0553] - A multivalent protein scaffold comprising an oligomeric core, the oligomeric core comprising a plurality (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 3; wherein each monomer comprises a first binding site and a second binding site, wherein the first binding site is orthogonal to the second binding site, and wherein each of the first binding site and the second binding site independently has at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with any one of SEQ ID NOs: 4-9, 11-13, 23 or 15-18; and wherein each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first binding site and the second binding site has at least 50% amino acid identity with SEQ ID NO:4, 6 or 8, and the other has at least 50% amino acid identity with SEQ ID NO:12.
[0554] - A protein complex comprising the multivalent protein scaffold of the third aspect, wherein the first binding site binds to a first polypeptide target linked to a first effector moiety; the second binding site binds to the first polypeptide target linked to a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and wherein each of the first binding site / polypeptide target pair and the second binding site / polypeptide target pair is independently selected from the following combinations: (i) any one of SEQ ID NO: 4, 6 or 8 and any one of SEQ ID NO: 5, 7 or 9; (ii) SEQ ID NO: 12 and SEQ ID NO: 13 or 15; (iii) SEQ ID NO: 5 and SEQ ID NO: 11; (iv) SEQ ID NO: 15 and SEQ ID NO: 16; (v) SEQ ID NO: 17 and SEQ ID NO: 18; or (vi) SEQ ID NO: 23 and SEQ ID NO: 16.
[0555] - A screening platform comprising a library comprising a plurality of sets of protein complexes according to the third aspect, wherein each set comprises a different combination of a first effector moiety and a second effector moiety.
[0556] - A therapeutic drug candidate, comprising an oligomeric core, wherein the oligomeric core comprises a plurality of (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 3; wherein each monomer is directly linked to a first effector moiety and a second effector moiety (e.g., genetically fused); wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and, preferably, wherein each monomer is directly linked to the first effector moiety and the second effector moiety to which it is linked via a polypeptide linker.
[0557] In a fourth preferred aspect, the present application specifically provides:
[0558] - A multivalent protein scaffold comprising an oligomeric core, the oligomeric core comprising a plurality (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 19; wherein each monomer comprises a first binding site and a second binding site, wherein the first binding site is orthogonal to the second binding site, and wherein each of the first binding site and the second binding site independently has at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with any one of SEQ ID NOs: 4-9, 11-13, 23 or 15-18; and wherein each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first binding site and the second binding site has at least 50% amino acid identity with SEQ ID NO:4, 6 or 8, and the other has at least 50% amino acid identity with SEQ ID NO:12.
[0559] - A protein complex comprising the multivalent protein scaffold of the fourth aspect, wherein the first binding site binds to a first polypeptide target linked to a first effector moiety; the second binding site binds to the first polypeptide target linked to a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and wherein each of the first binding site / polypeptide target pair and the second binding site / polypeptide target pair is independently selected from the following combinations: (i) any one of SEQ ID NO: 4, 6 or 8 and any one of SEQ ID NO: 5, 7 or 9; (ii) SEQ ID NO: 12 and SEQ ID NO: 13 or 15; (iii) SEQ ID NO: 5 and SEQ ID NO: 11; (iv) SEQ ID NO: 15 and SEQ ID NO: 16; (v) SEQ ID NO: 17 and SEQ ID NO: 18; or (vi) SEQ ID NO: 23 and SEQ ID NO: 16.
[0560] - A screening platform comprising a library comprising a plurality of sets of protein complexes according to the fourth aspect, wherein each set comprises a different combination of a first effector moiety and a second effector moiety.
[0561] - A therapeutic drug candidate, comprising an oligomeric core, wherein the oligomeric core comprises a plurality of (e.g., 3 to 6, preferably 3) monomers, each monomer comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99% or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 4; wherein each monomer is directly linked (e.g., genetically fused) to a first effector moiety and a second effector moiety; wherein the first effector moiety and the second effector moiety may be the same or different, and are preferably different; and, preferably, wherein each monomer is directly linked to the first effector moiety and the second effector moiety to which it is linked via a polypeptide linker.
[0562] In a fifth preferred aspect, the present application specifically provides:
[0563] - A polypeptide comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain, and wherein the first antigen binding domain and the second antigen binding domain are capable of binding to their targets when the target molecules are expressed on a single cell or immobilized on a plate or a single bead.
[0564] - A polypeptide oligomer, wherein each polypeptide within the oligomer comprises or consists of a polypeptide comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain, and wherein the first antigen-binding domain and the second antigen-binding domain are capable of binding to their targets when the target molecule is expressed on a single cell or immobilized on a plate or a single bead.
[0565] -A polypeptide comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain, wherein the first antigen binding domain and the second antigen binding domain are capable of binding to their targets when the target molecule is expressed on a single cell or immobilized on a plate or a single bead. The first binding domain and the second binding domain are catcher domains, each of which is capable of forming an isopeptide bond with a cognate peptide. Such cognate peptides are generally referred to as tag peptides, for example, as is known in the art and as described above, SpyTag forms an isopeptide bond with a SpyCatcher domain. The cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain.
[0566] -A polypeptide oligomer, wherein each polypeptide within the oligomer comprises or consists of a polypeptide comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein the first binding domain and the second binding domain are separated by a structural domain, wherein the first antigen binding domain and the second antigen binding domain are capable of binding to their targets when the target molecule is expressed on a single cell or immobilized on a plate or a single bead. The first binding domain and the second binding domain are catcher domains, each of which is capable of forming an isopeptide bond with a cognate peptide. Such cognate peptides are generally referred to as tag peptides, for example, as is known in the art and as described above, SpyTag forms an isopeptide bond with a SpyCatcher domain. The cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain.
[0567] Other aspects of the present disclosure
[0568] The present application also provides a polynucleotide encoding at least one monomer of the oligomeric core of the multivalent protein scaffold described in further detail herein. The present application also provides a polynucleotide encoding a multi-domain polypeptide construct comprising a first binding domain, a second binding domain, and a structural domain as described in further detail herein.
[0569] Also provided are a vector comprising a polynucleotide, a cell comprising the vector, and a method for producing a monomeric, oligomeric core and / or multivalent protein scaffold, comprising: culturing the cell in a culture medium to produce the protein scaffold.
[0570] Also provided are a vector comprising the polynucleotide, a cell comprising the vector, and a method of producing a multi-domain polypeptide construct comprising: culturing the cell in a culture medium to produce the multi-domain polypeptide.
[0571] The selection of a suitable polynucleotide sequence encoding at least one monomer, a suitable expression vector, and suitable cells for expressing the monomer, oligomeric core and / or multivalent protein scaffold is routine for those skilled in the art.
[0572] therapeutic effects
[0573] The protein complexes, therapeutic analogs, and therapeutic drug candidates provided herein can be used therapeutically. Multi-domain polypeptide constructs can generally be used therapeutically. Such substances provided herein are also referred to as "therapeutic protein complexes."
[0574] Thus, the present invention provides therapeutic protein complexes and constructs as described herein for use in medicine. The present invention provides therapeutic protein complexes as described herein for use in the treatment of a human or animal body. The present invention provides therapeutic protein constructs as described herein for use in the treatment of a human or animal body.
[0575] The present invention provides a method for treating a human or animal in need of such treatment, comprising administering to the human or animal a protein complex, a multi-domain polypeptide construct (in monomeric or oligomeric form), a therapeutic drug analog, a therapeutic drug candidate or a therapeutic drug as described herein.
[0576] Also provided is a pharmaceutical composition comprising one or more therapeutic protein complexes as described herein and a pharmaceutically acceptable carrier or diluent. Typically, the composition contains no more than 85% by weight of the therapeutic protein complexes of the invention. More typically, it contains no more than 50% by weight of the therapeutic protein complexes of the invention. Preferably, the pharmaceutical composition is sterile and pyrogen-free.
[0577] Also provided is a pharmaceutical composition comprising one or more multi-domain polypeptide constructs as described herein and a pharmaceutically acceptable carrier or diluent. Typically, the composition contains no more than 85% by weight of the therapeutic protein complex of the invention. More typically, it contains no more than 50% by weight of the therapeutic multi-domain polypeptide construct of the invention. Preferably, the pharmaceutical composition is sterile and pyrogen-free.
[0578] The compositions of the present invention may be provided as a kit including instructions for use of the kit in the methods described herein, or details regarding the subjects for which the methods are applicable.
[0579] As described above, the therapeutic protein complexes and constructs provided herein can be used to treat or prevent various conditions. Conditions that can be treated with the therapeutic p...
Claims
1. A polypeptide comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, wherein: The first binding domain and the second binding domain are separated by a CutA1 domain containing one or more substitutions or deletions relative to wild-type CutA1, and wherein the first binding domain and the second binding domain are the same or different.
2. The polypeptide according to claim 1 or 2, wherein The CutA1 domain is human CutA1.
3. The polypeptide according to claim 1 or 2, wherein The substitution is of one or more cysteine residues, or wherein the deletion is of one or more residues at the N-terminus and / or C-terminus.
4. The polypeptide according to any one of claims 1 to 3, wherein One or more cysteine residues in CutA1 are substituted with one or more alanine, valine or serine residues.
5. The polypeptide according to any one of claims 1 to 4, wherein The substitution comprises or consists of two cysteines being replaced by two alanines.
6. The polypeptide according to any one of claims 1 to 4, wherein The substitution comprises or consists of a substitution of one cysteine by valine and one cysteine by serine. 7 . The polypeptide according to claim 1 , wherein the cysteine residues at positions 75 and 96 of wild-type human CutA1 (SEQ ID NO: 19) are substituted by different residues.
8. The polypeptide according to claim 7, wherein The cysteine substitution comprises or consists of (i) C75A, C96A or (ii) C75V, C96S.
9. A polypeptide according to any one of the preceding claims, wherein The CutA1 domain comprises a cysteine residue at a position other than a cysteine residue in the wild-type sequence.
10. The polypeptide according to claim 9, wherein One or more of the following residues of wild-type human CutA1 (SEQ ID NO: 19) are substituted with cysteine: V64, E78, K79, K82, E83, K91, Q102, K110, E114, F136, S139, F158, and Q166.
11. A polypeptide according to any one of the preceding claims, wherein 5-60 residues are deleted from the N-terminus of CutA1, optionally wherein 10-59 residues are deleted.
12. A polypeptide according to any one of the preceding claims, wherein 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 residues are deleted from the N-terminus of CutA1.
13. A polypeptide according to any one of the preceding claims, wherein CutA1 is a truncated human CutA1 that begins at residue 33 of SEQ ID NO:19, or begins at residue 44 of SEQ ID NO:19, or begins at residue 60 of SEQ ID NO:19, or begins at residue 24 of SEQ ID NO:
19.
14. A polypeptide according to any one of the preceding claims, wherein From 5 to 20 residues are deleted from the C-terminus of CutA1, optionally wherein 6 to 12 residues are deleted.
15. A polypeptide according to any one of the preceding claims, wherein CutA1 is a truncated CutA1 that starts at any of residues 30-67 and ends at any of residues 165-179 of SEQ ID NO:
19.
16. A polypeptide according to any one of the preceding claims, wherein CutA1 is a truncated CutA1 consisting of residues 44-179 or residues 60-171 of SEQ ID NO:
19.
17. A polypeptide according to any one of the preceding claims, wherein The first binding domain and the second binding domain are different antigen binding domains, optionally wherein one or both antigen binding domains are antigen binding fragments of an antibody, optionally scFv or Fab, or single domain antibodies (sdAbs), or other antibody mimetics or scaffolds selected to bind to a specific target, or other proteins or peptides capable of specific binding to a biomolecule; and / or wherein one or both of the first binding domain and the second binding domain are agonists of a TNF receptor superfamily member, optionally wherein the TNF receptor superfamily member is a TRAIL receptor, such as death receptor 5 (DR5).
18. The polypeptide according to any one of claims 1 to 16, wherein The first binding domain and the second binding domain are catcher domains, each catcher domain being capable of forming an isopeptide bond with a cognate peptide, optionally wherein the cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain, or Optionally, wherein the cognate peptide of the first binding domain is the same as the cognate peptide of the second binding domain, optionally, selective binding or coupling is achieved by temporal or sequential control such as activation or inactivation of the second binding domain and / or by competitive binding.
19. The polypeptide according to claim 18, wherein Each homologous peptide is linked to an antigen binding domain, optionally wherein one or two homologous peptides are linked to the first catcher domain and / or the second catcher domain via an isopeptide bond.
20. The polypeptide according to claim 18, wherein One or both of the homologous peptides are linked to an agonist of a TNF receptor superfamily member, optionally wherein the TNF receptor superfamily member is a TRAIL receptor, such as death receptor 5 (DR5).
21. The polypeptide according to any one of claims 1 to 20, wherein When expressed on a single cell or immobilized on a plate or a single bead, the first and second binding domains are capable of binding to their targets.
22. A polypeptide according to any one of the preceding claims, comprising an effector molecule, optionally a drug molecule or a dye molecule, optionally wherein the drug molecule comprises or consists of: (a) an anticancer drug, optionally a cytotoxic drug; (b) a tubulin inhibitor, optionally a maytansinoid, an auristatin or a paclitaxel derivative; (c) monomethyl auristatin E (MMAE) or monomethyl auristatin F (MMAF); (d) compounds derived from dolastatin 10; (e) tubulysins, such as tubulysin A; (f) DNA damaging agents; (g) Ducarcin; (h) calicheamicin; (i) Pyrrolobenzodiazepine (j) SN-38 (active metabolite of irinotecan); (k) immunomodulators, such as TLR agonists or STING agonists; (l) Delutec; (m)Metansin; or (n) an immunotoxin, optionally Pseudomonas exotoxin A (PE).
23. The polypeptide of claim 22, wherein the drug molecule or dye molecule is chemically coupled to the polypeptide, optionally in the CutA1 domain, and optionally to a cysteine residue in the CutA1 domain, optionally wherein the coupled cysteine residue is not present in the native CutA1 sequence.
24. An oligomer comprising two or more polypeptides according to any one of the preceding claims.
25. The polypeptide according to any one of claims 1 to 23 or the oligomer according to claim 24, wherein The polypeptide or oligomer comprises the features of any one of embodiments A1 to A26.
26. A polypeptide comprising: an N-terminal first binding domain and / or a C-terminal second binding domain of a CutA1 domain or a cytokine domain, wherein when both the first binding domain and the second binding domain are present, they are separated by the CutA1 domain or the cytokine domain, optionally wherein the CutA1 domain or the cytokine domain contains one or more substitutions or deletions relative to the corresponding wild-type CutA1 or cytokine, and wherein the first binding domain and the second binding domain are the same or different; or A first binding domain linked to a second binding domain, wherein the second binding domain is linked to a CutA1 domain or a cytokine domain, optionally wherein the CutA1 domain or the cytokine domain contains one or more substitutions or deletions relative to the corresponding wild-type CutA1 or cytokine, and wherein the first binding domain and the second binding domain are the same or different.
27. The polypeptide of claim 26, wherein the cytokine is TNF, TL1A, OX40L or CD40L, SEQ ID NO: 80 or a modified form thereof, SEQ ID NO: 31 or a modified form thereof, SEQ ID NO: 78 or a modified form thereof, or SEQ ID NO: 58 or a modified form thereof.
28. A CutA1 protein comprising one or more substitutions or deletions relative to wild-type CutA1, optionally linked to a different polypeptide at the N-terminus and / or C-terminus of the CutA1 protein via a peptide bond as a fusion protein.
29. The CutA1 protein according to claim 28, wherein at least one native cysteine residue is substituted by a different amino acid residue, and / or wherein at least one non-cysteine native residue is substituted by a cysteine residue.
30. A single polypeptide chain comprising a plurality of CutA1 sequences, optionally wherein a linker polypeptide of 3-30 amino acids in length is located between some or all of the CutA1 sequences, optionally wherein at least one and optionally all of the CutA1 sequences are as defined in any one of claims 1 to 16.
31. A single polypeptide according to claim 30, comprising three human CutA1 sequences, each separated by a linker polypeptide, optionally comprising an isopeptide bond forming domain at the N-terminus and / or C-terminus.
Citation Information
Patent Citations
Recombinant antibodies specific for a growth factor receptor
US5571894A
Anti-erbB-2 antibodies, combinations thereof, and therapeutic and diagnostic uses thereof
US5587458A
Immunoglobulins devoid of light chains
WO1994004678A1
Methods and products for fusion protein synthesis
WO2016193746A1
Multivalent and multispecific dr5-binding fusion proteins
WO2017011837A2