Multivalent proteins and screening methods
Multivalent protein scaffolds with engineered termini enable efficient and scalable screening of novel therapeutics, addressing production and targeting challenges of antibody-based drugs, facilitating diverse therapeutic applications.
Patent Information
- Application Number
- JP2025518049
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-16
- Filing Date
- 2023-09-28
- Publication Date
- 2025-10-15
AI Technical Summary
Existing antibody-based therapeutics face challenges such as high production costs, complex post-translational modification chemistry, poor tumor targeting, and limited modularity in screening multiple antigen-binding domains, which hinder their widespread use and therapeutic potential.
Development of multivalent protein scaffolds with customizable, reproducible, and scalable constructs that allow for cis-oriented bispecific binding, using engineered polypeptides with modified N- and C-termini, enabling efficient screening and identification of novel therapeutic agents.
The multivalent protein scaffolds provide a modular platform for rapid and scalable identification of therapeutic agents, overcoming production costs and targeting limitations, while minimizing immune responses and enabling diverse therapeutic applications.
Smart Images

Figure 2025534309000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to multivalent protein scaffolds and their use as a modular system for phenotypic screening of combinations of target molecules and as therapeutic agents. The present invention also relates to multi-domain polypeptide constructs comprising multiple binding domains and a structural domain. The present invention also relates to methods for identifying new therapeutic agents using the protein scaffolds described herein, and the therapeutic agents that can be identified in this way. [Background technology]
[0002] There is a continuing need to identify new therapeutic agents for many pathological conditions.
[0003] Protein-based therapeutics offer an attractive approach to address many common diseases. Such therapeutics have proven highly successful clinically, with many protein therapeutics approved for clinical use by regulatory agencies worldwide.
[0004] Protein-based therapeutics can act in a variety of ways: for example, by replacing missing or abnormal proteins; by enhancing existing pathways; by providing a therapeutic utility with novel functions or activities; by interfering with molecules or organisms; and by delivering other compounds or proteins, such as radionuclides, cytotoxic drugs, or effector proteins. Therapeutic proteins can be grouped based on their physical and structural properties, such as antibody-based drugs, Fc fusion proteins, anticoagulants, blood factors, bone morphogenetic proteins, engineered protein scaffolds, enzymes, growth factors, hormones, interferons, interleukins, and thrombolytic agents. Therapeutic proteins can also be classified based on their molecular mechanism of activity. For example, monoclonal antibodies typically act by noncovalently binding to targets. Enzymes can affect covalent binding of targets. Other proteins, such as serum albumin, can exert their activity without any specific interaction.
[0005] One class of protein-based therapeutics that has achieved clinical success is therapeutic antibodies. Antibodies, also known as immunoglobulins (Ig), are being evaluated as promising treatments for many disease states. For example, monoclonal antibody therapy has been used to treat diseases including rheumatoid arthritis, multiple sclerosis, psoriasis, and various forms of cancer. Commercially available antibody therapeutics include muromomab, abciximab, rituximab, daclizumab, basiliximab, palivizumab, infliximab, trastuzumab, etanercept, gemtuzumab, alemtuzumab, ibritomomab, adalimumab, alefacept, omalizumab, tositumomab, efalizumab, cetuximab, bevacizumab, natalizumab, ranibizumab, panitumumab, eculizumab, and certolizumab.
[0006] An antibody typically comprises four polypeptide chains that form one Fc region and two antigen-binding (Fab) regions. Each Fab region contains a variable region (Fv) that forms a paratope and contacts the antigen. Naturally occurring antibodies typically exhibit symmetric binding in the variable regions. However, this limitation typically means that a given antibody can only target a single receptor type (or other target) at the variable regions.
[0007] To address this, there has been much interest recently in bispecific antibodies. Bispecific antibodies differ from conventional monospecific antibodies in that each of the two Fab regions binds to two different antigens. Bispecific antibodies are often classified as either Ig-like or non-Ig-like, the latter of which may consist of chemically linked Fab regions.
[0008] Bispecific antibodies are being actively investigated for clinical use. Two examples of commercially available bispecific antibodies include blinatumomab (sold under the trade name Blincyto), which contains both a CD3 site for targeting T cells and a CD19 site for targeting B cells and is useful for treating Philadelphia chromosome-negative relapsed or refractory acute lymphoblastic leukemia; and emicizumab (sold under the trade name Hemlibra), which targets both blood coagulation factors IXa and X and is used to treat hemophilia A. Bispecific antibodies are generally used to simultaneously bind to multiple cell types, for example, by simultaneously binding to tumor cell receptors and recruiting cytotoxic immune cells.
[0009] Despite the promise of some bispecific antibodies, problems remain. Antibody therapeutics are associated with high production costs, at least due to their size and complex post-translational modification chemistry, including complex glycosylation patterns. Antibody production requires the use of very large-scale mammalian cell cultures followed by extensive purification steps, which leads to extremely high production costs and limits the widespread use of these drugs. Antibodies also have poor tumor targeting, limiting their use in cancer therapy (e.g., studies in mouse xenograft models have shown that typically less than 20% of administered antibodies interact with tumors). The Fc portion of antibodies, such as IgG antibodies, can interact with various receptors expressed on the surface of several cell types, thereby increasing their persistence in the circulation. The large size of antibodies also leads to slow diffusion in vivo. IgG-like antibodies are immunogenic and can trigger harmful downstream immune responses via Fc receptor activation. With regard to bispecific antibodies in particular, the previously described "knobs-into-holes" Ig-like approach lacks modularity and is therefore not easily applicable to screening a large number of antigen-binding domains. Furthermore, the bispecific antibody approach is practically limited to screening the Fv / Fab region and therefore cannot be used to explore the therapeutic potential of other non-immunoglobulin protein domains. The non-Ig-like approach of tandem fusions is more applicable but is not easily scaled up.
[0010] Blanco-Toribio et al. (MAbs. 2013 Jan 1;5(1):70-79) described the generation and characterization of monospecific and bispecific hexavalent trimerbodies. These molecules, termed "trimerbodies," use a modified N-terminal trimerization region of the noncollagenous 1 (NC1) domain of human collagen XVIII, flanked by two flexible linkers, as a trimerization scaffold. By fusing single-chain variable fragments (scFvs) with the same or different specificities to both the N- and C-termini of the trimerization scaffold domain, the authors generated monospecific or bispecific hexavalent molecules that were efficiently secreted as soluble proteins by transfected mammalian cells. The bispecific anti-laminin × anti-CD3 N- / C-trimerbodies were found to be trimeric in solution. One drawback of this method is that the required transfection process is not always a convenient production method.
[0011] WO 2020 / 0188346 describes a bispecific antigen-binding protein in which two antigen-binding domains ("ABD") are covalently linked to a fusion protein formed by two or more domains that form isopeptide linkages with the antigen-binding protein. The isopeptide linkage-forming domains are typically catcher domains such as Spycatcher ("SC"), and the resulting bispecific protein is of the format ABD-SC-SC-ABD.
[0012] Brune et al. (2017) (Bioconjugate Chem. 2017, 28, 5, 1544-1551) described a plug-and-display synthetic assembly using orthogonally reactive proteins for twin antigen immunization. The authors prepared dual-addressable synthetic nanoparticles by genetically engineering the multimerizing coiled-coil IMX313 and two orthogonally reactive split proteins. This construct, in the SpyCatcher-IMX-SnoopCatcher format, provides a modular platform. This allows for the multimerization of SpyTag-antigen and SnoopTag-antigen on opposite faces of the particle by simply mixing them.
[0013] Thus, new paradigms are needed that enable the rapid, scalable, and applicable identification and design of new protein therapeutics that overcome some or all of these challenges, particularly new platform technologies that rival traditional antibody platforms. Summary of the Invention
[0014] The present inventors have recognized the problems described above. It is now recognized that a non-antibody assembly platform can be used to provide a customizable, reproducible, scalable, and highly applicable scaffold for screening, identifying, and developing novel therapeutics. The approach described herein allows the geometry, valency, and / or functionality of multiple different proteins to be evaluated for potential therapeutic effects.
[0015] Part of the inventors' approach was to develop protein constructs with desirable properties. Such constructs may be prepared recombinantly by expression as fusion proteins, or the component domains may be joined by other means known in the art, such as chemical conjugation. In particular, the inventors have identified that both the N-terminus and C-terminus of a polypeptide can be advantageously modified to provide a polypeptide with two modified termini. These modifications typically involve the addition of polypeptide domains, each capable of binding to a target molecule, such as an antigen-binding region or an isopeptide bond-forming region. Different target molecules can be attached to the N-terminus and C-terminus to provide so-called bispecific binding constructs. The resulting protein constructs are capable of binding to target molecules at their modified N-terminus and modified C-terminus. The inventors have specifically engineered protein constructs in which the N-terminus and C-terminus exhibit the same basic orientation, such that the resulting constructs are capable of binding to binding partners at each terminus when the binding partners are present in the same basic spatial and orientation, for example, when attached to a solid surface such as a plate or bead, or to the surface of a cell. This can be referred to as providing the modified N- and C-termini of a single polypeptide chain in a "cis" orientation. Typically, a cis-oriented bispecific construct is provided. Two or more of these protein constructs can be combined to form an oligomeric protein.
[0016] Such protein constructs allow for the creation of combinatorial systems that can be used to screen for useful combinations of effector moieties, such as binding regions (e.g., antigen-binding regions). Furthermore, once useful combinations are identified, the constructs can be modified to remove (or replace with, for example, linkers) features required for combinatorial screening, thereby providing simpler protein constructs with the identified preferred combination of binding regions. Such constructs, as monomers or oligomers of more than one construct, can be particularly useful as therapeutic, diagnostic, or analytical agents.
[0017] Therefore, in this specification, - an oligomeric core comprising a plurality of subunit monomers; and - at least two first binding sites that are orthogonal to at least two second binding sites; 1. A multivalent protein scaffold comprising: the first binding site and the second binding site are located on the same face of the scaffold; A multivalent protein scaffold is provided.
[0018] Also, - an oligomeric core containing multiple subunit monomers; - at least one first binding site that is orthogonal to at least one second binding site; 1. A multivalent protein scaffold comprising: the first binding site and the second binding site are located on the same face of the scaffold; the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target; A multivalent protein scaffold is provided.
[0019] moreover, - an oligomeric core comprising a plurality of subunit monomers; - at least one first binding site that is orthogonal to at least one second binding site; 1. A multivalent protein scaffold comprising: the first binding site and the second binding site are located on the same face of the scaffold; The oligomer core does not include the Fc region of an antibody. A multivalent protein scaffold is provided.
[0020] Preferably, the oligomer core comprises at least three subunit monomers, and more preferably, the oligomer core comprises three to six subunit monomers.
[0021] Preferably, in one embodiment, the subunit monomers are non-covalently attached together. Preferably, in another embodiment, the subunit monomers are covalently attached together. Preferably, when the subunit monomers are covalently attached together, the subunit monomers are genetically fused together. In some embodiments, the subunit monomers are expressed as a single polypeptide chain from a recombinant nucleic acid.
[0022] In one embodiment, the oligomer core is preferably a homo-oligomer core. In such an embodiment, each monomer of the oligomer core preferably comprises at least one first binding site and at least one second binding site, and the at least one first binding site is orthogonal to the at least one second binding site. Preferably, in one aspect, each monomer comprises a first binding site attached to a first end of the monomer and a second binding site at a second end of the monomer. Preferably, the first and second ends of each monomer are located on the same face of the monomer. Preferably, in another aspect, each monomer comprises a first binding site attached to a first end of the monomer and a second binding site attached to the first binding site.
[0023] In another embodiment, the oligomeric core is preferably a hetero-oligomeric core. Preferably, in such an embodiment, said core comprises at least one first subunit monomer comprising a first binding site and at least one second subunit monomer comprising a second binding site, wherein the first binding site is orthogonal to the second binding site.
[0024] Preferably, in the present invention, the protein scaffolds provided herein typically have each subunit monomer comprising less than 300 amino acids, preferably less than 200 amino acids, more preferably less than 150 amino acids. Preferably, the oligomeric core has a molecular weight of less than about 150 kDa, preferably less than about 100 kDa, more preferably less than about 70 kDa.
[0025] In some embodiments, the oligomer core does not comprise an Fc region of an antibody. In some embodiments, the oligomer core does not comprise a CH2 domain. In some embodiments, the oligomer core does not comprise a CH3 domain. In some embodiments, the oligomer core does not comprise a CH2 domain and does not comprise a CH3 domain.
[0026] Preferably, the oligomeric core and / or scaffold does not elicit an immune response when administered to a human subject, as described in further detail herein.
[0027] In some embodiments, the oligomeric core and / or scaffold or structural domain does not elicit an adverse immune response when administered to a human subject, e.g., it does not elicit an active B cell or T cell response against the structural domain, and / or it does not specifically bind to an immunoglobulin receptor or activate antibody-dependent cell-mediated cytotoxicity (ADCC).
[0028] Preferably, the oligomeric core comprises a soluble multimerization structural element of a multimeric protein, such as a collagen NC (non-collagenous) domain (e.g., an NC1 domain), CutA1, a C1q head domain, TNF, p53, fibrinogen, C4, Bacillus subtilis AbrB, or a homolog or paralog thereof.
[0029] Preferably, the multimeric protein comprises type VIII collagen NC1 (non-collagenous) domain, type X collagen NC1 (non-collagenous) domain, C1q head domain, CutA1 protein, macrophage migration inhibitory factor (MIF) or macrophage migration inhibitory factor 2 (MIF-2), tumor necrosis factor (TNF), a TNF family member including TL1A, CD40L or OX40L, or a homolog or paralog thereof.
[0030] Preferably, the multimerization structural element comprises a polypeptide having at least 30% or at least 50% amino acid identity with SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:29, SEQ ID NO:60, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:42, SEQ ID NO:31, SEQ ID NO:58, SEQ ID NO:78, SEQ ID NO:80 or SEQ ID NO:19.
[0031] The present inventors have particularly identified CutA1, typically human CutA1, as a preferred structural element. In some embodiments, the CutA1 protein is genetically engineered to comprise one or more substitutions or deletions compared to wild-type CutA1. Optionally, genetically engineered CutA1 domains are particularly useful as scaffolds. Thus, in some embodiments, the CutA1 domain is connected to different polypeptide domains at the N-terminus and / or C-terminus of CutA1 by peptide linkage as a fusion protein.
[0032] In certain embodiments, the multimerization structural element is human CutA1 (SEQ ID NO: 19). In certain embodiments, the multimerization structural element is a truncated form of SEQ ID NO: 19 (human CutA1) as described elsewhere herein. In certain embodiments, the multimerization structural element is a modified form of SEQ ID NO: 19 (human CutA1), in which at least one cysteine residue of SEQ ID NO: 19 has been replaced with a different amino acid residue as described elsewhere herein. In certain embodiments, the multimerization structural element is a modified and truncated form of SEQ ID NO: 19 (human CutA1), in which at least one N-terminal and / or C-terminal amino acid residue has been removed and at least one cysteine residue of SEQ ID NO: 19 has been replaced with a different amino acid residue as described elsewhere herein.
[0033] In certain embodiments, the multimerization structural element is a cytokine. In certain embodiments, the multimerization structural element belongs to the TNF superfamily. In certain embodiments, the multimerization structural element is TNF (SEQ ID NO: 80) or OX40L (SEQ ID NO: 78) or CD40L (SEQ ID NO: 58) or TL1A (SEQ ID NO: 31), or is derived from TNF or OX40L or CD40L or TL1A (e.g., a truncated and / or modified form thereof), as described elsewhere herein. Thus, the multimerization structural element may comprise or consist of a polypeptide having at least 30%, for example at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity with the amino acid sequence of human CutA1 (SEQ ID NO: 19), OX40L (SEQ ID NO: 78), CD40L (SEQ ID NO: 58), TL1A (SEQ ID NO: 31), or TNF (SEQ ID NO: 80).
[0034] In certain embodiments, the multimerization structural element is a modified and truncated form of SEQ ID NO:78, SEQ ID NO:58, SEQ ID NO:80, or SEQ ID NO:31, in which at least one N-terminal and / or C-terminal amino acid residue has been removed and at least one cysteine residue has been replaced with a different amino acid residue as described elsewhere herein.
[0035] CutA1 is structurally similar to the non-human PII protein family (Bagautdinov et al., Acta Crystallogr Sect F Struct Biol Cryst Commun. 2008 May 1; 64(Pt 5): 351-7). In some embodiments, the structural element comprises or consists of a protein of the PII protein family, or a truncated or modified version thereof. Examples of PII proteins include the signal transduction protein PII GlnB (SEQ ID NO: 81) from Synechococcus elongatus PCC7942 (Xu et al., Acta Crystallogr D Biol Crystallogr. 2003; 59: 2183-2190), GlnK (SEQ ID NO: 82) from Escherichia coli (Xu et al., J Mol Biol. 1998; 282: 149-165), the PII-like domain of YqfO from Bacillus cereus (SEQ ID NO: 83) (Godsey et al., Protein Sci. 2007 Jul; 16: 1285-1293), or the PII-like domain of Staphylococcus aureus (SEQ ID NO: 84). Examples include the putative protein SA13888 (SEQ ID NO: 84) from S. aureus (Saikatendu et al., BMC Struct Biol. 2006; 6: 27), or truncated or modified versions thereof described by Luddecke et al., Sci Rep.
[0036] Preferably, the first binding site and / or the second binding site comprise a protein domain. Typically, the first binding site comprises a first protein domain and the second binding site comprises a second protein domain. Preferably, the first binding site and / or the second binding site are genetically fused to the subunit monomer to which they are attached, forming a single polypeptide chain.
[0037] Preferably, the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target. Preferably, the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target. More preferably, the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target. Preferably, the first protein domain is capable of forming an isopeptide bond with the first polypeptide target, and the second protein domain is capable of forming an isopeptide bond with the second binding target.
[0038] Preferably, the first binding site and the second binding site each comprise a different split ligand-binding protein domain. More preferably, one of the first binding site and the second binding site comprises a split Streptococcus pyogenes fibronectin-binding protein domain, and the other of the first binding site and the second binding site comprises a split Streptococcus pneumoniae adhesin domain.
[0039] Preferably, the first and second binding sites each independently have at least 50% amino acid identity to any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18. In some embodiments, the first and second binding sites each independently have at least 60%, at least 70%, at least 80%, or at least 90% amino acid identity to any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18.
[0040] Also provided herein is a protein complex comprising a protein scaffold as described herein, wherein the first binding site binds to a first polypeptide target attached to a first effector moiety, and the second binding site binds to a second polypeptide target attached to a second effector moiety.
[0041] Preferably, in the complex, the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are each independently selected from the following: (i) a combination of any one of SEQ ID NOs: 4, 6, or 8 with any one of SEQ ID NOs: 5, 7, or 9; (ii) a combination of SEQ ID NO: 12 with SEQ ID NO: 13 or 15; (iii) a combination of SEQ ID NO: 5 with SEQ ID NO: 11; (iv) a combination of SEQ ID NO: 15 with SEQ ID NO: 16; (v) a combination of SEQ ID NO: 17 with SEQ ID NO: 18, or (vi) a combination of SEQ ID NO: 23 with SEQ ID NO: 16.
[0042] Also provided is a screening platform comprising a library, said library comprising a plurality of populations of protein complexes described herein, each population of protein complexes comprising a different combination of a first effector moiety, a second effector moiety, and / or an oligomeric core.
[0043] Also provided is a method for identifying a therapeutic drug or drug analog, comprising: Providing a protein complex as described herein; contacting the protein complex with a biological system; and Determining whether a protein complex induces a desired change in a property of a biological system Including, Optionally, methods are provided that further comprise selecting a protein complex that induces a desired change in a property of the biological system.
[0044] The method may further comprise synthesizing a therapeutic drug or drug candidate comprising an oligomeric core of the identified therapeutic drug analog protein complex scaffold attached to first and second effector moieties of said protein complex.
[0045] Also provided are therapeutic drug candidates obtainable according to the methods described herein.
[0046] Also provided are therapeutic drugs obtainable according to the methods described herein.
[0047] Also provided are therapeutic drugs or drug candidates comprising or consisting of one or more of the constructs or polypeptides described herein.
[0048] Also provided herein is a therapeutic drug or drug candidate comprising an oligomer core, the oligomer core comprising a plurality of subunit monomers attached to one or more first effector moieties and one or more second effector moieties, wherein the one or more first effector moieties and the one or more second effector moieties are located on the same face of the oligomer core, and wherein (i) the one or more first effector moieties comprise two or more first effector moieties and the one or more second effector moieties comprise two or more second effector moieties; and / or (ii) the oligomer core does not comprise an antibody or antibody fragment. Preferably, the oligomer core of the therapeutic drug counterpart is as described in more detail herein.
[0049] Preferably, in the provided therapeutic drug candidates, the oligomer core comprises multiple subunit monomers, and (i) each subunit monomer comprises a collagen NC1 domain, CutA1, C1q domain, TNF, p53, fibrinogen, C4, Bacillus subtilis AbrB, or a homolog or paralog thereof; and / or (ii) each subunit monomer comprises a multimerization structural element comprising a polypeptide having at least 50% amino acid identity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:19.
[0050] Preferably, in the provided therapeutic drugs or drug candidates, the oligomer core comprises a plurality of subunit monomers, and (i) each subunit monomer comprises a type VIII collagen NC1 (non-collagenous) domain, a type X collagen NC1 (non-collagenous) domain, a C1q head domain, a CutA1 protein, macrophage migration inhibitory factor (MIF) or macrophage migration inhibitory factor 2 (MIF-2), tumor necrosis factor (TNF), a TNF family member including TL1A or CD40L or OX40L, or a homolog or paralog thereof; and / or (ii) each subunit monomer comprises a multimerization structural element comprising a polypeptide having at least 50% amino acid identity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:29, SEQ ID NO:60, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:42, SEQ ID NO:31, SEQ ID NO:58, SEQ ID NO:78, SEQ ID NO:80, or SEQ ID NO:19.
[0051] In certain embodiments, the present invention provides a polypeptide comprising a first binding domain at the N-terminus and / or a second binding domain at the C-terminus, wherein the first and second binding domains are separated by a structural domain. The polypeptide is typically a single genetically engineered polypeptide chain expressed as a fusion protein from a recombinant nucleic acid. Typically, the first binding domain and the second binding domain can bind to their target when the target molecule is expressed on a single cell or immobilized on a plate or single bead. This is sometimes described herein as providing the first and second binding domains in a "cis" orientation.
[0052] Thus, in some embodiments, a first binding domain is present at the N-terminus of the polypeptide and a second binding domain is present at the C-terminus, with the first and second binding domains separated by a structural domain. However, in some embodiments, the first binding domain is connected to the second binding domain, which is connected to the structural domain. Typically, the second binding domain is connected to the N-terminus or C-terminus of the structural domain. As noted elsewhere herein, the first and second binding domains are typically different, but in some embodiments, they may be the same. A linker, e.g., a linker peptide of 2 to 30 amino acid residues, typically 5 to 25 amino acid residues, may be included between any two domains, or between all domains of the construct, or between any selected domains of the construct.
[0053] In some embodiments, the first binding domain and the second binding domain are different antigen binding domains, in which case the construct is a bispecific construct.
[0054] In some embodiments, the first binding domain and / or the second binding domain is a protein or peptide capable of specifically binding to a biological molecule, which may be a signaling molecule capable of specifically interacting with a binding partner, such as a ligand or receptor of the protein or peptide, e.g., a cytokine or cell surface receptor.
[0055] In other embodiments, the first binding domain and the second binding domain are each catcher domains (i.e., split ligand-binding protein domains) capable of forming an isopeptide bond with a cognate peptide. Such cognate peptides are often referred to as tag peptides; for example, SpyTag forms an isopeptide bond with a SpyCatcher domain, as known in the art and discussed below. Typically, the cognate peptide for the first binding domain is different from the cognate peptide for the second binding domain. However, in some embodiments, the cognate peptides may be the same for both the first and second binding domains. A polypeptide having a catcher at each end can be provided separately from a tag-containing molecule (e.g., a protein), for example, as part of a kit. In some embodiments, each tag peptide is covalently attached to its cognate catcher domain, and optionally, one or both cognate peptides are linked to the first and / or second catcher domains by an isopeptide bond. In some embodiments, one or both cognate peptide tags are present as a fusion polypeptide with an effector moiety, typically an antigen-binding domain. Thus, the linkage of a catcher to its cognate peptide tag links an effector moiety (eg, an antigen binding domain) to its binding domain.
[0056] In embodiments in which the cognate peptides of both the first and second binding domains are the same, preferential, selective, or sequential binding or conjugation can be achieved by temporal or sequential control, spatial control, or spatiotemporal control, which may be, for example, by activation or inactivation of one binding domain and / or by competitive binding. Without limitation, activation can occur, for example, by enzymatic cleavage of an inhibitory peptide (e.g., as described in Driscoll et al, September 2023, bioRxiv 2023.08.31.555700; doi: https: / / doi.org / 10.1101 / 2023.08.31.555700, which is incorporated herein by reference in its entirety), or by photoactivation using light triggering of covalent bond formation in a photocaged binding domain, for example, created by site-specific incorporation of a non-natural residue such as coumarin-lysine at the reactive site (e.g., Rahikainen et al, 2023 "Visible light-induced specific protein reaction delineates early stages of cell adhesion", bioRxiv 2023.07.21.549850; doi: https: / / doi.org / 10.1101 / 2023.07.21.549850 for SpyCatcher003, which is incorporated by reference in its entirety. Further photoactivation techniques using photocaged glutamic acid analogs are described in "Photoactivatable Protein with Genetically Encoded Photocaged Glutamic Acid," by Yang et al., Angewendte Chemie Volume 62, Issue 40, October 2, 2023 (published online August 16, 2023). Alternatively, control over the availability of reactive sites can be achieved by removing any inhibitory domains, for example, by steric hindrance, competing for binding, or blocking reactive or catalytic residues.
[0057] In some embodiments, two antigen-binding domains, each containing the same isopeptide bond-forming tag, can be precisely and selectively conjugated to a construct of the present invention containing two identical isopeptide bond-forming catcher domains (to which the tags are attached), one of which is unreactive until reactivity is revealed using photoactivation. Typically, one of the catcher domains is caged at the reactive isopeptide bond-forming site by the presence of a non-natural photoreactive residue, such as coumarin-lysine (7-hydroxycoumarin lysine, "HCK") amino acid or a photocaged glutamic acid analog, making it unreactive. Once a first tagged peptide is attached to the first catcher by forming an isopeptide bond between the first catcher and the tag, the photocage residue is released from the second catcher by irradiation with appropriate light (e.g., 405 nm light). The second catcher is then available for binding, allowing the addition of a second tagged peptide.
[0058] In some embodiments, two antigen-binding domains each containing the same isopeptide bond-forming tag can be precisely and selectively conjugated to a construct of the present invention containing two identical isopeptide bond-forming catcher domains (to which the tags bind), one of which is non-reactive until reactivity is revealed using a site-specific protease. Typically, one of the catcher domains is non-reactive due to the presence of a non-reactive tag mutant sequence fused to one of the catcher domains via a flexible linker containing a protease cleavage site. Once a first tagged peptide is attached to the first catcher by forming an isopeptide bond between the first catcher and the tag, a protease is added to release the non-reactive tag mutant from the second catcher. The second catcher is then available for binding, allowing the addition of a second tagged peptide.
[0059] In certain embodiments, the non-reactive tag mutant is SpyTag003 D117A (SpyTag003DA) fused to the C-terminus of a second SpyCatcher003 via a flexible linker containing a tobacco etch virus (TEV) protease cleavage site (Keeble et al., 2019 Proc. Natl. Acad. Sci. USA 116, 26523-26533, which is incorporated herein by reference in its entirety). After cleavage at the TEV site, the SpyTag003DA peptide is free to dissociate, exposing the reactive Lys of SpyCatcher003 and allowing it to react with a second provided SpyTag-linked conjugate. Cleavage can be performed using any suitable protease, such as a super-TEV protease.
[0060] In some embodiments, the polypeptide comprises a first binding domain at its N-terminus and a second binding domain at its C-terminus, the first and second binding domains being separated by a structural domain, the first binding domain and the second binding domain being catcher domains each capable of forming an isopeptide bond with a cognate peptide, the first catcher domain being linked to its cognate peptide tag by an isopeptide bond, and the second catcher domain being linked to its cognate peptide tag by an isopeptide bond. In some embodiments, each peptide tag is attached to an antigen-binding domain.
[0061] The polypeptides of the above four paragraphs are useful as monomers or can be combined to form oligomers.
[0062] In some aspects, there are provided oligomers comprising two or more polypeptides as defined above and elsewhere herein.
[0063] In some embodiments, the polypeptide or oligomer defined in the preceding paragraph comprises the features described elsewhere herein. For example, the structural domain of the polypeptide construct is typically a "subunit monomer" as broadly described herein, and therefore the definition and description of the subunit monomer also applies to the structural domain. Similarly, the first binding domain and the second binding domain are typically the first binding site and the second binding site as described elsewhere herein, and therefore the definition and description of the first and second binding site also apply to the first binding domain and the second binding domain of the polypeptide construct.
[0064] In certain embodiments, the structural domain (or subunit monomer of the oligomeric core) of the polypeptide construct is a CutA1 protein, typically a human CutA1 protein. The data presented below, particularly Figure 23, demonstrate various fusions of HsCutA1 ("Homo sapiens CutA1") and an effector protein, as well as the ability of such direct fusions to exert a biological effect. Thus, in certain embodiments, the present invention provides a polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, wherein the first and second binding domains are separated by a human CutA1 structural domain. In certain embodiments, CutA1, typically human CutA1, has been genetically engineered to remove one or more cysteine residues from the native sequence (shown herein as SEQ ID NO: 19). Typically, this is by replacing one or more cysteine residues with one or more non-cysteine residues. Removal of one or more cysteine residues is beneficial because it allows for targeted cysteine conjugation at non-native sites or in fusion proteins. Without being bound by theory, it may be beneficial to remove unpaired cysteines because unpaired cysteines often interfere with stability and downstream applications. In some embodiments, it is beneficial to remove one or more unpaired cysteines and reintroduce one or more cysteine residues only at optimized target positions.
[0065] In some embodiments, one or more cysteine residues of CutA1 are substituted with one or more alanine residues. In some embodiments, one or more cysteine residues of CutA1 are substituted with one or more valine residues. In some embodiments, one or more cysteine residues of CutA1 are substituted with one or more serine residues. In some embodiments, the substitution comprises or consists of two cysteines being substituted with two alanines, referred to herein as a "CACA" substitution. In some embodiments, the substitution comprises or consists of one cysteine being substituted with a valine and one cysteine being substituted with a serine, referred to herein as a "CVCS" substitution. In some embodiments, the cysteine residues at positions 75 and 96 of wild-type human CutA1 (e.g., SEQ ID NO: 19) are substituted with different residues. In some embodiments, human CutA1 is engineered to have two cysteine residues substituted, the cysteine substitutions comprising or consisting of (i) C75A, C96A, or (ii) C75V, C96S. Thus, in some embodiments, the structural domain (or subunit monomer of the oligomer core) of the polypeptide construct is a human CutA1 protein engineered to have two cysteine residues substituted, the cysteine substitutions comprising or consisting of (i) C75A, C96A, or (ii) C75V, C96S. The generation and biological effects of polypeptide constructs comprising such cysteine-substituted human CutA1 domains are illustrated herein ( FIG. 25 ).
[0066] In some embodiments, it is beneficial to remove one or more unpaired cysteines and reintroduce one or more cysteine residues only at optimized target positions. The reintroduced cysteines may be useful as conjugation sites for drugs or dyes, for example, when producing labeled antibodies (or antibody-type molecules) or antibody-drug conjugate ("ADC")-type molecules. Example 13 illustrates the substitution of non-cysteine residues with cysteine residues, and it was observed that CutA1 with the newly introduced Cys residues exhibited higher conjugation efficiency when conjugating dyes or drugs to CutA1 compared to the wild type and the negative control (CC041, i.e., CutA1 CACA engineered to remove the native cysteines). The newly introduced one or more cysteines can be substituted at preferred positions within the sequence. As illustrated in Example 13, exemplary positions in human CutA1 include one or more of V64, E78, K79, K82, E83, K91, Q102, K110, E114, F136, S139, F158, and Q166. In some embodiments, human CutA1 (e.g., SEQ ID NO: 19) has substituted-in cysteine residues at two, three, four, five, or six of residues E78, K82, Q102, E114, F136, and Q166.
[0067] In some embodiments, CutA1 is conjugated to an effector molecule, which is any molecule that produces a desired effect. Effector molecules are sometimes referred to in the art as "payloads," especially in the context of conjugating the effector molecule to an antibody to form an antibody-drug conjugate (i.e., ADC). Thus, such payloads or effector molecules are well known in the art. Typically, the effector molecule comprises or consists of a label, a dye molecule, or a drug, i.e., a substance that exhibits a therapeutic or preventive physiological effect on the human or animal body when consumed internally. The drug is typically a pharmaceutical.
[0068] In some embodiments, the effector molecule is a nucleic acid, polynucleotide, or oligonucleotide. The nucleic acid, polynucleotide, or oligonucleotide may be DNA, RNA, XNA, LNA, or a mixture thereof. Typically, it is DNA or RNA. The nucleic acid, polynucleotide, or oligonucleotide may be single-stranded or double-stranded. In some embodiments, the oligonucleotide comprises or consists of 3 to 50 nucleotides (or nucleotide pairs if double-stranded), for example, 5 to 30 nucleotides (or pairs). Antibody-oligonucleotide conjugates are often known in the art as a subset of ADCs.
[0069] In some embodiments, CutA1 is conjugated with a drug to form a CutA1-drug conjugate.As will be clear to those skilled in the art, the drug can be of any type.The drug types that can be conjugated with CutA1 include but are not limited to analgesics, antibiotics, anticancer drugs, anticoagulants, antidepressants, antidiabetic drugs, antiepileptic drugs, antipsychotic drugs, anticonvulsants, antiviral drugs, cardiovascular drugs, depressants, sedatives and stimulants.
[0070] In some embodiments, two or more different drugs are conjugated to a single CutA1 molecule.An antibody-drug conjugate with dual payload is described in Yamazaki et al., Nature Communications volume 12, Article number: 3528 (2021), in which both MMAE and MMAF are conjugated to address breast tumor heterogeneity and drug resistance.
[0071] In some embodiments, the drug is an anticancer drug, e.g., a cytotoxic drug. In some embodiments, the anticancer drug is a tubulin inhibitor, e.g., a maytansinoid such as mertansine (also known as DM1), or an auristatin, or a taxol derivative. Examples of drugs that inhibit tubulin polymerization are auristatins such as tissue factor-directed monomethyl auristatin E (MMAE) and monomethyl auristatin F (MMAF), compounds derived from dolastatin 10 (e.g., TZT-1027, described in Kobayashi et al. Jpn J Cancer Res. 1997 Mar;88(3):316-27), and tubulysins such as tubulinsin A. In some embodiments, the anticancer drug is a DNA-damaging agent that causes cell death by cleaving DNA and / or causing DNA alkylation. Examples of DNA-damaging molecules are duocarmycin, calicheamicin, and pyrrolobenzodiazepines. In some embodiments, the drug is a topoisomerase I inhibitor, which binds to and stabilizes the complex between topoisomerase I and DNA, thereby preventing DNA religation and resulting in DNA damage. Examples of drugs that inhibit topoisomerase I are DXd and SN-38. In some embodiments, the drug is an RNA polymerase II inhibitor, such as α-amanitin. In some embodiments, the drug is an immunomodulator, such as a TLR agonist or a STING agonist (e.g., as described in Table A of Fu et al., Signal Transduction and Targeted Therapy, volume 7, Article number: 93 (2022)).
[0072] In some embodiments, drug is a small molecule drug.Usually, small molecule drug has molecular weight of 1000Da or less, more typically 750Da or less, or 500Da or less.This molecular weight is the molecular weight of drug molecule itself, and does not include any linker.
[0073] In some embodiments, the drug is a cytotoxic drug, typically a small molecule cytotoxic drug, such as DXd, mertansine, monomethyl auristatin E (MMAE), or monomethyl auristatin F (MMAF). Cytotoxic small molecule drugs are often referred to as chemotherapeutic drugs or chemotherapeutic agents, particularly in the context of cancer treatment.
[0074] Conjugation of a drug to CutA1 can be carried out by any known method. Typically, a drug is conjugated to CutA1 at a cysteine residue, more typically at a cysteine residue not present in the native sequence. When conjugating to a cysteine, the molecule to be conjugated may contain a maleimide. Cysteine residues readily react with maleimides to form succinimidyl thioether conjugates.
[0075] In some embodiments, the drug has a molecular weight greater than 1000 Da. In some embodiments, the drug is or comprises an oligopeptide or polypeptide comprising two or more, e.g., 3, 4, 5, 6, 7, 8, 9, 10 or more, or 20 or more, or 50 or more, or 100 or more amino acid residues covalently linked by peptide bonds. In some embodiments, the drug is an immunotoxin, such as Pseudomonas exotoxin A (PE), which has been described as an anticancer agent by Wolf and Beile (Int J Med Microbiol. 2009 Mar;299(3):161-76. doi: 10.1016 / j.ijmm.2008.08.003).
[0076] In some embodiments, CutA1 is conjugated to a dye molecule. In some embodiments, the dye molecule is a small molecule dye. Typically, small molecule dyes have a molecular weight of 1000 Da or less, more typically 750 Da or less, or 500 Da or less. In some embodiments, the dye is fluorescent, such as fluorescein. Conjugation can be performed by any known method. Typically, the dye is conjugated to CutA1 at a cysteine residue, more typically at a cysteine residue not present in the native sequence. When conjugated to cysteine, the molecule to be conjugated may contain a maleimide. Cysteine residues readily react with maleimides to form succinimidyl thioether conjugates.
[0077] In some embodiments, the effector molecule is a nanoparticle, such as a gold nanoparticle, a silica nanoparticle, a lipid nanoparticle, or a lipid-polydopamine hybrid nanoparticle ("LPN") as described in Yang et al., Acta Pharmaceutica Sinica B Volume 10, Issue 11, November 2020, pages 2212-2226. Typically, the nanoparticle is conjugated to one or more cysteine residues in CutA1 by site-specific conjugation to maleimide groups on the lipid nanoparticle or polydopamine (PDA) hybrid nanoparticle.
[0078] In certain aspects, CutA1, typically human CutA1 (e.g., SEQ ID NO: 19), has been genetically engineered to delete one or more residues from either or both ends of the native sequence to form a truncated CutA1 domain that is incorporated into a construct of the invention. In some embodiments, one or more residues are deleted from the N-terminus of CutA1, typically human CutA1. Typically, 5 to 70 residues, e.g., 10 to 59 residues, are deleted from the N-terminus of CutA1, typically human CutA1. In some embodiments, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 59, 60, 61, 62, 63, 64, 65, or 66 residues are deleted from the N-terminus of CutA1. In some embodiments, the truncated CutA1 begins at residue 30 (i.e., 29 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 33 (i.e., 32 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 44 (i.e., 43 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 60 (i.e., 59 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 67 (i.e., 66 N-terminal residues are deleted), e.g., "HsCutA1" illustrated in Figure 23b. 67~171 CACA" direct fusion construct.
[0079] In some embodiments, one or more residues are deleted from the C-terminus of CutA1, typically human CutA1. C-terminal deletions may be in place of or in addition to N-terminal deletions. Typically, 5 to 20 residues are deleted from the C-terminus of CutA1, typically human CutA1, e.g., 6 to 12 residues are deleted. In some embodiments, approximately 8 residues are deleted. In human CutA1, deleting 8 residues results in a truncated protein having a C-terminus at residue 171 (valine) of the wild-type sequence shown in FIG. 24 and SEQ ID NO: 19. In some embodiments, the truncated protein has a C-terminus at residue 168 (threonine) of the wild-type sequence shown in FIG. 24 and SEQ ID NO: 19. In some embodiments, the truncated protein has a C-terminus at residue 169, 170, 172, 173, 174, 175, or 176 of the wild-type sequence shown in FIG. 24 and SEQ ID NO: 19.
[0080] In some embodiments, the truncated CutA1 begins at any of residues 30-67 of SEQ ID NO:19. In some embodiments, the truncated CutA1 begins at any of residues 44-67 of SEQ ID NO:19. In certain embodiments, the truncated CutA1 consists of residues 44-179 of SEQ ID NO:19 (SEQ ID NO:29), residues 61-168, or residues 60-171 of SEQ ID NO:19, as shown in FIG. 24. In some embodiments, the truncated CutA1 begins at any of residues 44-67 of SEQ ID NO:19 and ends at any of residues 168-179. In other embodiments, the truncated CutA1 begins at any of residues 30-65 of SEQ ID NO:19 and ends at any of residues 165-179. In some embodiments, the truncated CutA1 begins at any of residues 44-67 of SEQ ID NO:19 and ends at any of residues 171-179.
[0081] Truncating CutA1 can improve the precision with which fusion constructs can be constructed. Figure 24 shows the generation, production, and biological effects of polypeptide constructs containing such truncated human CutA1 domains.
[0082] In certain aspects, CutA1 is human CutA1 that has been genetically engineered to remove one or more cysteine residues from the native sequence, resulting in truncation at the N-terminus and / or C-terminus. In some embodiments, the truncated CutA1 consists of residues 44-179 or residues 60-171 of SEQ ID NO: 19 as shown in Figure 24, with cysteine substitutions that include or consist of (i) C75A, C96A, or (ii) C75V, C96S. In other embodiments, the truncated CutA1 begins at any of residues 30-65 and ends at any of residues 165-179 of SEQ ID NO: 19, and also has cysteine substitutions that include or consist of (i) C75A, C96A, or (ii) C75V, C96S.
[0083] In certain aspects, CutA1 is human CutA1 that has been genetically engineered to remove one or more cysteine residues from the native sequence, truncated at the N-terminus and / or C-terminus, and substituted at least one non-cysteine residue with a cysteine residue. In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19), with a cysteine residue substitution removal from the native CutA1 sequence comprising or consisting of (i) C75A, C96A, or (ii) C75V, C96S, and with at least one cysteine at a residue position that is not a cysteine residue in the native sequence.
[0084] In some embodiments, the truncated CutA1 consists of residues 44 to 179 of human CutA1 (SEQ ID NO: 19), with cysteine substitution deletions including or consisting of C75A, C96A, and cysteine residue substitution incorporation at one, two, or three of residues K82, E114, and F136.
[0085] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19), with a cysteine substitution deletion including or consisting of C75A, C96A, and an incorporation of a cysteine residue substitution at one or more of residues V64, E78, K79, E83, K91, Q102, K110, S139, F158, and Q166. In some embodiments, this variant CutA1 has an incorporation of a cysteine residue substitution at 2, 3, 4, 5, 6, 7, 8, 9, or 10 of residues V64, E78, K79, E83, K91, Q102, K110, S139, F158, and Q166.
[0086] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19), with a cysteine substitution deletion including or consisting of C75A, C96A, and an incorporation of a cysteine residue substitution at one or more of residues E78, Q102, and Q166. In some embodiments, this variant CutA1 has an incorporation of a cysteine residue substitution at two or three of residues E78, Q102, and Q166.
[0087] In some embodiments, the truncated CutA1 consists of residues 44-179 of human CutA1 (SEQ ID NO: 19), with a cysteine substitution deletion including or consisting of C75A, C96A, and an incorporation of a cysteine residue substitution at one or more of residues E78, K82, Q102, E114, F136, and Q166. In some embodiments, this variant CutA1 has an incorporation of a cysteine residue substitution at two, three, four, five, or six of residues E78, K82, Q102, E114, F136, and Q166.
[0088] In some embodiments, a variant CutA1 comprises or consists of any of the sequences defined above or the specific sequences described in the Examples, but contains 1 to 10 amino acid substitutions at residues not designated as cysteine residue substitution deletions or incorporations. For example, some embodiments provide variant human CutA1 sequences containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions at positions other than 95, 64, 76, 78, 79, 82, 83, 91, 95, 102, 110, 114, 136, 139, 158, and 166. These 1 to 10 substitutions may be for any of the 20 standard amino acids, but typically will not replace non-cysteine residues with cysteines. Typically, these 1 to 10 substitutions will be conservative substitutions. In some embodiments, one, two, three, or more of these 1-10 substitutions can be unnatural amino acids (UAAs), also known as non-proteinogenic amino acids. Unnatural amino acids can be useful for enabling conjugation of molecules of interest via bioorthogonal chemistry, e.g., click chemistry. Unnatural amino acids are known in the art and include D-amino acids, homoamino acids, N-methylamino acids, hydroxyproline (Hyp), beta-alanine, citrulline (Cit), ornithine (Orn), norleucine (Nle), 3-nitrotyrosine, nitroarginine, and pyroglutamic acid (Pyr). For example, propargyl lysine is an unnatural amino acid that, when incorporated into a protein, can be utilized to attach commercially available fluorescent azide dyes via copper-catalyzed alkyne-azide cycloaddition (click) reaction (also known as click reaction). Other UAAs suitable for site-specific modification of polypeptide sequences include: 1: 3-(6-acetylnaphthalen-2-ylamino)-2-aminopropanoic acid (Anap); 2: (S)-1-carboxy-3-(7-hydroxy-2-oxo-2H-chromen-4-yl)propan-1-aminium (CouAA); 3: 3-(5-(dimethylamino)naphthalene-1-sulfonamido)propanoic acid (dansylalanine), 4: Nε-p-azidobenzyloxycarbonyl lysine (PABK), 5: propargyl-L-lysine (PrK), 6: Nε-(1-methylcycloprop-2-enecarboxamido) lysine (CpK), 7: Nε-acryl lysine (AcrK), 8: Nε-(cyclooct-2-yn-1-yloxy)carbonyl)L-lysine (CoK), 9: bicyclo[6.1.0]non-4-yn-9-ylmethanol lysine (BCNK), 10: trans-cyclooct-2-ene lysine (2'-TCOK), 11: trans-cyclooct-4-ene lysine (4'-TCOK), 12: di These include oxo-TCO lysine (DOTCOK), 13: 3-(2-cyclobuten-1-yl)propanoic acid (CbK), 14: Nε-5-norbornen-2-yloxycarbonyl-L-lysine (NBOK), 15: cyclooctyne lysine (SCOK), 16: 5-norbornen-2-ol tyrosine (NOR), 17: cyclooct-2-ynol tyrosine (COY), 18: (E)-2-(cyclooct-4-en-1-yloxyl)ethanol tyrosine (DS1 / 2), 19: azidohomoalanine (AHA), 20: homopropargylglycine (HPG), 21: azidonorleucine (ANL), and 22: Nε-2-azideoethyloxycarbonyl-L-lysine (NEAK).
[0089] Without being bound by theory, trimerization or multimerization of binder proteins may be a useful property for enhancing or enabling biological effects, and novel components for trimerization or multimerization have been extensively investigated (Cuesta, AM et al., Trends in Biotechnology, 2010). CutA1 is suitable for fusing effector proteins or binding domains to both the N-terminus and C-terminus of CutA1, and is also a novel component for stable multimerization of binder proteins. The data presented below, particularly Figures 22 and 23, demonstrate that CutA1 may be a suitable domain for conjugation or fusion with binder proteins or effector proteins, generally promoting trimerization to exert biological effects, for example, after conjugation with a binder protein via a catcher domain or as a direct fusion with a binder protein. Thus, the present invention also provides polypeptides comprising a binding domain or effector protein fused to the N-terminus or C-terminus of CutA1. In some embodiments, CutA1, typically human CutA1, has been modified as described elsewhere herein.
[0090] In some embodiments, the structural domain (or subunit monomer of the oligomer core) of the polypeptide construct is a TNF family protein, including a TNF protein, a TL1A protein, an OX40L protein, or a CD40L protein, typically a human protein. In certain embodiments, the TNF, TL1A, OX40L, or CD40L protein, typically a human protein, is genetically engineered to remove one or more cysteine residues from the native sequence. Typically, this is the replacement of one or more cysteine residues with one or more non-cysteine residues. Removal of one or more cysteine residues is beneficial because it allows targeted cysteine conjugation at non-native sites or in fusion proteins. Without being bound by theory, it may be beneficial to remove unpaired cysteines, as unpaired cysteines often interfere with stability and downstream applications. In certain aspects, a TNF, TL1A, OX40L, or CD40L protein, typically a human protein, has been engineered to delete one or more residues from either or both ends of the native sequence to form a truncated TNF, TL1A, OX40L, or CD40L domain that is incorporated into a construct of the invention. In some embodiments, one or more residues, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more, have been deleted from the N-terminus of the native sequence. In some embodiments, one or more residues, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more, have been deleted from the C-terminus of the native sequence. In some embodiments, one or more residues, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more, have been deleted from the N-terminus of the native sequence, and one or more residues, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more, have been deleted from the C-terminus of the native sequence.
[0091] In some aspects, as shown in the examples, the modular assemblies provided by the present invention can be combined with a simple post-assembly cleanup. This is particularly useful for producing drug candidates for downstream analysis. In some embodiments, such methods can be automated. In some embodiments, the cleanup uses beads, e.g., paramagnetic beads, that can bind to the appropriate catcher and quench unconjugated tagged proteins from the tagged conjugate-catcher core assembly (see Figure 37). This allows for cost-effective, rapid, and easily scalable purification of the assembly (by removing unbound tagged conjugates).
[0092] Cleanup methods typically involve contacting the post-conjugation reaction mixture with a cleanup binding domain capable of binding to unbound conjugates, and the cleanup binding domain bound to the previously unbound conjugates can then be removed from the reaction mixture.
[0093] This cleanup method is particularly suitable for plate-based formats that can be automated using liquid-handling robots. For example, GST-catchers can be loaded onto glutathione-conjugated beads (e.g., Figure 35), and unbound excess can be removed by washing with PBS (e.g., Figure 36). The GST-catcher-bound beads can then be added to each conjugation reaction, typically in excess relative to the corresponding unconjugated tagged conjugate (e.g., a two-fold excess of GST-catcher). The beads can then be removed by appropriate means. For example, if the beads are magnetic, they can be removed by magnetic particle capture, leaving behind a supernatant containing only fully assembled molecules (e.g., Figure 38). This method is advantageous over methods in which catchers are covalently conjugated to beads because it can be performed on demand at various scales. Furthermore, no complex chemical processes or reducing conditions are required. Furthermore, the design of the catcher construct, the ratio of catcher proteins, or the resin can be easily modified as needed.
[0094] This cleanup technique is a general improvement to the art and may be applicable to other modular assembly methods, particularly those described in Driscoll et al., September 2023, bioRxiv 2023.08.31.555700; doi: https: / / doi.org / 10.1101 / 2023.08.31.555700.
[0095] This cleanup technique can be adapted for techniques such as sortase-based assembly, as described in Andres et al., Mol Cancer Ther (2020) 19 (4): 1080-1088 "High-Throughput Generation of Bispecific Binding Proteins by Sortase A-Mediated Coupling for Direct Functional Screening in Cell Culture."
[0096] The cleanup step typically uses a fusion protein capable of binding excess binders in the reaction mixture after assembly of the catcher and tagged binders. Such cleanup fusion proteins typically comprise a protein useful in the immediately preceding cleanup method fused to at least one binding domain described herein, typically a catcher domain capable of binding excess binders in the reaction mixture. Proteins useful in the cleanup step typically form a protein domain-based connection useful in the purification process. This connection may be a covalent bond (e.g., a halotag, a modified haloalkane dehalogenase designed to bind covalently or noncovalently (e.g., maltose-binding protein [MBP]) to a synthetic ligand containing a chloroalkane linker attached to various useful molecules such as fluorescent dyes, affinity handles, or solid surfaces). As noted above, this construct is typically a glutathione-S-transferase "GST" sequence capable of binding GSH on beads. In some embodiments, the cleanup fusion protein is attached to a solid support, such as beads. In some embodiments, the beads may be magnetic or paramagnetic.
[0097] GST / GSH is a particularly suitable system due to the commercially available paramagnetic beads with high protein capacity at relatively low cost.
[0098] The cleanup fusion protein may contain a single type of binding domain, allowing it to clean up one excess binder. This single binding domain may be provided in a single copy (e.g., GST-SpyCatcher) or, optionally, in multiple copies (e.g., in the form of GST-SpyCatcher-SpyCatcher). If there are two or more excess binders to be cleaned up after assembly, two or more cleanup fusion proteins, each with a different single type of binding domain, may be used in combination. GST-SpyCatcher and GST-DogCatcher are examples of cleanup fusion proteins. In some exemplary embodiments, for example, the first cleanup fusion protein is GST-SpyCatcher and the second cleanup fusion protein is GST-DogCatcher. Specific sequences are provided in the Examples for GST-SpyCatcher003 and GST-DogCatcher. In another exemplary embodiment, the cleanup fusion protein comprises GST-SpyCatcher002 or GST-SpyCatcher-003. In another exemplary embodiment, the cleanup fusion protein comprises GST-SpyCatcher002 and GST-SpyCatcher-003 used in combination. Other cleanup fusion proteins can include inactive variants of SpyCatcher, such as SpyCatcher002KA, SpyDock, or SpySwitch (e.g., as described in Khairil Anuar et al, Nature Communications volume 10, Article number: 1734 (2019), particularly Figure 7).
[0099] MBP-SpyCatcher and MBP-DogCatcher are further examples of cleanup fusion proteins. In some exemplary embodiments, for example, the first cleanup fusion protein is MBP-SpyCatcher and the second cleanup fusion protein is MBP-DogCatcher. In other exemplary embodiments, the cleanup fusion protein comprises MBP-SpyCatcher002 or MBP-SpyCatcher-003. In other exemplary embodiments, the cleanup fusion protein comprises MBP-SpyCatcher002 and MBP-SpyCatcher-003 used in combination.
[0100] Halotag-SpyCatcher and Halotag-DogCatcher are further examples of cleanup fusion proteins. In some exemplary embodiments, for example, the first cleanup fusion protein is Halotag-SpyCatcher and the second cleanup fusion protein is Halotag-DogCatcher. In other exemplary embodiments, the cleanup fusion protein comprises Halotag-SpyCatcher002 or Halotag-SpyCatcher-003. In other exemplary embodiments, the cleanup fusion protein comprises Halotag-SpyCatcher002 and Halotag-SpyCatcher-003 used in combination.
[0101] In some embodiments, two or more binding domains are attached to a single protein useful in the cleanup methods described above. Multiple binding domains, typically multiple catcher domains, can be attached to any suitable position of the protein useful for cleanup (e.g., GST, MBP, Halotag). Multiple binding (e.g., catcher) domains can be arranged in tandem, for example, at one end of the cleanup protein, or at different ends. As with other embodiments described herein, a linker peptide (e.g., 2-30 amino acid residues) can be included between the separate components of the fusion, if desired. Some exemplary formats in which each Catcher is different include Catcher1-GST-Catcher2, Catcher1-Catcher2-GST, GST-Catcher1-Catcher2, Catcher1-MBP-Catcher2, Catcher1-Catcher2-MBP, MBP-Catcher1-Catcher2, Catcher1-halotag-Catcher2, Catcher1-Catcher2-halotag, or halotag-Catcher1-Catcher2.
[0102] In some embodiments, the present invention provides polypeptides comprising a first binding domain at the N-terminus and a second binding domain at the C-terminus, the first and second binding domains being separated by a structural domain useful for the cleanup methods described immediately above, where the first and second binding domains are the same or different. In such embodiments, the structural domain is typically capable of forming a protein domain-based connection useful for the purification process. This connection can be covalent (e.g., halotag) or non-covalent (e.g., maltose binding protein [MBP]). As noted above, this construct is typically a glutathione-S-transferase "GST" sequence capable of binding to GSH on beads.
[0103] In some embodiments, the present disclosure provides a TNF receptor superfamily (TNFRSF) agonist, such as a TRAIL receptor agonist, e.g., a DR5 agonist, comprising CutA1 or a variant thereof described herein. TNFRSF agonists are known in the art. Typically, agonism is provided by one or more binding domains (e.g., Fab, ScFv, nanobody) attached to CutA1, for example, at the N-terminus and / or C-terminus. Agonistic antibodies against TRAIL receptors such as DR5 are known in the art, such as IGM8444 and TAS266 (see, for example, (i) Wang et al., ASCO 2020 Annual Meeting Abstract: "IGM-8444 as a potent agonistic Death Receptor 5 (DR5) IgM antibody: Induction of tumor cytotoxicity, combination with chemotherapy, and in vitro safety profile."; (ii) Huet et al., Cancer Res (2012) 72 (8_Supplement): 3853. "Abstract 3853: TAS266, a novel tetrameric nanobody agonist targeting death receptor 5 (DR5), elicits superior antitumor efficacy than conventional DR5-targeted approaches"; and (iii) WO 2017 / 011837).
[0104] The DR5 conjugate itself may not be agonizing. Typically, without wishing to be bound by theory, anti-DR5 antibodies (including antigen-binding fragments and genetically engineered forms thereof, as disclosed herein) must have a specific valency state, e.g., bivalent, trivalent, or tetravalent, to exert an agonistic effect on DR5. The valency of the conjugate can be increased by connecting it to a multimerization scaffold (e.g., CutA1). The data in at least Examples 11 and 14 indicate that constructs of the present invention can achieve effective agonism if they are at least trivalent. This indicates that CutA1 is particularly useful for targeting TNFRSF, likely due to its small size and geometric shape (including its trimeric nature) similar to the natural ligand of the TNFRSF receptor. Figure 11 provides an example of a ligand of the TNFSF family for reference. Thus, in one embodiment, a trivalent CutA1-DR5 agonist is provided. In some embodiments, a hexavalent CutA1-DR5 agonist is provided.
[0105] The C- and N-terminal positions can differ significantly between the two trimers, and one may be preferred over the other, e.g., in the case of DR5, a CutA1 C-terminal fusion may sometimes be preferred (see Figures 22, 23c, and 32).
[0106] In some embodiments, provide CutA1-DR5 agonist multimer.In some embodiments, provide the trimer of CutA1-DR5 agonist construction.In some embodiments, multimer can be dimer, or can be higher order, for example, tetramer, pentamer, hexamer, heptamer or even higher order.
[0107] Typically, the CutA1-DR5 agonist construct includes a second binding domain for a different target, i.e., it is a bispecific construct. The second binding domain is typically provided at a different position on CutA1, most typically at the opposite end. However, as described elsewhere herein, it may also be provided in tandem with the DR5 agonist at the same position on CutA1. Thus, suitable construct formats include: binding domain 2--CutA1--DR5 conjugate, DR5 conjugate--CutA1--binding domain 2, DR5 conjugate--binding domain 2--CutA1, CutA1--DR5 conjugate--binding domain 2, binding domain 2--DR5 conjugate--CutA1, or CutA1--binding domain 2--DR5 conjugate (shown in the conventional N to C orientation). Each of these constructs is provided in, for example, multivalent (eg, trivalent, hexavalent) and / or multimeric (eg, trimeric or hexameric) formats, as described above.
[0108] TAS266 is a tetrameric nanobody agonist targeting death receptor 5 (DR5). As discussed in the Examples, the inventors hypothesized that the severe hepatotoxicity associated with the clinical failure of TAS266 was the result of anti-drug antibody-dependent DR5 hyperclustering and consequent excessive induction of apoptosis in liver cells (Inhibrx, WO 2017 / 011837). Such antibodies are present in commercially available intravenous immunoglobulin (IVIG) solutions (Inhibrx, WO 2017 / 011837). In contrast to TAS266, a CutA1-based formulation (exemplified as L11-HsCutA1-L7) did not increase IVIG-dependent toxicity in the HepG2 liver cell line (see, e.g., Figure 41). Without wishing to be bound by theory, the inventors observed that by reducing the valency of DR5-targeting conjugates from 4 to 3 and simultaneously introducing bispecificity, CutA1-based constructs, exemplified by L11-HsCutA1-L7, are able to induce milder adverse effects and exhibit increased tumor specificity.
[0109] In some embodiments, a construct is provided comprising CutA1 or a variant thereof and a DR5-binding domain. Typically, the construct will also comprise a second binding domain for another target. In some embodiments, the DR5-binding domain competes with a TAS266 tetramer construct for binding to DR5. In some embodiments, the DR5-binding domain comprises a TAS266 "VHH" binding domain (e.g., SEQ ID NO: 106). In some embodiments, the construct comprises CutA1 or a variant thereof described herein and a TAS266 "VHH" binding domain (e.g., as exemplified by SEQ ID NO: 102), and optionally a second binding domain (e.g., scFv, VHH, or Fab) that binds to a different target protein. A linker sequence may be included between the anti-DR5 VHH and CutA1 and / or between the CutA1 and the second binding domain. Typically, the two binding domains are at opposite ends of CutA1, but may also be provided contiguously at one end of CutA1.
[0110] SEQ ID NO: 106 provides an example of a DR5 binder useful for constructing multivalent DR5 agonists, such as tetravalent TAS266 or trimeric CutA1-DR5. In some embodiments, the DR5 binder competes with SEQ ID NO: 106 for binding to DR5. This competition can be determined, for example, by a surface plasmon resonance assay, such as a Biacore assay. In some embodiments, the DR5 agonist has at least 70% identity, at least 80% identity, at least 90% identity, e.g., at least 95% identity, or at least 99% identity, to SEQ ID NO: 106. Typically, any mutations are present only in framework residues, with no mutations tolerated in CDR residues.
[0111] SEQ ID NO: 102 provides an example of an agonist hsCutA1-DR5 conjugate. Sequences having at least 70% identity, at least 80% identity, at least 90% identity, for example at least 95% identity, or at least 99% identity to SEQ ID NO: 102 can be used. Typically, any mutations are present only in framework residues, with no mutations allowed in CDR residues.
[0112] In a further aspect, as exemplified in Example 14, the present invention provides a further engineered variant of CutA1, single-chain CutA1, which is an engineered variant of CutA1 comprising multiple CutA1 molecules (whether full-length or truncated, and whether optionally substituted or otherwise engineered as detailed herein), wherein all "monomers" of CutA1 are optionally covalently connected by a linker.
[0113] In one embodiment of this aspect, the single polypeptide chain comprises multiple CutA1 sequences, optionally with a linker polypeptide 3-30 amino acids in length interposed between some or all of the CutA1 sequences, optionally with at least some and optionally all of the CutA1 sequences as defined elsewhere herein. In a further embodiment, the single polypeptide chain comprises three human CutA1 sequences, each separated by a linker polypeptide (optionally comprising an isopeptide bond-forming domain at the N-terminus and / or C-terminus).
[0114] Typically, single-chain CutA1 is expressed as a fusion protein containing two or more CutA1 molecules, e.g., three or more CutA1 molecules, e.g., four, five, six, seven, eight, or more CutA1 molecules. Three CutA1 molecules is typical. Some or all of the CutA1 molecules in the chain are typically separated by a polypeptide linker, which is typically 3 to 30 amino acid residues in length, more typically 10 to 20 amino acid residues in length, e.g., 12, 13, 14, 15, 16, or 17 amino acids in length. A typical, non-limiting linker is 15 amino acids in length, e.g., (GGGGS)3.
[0115] For example, in a trimeric single-chain CutA1 construct, each of the three CutA1 components can be connected by a linker to generate the construct. The construct may also include a binding moiety of interest, typically (but not necessarily) at its terminus. An exemplary single-chain CutA1 construct is SpyCatcher-linker-CutA1-linker-CutA1-linker-CutA1-DogCatcher. This single-chain construct may be particularly useful for providing a 1+1 bivalent control with the same geometry as the 3+3 molecule. This single-chain CutA1 has numerous potential uses, including, but not limited to, as an in vitro or in vivo control scaffold to compare with the 3+3 CutA1, as a scaffold for evaluating 1+1 molecules with fixed geometries, and in vivo applications when combined with other CutA1 technologies, such as ADCs (e.g., to create ADCs with high drug-antibody ratios (DARs)).
[0116] The present invention allows for highly versatile screening of large numbers of effector moieties. The effector moiety may be any protein domain, including but not limited to Fab / Fv regions or other antigen-binding domains. The present invention can also be used to investigate molecular effects achieved through higher valency interactions that would not be observed using conventional bispecific antibodies or other approaches. Thus, the present invention provides a system for high-throughput screening of bispecific molecular combinations that can uncover effects observed only through higher valency interactions. The present invention also provides novel therapeutic drug candidates that can be identified according to the methods provided herein. The therapeutic drug candidates offer the benefits of multiple functionalities and increased valency.
[0117] The art has described limited uses of multivalent protein scaffolds. One attempt described by Brune et al., Bioconjugate Chemistry, 28(5), pp.1544-1551, concerns the use of heptameric assemblies in which antigens are assembled on opposite faces of the IMX313 heptameric core as a potential malaria vaccine. However, because such constructs present the attached antigens on opposite sides, they are not suitable for engaging multiple receptors on the same cell, limiting their application in same-cell binding to treat disorders such as cancer and autoimmune diseases.
[0118] Thus, there is a need for new and / or improved methods for phenotypic screening of combinations of effector moieties (also referred to herein as ligands), and for developing new therapeutic agents. Previous methods have not been able to increase the valency of the functional combinations investigated and / or have not recognized the synergistic benefits of presenting antigens on the same face of a protein scaffold. There is also a need for new and / or improved therapeutic agents, including those provided herein, that can be designed and identified in accordance with the present disclosure. [Brief explanation of the drawings]
[0119] [Figure 1] FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a first binding site and a second binding site are located on the same face of the oligomeric core and therefore the same face of the multivalent protein scaffold. [Figure 2] FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a first binding site and a second binding site are located on opposite faces of the oligomeric core and thus on opposite faces of the multivalent protein scaffold. [Figure 3] FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a plurality of first binding sites and a plurality of second binding sites are located on the same face of the oligomeric core, and therefore the same face of the multivalent protein scaffold, for engagement with a surface. [Figure 4] FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a plurality of first binding sites and a plurality of second binding sites are located on opposite faces of an oligomeric core and thus on opposite faces of the multivalent protein scaffold, and therefore cannot simultaneously engage a surface. [Figure 5] FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a first binding site and a second binding site are located on the same face of the oligomeric core, and thus on the same face of the multivalent protein scaffold, in a face-to-face orientation. [Figure 6] FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a first binding site and a second binding site are located on the same face of the oligomeric core, and thus on the same face of the multivalent protein scaffold, in a front-to-side orientation. [Figure 7]FIG. 1 is a schematic diagram illustrating a multivalent protein scaffold as described herein, wherein a first binding site and a second binding site are located on the same face of the oligomeric core, and thus on the same face of the multivalent protein scaffold, in an orientation intermediate between a front-to-front orientation and a front-to-side orientation. [Figure 8] FIG. 1 is a schematic diagram showing the angle (X) formed between a first binding site and a second binding site attached to a subunit monomer of an oligomeric core of a multivalent protein scaffold described herein. [Figure 9] 1 is a schematic diagram illustrating an embodiment of the invention in which tandem fusions of first and second binding sites are attached to an oligomeric core as described herein, allowing for the production of multivalent protein scaffolds. A variety of different geometries and stoichiometries can be produced using the methods disclosed herein. [Figure 10]Figure 2 shows a 2D depiction of fusion sites in the "cis" position. a) For a given target plane, the longest cross-section of the core protein is determined by an orthogonal line drawn from the target plane, characterized as the distance dc. Parallel planes intersecting the midpoint of this cross-section are shown. In this case, sites of protein conjugation (stars, circles) are considered to be preferentially in the "cis" position if, for all conjugation sites, the distance of the shortest path from the conjugation site to the target plane that does not intersect with the protein surface is less than a certain percentage of the distance dc, e.g., less than 50% of dc. b) An example of a protein in which all binding sites are in the cis position. c) An example of a protein in which all binding sites are in the cis position relative to each other according to a), such that all secondary binding sites (circles) are characterized by a minimum path length just below the 50% dc threshold. d) The projected distance (across the protein) from all binding sites to the target plane is the same as in c), but the protein is characterized by a geometry that obstructs the shortest path from all secondary binding sites (circles), so that these binding sites are not accessible and / or do not lie on the same side of the scaffold. e, f) The binding sites are too far apart to be considered cis to any target plane. [Figure 11]Figure 1 provides an overview of the protein structures referenced herein. The protein structure of a given PDB ID is visualized as a schematic diagram using chains of different colors. The N- and C-termini of a single monomer are each annotated and point approximately toward the binding surface. The symmetry of the protein structure, e.g., circular C3 symmetry, circular C4 symmetry, and dihedral D2 symmetry, is indicated in parentheses. *: Protein symmetry is inferred from the NMR structure. †1PK6 is a heteromer characterized by C1 symmetry, but the domains are homomeric and their arrangement resembles C3 symmetry. By selecting a multimeric protein core with a suitable monomer configuration, multiple binding sites can be projected per monomer to form a single binding surface. In addition to recombinant fusion, such proteins can then be used for modular assembly, e.g., by recombinant fusion of SpyCatcher and SnoopCatcher or DogCatcher at the N- and C-termini, to rapidly confer multivalency or other properties to suitably modified peptides or proteins, e.g., via recombinant fusion of SpyCatcher at the N-terminus and SnoopCatcher at the C-terminus of a monomer of an oligomeric core protein. Examples include: (1) HsCutA1 (PDB ID 2ZFH), a highly thermostable human copper-binding protein characterized by the close proximity of both the N- and C-termini of each monomer, resulting in a C3 structure when projected onto a single plane of the assembled trimer; and (2) PhCutA1 (PDB ID 4NYO), a hyperthermostable homologue from Pyrococcus horikoshii that shares high structural similarity with HsCutA1. (3) NC1 domain of type X collagen (PDB ID: 1GR3), (4) NC1 domain of type VIII collagen (PDB ID: 1O91), (5) macrophage migration inhibitory factor 2 (PDB ID: 7MSE), (6) tumor necrosis factor (PDB ID: 1TNF), and (7) TNF-like protein TL1A (PDB ID: 2RE9). [Figure 12]Figure 1 shows that cis-oriented multimeric protein complexes can be easily expressed and prepared by standard protein purification methods. a) Ni-NTA purification of H6-SpC-PhCutA1-SnC [SEQ ID NO: 21]. H6-SpC-PhCutA1-SnC was easily expressed in Escherichia coli (E. coli) BL21(DE3) and retained the intact trimeric structure characteristic of the ultrastable PhCutA1 even after boiling in SDS-PAGE loading buffer. Washes 1 and 2 were performed with 10 column volumes of equilibration buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 10 mM imidazole). Washes 3 and 4 were performed with 10 column volumes of wash buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 30 mM imidazole). All elution steps were performed using two column volumes of elution buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 200 mM imidazole). Samples were analyzed using a 12% SDS-PAGE gel with Coomassie staining. P - lysate pellet; CL - clarified lysate; FT - flow-through; W - wash; E - elution. b, c) Size-exclusion chromatography of H6-SpC-PhCutA1-SnC using HiLoad 16 / 600 Superdex 200 pg. b) UV A280 absorbance chromatogram with the H6-SpC-PhCutA1-SnC peak highlighted. c) SDS-PAGE of a 2 mL fraction of H6-SpC-PhCutA1-SnC from the highlighted region of the chromatogram above. [Figure 13]Transition from a dihedral hexamer to a circularly symmetric trimeric core protein for cis-oriented display. a) Homohexameric antiparallel coiled-coils are characterized by the N- and C-termini being on two opposite sides of the protein assembly (PDB ID: 5W0J). b) Point mutagenesis can induce heteromeric assemblies from homomeric assemblies by, for example, introducing or modifying salt bridges to "lock" the assembly in one direction (PDB ID: 5VTE). c) Linking the ends of heteromeric assemblies induces homomeric assemblies suitable for cis-oriented display (compare with HIV GP41, PDB ID: 1I5Y). Black-to-white direction (structure, schematic) and arrow direction (schematic) are from N- to C-terminus. Structures were visualized with PyMOL. [Figure 14] SpC-PhCutA1-SnC is a highly stable trimeric protein. a) Samples of SpC-PhCutA1-SnC were heated at 97°C for 2 hours in either 0%, 0.5%, or 1% SDS in PBS before adding SDS loading dye. Samples were separated on 12% SDS-PAGE and stained with Coomassie gel. The unheated control sample without SDS is shown in the first lane. As the SDS concentration increased, partial monomerization of the trimer was observed, confirming that the trimer was not covalently crosslinked. b) SpC-PhCutA1-SnC was maintained even after long-term storage. Protein aliquots were stored at 4°C, ambient temperature (21°C), or 37°C for 7 days, after which samples were prepared in SDS loading buffer and separated by SDS-PAGE. The protein showed little sign of degradation at 21°C to 37°C compared to storage at 4°C. Optionally, the protease inhibitor PMSF was added, which had a similar effect. [Figure 15]Figure 1. Preparation of SpyTagged and SnoopTagged ligand components for modular assembly into platform proteins. a) Purification of SnT-L1 from a 200 mL BL21(DE3) culture using Ni-NTA chromatography. Each wash step was performed using 5 mL of Ni-NTA wash buffer. Each elution step was performed using 2 mL of Ni-NTA elution buffer. Samples were analyzed using a 12% SDS-PAGE gel with Coomassie staining. pel. - lysate pellet; FT - flow-through; W - wash; E - elution; L - ladder. b) Purification of L2-SpT as in a). c) Size-exclusion chromatography of SnT-L1 after Ni-NTA purification. UV A280 absorbance chromatogram of SnT-L1 on an AKTA Pure 25 with a HiLoad Superdex 16 / 600 75 pg column. Inset: SDS-PAGE of a 2 mL fraction of SnT-L1 from the highlighted region of the chromatogram. d) Size exclusion chromatography of L2-SpT as in c). [Figure 16]Figure 1 shows that SpC-PhCutA1-SnC promotes stable trimerization of Spy- and Snoop-tagged proteins. a) H6-SpC-PhCutA1-SnC was conjugated with a 1:2:2 molar excess of SnT-L1 or L2-SpT for the indicated times. After conjugation, samples were supplemented with SDS loading buffer and denatured by boiling all samples at 95°C for 5 min. Samples were separated on 8% and 16% SDS-PAGE gels and subsequently Coomassie-stained. We observed time-dependent conjugation of SpC-PhCutA1-SnC with SnT-L1 or L2-SpT, with the ligand components being consumed. b) Conjugation of SpC-PC-SnC with SnT-L1 and L2-SpT, as in a). c) Conjugation of SpC-PhCutA1-SnC with SnT-L1 and L2-SpT. SpC-PhCutA1-SnC was incubated with SnT-L1 and / or L2-SpT in a 1:2:2 molar excess, and the samples were incubated at 25°C for 64 hours. Samples were analyzed using a 16% SDS-PAGE gel with Coomassie staining. Notably, SpC-PhCutA1-SnC and SnT-L1 / L2-SpT conjugated to completion and retained the characteristic hyperthermostability of PhCutA1. d) Conjugation of SpC-PC-SnC with SnT-L1 and L2-SpT, as in c). SpC-PC-SnC and SnT-L1 / L2-SpT conjugation was carried out to completion. [Figure 17]High molecular weight scaffolds enable post-assembly purification by dialysis. a) Ni-NTA-purified SpC-PhCutA1-SnC and b) SpC-PhCutA1-SnC:SnT-L1:L2-SpT assemblies were dialyzed using a 96-well format. a) SpC-P2-SnC was purified using Ni-NTA chromatography, and the elution fractions were combined and concentrated. b) Protein conjugation of SpC-PhCutA1-SnC, SnT-L1, and L2-SpT was performed at 25°C for 2 hours. Dialysis was performed in an HTDialysis 96-well block using a 100 kDa MWCO membrane with a 1:1 sample-to-dialysis buffer ratio. Dialysis was performed at ambient temperature with orbital shaking. PBS dialysate was replaced every 30 minutes for up to 90 minutes. Dialysis for 24 hours shows the removal of protein impurities (a-b) and unconjugated ligand (b) to equilibrium (sample to dialysis buffer ratio 1:1). Samples were boiled with reducing SDS loading dye, separated on 12% SDS-PAGE, and visualized by Coomassie staining. c) Alternatively, the SpC-PhCutA1-SnC:SnT-L1:L2-SpT conjugate was purified by dialysis in a 12-well plate format, characterized by a sample to dialysis buffer ratio of 1:30. The 1:1 conjugation of SpC-PhCutA1-SnC, SnT-L1, and L2-SpT was set up at 25°C for 24 hours. Dialysis was performed without stirring using an HTDialysis 12-well block and a 100 kDa MWCO cellulose membrane at ambient temperature for 16 hours. A sample volume of 100 μL and a PBS dialysate volume of 3 mL were used. Both contained 1x PMSF. Samples and dialysate were collected at 2, 4, 8, and 16 hours. Samples were analyzed using 14% SDS-PAGE gels and Coomassie staining. S = dialysate sample, D = dialysate (PBS). [Figure 18]Figure 1 shows the variation of the seed core component (PhCutA1 to MIF2m or HsCutA1), the protein component for conjugation (SpC / SnC to SpC3 / DgC), and the variable linker length (GGGGSGGGGSGGGGGS for MIF2m and GGGGS for HsCutA), highlighting the potential for rapid prototyping. a, b) Samples derived from Ni-NTA purification of H6-SpC3-HsCutA1-DgC or H6-SpC3-MIF2m-DgC. TL - total lysate; P - lysate pellet; CL - clarified lysate; FT - flow-through; W - wash; E - elution. Samples were analyzed using SDS-PAGE gels with Coomassie staining. c) Both SpC3-MIF2m-DgC and SpC3-HsCutA1-DgC can be rapidly conjugated to DogTagged or SpyTagged proteins. The platform proteins were incubated at a 1:1.5:1.5 molar ratio for 16 hours at 25°C. d) Incubation of H6-SpC3-HsCutA1-DgC with 0.1% glutaraldehyde demonstrates crosslinking of the trimeric protein in solution. 10 μM H6-SpC3-HsCutA1-DgC was crosslinked for 0-20 minutes using 0.1% glutaraldehyde. The reaction was incubated at 37°C and stopped by the addition of 100 mM Tris, pH 8.8. The trimeric crosslinked species formed rapidly, and prolonged incubation resulted in the consumption of the monomeric H6-SpC3-HsCutA1-DgC and the formation of a crosslinked species of the molecular weight expected for a crosslink between two trimers. e) SpC3-HsCutA1-DgC was reacted with L1-SpT and L3-DgT in a 1:2:2 molar ratio for 16 h at 25°C. The SpC3-HsCutA1-DgC alone, L1-SpT:SpC3-HsCutA1-DgC, SpC3-HsCutA1-DgC:L3-DgT, and L1-SpT:SpC3-HsCutA1-DgC:L3-DgT samples were then subjected to size-exclusion chromatography using a Superose6 Increase5 / 150 GL column, which demonstrated an increase in the hydrodynamic radius of each protein scaffold or protein assembly.The peak fractions of each chromatogram peak in the shaded region were loaded onto an SDS page gel to remove excess ligand, revealing only the scaffold and assembly proteins. These samples were used as input for Figure 20c. [Figure 19] This diagram shows how Alphafold v2.0 was used to predict the cis orientation of the fusion proteins. The highest-ranked model for each SpC3-scaffold-DgC assembly was visualized in PyMOL. Single chains are highlighted, showing SpyCatcher3 (medium gray), the scaffold (dark gray), and DogCatcher (light gray). Additional chains are shown as transparent surfaces, depicted schematically in white. In all simulations, a GSGS linker was used between the catcher and the scaffold. In the case of type XV collagen NC1, the predicted structure collapsed when the GSGS linker was used. Therefore, the prediction was repeated using a (GGGGS)2 linker, and this structure is shown here. Monomer sequences used for Alphafold prediction: TL1A - SEQ ID NO: 32; Col XV NC1 - SEQ ID NO: 33; MIF2m - SEQ ID NO: 34; HsCutA1 - SEQ ID NO: 54; PhCutA1 - SEQ ID NO: 55; Col X NC1 - SEQ ID NO: 56; TNF - SEQ ID NO: 57. [Figure 20]Figure 1 shows the suitability of the scaffold for in vitro testing. a) H6-SpC-PhCutA1-SnC was conjugated to two different ligands, SnT-L1 and L2-SpT. Serum-starved NCI-N87 cells were grown for 7 days in the presence of relevant growth factors and the dual conjugate assembly (H6-SpC-PhCutA1-SnC:SnT-L1:L2-SpT), single conjugate assemblies (H6-SpC-PhCutA1-SnC:SnT-L1, H6-SpC-PhCutA1-SnC:L2-SpT), and ligand-only controls (SnT-L1, L2-SpT, SnT-L1+L2-SpT), followed by MTT cell viability measurements. Control antibodies against L1, L2, and L1+L2 at 10 nM were used as controls (data not shown), yielding similar results compared to single and dual conjugate assembly samples. Error bars indicate n=3 technical replicates. b) The fully assembled scaffold H6-SpC-PhCutA1-SnC:SnT-L1:L2-SpT inhibits Akt and Erk1 / 2 activation. NCI-N87 cells were treated for 1 hour with scaffold alone (S; H6-SpC-PhCutA1-SnC), ligand alone (L1; SnT-L1 and L2; L2-SpT), and single-conjugate (S×L1, S×L2) and dual-conjugate (S×L1×L2) assemblies, demonstrating the inhibition of downstream activation of Akt / ERK signaling by the complete assembly H6-SpC-PhCutA1-SnC:SnT-L1:L2-SpT. c) H6-SpC-HsCutA1-SnC was conjugated with two different ligands, SnT-L1 and L3-SpT. Serum-starved NCI-N87 cells were treated with the dual conjugate assembly (H6-SpC-HsCutA1-SnC:SnT-L1:L3-SpT), single conjugate assemblies (H6-SpC-HsCutA1-SnC:SnT-L1, H6-SpC-HsCutA1-SnC:L3-SpT), and ligand-only controls (SnT-L1, L3-SpT, SnT-L1+L3-SpT) for 2 days, followed by MTT cell viability measurements. Error bars indicate n=3 technical replicates. [Figure 21]Figure 1 shows the purification of L1-PhCutA1-L2 as a direct fusion multidomain polypeptide by both Ni-NTA and size-exclusion chromatography. a) L1-PhCutA1-L2 was readily expressed in E. coli BL21(DE3) and purified by Ni-NTA chromatography using HisPur resin (ThermoFisher). Washes 1 and 2 were performed with 10 column volumes of equilibration buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 10 mM imidazole). Washes 3 and 4 were performed with 2 column volumes of wash buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 30 mM imidazole). All elution steps were performed with 2 column volumes of elution buffer (50 mM Tris, pH 7.8; 300 mM NaCl; 200 mM imidazole). Samples were analyzed on a 12% SDS-PAGE gel and stained with Coomassie. P - lysate pellet; CL - clarified lysate; FT - flow-through; W - wash; E - elution. b) Size-exclusion chromatography of L1-PhCutA1-L2 using HiLoad 16 / 600 Superdex 200 pg. b) UV A280 absorbance chromatogram with the L1-PhCutA1-L2 peak highlighted. Inset: SDS-PAGE of a 2 mL fraction of L1-PhCutA1-L2 from the highlighted region of the above chromatogram. [Figure 22] Figure 1 shows cell viability upon treatment of two cancer cell lines with various ligands (L1, L2, L4, L5, L6, L7) and an assembly consisting of SpC3-HsCutA144-179-DgC (SEQ ID NO: 24). Cell viability of Colo205 and NCI-N87 cells is shown as a heat map. Cells were treated with single assemblies (SpyTag or DogTag ligands on the scaffold only) and the fully conjugated assembly. Colo205 and NCI-N87 were treated with doses ranging from 0.16 to 100 nM and 4 to 100 nM, respectively. Data shown represent cell viability at the highest dose used. The bottom panel shows a measure of relative cell viability. [Figure 23]Figure 1 shows HsCutA1 direct fusions (without modular protein binding (catcher)). a) Structural models comparing the lengths of SpyCatcher003 (SpC3) with SpyTag003 (SpT3), and DogCatcher (DgC) with DogTag (DgT), compared to alternative linkers, e.g., (G4S)n, where n=1, 2, 3, (EAAAK)n, where n=1, 2, 3, and the ribosomal L9 linker mediating the link between the polypeptide (protein binder) and the multivalent scaffold instead of the modular binding site (SpyCatcher003 and DogCatcher). The lengths between the tag and the catcher are shown when the protein binder is fused to the tag at either the N- or C-terminus. b) Cytotoxicity of SpC3-(G4S)-HsCutA1-(G4S)-DgC (SEQ ID NO: 24) assembled with the indicated protein pairs (e.g., L5 / L4, where L5 is Spy-tagged and L4 is Dog-tagged) was compared to direct fusions in which the modular display platform was replaced as follows: L5 / L4 to L5-GGGGS-[C75A,C96A]HsCutA167-171-GGGGS-L4, L1 / L2 to L1-(GGGGS)3-[C75A,C96A]HsCutA167-171-(GGGGS)3-L2, and L2 / L4 to L2-(GGGGS)4-HsCutA144-179-(GGGGS)4-L4. Cell viability of NCI-N87 cells was measured after 7 days of treatment with 20 nM of the assembly and direct fusions. c) Cytotoxicity induced by SpC3-[C75V, C96S]HsCutA144-179-DgC (SEQ ID NO: 73), assembled with L7 to form trivalent monospecific molecules at the N- or C-terminus, was compared with their direct fusion counterparts. The direct fusion equivalents were L7-(GGGGS)5-[C75V, C96S]HsCutA144-179 and [C75V, C96S]HsCutA144-179-(GGGGS)8-L7. The percentage of viable Colo205 cells was determined 48 hours after treatment with the indicated doses. The same assays were performed for d) and e).d) Levels of apoptosis induced by SpC3-[C75V,C96S]HsCutA144-179-DgC (SEQ ID NO: 73) and SpC3-[C75V,C96S]HsCutA160-171-DgC (SEQ ID NO: 76), both assembled with L7 to form hexavalent monospecific molecules, and comparison with their direct fusion equivalents, L7-(GGGGS)5-[C75V,C96S]HsCutA144-179-(GGGGS)8-L7 and L7-(GGGGS)5-[C75V,C96S]HsCutA160-171-(GGGGS)8-L7. e) Levels of apoptosis induced by several direct fusion molecules of different Cores fused at the N- and C-termini to L7 in the format L7-(GGGGS)5-[Core]-(GGGGS)8-L7, where [Core] was [C75V,C96S]HsCutA144-179 (SEQ ID NO: 68), [C75V,C96S]HsCutA160-171 (SEQ ID NO: 70), OX40L (SEQ ID NO: 78), CD40L (SEQ ID NO: 79), and full-length soluble TNF (SEQ ID NO: 80). [Figure 24]Figure 1 shows truncations of HsCutA1. a) Schematic representation of the coding sequences of different truncated HsCutA1s fused to SpyCatcher003 (SpC3) and DogCatcher (DgC) via exemplary linkers L1 and L2. The truncations are represented as 44-179 (SEQ ID NO: 29) and 60-171 (SEQ ID NO: 66), respectively. The sequences are compared in a multiple sequence alignment between WT HsCutA1 (33-179) and HsCutA1 truncated to 44-179 and 60-171. b) Schematic representation of the truncated sites of HsCutA1. Residues 44-59 are shown at the N-terminus, and residues 172-179 are shown at the C-terminus. Schematic representation modeled from PDB ID: 2ZFH. c) Comparison of expression of HsCutA144-179 and HsCutA160-171. Comparison of E. coli cell lysates demonstrated the changes in protein expression mediated by truncation. d) Cytotoxicity of SpC3-(GGGGS)3-[C75V, C96S]HsCutA144-179-(GGGGS)3-DgC (SEQ ID NO: 73) and SpC3-(GGGGS)3-[C75V, C96S]HsCutA160-171-(GGGGS)3-DgC (SEQ ID NO: 76), which variably assembled with L7 on either SpC3, DgC, or both to form trivalent or hexavalent monospecific molecules. The percentage of viable Colo205 cells was determined 48 hours after treatment at the indicated doses. [Figure 25]Figure 1. Cysteine removal from HsCutA1. a) Schematic representation of the naturally occurring cysteine sites in HsCutA1 targeted for removal. Schematic modeled from PDB ID: 2ZFH. b) Table of amino acid frequencies at sites similar to C75 and C96 in HsCutA1 after multiple sequence alignment (MSA) of 406 homologs. Amino acids with zero frequency are not shown. c) Comparison of expression of HsCutA144-179 and HsCutA160-171 with native WT cysteines (C75, C96), alanine-mutated cysteines (C75A, C96A), and MSA-informed mutations (C75V, C96S). E. coli cell lysates were compared to demonstrate the changes in protein expression induced by mutagenesis. d) SpC3-GGGGS-HsCutA144-179-GGGGS-DgC (SEQ ID NO: 24), SpC3-(GGGGS)3--[C75V, C96S]HsCutA144-179-(GGGGS)3-DgC (SEQ ID NO: 73), SpC3-(GGGGS)3-HsCu, assembled with L7 on both SpC3 and DgC to form hexavalent monospecific molecules. Cytotoxicity induced by tA160-171-(GGGGS)3-DgC (SEQ ID NO: 74), SpC3-(GGGGS)3--[C75A, C96A]HsCutA160-171-(GGGGS)3-DgC (SEQ ID NO: 75), and SpC3-(GGGGS)3--[C75V, C96S]HsCutA160-171-(GGGGS)3-DgC (SEQ ID NO: 76). The percentage of viable Colo205 cells was determined 48 hours after initiating treatment at the indicated doses. [Figure 26]
[0023] Figure 1 illustrates a typical thiol-maleimide reaction with a monomeric protein, as well as various linker-payload structures. A) The thiol group of the monomeric protein reacts with the maleimide moiety of the payload-linker compound to generate a thiosuccinimide payload-linker-protein adduct. B) Examples of linker-payload compounds conjugated to catcher core proteins: 1-fluorescein-5-maleimide, 2-deruxtecan, 3-maleimidocaproyl-Val-Cit-PAB-DM1, 4-maleimidocaproyl-Val-Cit-PAB-MMAF. [Figure 27] Schematic representation of the CutA1 structure (modeled from PDB ID: 2ZFH) and the location of Cys mutations in a model of the fluorescein-5-thiosuccinimide-CutA1 conjugate. A) Spatial location of the Cys mutations in one monomer of the CutA1 structure. The Cys thiol is shown as a sphere. B) Model of the fluorescein-5-thiosuccinimide conjugate to the K82C mutant of CutA1 (SEQ ID NO: 86). All images were generated with PyMOL. [Figure 28]Figure 1 shows deconvoluted ESI mass spectra of fluorescein-5-maleimide conjugates of all Cys mutant CutA1 Catcher core variants (SEQ ID NOs: 95, 93, 96, 92, 94, 97) and wild-type (SEQ ID NO: 24) and cysteine-free (SEQ ID NO: 91) controls, as well as unconjugated CC042 and fluorescein-5-thiosuccinimide-CC042 conjugates. A) The dye / protein ratios of all mutants were calculated as described in the methods for conjugated samples after PD-10 purification. B) 12% SDS-PAGE gel of unconjugated and conjugated CutA1 Catcher core mutants. Little to no conjugation is observed in the CC041 negative control, while all other mutants conjugate with fluorescein-5-maleimide. CC refers to the Catcher core protein gel band. C) ESI deconvolution mass spectrum of unconjugated Catcher core: the main peak is within 1 mass unit of the expected molecular weight of the CC042 monomer. D) Mass spectrum of the fluorescein-5-thiosuccinimide-CC042 conjugate: the main peak is approximately 18 Da larger than the expected molecular weight of the monomer, likely due to thiosuccinimide ring hydrolysis after conjugation. [Figure 29] SDS-PAGE (4-20% gradient) of fluorescein-5-maleimide conjugates and catcher / tag assembly of CC042 (SEQ ID NO: 93) and the CC041 cysteine-free negative control (SEQ ID NO: 91). Only CC042 conjugates efficiently to fluorescein-5-maleimide. A low level of background signal indicating conjugation is observed with CC041, suggesting that maleimide conjugation is primarily specific to Cys residues. Conjugate CC042 exhibits a clear shift in migration within the gel when reacted with the Dog-tagged conjugate (L7), the Spy-tagged conjugate (L8), or both conjugates simultaneously. A similar band profile is observed with the CC041 control, suggesting that maleimide conjugation does not impair assembly efficiency. [Figure 30] Figure 1 shows specific membrane binding and internalization by RTK conjugates, single assemblies, and full assemblies. HsCutA1 refers to CC7, and HsCutA1* refers to CC042 conjugated to fluorescein-5-maleimide. A) His-tagged conjugates L1 and L2 and the full assembly with His-tagged CC7 (SEQ ID NO: 76) stained with anti-His antibody show that all compounds bound to the membrane and were internalized. (Note the different ratios of His tags.) B) CutA1 staining of the full assembly of CC7 with conjugates L1 and L2 confirms membrane binding and internalization. C) Fluorescein-5-maleimide (F-5-M)-conjugated CC042 (SEQ ID NO: 93) alone as a single assembly and full assembly shows RTK-specific internalization of all assemblies without off-target binding of HsCutA1*. [Figure 31] Figure 28 shows the structures of cytotoxic drug-linker compounds and the corresponding deconvoluted ESI mass spectrum of CC042 (SEQ ID NO: 93) conjugated to such compounds. For reference, the mass spectrum of unconjugated CC042 can be found in Figure 28C. A) Deruxtecan: The main peak of the conjugate is within one mass unit of the expected molecular weight. B) Maleimidocaproyl-Val-Cit-PAB-DM1: The main peak of the conjugate is within one mass unit of the expected molecular weight. C) Maleimidocaproyl-Val-Cit-PAB-MMAF: The main peak of the conjugate is within one mass unit of the expected molecular weight. [Figure 32]Figure 1 shows cell viability of two cancer cell lines upon treatment with Spy-tagged ligand L7 and / or Dog-tagged ligand L7 and assemblies consisting of SpC3-(GGGGS)3-[C75A,C96A]-HsCutA1-(GGGGS)3-DgC (SEQ ID NO: 75), SpC3-(GGGGS)-scHsCutA1-(GGGGS)-DgC (SEQ ID NO: 99), or SpC3-(GGGGS)3-L9-(GGGGS)3-DgC (SEQ ID NO: 98). Colo205 and HCT116 cells were treated at doses ranging from 0.0008 to 2 nM. Data shown are normalized to mock-treated controls. [Figure 33] Figure 1 shows cell viability of the cancer cell line Colo205 upon treatment with the assembly consisting of DogTag ligand L7 and SpyTag ligands L8 or -L10, both in scFv and Fab formats, and SpC3-GGGGS--[C75A, C96A, K82C]-HsCutA44~179-GGGGS-DgC (SEQ ID NO: 93). The hexavalent nanobody-based L7 assembly served as a control. Colo205 cells were treated with doses ranging from 0.0008 to 2 nM. Data shown are normalized to mock-treated controls. [Figure 34] Figure 1 shows cell viability of the cancer cell line HCT116 upon treatment with [C75A, C96A]HsCutA1160-171-based multivalent bispecific fusions containing L7 and L11 conjugates, or a tetrameric L7 nanobody fusion, a failed clinical candidate. Monovalent nanobody-based L7 conjugates and [C75A, C96A]HsCutA160-171 were used as controls. HCT116 cells were treated with doses ranging from 0.000256 to 0.8 nM. Data shown are normalized to mock-treated controls. [Figure 35]Schematic diagram of the coupling of SpyCatcher003 and DogCatcher to magnetic beads used during the assembly purification pipeline. Glutathione (GSH)-conjugated magnetic beads are mixed with equimolar concentrations of GST-SpC3 and GST-DgC. Binding of GST-catcher to the beads is incubated at 25°C for 1 hour. The beads are then precipitated using a magnetic block and washed 3x with PBS, or until a baseline absorbance reading at 280 nm is achieved. The GST-bound beads containing both catchers are retained for use during assembly purification. [Figure 36] SDS-PAGE (4-20% gradient) of a preparation of glutathione-conjugated beads carrying GST-SpC3 and GST-DgC. MagneGST resin was prepared using SDS loading buffer to confirm the presence of contaminating proteins in the stock. W1-3: PBS washes. [Figure 37] Figure 35 shows a schematic diagram of another variation of catcher-based protein assembly and cleanup. SpC3-core-DgC is combined with SpT and DgT proteins at a 1.4x excess of tagged protein relative to catcher-core. The conjugation reaction is incubated at 25°C for 1 hour to allow complete capture of the tagged protein. Pre-prepared paramagnetic glutathione-conjugated beads bound to GST-SpC3 and GST-DgC as shown in Figure 35 are added at a 2x excess of each GST catcher relative to the corresponding tagged protein. The sample is transferred to a magnetic block to capture the bead-GST-catcher-tag conjugates, and the supernatant containing the core-catcher-tag conjugates is retained for further analysis and downstream assays. [Figure 38]SDS-PAGE (4-20% gradient) of protein assembly and MagneGST purification. The L7SpT:CC7:L2DgT protein assembly before purification (L7SpT:CC7:L2DgT(pre)) shows excess of both conjugates, while the L7SpT:CC7:L2DgT protein assembly after purification (L7SpT:CC7:L2DgT(post)) shows only fully assembled molecules. GST-catcher+ resin indicates that excess conjugates were captured during MagneGST purification. (CC7 corresponds to SEQ ID NO: 76). [Figure 39] Design and production of single-chain SpC3-HsCutA1-DgC. (A) Schematic diagram of SpC3-HsCutA1-DgC as a linear construct and tertiary structure configuration. The predicted structure of the single-chain HsCutA1 core generated by AlphaFold2 is also shown. (B) UV A280 absorbance of single-chain SpC3-HsCutA1-DgC from size-exclusion chromatography using HiLoad16 / 600 Superdex 200 pg. 2 mL fractions of the highlighted peak were separated by SDS-PAGE, and fractions B6–B11 were retained as purified single-chain SpC3-HsCutA1-DgC. [Figure 40] Figure 1 shows the induction of apoptosis by trivalent and hexavalent L7 assemblies, measured as activation of caspase 3 / 7. L7 was conjugated to SpC3-HsCutA44~179-DgC (SEQ ID NO: 24) to generate trivalent and hexavalent assembly molecules. Colo205 cells were treated with the assemblies, and caspase 3 / 7 activity was measured 1 hour (A) and 24 hours (B) after treatment began. After 1 hour, only the hexavalent molecules, but not the trivalent molecules or the combination of both trivalent molecules, induced apoptosis. Caspase 3 / 7 activity was observed after 24 hours when treated with the trivalent assemblies. [Figure 41]Figure 1 shows cell viability of the liver cancer cell line HepG2 upon treatment with the L11-HsCutA1-L7 direct fusion, the failed clinical asset TAS266, and HsCutA1 in the presence or absence of intravenous immunoglobulin (IGIV 0.2-10 mg / ml). Untreated cells growing in the absence of IVIG were considered 100% viable. In contrast to TAS266, L11-HsCutA1-L7 did not cause an increase in ADA-associated cytotoxicity. DETAILED DESCRIPTION OF THE INVENTION
[0120] The present invention will be described with respect to particular embodiments and with reference to certain drawings, but the present invention is not limited thereto, but only by the claims. Any reference signs in the claims should not be construed as limiting the scope. Of course, it should be understood that not necessarily all aspects or advantages will be achieved in accordance with any particular embodiment of the invention. Thus, for example, one skilled in the art will recognize that the present invention can be embodied or performed in a manner that achieves or optimizes one advantage or group of advantages taught herein, but does not necessarily achieve other aspects or advantages that may be taught or suggested herein.
[0121] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to a "scaffold" includes two or more scaffolds, reference to an "oligomer" includes two or more such oligomers, and so forth.
[0122] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.
[0123] definition The following terms or definitions are provided merely to aid in understanding the present invention. Unless otherwise defined herein, all terms used herein have the same meaning as those of ordinary skill in the art of the present invention. For definitions and terms in the art, experts can refer to, among others, Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be interpreted as having a narrower scope than those understood by those skilled in the art.
[0124] As used herein, "about" when referring to measurable values such as amounts and durations is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and even more preferably ±0.1% from the particular value, where such variations are appropriate for practicing the methods of the present disclosure.
[0125] The term "amino acid" in the context of this disclosure is used in its broadest sense and is meant to include organic compounds containing an amine (NH) functional group and a carboxyl (COOH) functional group, along with a side chain (e.g., an R group) specific to each amino acid. In some embodiments, amino acid refers to a naturally occurring L α-amino acid or residue. Commonly used one-letter and three-letter abbreviations for naturally occurring amino acids are used herein: A = Ala; C = Cys; D = Asp; E = Glu; F = Phe; G = Gly; H = His; I = Ile; K = Lys; L = Leu; M = Met; N = Asn; P = Pro; Q = Gln; R = Arg; S = Ser; T = Thr; V = Val; W = Trp; and Y = Tyr (Lehninger, AL, (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes chemically modified amino acids such as D-amino acids, retro-inverso amino acids, and amino acid analogs, naturally occurring amino acids that are not normally incorporated into proteins, such as norleucine, and chemically synthesized compounds having properties known in the art to be characteristic of amino acids, such as β-amino acids. For example, analogs or mimetics of phenylalanine or proline that allow the same conformational restriction of peptide compounds as natural Phe or Pro are included within the definition of amino acid. Such analogs and mimetics are referred to herein as "functional equivalents" of the corresponding amino acids. Other examples of amino acids are listed in Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5, p. 341, Academic Press, Inc., NY 1983. This document is incorporated herein by reference.
[0126] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to a polymer of amino acid residues, as well as variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are synthetic, non-naturally occurring amino acids, such as chemical analogs of corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. Polypeptides may also undergo maturation or post-translational modification processes, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, and phosphorylation. Peptides can be produced using recombinant techniques, e.g., by expression of recombinant or synthetic polynucleotides. Recombinantly produced peptides are typically substantially free of culture medium; e.g., culture medium represents less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of the protein preparation.
[0127] The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. A protein may be composed of a single polypeptide or may include multiple polypeptides that assemble to form a multimer. A multimer may be a homo- or hetero-oligomer. A protein may be a naturally occurring protein, a wild-type protein, or a modified or non-naturally occurring protein. A protein may differ from a wild-type protein by, for example, the addition, substitution, or deletion of one or more amino acids.
[0128] Protein "variants" include peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question, and have biological and functional activity similar to that of the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the degree to which sequences are identical on an amino acid-by-amino acid basis over a comparison window. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions where identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) are present in both sequences to calculate the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window (i.e., window size), and multiplying the result by 100 to calculate the percentage of sequence identity.
[0129] In all aspects and embodiments of the invention, a "variant" typically has at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% complete sequence identity with the amino acid sequence of the corresponding wild-type protein. Sequence identity may also be relative to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, while a sequence may exhibit only 50% overall sequence identity with a full-length reference sequence, the sequence of a particular region, domain, or subunit may share 80%, 90%, or even 99% sequence identity with the reference sequence.
[0130] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is the gene most frequently observed in a population and is therefore an arbitrarily designed "normal" or "wild-type" form of the gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence alterations (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., altered characteristics) compared to the wild-type gene or gene product. Note that naturally occurring mutants may be isolated. They are identified by the fact that they have altered characteristics compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted for arginine (R) by replacing the methionine codon (ATG) with an arginine codon (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the mutant monomer. Alternatively, non-naturally occurring amino acids can be introduced by expressing mutant monomers in E. coli that are auxotrophic for a particular amino acid in the presence of a synthetic (i.e., non-naturally occurring) analog of that particular amino acid. Mutant monomers can also be generated by naked ligation when generated using partial peptide synthesis. In a conservative substitution, an amino acid is replaced with another amino acid having a similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acid may have a similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the amino acid being replaced. Alternatively, a conservative substitution can replace an existing aromatic or aliphatic amino acid with another amino acid that is aromatic or aliphatic.Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 major amino acids defined below in Table 1. If the amino acids have similar polarity, this can also be determined by reference to the hydropathic index scale of amino acid side chains in Table 2.
[0131] [Table 1]
[0132] [Table 2]
[0133] Mutant or modified proteins, monomers, or peptides can also be chemically modified in any manner and at any site. Mutant or modified monomers or peptides can be chemically modified by attaching a molecule to one or more cysteines (cysteine ligation), by attaching a molecule to one or more lysines, by attaching a molecule to one or more non-naturally occurring amino acids, by enzymatic modification of an epitope, or by terminal modification. Suitable methods for performing such modifications are well known in the art. Mutants of modified proteins, monomers, or peptides can also be chemically modified by attaching any molecule. For example, mutants of modified proteins, monomers, or peptides can be chemically modified by attaching a dye or fluorophore.
[0134] Polypeptide Constructs The present invention relates in part to multi-domain polypeptide constructs, which are useful when two or more are combined to form oligomeric proteins, and in some embodiments, can also be used as monomers. Multi-domain polypeptides are typically engineered to combine domains that do not exist together in nature. In some embodiments, 3, 4, 5, or 6 polypeptide constructs are combined to form oligomers. In some embodiments, 3 constructs are combined to form trimers, such as homotrimers.
[0135] The description provided herein is of an oligomeric core of subunit monomers. The subunit monomers are typically structural domains of a multi-domain polypeptide construct. The oligomerization of these structural domains can then form the core of a multivalent protein scaffold.
[0136] Thus, the considerations regarding the characteristics of subunit monomers also apply to the disclosure and definition of structural domains of individual polypeptide constructs.
[0137] Similarly, the first and second binding domains of a polypeptide construct may form a first and second binding site, as described elsewhere herein, or may form a first and second effector moiety, as described elsewhere herein, as the case may be. For example, if the binding domain is an isopeptide bond-forming "catcher" domain (or other binding site described herein), it is a binding site as described elsewhere below. If the binding domain is, for example, an antibody, antigen-binding fragment, antibody mimetic, protein or peptide ligand, protein or peptide signaling molecule (e.g., cytokine), biological receptor, or other molecule described herein as an effector moiety, the binding domain is an effector moiety as described elsewhere herein.
[0138] Thus, the definitions and descriptions of first and second binding sites, or first and second effector moieties, apply accordingly to the first binding domain and second binding domain of the polypeptide construct.
[0139] In some embodiments, the polypeptide construct comprises a first binding domain at the N-terminus and a second binding domain at the C-terminus, the first and second binding domains being separated by a structural domain. The N-terminus refers to the terminal amino acid residue at the amino terminus of a polypeptide. The C-terminus refers to the terminal amino acid residue at the carboxy terminus of a polypeptide. Typically, the first binding domain and the second binding domain can bind to their targets when the target molecules are expressed on a single cell or immobilized on a plate or single bead. This is sometimes described herein as providing the first and second binding domains in a "cis" orientation. Typically, the first binding domain and the second binding domain bind to a cellular target on the surface of a single cell, and both targets can be clustered in the cell membrane. Therefore, cis orientation may be preferred for some cis-acting agents (i.e., those that can act on a single cell). The cis orientation of bispecific antibodies, along with the inverse "trans" orientation, is discussed in Dickopf et al. (Computational and Structural Biotechnology Journal, Volume 18, 2020, Pages 1221-1227).
[0140] "Cis orientation" in the context of geometry, as with cis-trans isomerization and as previously used to describe bispecific antibody architectures, is used herein to refer to a spatial arrangement in which two components are on the "same side" of a plane (derived from Latin), as opposed to "trans" or "trans orientation" in the context of geometry in which the two components "cross" (derived from Latin), i.e., are on different sides of a plane. This geometric definition differs from "cis-acting" and "trans-acting" in the context of biological effects in which a single bispecific molecule acts on a single or adjacent cell (cis) or on different cell populations (trans), e.g., recruiting effector cells to target cells. There is active interest and effort in exploring various forms of geometry (see, e.g., Dengl et al. Nat Commun. 2020;11:4974). "Cis orientation" can also be beneficial in certain "trans-acting" bispecifics, for example, by reducing intermolecular distance (Dickopf et al., Computational and Structural Biotechnology Journal, Volume 18, 2020, pp. 1221-1227). Cis orientation is particularly interesting for multivalent cis orientations for cis-acting bispecifics due to multivalent binding or clustering of targets on a single cell (e.g., compared to the bispecific tandem fusions described in Veggiani et al., Biochemistry January 19, 2016, 113 (5) 1202-1207, or higher order monospecific clustering, Khairil Anuar et al., Nature Communications volume 10, Article number: 1734 (2019)).
[0141] The function of the structural domain is to provide structurally defined support for the binding domain. Advantageously, the structural domain can ensure that the binding domain has a desired orientation, typically such that both binding domains can bind to the target in a cis orientation. Thus, the construct can provide a single binding surface.
[0142] In certain embodiments, the attachment site of the binding domain to the structural domain (oligomeric core) allows for attachment even with a short linker.
[0143] A structural domain may be any polypeptide domain that comprises a defined secondary structure, typically an alpha helix or a beta sheet. In some particularly advantageous embodiments, a structural domain has its N-terminus and C-terminus in the same spatial region, for example, substantially adjacent or adjacent to each other. Attaching a binding domain to the end of a structural domain provides two binding domains that are substantially adjacent in a three-dimensional structure. In some embodiments, the N-terminus and C-terminus are oriented so that they face substantially the same direction. As described elsewhere herein, providing spatially adjacent N-terminus and C-terminus results in binding domains located on the same surface of a construct or on the same surface of an oligomer containing multiple constructs. Thus, a construct typically provides a single binding surface. A construct typically provides a cis-oriented binding region.
[0144] In some embodiments, the structural domain presents the binding domain in a preferred orientation due to the tertiary structure of the monomer (e.g., preferred relative positioning of the N- and C-termini of a single monomer). In some embodiments, the structural domain presents the binding domain in a preferred orientation due to the connection of the monomer to the quaternary structure (e.g., preferred relative positioning of the N- and / or C-termini between monomers). In some preferred embodiments, the structural domain presents the binding domain in such a preferred orientation due to the tertiary structure of the monomer combined with the connection of the monomer to the quaternary structure (e.g., preferred relative positioning of the N- and C-termini within and between monomers).
[0145] A structural domain may comprise a single polypeptide chain, or may be formed of two or more separate polypeptide chains that connect to form a single structural domain, e.g., two antiparallel (NC CN) alpha helices or two or more beta strands that connect to form a beta sheet. In some embodiments, two or more polypeptide chains with appropriate characteristics are identified and then fused to form a single polypeptide chain (i.e., a fusion protein), typically by recombinant means, but also by chemical conjugation or bonding to form a single covalent molecule.
[0146] The structural domain is different from the two binding domains. Thus, if the binding domain is a catcher polypeptide, such as SpyCatcher, DogCatcher, or SnoopCatcher, the structural domain is not a catcher polypeptide.
[0147] Typically, the structural domain does not include a CH2 domain. Typically, the structural domain does not include a CH3 domain. In some embodiments, the structural domain does not include a CH2 domain and does not include a CH3 domain.
[0148] As discussed below and in the Examples, the inventors have identified numerous exemplary structural domains. In some embodiments, the structural domain comprises or consists of the type X collagen NC1 domain (SEQ ID NO: 2), or a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, e.g., at least 90% or at least 95% identity thereto. In some embodiments, the structural domain comprises or consists of the type VIII collagen NC1 domain (SEQ ID NO: 3), or a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, e.g., at least 90% or at least 95% identity thereto. In some embodiments, the structural domain comprises or consists of the CutA1 polypeptide (e.g., SEQ ID NO: 1 or SEQ ID NO: 19), or a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, e.g., at least 90% or at least 95% identity thereto.
[0149] In certain embodiments, the structural element comprises or consists of a polypeptide having at least 50% amino acid identity, such as at least 90% identity, to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:19, SEQ ID NO:29, SEQ ID NO:60, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:42, SEQ ID NO:31, SEQ ID NO:78, SEQ ID NO:80 or SEQ ID NO:58.
[0150] A suitable structural domain is the type IV collagen C4 domain (also known as the type IV collagen NC1 domain) (PDB ID: 1m3d, SEQ ID NO: 49) or a polypeptide having at least 50%, at least 60%, at least 70%, or at least 80%, such as at least 90% or at least 95% identity thereto.
[0151] Generally, collagen NC1 domains, including NC1 domains from collagen types IV, VIII, and X, can be used as structural domains according to the present invention.
[0152] However, not all collagen NC1 domains are suitable as structural domains, and in particular, the NC1 domains from collagen types XV and XVIII do not have the required orientation.
[0153] In certain embodiments, the structural domain comprises human macrophage migration inhibitory factor (MIF) (PDB ID: 1CA7, or SEQ ID NO: 25, or PDB ID: 6OY8 with Y99G mutation) or human macrophage migration inhibitory factor 2 (MIF2) (PDB ID: 7MSE, or SEQ ID NO: 26, or SEQ ID NO: 27 with S62A and F99A mutations), or a homolog or paralog thereof.
[0154] In certain embodiments, the structural domain comprises a TNF family protein, including TNF (PDB ID: 1TNF, SEQ ID NO: 42, full length soluble domain SEQ ID NO: 80), TL1A (PDB ID: 2RE9, SEQ ID NO: 31), OX40L (SEQ ID NO: 78), or CD40L (PDB ID: 3LKJ, SEQ ID NO: 58).
[0155] Further examples of structural domains include a suitably modified antiparallel coiled-coil hexamer (PDB ID: 5W0J, see Example 4, SEQ ID NO: 43), HIV-1 GP41 core (PDB ID: 1I5Y or SEQ ID NO: 44), cytochrome c555 (PDB ID: 5Z25 or SEQ ID NO: 45), MHC class II-associated chaperonin and targeting protein invariant chain (Ii) (PDB ID: 1iie or SEQ ID NO: 46), p53 (PDB ID: 1C26 or SEQ ID NO: 47); fibrinogen-like domain (PDB ID: 4M7F or SEQ ID NO: 48), Bacillus subtilis AbrB (PDB ID: 1YFB or SEQ ID NO: 50), bacteriophage lambda head protein D (e.g., PDB ID: 1C5E or PDB ID: 1C5E or SEQ ID NO: 51); a domain-swap trimeric variant of HCRBPII (PDB ID: 1C5F or SEQ ID NO: 52); ID: 6VIS or SEQ ID NO: 52); T1L reovirus attachment protein sigma 1 (chains A, B, C of PDB ID: 4ODB, or SEQ ID NO: 53).
[0156] In certain embodiments, the structural domain (or subunit monomer of the oligomeric core) of the polypeptide construct is a CutA1 protein, typically a human CutA1 protein. The data presented below, particularly Figure 23, demonstrates various fusions of HsCutA1 ("Homo sapiens CutA1") with effector proteins and the ability of such direct fusions to exert a biological effect. Thus, in certain embodiments, the present invention provides a polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, wherein the first and second binding domains are separated by a human CutA1 structural domain.
[0157] In certain embodiments, CutA1, typically human CutA1, has been genetically engineered to remove one or more cysteine residues from the native sequence (presented herein as SEQ ID NO: 19). Typically, this is a substitution of one or more cysteine residues with one or more non-cysteine residues. Removal of one or more cysteine residues is beneficial because it allows for targeted cysteine conjugation at non-native sites or in fusion proteins. Without being bound by theory, it may be beneficial to remove unpaired cysteines, as unpaired cysteines often interfere with stability and downstream applications.
[0158] In some embodiments, one or more cysteine residues of CutA1 are substituted with one or more alanine residues. In some embodiments, one or more cysteine residues of CutA1 are substituted with one or more valine residues. In some embodiments, one or more cysteine residues of CutA1 are substituted with one or more serine residues. In some embodiments, the substitution comprises or consists of two cysteines being substituted with two alanines, referred to herein as a "CACA" substitution. In some embodiments, the substitution comprises or consists of one cysteine being substituted with a valine and one cysteine being substituted with a serine, referred to herein as a "CVCS" substitution. In some embodiments, the cysteine residues at positions 75 and 96 of wild-type human CutA1 (e.g., SEQ ID NO: 19) are substituted with different residues. In some embodiments, human CutA1 is engineered to have two cysteine residues substituted, the cysteine substitutions comprising or consisting of (i) C75A, C96A, or (ii) C75V, C96S. Thus, in some embodiments, the structural domain (or subunit monomer of the oligomer core) of the polypeptide construct is a human CutA1 protein engineered to have two cysteine residues substituted, the cysteine substitutions comprising or consisting of (i) C75A, C96A, or (ii) C75V, C96S. The generation, production, and biological effects of polypeptide constructs comprising such cysteine-substituted human CutA1 domains are described below (Figure 25).
[0159] In certain aspects, CutA1, typically human CutA1 (e.g., SEQ ID NO: 19), has been genetically engineered to delete one or more residues from either or both ends of the native sequence to form a truncated CutA1 domain that is incorporated into the constructs of the invention. In some embodiments, one or more residues are deleted from the N-terminus of CutA1, typically human CutA1. Typically, 5 to 70 residues, e.g., 10 to 59 residues, are deleted from the N-terminus of CutA1, typically human CutA1. In some embodiments, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 59, 60, 61, 62, 63, 64, 65, or 66 residues are deleted from the N-terminus of CutA1. In some embodiments, the truncated CutA1 begins at residue 33 (i.e., 32 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 44 (i.e., 43 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 60 (i.e., 59 N-terminal residues are deleted). In some embodiments, the truncated CutA1 begins at residue 67 (i.e., 66 N-terminal residues are deleted), for example, the "HsCutA1 67-171 CACA" direct fusion construct illustrated in Figure 23b.
[0160] In some embodiments, one or more residues are deleted from the C-terminus of CutA1, typically human CutA1. C-terminal deletions may be made instead of or in addition to N-terminal deletions. Typically, 5 to 20 residues are deleted from the C-terminus of CutA1, typically human CutA1, e.g., 6 to 12 residues are deleted. In some embodiments, approximately 8 residues are deleted, resulting in a truncated protein having a C-terminus at residue 171 (valine) of the wild-type sequence shown in FIG. 24 and SEQ ID NO: 19 for human CutA1. In some embodiments, the truncated protein has a C-terminus at residue 168 (threonine) of the wild-type sequence shown in FIG. 24 and SEQ ID NO: 19. In some embodiments, the truncated protein has a C-terminus at residue 169, 170, 172, 173, 174, 175, or 176 of the wild-type sequence shown in FIG. 24 and SEQ ID NO: 19.
[0161] In certain embodiments, the truncated CutA1 consists of residues 44-179 or residues 60-171 of SEQ ID NO:19 as shown in Figure 24. In some embodiments, the truncated CutA1 begins at any of residues 44-67 of SEQ ID NO:19 and ends at any of residues 168-179. In other embodiments, the truncated CutA1 begins at any of residues 30-65 of SEQ ID NO:19 and ends at any of residues 165-179. In some embodiments, the truncated CutA1 begins at any of residues 44-60 of SEQ ID NO:19 and ends at any of residues 171-179.
[0162] Truncating CutA1 can improve the precision with which fusion constructs can be constructed. Figure 24 shows the generation, production, and biological effects of polypeptide constructs containing such truncated human CutA1 domains.
[0163] In certain aspects, CutA1 is human CutA1 that has been genetically engineered to remove one or more cysteine residues from the native sequence, resulting in truncation at the N-terminus and / or C-terminus. In some embodiments, the truncated CutA1 consists of residues 44-179 or residues 60-171 of SEQ ID NO: 19 as shown in Figure 24, with cysteine substitutions that include or consist of (i) C75A, C96A, or (ii) C75V, C96S. In other embodiments, the truncated CutA1 begins at any of residues 30-65 and ends at any of residues 165-179 of SEQ ID NO: 19, and also has cysteine substitutions that include or consist of (i) C75A, C96A, or (ii) C75V, C96S.
[0164] The polypeptide construct according to the invention comprises, in addition to the structural domain, a first binding domain and a second binding domain.
[0165] In certain embodiments, the binding domain is capable of forming an isopeptide bond with a cognate peptide, such as various catcher domains known in the art. Constructs comprising such isopeptide bond-forming domains are particularly well suited for screening different pairs of effector molecules, such as antigen-binding proteins. As discussed elsewhere herein, multiple combinations of effector molecules can be linked to constructs comprising isopeptide bond-forming binding domains via isopeptide-forming peptide tags. Thus, constructs comprising binding domains capable of forming an isopeptide bond with a cognate peptide are particularly useful as drug discovery platforms.
[0166] Aspects of the invention relating to isopeptide bond formation are generally illustrated with reference to a larger molecule (domain), typically called a catcher, attached to a structural domain, and a smaller polypeptide or peptide, typically called a tag, that forms part of the binding region of interest (e.g., antigen-binding domain). However, all aspects and embodiments can be practiced in the reverse orientation, where the larger (e.g., catcher) molecule forms part of the binding region of interest (e.g., antigen-binding domain) and the smaller tag peptide forms the binding domain attached to the structural domain.
[0167] In some embodiments, the first and second binding domains in the polypeptide construct are effector molecules, such as antigen-binding domains, and the construct is particularly suitable for use as a diagnostic, analytical, or therapeutic agent.
[0168] In some embodiments, interesting or effective pairs of antigen-binding regions are identified using the drug discovery platform of the invention (e.g., where the construct comprises an isopeptide bond-forming binding domain), and then constructs are expressed having the identified combination of antigen-binding domains (or other effector moieties) directly connected without the isopeptide bond-forming binding regions, without an intermediate catcher domain in the structural domain, and without a peptide tag in the antigen-binding domain of the structural domain. For the avoidance of doubt, such direct fusion constructs may still include a linker region between the terminal residue of the structural domain and the terminal residue of the or each effector moiety (e.g., antigen-binding region), as described in detail elsewhere herein.
[0169] Thus, one aspect of the present invention provides a system for large-scale, high-throughput screening of large numbers of possible combinations of effector molecules using combinatorial pairs of tagged effector proteins; combinations identified as useful can be used in the form in which they are provided in the screening construct, or can be converted into a more convenient format (e.g., for therapeutic drug candidates) by creating direct fusions with effector molecules (e.g., antigen-binding regions) at the same structural domains as those used in the drug discovery platform. This provides a simple, rapid, and reliable technique for identifying and developing bispecific and multispecific agents.
[0170] Antigen-binding domains are exemplary domains that can be used and applied in accordance with the present invention. In certain embodiments of the present invention, the antigen-binding domain comprises a peptide tag capable of forming an isopeptide bond, such as SpyTag or SnoopTag, and can be linked via an isopeptide bond to a construct comprising a cognate catcher domain, for example, to create a combinatorial or modular screening platform. In other embodiments of the present invention, the constructs of the present invention comprise a first antigen-binding domain at the N-terminus and a second antigen-binding domain at the C-terminus, the first and second binding domains being separated by a structural domain. Optionally, a linker sequence may be present between one or each of the antigen-binding domains and the structural domain. As discussed below, suitable peptide linkers for use in connecting binding domains (binding sites) to structural domains (monomer subunits) are typically 1 to 150, 1 to 100, 1 to 50, 1 to 25, 1 to 20, 1 to 15, or 1 to 10 amino acids in length. Examples of linker sequences are GSGS, GGGGS, GGGGSGGGGS, or GGGGSGGGGSGGGGGS. Other suitable linkers include linkers that form alpha-helical secondary structures, such as EAAAK and repeats thereof, such as (EAAAK)2, (EAAAK)5, or (EAAAK)7, or an alpha-helical linker derived from the ribosomal L9 protein having the sequence PANLKALEAQKQKEQRQAAEELANAKKLKEQLEK (Kuhlman et al 1997 J Mol Biol 270, 5, 640-647) or derivatives thereof (e.g., derivatives containing 1, 2, 3, 4, 5, or up to 10 substitutions, insertions, or deletions) or repeats thereof, such as 2, 3, 4, 5, or up to 10 repeats of the L9 linker.Other suitable linkers include linkers that confer variable functionality to the fusion protein, such as those discussed in Chen et al 2013 (Adv Drug Deliv Rev. 2013, 65, 10, 1357-1369), including, but not limited to, PAPAP, CC, AP, VSQTSKLTRAETVFPDV, PLGLWA, RVLAEA, EDVVCCSMSY, GGIEGRGS, TRHRQPRGWE, AGNRVRRSVG, RRRRR, RRRRRRRRR, GFLG, or LE, or derivatives thereof (e.g., derivatives containing 1, 2, 3, 4, 5, or up to 10 substitutions, insertions, or deletions) or repeats thereof, e.g., 2, 3, 4, 5, or up to 10 repeats of the PAPAP linker. The catcher domain shares structural similarity with antibody domains and has an immunoglobulin fold (Kang et al. 2007, Science. 318, 5856, 1625-1628). Substituting the catcher domain with a human antibody domain allows for an alternative route to generate a humanized direct fusion structure. Therefore, further suitable linkers include antibody fragment domains, such as CH1, CH2, or CH3, derived from various antibody isotypes or various species, preferably of human origin, and modified versions thereof.Examples include a CH1 domain from a human IgG1 isotype having the sequence ASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKKV, a CH2 domain from a human IgG1 isotype having the sequence PCPAPELLGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVEVHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPIEKTISKAK, or a CH3 domain from a human IgG1 isotype having the sequence GQPREPQVYTLPPSRDELTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK.
[0171] The antigen-binding domain is typically an antigen-binding fragment of an antibody. Antigen-binding antibody fragments are not full-length intact antibodies and typically lack at least the CH2 and / or CH3 domains. Such antigen-binding fragments are well known and include Fab, F(ab')2, Fv, or single-chain Fv fragments (scFv). Antigen-binding fragments typically contain the CDRs (typically six CDRs) necessary for antigen binding and framework residues necessary for correct CDR structure. In some embodiments, the antigen-binding domain comprises a heavy (H) chain variable domain sequence (VH) and a light (L) chain variable domain sequence (VL).
[0172] In some embodiments, the antigen-binding region may be a single-domain antibody. Single-domain antibodies (sdAbs) can include antibodies whose complementarity-determining regions are part of a single-domain polypeptide. Examples include, but are not limited to, heavy-chain antibodies, antibodies naturally lacking light chains, single-domain antibodies derived from conventional four-chain antibodies, genetically engineered antibodies, and single-domain scaffolds other than those derived from antibodies. Single-domain antibodies may be any of those in the art or any future single-domain antibodies. Single-domain antibodies may be derived from species including, but not limited to, mouse, human, camel, llama, fish, shark, goat, rabbit, and cow. Single-domain antibodies may also be naturally occurring single-domain antibodies known as heavy-chain antibodies lacking light chains. Such single-domain antibodies are disclosed, for example, in WO 94 / 04678. For clarity, this variable domain derived from a heavy-chain antibody naturally lacking light chains is sometimes referred to as a VHH or nanobody to distinguish it from the conventional VH of four-chain immunoglobulins. Such VHH molecules may be derived from antibodies raised in Camelidae species, such as camel, llama, dromedary, alpaca, and guanaco. Other non-Camelidae species can produce heavy chain antibodies that naturally lack light chains, and such VHHs are within the scope of the present invention.
[0173] The antigen-binding domain may also comprise or consist of an antibody mimetic, such as an affibody or DARPin. As known in the art, an affibody is a small polypeptide typically containing about 58 amino acids and three alpha helices with a molecular mass of about 6 kDa. As known in the art, a DARPIN (designed ankyrin repeat protein) is a genetically engineered antibody-mimetic protein that typically exhibits highly specific and high-affinity target protein binding.
[0174] The binding domain may comprise a naturally occurring ligand such as a cytokine, as an alternative to an antigen binding domain.
[0175] When two different antigen-binding domains are present in a bispecific molecule, they typically bind to different epitopes. This may be different epitopes on the same target molecule, or different epitopes on different target molecules. In some embodiments, both epitopes are on a therapeutic target, and the binding of the antigen-binding domains to the therapeutic target modifies a biological mechanism, typically a pathological mechanism, for therapeutic benefit.
[0176] In some embodiments, each antigen-binding domain can agonize a biological target. In some embodiments, each antigen-binding domain can antagonize a biological target. In some embodiments, one antigen-binding domain can agonize a first target and the other antigen-binding domain can antagonize a second target.
[0177] In some embodiments, the construct may comprise two binding regions that bind to the same epitope but have different affinities for that epitope, hi some embodiments, the construct may comprise two binding regions that bind to the same epitope, optionally with different affinities, and the two binding regions have different formats, for example, one antigen-binding region is an scFv and the second antigen-binding region is a Fab.
[0178] Multi-domain polypeptide constructs of the invention typically have the following format, shown in N to C orientation according to conventional convention: (binding domain 1)-linker 1-structural domain-linker 2-(binding domain 2) where Linker1 and Linker2 are optional linker sequences, optionally 1 to 20 amino acids, e.g., GSGS. An optional purification tag, e.g., a His tag, e.g., 6xHis, can be incorporated at either end of the construct.
[0179] The binding domains may be the same or, typically, different. A binding domain may typically be a catcher polypeptide or an antigen-binding domain. A number of exemplary configurations are provided below. SpC-Linker1-CutA1-Linker2-SpC SnC-Linker1-CutA1-Linker2-SnC SnC-Linker1-CutA1-Linker2-SpC SpC-Linker1-CutA1-Linker2-SnC SpC3-Linker1-CutA1-Linker2-DgC scFv-Linker1-CutA1-Linker2-ScFv Fab-Linker1-CutA1-Linker2-Fab ScFv-Linker1-CutA1-Linker2-Fab Fab-Linker1-CutA1-Linker2-ScFv Nanobody 1-Linker 1-CutA1-Linker 2-Nanobody 2 Nanobody-Linker1-CutA1-Linker2-DgC SpC3-Linker1-CutA1-Linker2-Nanobody where SpC is SpyCatcher, SpC3 is SpyCatcher003, SnC is SnoopCatcher, Fab is antibody "fragment antigen binding," and scFv is single-chain Fv. In each case, the linker 1 and linker 2 linkers are optional. The CutA1 sequence may be derived from human or Pyrococcus horikoshii, or may be a homologue from another species, or may have at least 30%, at least 50%, at least 70%, or at least 90% identity to the human or Pyrococcus horikoshii sequence.
[0180] Further exemplary constructs are as follows: SpC-Linker1-NC1-Linker2-SpC SnC-Linker1-NC1-Linker2-SnC SnC-Linker1-NC1-Linker2-SpC SpC-Linker1-NC1-Linker2-SnC scFv-Linker1-NC1-Linker2-ScFv Fab-Linker1-NC1-Linker2-Fab ScFv-Linker1-NC1-Linker2-Fab Fab-Linker1-NC1-Linker2-ScFv Nanobody 1-Linker 1-NC1-Linker 2-Nanobody 2 Nanobody-Linker1-NC1-Linker2-DgC SpC3-Linker1-NC1-Linker2-Nanobody where NC1 is the collagen NC1 domain derived from collagen type VIII or collagen type X. In either case, the linkers Linker 1 and Linker 2 are optional.
[0181] Further exemplary constructs include macrophage migration inhibitory factor (MIF) (SEQ ID NO: 25) or macrophage migration inhibitory factor 2 (MIF2) (SEQ ID NO: 26) or the S62A F99A mutant of MIF2 (MIF2m, SEQ ID NO: 27) as structural domains, including: SpC-linker1-MIF2-linker2-SpC SnC-linker1-MIF2-linker2-SnC SnC-linker1-MIF2-linker2-SpC SpC-linker1-MIF2-linker2-SnC scFv-linker1-MIF2-linker2-ScFv Fab-linker1-MIF2-linker2-Fab ScFv-linker1-MIF2-linker2-Fab Fab-linker1-MIF2-linker2-ScFv Nanobody 1-Linker 1-MIF2-Linker 2-Nanobody 2 Nanobody-linker1-MIF2-linker2-DgC SpC3-linker1-MIF2-linker2-nanobody
[0182] Any suitable structural domain, particularly any of the structural domains described herein, can be used in such constructs in place of the exemplary structural domains provided above, and therefore, such exemplary formats are described for use in structural domains generally and in the structural domains described herein.
[0183] The orientation of the binding domains of a multidomain construct can be functionally assessed in either monomeric or oligomeric form using a variety of assays, a selection of exemplary assays being described below.
[0184] FRET: FRET assays can be performed to demonstrate the cis orientation of a selected scaffold compared to a non-cis-oriented protein. A scaffold polypeptide with a catcher moiety, such as SpC3-HsCutA1-DgC, can be conjugated to a fluorescent protein FRET pair fused to the respective tag pairs, such as mCherry(6+)-SpT3-H6 and H6-DgT-mCitrine(4-). Once the tagged FRET pair is conjugated to the scaffold protein, the emission of the acceptor FRET protein can be measured using standard fluorescence reading methods and compared to the emission enhancement of the donor FRET protein. Protein scaffolds that exhibit a preferential cis orientation will exhibit higher acceptor emission, while protein scaffolds that exhibit a preferential trans orientation will exhibit higher donor emission enhancement.
[0185] SPR: SPR experiments can be performed to demonstrate that cis-oriented scaffolds conjugated to suitable ligands can preferentially bind to targets in the same plane compared to non-cis-oriented proteins. Target proteins for the scaffold-conjugated ligands, such as the L1 and L2 targets of SpC-PhCutA1-SnC:SnT-L1:L2-SpT, can be immobilized on the surface of an SPR sensor chip, either together or separately for control purposes, with only the L1 target or only the L2 target. Then, cis-oriented or non-cis-oriented scaffolds conjugated to L1 and L2 are loaded onto an SPR sensor chip with the L1 target and / or L2 target, and the binding of the conjugate assembly to the immobilized targets on the chip is determined. Assemblies with cis orientation are expected to show highly measurable binding to both the L1 target and the L2 target when both targets are immobilized on the same chip, while assemblies with non-cis orientation are expected not to show highly measurable binding to such a chip. Both assembly types are expected to exhibit highly measurable binding to immobilized L1 targets only or L2 targets only.
[0186] SEC-MALS: SEC-MALS experiments can be performed to demonstrate that the scaffold and assembly proteins in solution are in a native oligomeric state. The scaffold and assembly proteins can be prepared as described in the Methods section. The sample is then injected into an FPLC instrument coupled to a MALS instrument and detector to separate the sample by size, and the native protein mass is estimated by light scattering calculations. The oligomeric state of the protein can then be derived by dividing the native protein mass by the predicted monomer mass calculated using software such as ProtParam. Both scaffold and assembly proteins, such as SpC-PhCutA1-SnC and SpC-PhCutA1-SnC:SnT-L1:L2-SpT, are expected to exhibit an oligomeric state of 3.
[0187] Crosslinking and LC-MS / MS: To demonstrate simultaneous binding of cis-oriented conjugate assemblies to two targets, e.g., L1 and L2, target-expressing cells can be incubated with a biotinylated version of the conjugate assembly, biotin-SpC-PhCutA1-SnC:SnT-L1:L2-SpT. The binding between the target and the biotinylated assembly can then be crosslinked with BS3 (bis(sulfosuccinimidyl)suberate), followed by cell lysis and extraction of the crosslinked target-assembly complex with streptavidin. The complexes can then be trypsin-digested and subjected to LC-MS / MS to confirm binding of L1 and L2 to their respective targets. By performing the same methodology on non-cis-oriented assemblies, the LC-MS / MS data output can be compared between the two protein assemblies. Cis-oriented assemblies are expected to be able to bind preferentially to both L1 and L2 targets, whereas non-cis-oriented assemblies are expected to be able to bind preferentially to either L1 or L2 targets.
[0188] Multivalent Protein Scaffolds One aspect of the present invention relates to a modular system for screening target molecules. The system allows for multivalent display of target molecules. In one aspect, a multivalent protein scaffold is provided. The multivalent protein scaffold comprises an oligomeric core comprising multiple subunit monomers. The multivalent protein scaffold also comprises at least one first binding site that is orthogonal to at least one second binding site. Suitable binding sites are described in more detail herein.
[0189] The scaffold serves as a platform onto which other molecules can be attached. Various combinations of molecules can be attached to the scaffold in a modular manner. The scaffold allows for multivalent binding of molecules. Generally, molecules attached to the scaffold have potential therapeutic benefits, and the attached scaffold can be used to investigate whether multivalent assemblies of different molecules exhibit the desired effect. For example, various combinations of potential anti-cancer polypeptides can be attached to the scaffold, and the resulting assemblies can then be used in screening assays to determine whether the combination exhibits an effect, such as binding to and killing cancer cells. Once a combination of molecules is identified, therapeutic drug candidates can be generated by modifying the multivalent protein scaffold to directly attach the identified molecules rather than using a modular system. The multivalent protein scaffold "presents" the molecules on the same surface of the scaffold, allowing all of the molecules to potentially interact with target cells.
[0190] The provided scaffold typically comprises at least two first binding sites and at least two second binding sites. The existence of at least two first binding sites and at least two second binding sites allows this scaffold to be used for screening multivalent interactions that cannot necessarily be found when screening using, for example, a bispecific antibody format. It also allows various scaffolds to be tested as a "toolbox" for the same combination.
[0191] Therefore, in this specification, - an oligomeric core comprising a plurality of subunit monomers; and - at least two first binding sites that are orthogonal to at least two second binding sites; 1. A multivalent protein scaffold comprising: A multivalent protein scaffold is provided, wherein the first binding site and the second binding site are located on the same face of the scaffold.
[0192] The provided scaffolds typically contain binding sites capable of forming covalent bonds with their respective targets. The covalent bonds result in strong, irreversible connections. The complexes produced by covalently attaching the scaffolds of the present invention to the targets at the binding sites are physically robust and easy to produce. Such complexes can be produced with high yields and high homogeneity. Therefore, the biological responses that occur when such complexes are administered to a biological system, such as the subject described herein, are reproducible and controllable.
[0193] Therefore, in this specification, - an oligomeric core comprising a plurality of subunit monomers; - at least one first binding site that is orthogonal to at least one second binding site; 1. A multivalent protein scaffold comprising: Also provided is a multivalent protein scaffold, wherein the first binding site and the second binding site are located on the same face of the scaffold, the first binding site comprising a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprising a second protein domain capable of forming a covalent bond with a second polypeptide target.
[0194] As described herein, the scaffolds provided herein have significant advantages over conventional antibodies, including bispecific antibodies. The oligomer core of the provided scaffolds typically does not comprise the Fc region of an antibody. In some embodiments, the oligomer core does not comprise the CH2 region. In some embodiments, the oligomer core does not comprise the CH3 region. In some embodiments, the oligomer core does not comprise the CH2 region and does not comprise the CH3 region. The use of constant domains such as the immunoglobulin domain of an antibody, typically the Fc region, typically does not exhibit the advantages of the provided scaffolds described herein. For example, bispecific antibodies lack the modularity of the present invention and may not be useful for investigating multivalent interactions.
[0195] Therefore, in this specification, - an oligomeric core comprising a plurality of subunit monomers; - at least one first binding site that is orthogonal to at least one second binding site; 1. A multivalent protein scaffold comprising: Also provided is a multivalent protein scaffold, wherein the first binding site and the second binding site are located on the same face of the scaffold and the oligomeric core does not comprise an Fc region of an antibody.
[0196] Also, - an oligomeric core comprising a plurality of subunit monomers; - at least one first binding site that is orthogonal to at least one second binding site; 1. A multivalent protein scaffold comprising: Also provided are multivalent protein scaffolds, wherein the first binding site and the second binding site are located on the same face of the scaffold, and wherein the oligomeric core does not comprise a CH2 domain of an antibody, or does not comprise a CH3 domain of an antibody, or does not comprise a CH2 domain and does not comprise a CH3 domain.
[0197] The multivalent protein scaffold comprises an oligomeric core, at least one first binding site, and at least one second binding site, and may also include other features such as linkers, domain inserts, and / or functional groups, as described in more detail herein.
[0198] Preferably, the diameter of the multivalent protein scaffold is less than about 100 nm, such as less than about 50 nm, for example less than about 25 nm, for example less than about 10 nm. Preferably, the height of the multivalent protein scaffold is less than about 100 nm, for example less than about 50 nm, for example less than about 30 nm, for example less than about 20 nm, for example less than about 10 nm. The multivalent protein scaffold is preferably 1 to 500 Å in size, for example 2 to 250 Å, for example about 10 to about 100 Å, for example about 20 to about 80 Å.
[0199] Preferably, the multivalent protein scaffold itself preferably does not essentially induce an immune response in a biological system, cell culture, or subject, such as a human subject. In other words, typically, in the absence of effector moieties attached to the binding sites of the oligomeric core and / or the binding sites of the protein scaffold, administration of the protein scaffold to a biological system, such as a human subject, does not induce an immune response (or essentially does not induce an immune response, e.g., the immune response is not greater than that of a non-immunogenic protein). For example, administration of the protein scaffold (in the absence of effector moieties attached to the binding sites of the multivalent protein scaffold and / or the binding sites of the multivalent protein scaffold) to a biological system (e.g., a subject as defined herein) typically does not induce either the innate or adaptive immunity of the biological system. For example, the protein scaffold typically does not induce activation of the complement system, B cells, T cells, natural killer cells, mast cells, basophils, eosinophils, neutrophils, dendritic cells, or macrophages.
[0200] The multivalent protein scaffold preferably does not include an antibody or antibody fragment, although, as further described herein, antibodies and / or antibody fragments may be attached to the scaffold as effector moieties. The multivalent protein scaffold (e.g., no effector moieties whatsoever) more preferably does not include the Fc region of an antibody. The Fc region is the tail region of an antibody that interacts with cell surface receptors called Fc receptors and some proteins of the complement system. In some cases, the multivalent protein scaffold does not include an immunoglobulin constant region. In some embodiments, the multivalent protein scaffold does not include a CH2 domain. In some embodiments, the multivalent protein scaffold does not include a CH3 domain. In some embodiments, the multivalent protein scaffold does not include a CH2 domain and does not include a CH3 domain.
[0201] Preferably, the multivalent protein scaffold is thermodynamically stable, for example, at temperatures of about 0 to about 100°C, such as about 4°C to about 90°C, for example, about 10°C to about 50°C, for example, about 20 to about 38°C, for example, about 25 to about 37°C. In other words, preferably, the multivalent protein scaffold does not dissociate into its substituent subunit monomers and / or binding moieties do not dissociate from the oligomer core when present in aqueous solution at a temperature of about 0 to about 100°C (e.g., about 4°C to about 90°C, e.g., about 10°C to about 50°C, e.g., about 20 to about 38°C, e.g., about 25 to about 37°C), and e.g., at least 90%, such as at least 95%, such as at least 99%, such as at least 99.9%, e.g., at least 99.99% or 99.999% of the multivalent protein scaffold does not dissociate into its substituent subunit monomers and / or binding moieties do not dissociate from the oligomer core when present in aqueous solution at such a temperature. More preferably, the multivalent protein scaffold is stable at a temperature of about 0° C. to about 100° C., such as about 4° C. to about 90° C., for example, about 10° C. to about 50° C., for example, about 20° C. to about 38° C., for example, about 25° C. to about 37° C. The multivalent protein scaffold preferably has a shelf life of at least 10 minutes, more preferably at least 1 hour, for example, at least 1 day, for example, at least 1 week, for example, at least 1 month or at least 1 year, when determined at a temperature of about 0° C. to about 100° C., for example, about 4° C. to about 90° C., for example, about 10° C. to about 50° C., for example, about 20° C. to about 38° C., for example, about 25° C. to about 37° C.
[0202] Preferably, the interactions between the constituents of the multivalent protein scaffold are not weak, transient interactions. Weak, transient complexes exhibit a dynamic mixture of different oligomeric states in vivo, whereas strong, transient complexes change quaternary states only when triggered, for example, by ligand binding. Weak, transient interactions have dissociation constants (K) in the micromolar range. D ) and a lifetime of a few seconds. Strong, transient interactions stabilized by the binding of an effector molecule have a longer lifetime and a lower K in the nanomolar range. DThe constituent parts of the multivalent protein scaffold preferably interact with each other in at least a strong transient reaction, and more preferably, the constituent parts of the multivalent protein scaffold form permanent interactions. Permanent interactions mean that the multivalent protein scaffold does not dissociate into its constituent parts under normal conditions, e.g., 20°C to 40°C and pH 6 to 8. Multivalent protein scaffolds in which the constituent parts form permanent interactions typically dissociate only under denaturing conditions that denature the tertiary structure of the subunit monomers themselves.
[0203] Therefore, preferably, the constituents of the multivalent protein scaffold have a K of less than 1 μM, e.g., less than 100 nM, more preferably less than 10 nM, at a temperature of about 0° C. to about 100° C., e.g., about 4° C. to about 90° C., e.g., about 10° C. to about 50° C., e.g., about 20° C. to about 38° C., e.g., about 25° C. to about 37° C. D interact with each other.
[0204] Preferably, the multivalent protein scaffold and its constituent parts are stable to proteases. For example, the multivalent protein scaffold and its constituent parts can be exposed to a protease, such as trypsin, without losing the tertiary or quaternary structure of the scaffold. The multivalent protein scaffold and its constituent parts are stable to proteases, for example, at a temperature of about 10 to about 40° C., for example, about 20 to about 38° C., for example, about 25 to about 37° C., for at least 1 hour, for example, at least 2 hours, for example, at least 4 hours, for example, at least 8 hours, for example, at least 24 hours, or longer.
[0205] Oligomeric Core As explained above, the multivalent protein scaffolds provided herein comprise an oligomeric core comprising multiple subunit monomers, the subunit monomers of which are typically structural domains of the polypeptide constructs described elsewhere herein.
[0206] Any suitable number of subunit monomers can be used. For example, the oligomer core may contain from about 2 to about 20 subunit monomers, such as from about 2 to about 10 subunit monomers, more preferably from 3 to 7 subunit monomers, and more preferably from 3 to 6 subunit monomers. For example, the oligomer core may contain 2, 3, 4, 5, 6, 7, 8, 9, or 10 subunit monomers. Preferably, the oligomer core contains at least three subunit monomers. Most preferably, the oligomer core contains three subunit monomers. Preferably, the oligomer core does not contain or consist of seven subunit monomers.
[0207] Preferably, the subunit monomers, when polymerized, have rotational symmetry, such as 3-fold, 4-fold, 5-fold, 6-fold, or 7-fold. The oligomer core may have C2, C3, C4, D2, C5, C6, D3, C7, C8, D4, C9, C10, D5, C11, C12, D6, or T symmetry.
[0208] Some or all of the subunit monomers of the oligomer core may be non-covalently attached together. Some or all of the subunit monomers of the oligomer core may be covalently attached together. The oligomer core may contain a mix of covalently and non-covalently attached subunit monomers. For example, the oligomer core may contain a first monomer that is covalently bound to a second monomer to form a heterodimer. The oligomer core may contain at least two such heterodimers non-covalently attached together. For example, the oligomer core may contain three non-heterodimers non-covalently attached together, each heterodimer containing two monomers covalently attached together.
[0209] Some or all of the subunit monomers of the oligomer core may be attached by non-covalent interactions, including, but not limited to, electrostatic interactions such as ionic, hydrogen, and halogen bonds, van der Waals forces such as dipole-dipole interactions, π-π stacking, cation-π interactions, anion-π interactions, or polar-π interactions.
[0210] Some or all of the subunit monomers of the oligomer core may be covalently attached together. When the subunit monomers are covalently attached together, the monomers are typically of amino acid sequence corresponding to original or naturally occurring monomer domains.
[0211] Two or more subunit monomers may be covalently linked by disulfide bonds. Disulfide bonds are typically formed between cysteine residues in polypeptides. Artificial amino acids with free thiol groups can also participate in disulfide bond formation.
[0212] Two or more subunit monomers may be covalently attached by chemical crosslinking. Crosslinking reagents include homobifunctional crosslinking reagents, heterobifunctional crosslinking reagents, and photoreactive crosslinking reagents. Homobifunctional crosslinking reagents have the same reactive group at either end. Examples of homobifunctional crosslinking reagents include disuccinimidyl suberate (DSS), disuccinimidyl tartrate (DST), and succinimidyl dithiobispropionate (DSP). Some common examples of sulfhydryl-sulfhydryl crosslinkers include BMOE and DTME. Heterobifunctional crosslinking reagents have two different reactive groups and can be used to link different functional groups. Examples of heterobifunctional crosslinking reagents include MDS (m-maleimidobenzoyl-N-hydroxysuccinimide ester), GMBS (N-γ-maleimidobutyryloxysuccinimide ester), EMCS (N-(ε-maleimidocaproyloxy)succinimide ester), and sulfo-EMCS (N-(ε-maleimidocaproyloxy)sulfosuccinimide ester). Photoreactive crosslinking reagents are heterobifunctional crosslinkers that exhibit reactivity only upon exposure to ultraviolet or visible light. Two common classes of photoreactive chemical groups are aryl azides and diazirines. Aryl azides (N-((2-pyridyldithio)ethyl)-4-azidosalicylamides) are widely used. Upon exposure to 250-350 nm UV light, these reagents can promote the formation of nitrene groups, which can catalyze addition reactions with double bonds. Additionally, these crosslinkers can initiate the generation of C-H insertion products or react with nucleophiles. Some common cross-linking reagents in this group include ANB-NOS (N-5-azido-2-nitrobenzyloxysuccinimide) and sulfo-SANPAH. NHS ester diazirine or adipentanoate contains a photoactivatable diazirine ring and an N-hydroxysuccinimide (NHS) ester, which reacts efficiently with primary amino groups in neutral to basic buffers to form stable amide bonds.These exhibit better photostability compared to phenyl azide groups and can be readily activated with long-wavelength ultraviolet light (330-370 nm) to generate carbene intermediates that form covalent bonds with any peptide backbone or amino acid side chain within a spacer arm distance.
[0213] More preferably, two or more subunit monomers of the oligomer core may be genetically fused together. The subunit monomers are genetically fused together when they are encoded and expressed by a single polynucleotide sequence so as to be expressed as a single polypeptide chain. Thus, when the subunit monomers are genetically fused together, the oligomer core can comprise a single polypeptide chain.
[0214] Genetically fused subunit monomers may be genetically fused together via a peptide linker. A suitable peptide linker for use in linking subunit monomers is an amino acid sequence that can function as a hinge region between the subunit monomers, thus providing sufficient flexibility to allow the subunit monomers to fold independently of each other and retain the ability to multimerize. The length, flexibility, and hydrophilicity of the peptide linker are typically designed to allow the subunit monomers to easily assemble to form an oligomer core. Preferably, subunit monomers linked by a peptide linker can assemble to form an oligomer core in which the interactions between adjacent subunit monomers are substantially identical to the interactions between the same subunit monomers otherwise.
[0215] Peptide linkers suitable for use in connecting oligomeric core monomer subunits are typically 1-100, 1-50, 1-25, 1-20, 1-15, or 1-10 amino acids in length. Linkers may be composed of, for example, one or more of the following amino acids: lysine, serine, arginine, proline, glycine, and alanine. An example of a suitable flexible peptide linker is a stretch of 2-20, e.g., 4, 6, 8, 10, or 16, serine and / or glycine amino acids. An example of a rigid linker is a stretch of 2-30, e.g., 4, 6, 8, 16, or 24, proline amino acids. Examples of suitable linkers include, but are not limited to, the following: GGGS, PGGS, PGGG, RPPPPP, RPPPP, VGG, RPPG, PPPP, RPPG, PPPPPPPPP, PPPPPPPPPPPP, RPPG, GG, GGG, SGSG, SGSGSG, SGSGSGSG, SGSGSGSGSG, SGSGSGSGSG, and SGSGSGSGSGSGSGSG, where G is glycine, P is proline, R is arginine, S is serine, and V is valine. Additional exemplary linkers include GSGS, GGGGS, GGGGSGGGGS, and GGGGSGGGGSGGGGGS. Suitable linking groups can be designed using conventional modeling techniques. Linkers typically have sufficient flexibility to allow the monomers or their subunits to assemble into the corresponding protein oligomers.
[0216] Preferably, the total molecular weight of the oligomer core is less than about 1000 kDa, for example, less than about 500 kDa, for example, less than about 250 kDa. The total molecular weight of the oligomer core is preferably about 10 kDa to about 1000 kDa, for example, about 10 kDa to about 500 kDa, for example, about 10 kDa to about 250 kDa, for example, about 10 kDa to about 150 kDa. The total molecular weight of the oligomer core is more preferably about 20 kDa to about 150 kDa.
[0217] Preferably, the diameter of the oligomer core is less than about 100 nm, for example, less than about 50 nm, for example, less than about 30 nm, for example, less than about 20 nm, for example, less than about 10 nm. Preferably, the height of the oligomer core is less than about 100 nm, for example, less than about 50 nm, for example, less than about 30 nm, for example, less than about 20 nm, for example, less than about 10 nm. The oligomer core is preferably 1 to 50 nm, for example, 2 to 40 nm, for example, about 2 to about 20 nm, for example, about 5 to about 10 nm in size.
[0218] Preferably, the oligomer core is thermodynamically stable. Preferably, the oligomer core is stable at a temperature of about 0 to about 50°C. In other words, preferably, the oligomer core does not spontaneously dissociate into its substituent monomers when present in a solution at a temperature of about 0 to about 50°C. More preferably, the oligomer core is stable at a temperature of about 10 to about 40°C, for example, about 20 to about 38°C, for example, about 25 to about 37°C.
[0219] Preferably, the subunit monomers stably interact to form the oligomeric core. The interaction between the subunit monomers is preferably not a weak, transient interaction. Weak, transient complexes typically exhibit a dynamic mixture of different oligomeric states in vivo, whereas strong, transient complexes change quaternary state only when triggered, for example, by ligand binding, and exist in a single, predominant oligomeric state (e.g., at least 90%, e.g., at least 95%, e.g., at least 99%, e.g., at least 99.9%, e.g., at least 99.99%, or 99.999% of the complex can exist in a stable oligomeric state under standard conditions). Weak, transient interactions have dissociation constants (K) in the micromolar range. D ) and a lifetime of a few seconds. Strong, transient interactions stabilized by the binding of an effector molecule typically have a longer lifetime and a lower K in the nanomolar range. DThe subunit monomers more preferably interact in at least a strong transient reaction, and more preferably form a permanent interaction. Permanent interaction means that the oligomer does not or does not substantially dissociate into its constituent subunit monomers under normal conditions (e.g., when present in aqueous solution at a temperature of about 0 to about 100°C. Under such conditions, typically at least 90%, e.g., at least 95%, e.g., at least 99%, e.g., at least 99.9%, e.g., at least 99.99%, or 99.999% of the oligomer core does not dissociate), and typically the oligomer core dissociates only when denaturing conditions are used that denature the tertiary structure of the subunit monomers themselves. Thus, preferably, the subunit monomers have a K of less than 1 μM, e.g., less than 100 nM, more preferably less than 10 nM. D The oligomer core typically has a lifetime of at least 10 minutes, more preferably at least 1 hour, e.g., at least 1 day, e.g., at least 1 week, e.g., at least 1 month, or at least 1 year. The lifetime can be determined at any suitable temperature, such as from about 0°C to about 100°C, e.g., from about 4°C to about 90°C, e.g., from about 10°C to about 50°C, e.g., from about 20°C to about 38°C, e.g., from about 25°C to about 37°C, etc.
[0220] Preferably, the oligomer core is stable to proteases, for example, the oligomer core can be exposed to dilute concentrations of a protease, such as trypsin, for a limited period of time, e.g., 4 hours, without loss of the tertiary or quaternary structure of the oligomer core.
[0221] Preferably, the oligomer core is human or humanized. A human oligomer core is a multimer region of a human protein. A humanized oligomer core is an oligomer core that is a multimer region of a non-human protein that has been modified to more closely resemble the corresponding multimer region of a human protein. The humanized oligomer core may contain at least 50% amino acid identity, e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% amino acid identity, with the amino acid sequence of the multimer region of the corresponding human protein. The corresponding multimer region of the human protein is the multimer region of the human protein that has the highest amino acid sequence identity with the humanized oligomer core. This information can be identified using a Blast search (blast.ncbi.nlm.nih.gov) and limiting search results to Homo sapiens.
[0222] Preferably, the oligomer core of the multivalent protein scaffold does not itself induce an immune response in a biological system, cell culture, or subject (such as a non-human subject or a human subject). In other words, typically, in the absence of the binding site of the oligomer core and / or an effector moiety attached to the binding site of the oligomer core, administration of the oligomer core to a biological system does not induce an immune response. For example, administration of the oligomer core (in the absence of the binding site of the oligomer core and / or an effector moiety attached to the binding site of the oligomer core) to a biological system (e.g., a subject as defined herein) typically does not induce either the innate or adaptive immunity of the biological system. For example, the oligomer core typically does not induce activation of the complement system, B cells, T cells, natural killer cells, mast cells, basophils, eosinophils, neutrophils, dendritic cells, or macrophages. However, molecules attached to the oligomer core can be specifically designed to generate an immune response in a human subject.
[0223] Preferably, the oligomer core of the protein scaffold does not comprise an antibody or antibody fragment, although, as further described herein, antibodies and / or antibody fragments may be attached to the oligomer core as effector moieties. The oligomer core (e.g., without any effector moieties) more preferably does not comprise the Fc region of an antibody. In some cases, the oligomer core does not comprise an immunoglobulin constant region.
[0224] In some embodiments, the oligomeric core of a multivalent protein scaffold may be a homo-oligomeric core, i.e., the oligomeric core may comprise only one type of monomer. The homo-oligomeric core may comprise two or more, e.g., three or more, four or more, five or more, six or more, or seven or more identical subunit monomers.
[0225] In some other embodiments, the oligomer core of the multivalent protein scaffold may be a hetero-oligomer core, i.e., the oligomer core may comprise more than one type of monomer. Different types of subunit monomers can form the oligomer core. In other words, the subunit monomers can be attached together. The hetero-oligomer core may comprise two or more, for example, three or more, four or more, five or more, six or more, or seven or more different subunit monomers. For example, if the hetero-oligomer core comprises three monomers, the hetero-oligomer core may comprise two first monomers and one second monomer, or one first monomer, one second monomer, and one third monomer. When the hetero-oligomer core comprises four monomers, the hetero-oligomer core may comprise two first monomers and two second monomers; two first monomers, one second monomer, and one third monomer; or one first monomer, one second monomer, one third monomer, and one fourth monomer. Preferably, when the oligomer core is a hetero-oligomer core, the oligomer core comprises two types of subunit monomers, i.e., type A and type B monomers, and in the case of a hetero-oligomer core having a total of n subunit monomers, the hetero-oligomer core comprises (A a B b ), where a+b=n. When the oligomer core is a hetero-oligomer core, the subunit monomers contained therein may be modified such that a first type of subunit monomer preferentially binds to a second type of subunit monomer over another monomer of the first type (in other words, for a hetero-oligomer core comprising type A monomers and type B monomers, the hetero-oligomer core is of the form ABABAB··· rather than AAABBB···).
[0226] In yet further embodiments, the oligomer core may comprise multiple multimeric subunits. For example, two monomers may be fused (e.g., as a tandem fusion), and the monomer fusions may assemble to form the oligomer core. In this regard, the fused monomers may be the same or different.
[0227] For example, two or more identical monomers can be fused and the fusion product can be further assembled with identical fusion products to form a homo-oligomer core. The oligomer core can comprise multiple homodimers; for example, if the homodimer is "AA," the oligomer core can comprise "AA," "AAAA," or "AAAAAA," etc.
[0228] Alternatively, two or more different monomers can be fused together, and the fusion product can be assembled with additional identical fusion products to form an oligomeric core. As used herein, such a core is typically considered a homo-oligomeric core, with each fusion product being considered a monomeric subunit. The oligomeric core may also comprise multiple heterodimers; for example, if the heterodimer is "AB," the oligomeric core may comprise "AB," "ABAB," or "ABABAB," etc.
[0229] Alternatively, two or more identical monomers can be fused, and the fusion product can be assembled with additional non-identical fusion products to form a hetero-oligomeric core. The oligomeric core can comprise multiple homodimers; for example, if the first homodimer is "AA" and the second homodimer is "BB," the oligomeric core can comprise "AABB," etc.
[0230] Alternatively, two or more different monomers can be fused together and the fusion product can be assembled with additional non-identical fusion products to form a hetero-oligomeric core. The oligomeric core can comprise multiple heterodimers; for example, if the first heterodimer is "AB" and the second heterodimer is "CD," the oligomeric core can comprise "ABCD," etc.
[0231] The homo-oligomer core comprises multiple subunit monomers, each monomer comprising at least one first binding site and at least one second binding site. For example, if the homo-oligomer core comprises three subunit monomers, the oligomer core will comprise at least three first binding sites and at least three second binding sites. If the homo-oligomer core comprises four, five, six, or seven subunit monomers, the oligomer core will comprise at least four first binding sites and at least four second binding sites; at least five first binding sites and at least five second binding sites; at least six first binding sites and at least six second binding sites; or at least seven first binding sites and at least seven second binding sites. Thus, a multivalent protein scaffold preferably comprises at least two first binding sites and at least two second binding sites, i.e., all of the first binding sites are identical to other first binding sites and all of the second binding sites are identical to other second binding sites. A multivalent protein scaffold more preferably comprises at least three first binding sites and at least three second binding sites. In some embodiments, a multivalent protein scaffold comprises at least four, at least five, at least six, at least seven, or at least eight first and second binding sites, respectively.
[0232] A hetero-oligomer core comprises a plurality of subunit monomers, including at least two types of subunit monomers, where one type of subunit monomer comprises at least one first binding site and a second type of subunit monomer comprises at least one second binding site. For example, if the hetero-oligomer core comprises three subunit monomers, the hetero-oligomer core may comprise three different binding sites, or two first binding sites and one second binding site. If the hetero-oligomer core comprises four subunit monomers, the hetero-oligomer core may comprise four different binding sites, or two first binding sites, one second binding site, and one third binding site, or two first binding sites and two second binding sites.
[0233] Binding sites are described in more detail herein.
[0234] Monomer As explained above, the multivalent protein scaffolds provided herein comprise an oligomeric core comprising multiple subunit monomers, which are typically structural domains of the multi-domain polypeptide constructs described elsewhere herein.
[0235] Each subunit monomer (excluding any binding sites attached thereto, as described in more detail herein) preferably contains fewer than 300 amino acids, preferably fewer than 200 amino acids, more preferably fewer than 150 amino acids. For example, each subunit monomer (excluding any binding sites attached thereto) preferably has a molecular weight of less than 40 kDa, e.g., less than 30 kDa, e.g., less than 20 kDa. Protein scaffolds described herein comprising such monomers have a relatively low mass and are capable of efficient diffusion in vivo. They are typically expressed in bacterial or yeast cell expression systems and are capable of correct folding. Such expression systems can often produce much higher yields than mammalian cell culture, which is typically required for antibody production.
[0236] The subunit monomers preferably do not comprise or consist of an antibody or antibody fragment. The oligomer core or subunit monomers preferably do not comprise or consist of an Fc region of an antibody. In some embodiments, the subunit monomers do not comprise or consist of a CH2 domain. In some embodiments, the subunit monomers do not comprise or consist of a CH3 domain. In some embodiments, the subunit monomers do not comprise a CH2 domain or a CH3 domain.
[0237] Preferably, each monomer subunit of the oligomeric core is human or humanized. A human monomer is a monomer of a human oligomeric protein. A humanized monomer is a monomer of a non-human oligomeric protein that has been modified to more closely resemble the monomer of the corresponding human protein. Thus, a humanized monomer may contain at least 50% amino acid identity, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% amino acid identity with the amino acid sequence of the corresponding human protein. Typically, a human or humanized protein will not provoke an adverse immune response in a patient to whom it is administered.
[0238] The subunit monomers included in the oligomer core preferably each include a multimerization structural element, which is a structural and / or functional feature of the subunit monomer that allows the subunit monomer to multimerize.
[0239] The multimerization structural element can be a protein domain. The multimerization structural element of the oligomer core can be the multimerization domain of a naturally occurring multimeric protein or a de novo multimerization domain. A protein domain is an autonomous folding unit of a protein. A multimerization domain is typically a protein domain that is involved in protein-protein interactions with another protein domain. The multimerization structural element is preferably soluble, and therefore the monomer and oligomer core are soluble.
[0240] Considering the disclosure of this specification, those skilled in the art can identify suitable multimerization structural elements for use in the present invention.For example, those skilled in the art can identify multimeric proteins.For example, databases such as NCBI database (www.ncbi.nlm.nih.gov) and Protein Data Bank (PDB; www.rscb.org) list many multimeric proteins, and can search for multimeric proteins with rotational symmetry axis.
[0241] Preferably, the identified multimeric protein is a homo-oligomer, such as a homo-dimer, homo-trimer, homo-tetramer, homo-pentamer, homo-hexamer, or homo-heptamer. The multimeric protein may also be a hetero-oligomer, such as a hetero-dimer, hetero-trimer, hetero-tetramer, hetero-pentamer, hetero-hexamer, or hetero-heptamer. Functional and / or structural information can be used to identify which domain of the multimeric protein is involved in multimerization, i.e., the multimerization domain.
[0242] The multimerization structural element preferably comprises the multimerization interface of the multimerization domain (i.e., the structural or functional element of the multimerization domain that enables the domain to multimerize). Other aspects of the multimerization domain can be modified without affecting the multimerization of the subunit monomers. Thus, the subunit monomers of the oligomer core preferably comprise the multimerization structural element. The subunit monomers preferably comprise at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity with the multimerization domain from which they are derived. The subunit monomers retain the ability to form multimers (i.e., the oligomer core).
[0243] Preferably, the subunit monomers of the oligomer core comprise soluble multimerization structural elements of the multimeric protein. Soluble domains are preferred over multimerization domains found, for example, in membranes. Preferably, each subunit monomer of the oligomer core comprises a soluble multimerization structural element of the soluble multimeric protein.
[0244] The multimerization structural element may be derived from any multimeric protein with suitable symmetry (e.g., rotational or dihedral symmetry, as described in more detail herein), such as collagen (e.g., collagen NC1 domain), CutA, TNF, p53, fibrinogen, C4, Bacillus subtilis AbrB, or homologs or paralogs thereof.
[0245] Preferably, the subunit monomer may comprise a monomer or multimerization domain of a protein selected from the following: type X collagen (PDB ID: 1GR3) (e.g., its NC1 domain), type VIII collagen (PDB ID: 1o91) (e.g., its NC1 domain), C1q head domain (e.g., PDB ID: 1PK6 is the globular head of human C1q), or Pyrococcus horikoshii, Homo sapiens (PDB ID: 2ZFH), Thermus thermophiles (PDB ID: 1V6H); Oryza sativa (PDB ID: 2ZOM); or Shewanella sp. SIB1 (PDB ID: or a polypeptide having at least 60% amino acid sequence identity, at least 70%, at least 80%, at least 90%, at least 95%, or at least 97%, at least 98%, or at least 99% amino acid sequence identity to any one of the preceding polypeptides.
[0246] In some embodiments, the subunit monomer comprises or consists of: type X collagen (PDB ID: 1GR3, or SEQ ID NO: 2) (e.g., the NC1 domain thereof), type VIII collagen (PDB ID: 1o91, or SEQ ID NO: 3) (e.g., the NC1 domain thereof), heteromeric C1q head domain (e.g., PDB ID: 1PK6 is the globular head of human C1q; see SEQ ID NOs: 36-38), or a CutA (copper tolerance A) protein such as the CutA1 protein from Pyrococcus horikoshii (PDB ID: 4YNO, or SEQ ID NO: 1), Homo sapiens (PDB ID: 2ZFH, or SEQ ID NO: 19), Thermus thermophilus (PDB ID: 1V6H, or SEQ ID NO: 39), or rice (PDB ID: 2ZOM, or SEQ ID NO: 40); or Shewanella species SIB1 (PDB ID: 3AHP, or SEQ ID NO: 41); TNF-like protein TL1A (PDB ID: 2RE9, or SEQ ID NO: 31); TNF (PDB ID: 1TNF, or SEQ ID NO: 42, full length soluble domain SEQ ID NO: 80); TNF family protein CD40L (PDB ID: 3LKJ, or SEQ ID NO: 58); human macrophage migration inhibitory factor (MIF) (PDB ID: 1CA7, or SEQ ID NO: 25, or PDB ID: 6OY8 with Y99G mutation); human macrophage migration inhibitory factor 2 (MIF2) (PDB ID: 7MSE, or SEQ ID NO: 26, or SEQ ID NO: 27 with S62A and F99A mutations), or homologs or paralogs thereof.
[0247] Other multimerization domains include those of the antiparallel coiled-coil hexamer (PDB ID: 5W0J, see Example 4, SEQ ID NO: 43), HIV-1 GP41 core (PDB ID: 1I5Y or SEQ ID NO: 44), cytochrome c555 (PDB ID: 5Z25 or SEQ ID NO: 45), MHC class II-associated chaperonin and targeting protein invariant chain (Ii) (PDB ID: 1iie or SEQ ID NO: 46); p53 (PDB ID: 1C26 or SEQ ID NO: 47); fibrinogen-like domain (PDB ID: 4M7F or SEQ ID NO: 48), type IV collagen C4 (PDB ID: 1LI1 or SEQ ID NO: 49); Bacillus subtilis AbrB (PDB ID: 1L1 or SEQ ID NO: 49); ID:1YFB or SEQ ID NO:50); or a polypeptide having at least 50% amino acid sequence identity with any one of the preceding polypeptides; more preferably, a polypeptide having at least 60% amino acid sequence identity, for example at least 70%, at least 80%, at least 90%, at least 95%, or at least 97%, at least 98%, or at least 99% amino acid sequence identity with any one of the preceding polypeptides.
[0248] Other multimerization domains include the multimerization domains of bacteriophage lambda head domain D (e.g., PDB ID: 1C5E or SEQ ID NO: 51); a domain-swap trimer variant of HCRBPII (PDB ID: 6VIS or SEQ ID NO: 52); T1L reovirus attachment protein sigma 1 (chains A, B, C of PDB ID: 4ODB, or SEQ ID NO: 53); or a polypeptide having at least 50% amino acid sequence identity to any one of the preceding proteins; more preferably, at least 60% amino acid sequence identity to any one of the preceding proteins, for example, at least 70%, at least 80%, at least 90%, at least 95%, or at least 97%, at least 98%, or at least 99% amino acid sequence identity.
[0249] Preferably, in one embodiment, the oligomer core may comprise monomers derived from the multimerization structural elements of Pyrococcus horikoshii CutA1. Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or even 100% amino acid identity to the amino acid sequence of SEQ ID NO:1.
[0250] As discussed above and elsewhere herein, CutA1 (e.g., Pyrococcus horikoshii) is an exemplary structural domain of a multi-domain polypeptide construct.
[0251] In certain embodiments, the oligomer core may comprise monomers derived from the multimerization structural element of human CutA1 (SEQ ID NO: 19). Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 19.
[0252] In certain embodiments, the multimerization structural element is a truncated form of SEQ ID NO: 19 (human CutA1) as described elsewhere herein. In certain embodiments, the multimerization structural element is a modified form of SEQ ID NO: 19 (human CutA1), in which at least one cysteine residue of SEQ ID NO: 19 has been replaced with a different amino acid residue as described elsewhere herein.
[0253] Without wishing to be bound by theory, one advantage of CutA1, TNF, OX40L, CDL40L, or TL1A is that even when used as a monospecific assembly (e.g., as a "direct fusion" structure, such as a bispecific therapeutic agent), the use of a single surface display reduces the multimeric state. For example, the inventors have observed that CutA1-mediated Catcher-based display can achieve effective inhibition as hexavalent and trivalent assemblies with a trimeric core protein. Reducing the multimeric state is expected to result in a decrease in molecular weight and / or increased stability.
[0254] In certain embodiments, the oligomer core may comprise monomers derived from multimerization structural elements of human TNF (SEQ ID NO: 80) or OX40L (SEQ ID NO: 78) or CD40L (SEQ ID NO: 58) or TL1A (SEQ ID NO: 31). Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 80, 78, 58, or 31.
[0255] In certain embodiments, the multimerization structural element is a truncated form of SEQ ID NO: 80, 78, 58, or 31, as described elsewhere herein. An exemplary truncated sequence is SEQ ID NO: 79. In certain embodiments, the multimerization structural element is a modified form of SEQ ID NO: 80, 78, 58, or 31, in which at least one cysteine residue in the sequence has been substituted with a different amino acid residue as described elsewhere herein. In certain embodiments, the multimerization structural element is truncated, in which at least one cysteine residue has been substituted with respect to SEQ ID NO: 80, 78, 58, or 31.
[0256] Preferably, in another embodiment, the oligomer core may comprise monomers derived from the multimerization structural elements of type X collagen NC1. Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO:2.
[0257] As discussed above and elsewhere herein, type X collagen NC1 is a typical structural domain of a multi-domain polypeptide construct.
[0258] Preferably, in another embodiment, the oligomer core may comprise monomers derived from the multimerization structural elements of type VIII collagen NC1. Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO:3.
[0259] As discussed above and elsewhere herein, type VIII collagen NC1 is an exemplary structural domain of a multi-domain polypeptide construct.
[0260] Preferably, in another embodiment, the oligomer core may comprise monomers derived from the multimerization structural elements of CutA1 derived from Homo sapiens, and thus the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO: 19.
[0261] As discussed above and elsewhere herein, CutA1 (e.g., human) is a typical structural domain of a multi-domain polypeptide construct. Furthermore, as described in the Examples, the present inventors have genetically engineered human CutA1 to optimize and improve its characteristics and suitability as a structural domain.
[0262] Preferably, in another embodiment, the oligomer core may comprise monomers derived from the multimerization structural elements of MIF or MIF-2. Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, e.g., at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% amino acid identity to the amino acid sequence of SEQ ID NO:25 or SEQ ID NO:26 or SEQ ID NO:27.
[0263] As discussed above and elsewhere herein, MIF is a typical structural domain of a multi-domain polypeptide construct. As discussed above and elsewhere herein, MIF-2 is a typical structural domain of a multi-domain polypeptide construct.
[0264] Preferably, in another embodiment, the oligomer core may comprise monomers derived from multimerizing structural elements TNF family proteins, including TNF and TNF-like TL1A, OX40L or CD40L. Thus, the or each subunit monomer may comprise or consist of a polypeptide having at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% amino acid identity to the amino acid sequence of SEQ ID NO:42 or SEQ ID NO:31 or SEQ ID NO:58.
[0265] As discussed above and elsewhere herein, TNF is a typical structural domain of a multi-domain polypeptide construct. As discussed above and elsewhere herein, proteins of the TNF superfamily are typical structural domains of a multi-domain polypeptide construct. The TNF superfamily is known as a protein superfamily of type II transmembrane proteins that contain TNF homology domains and form trimers. This superfamily includes 19 members that bind to 29 members of the TNF receptor superfamily, including CD40L, OX40L, and TNF-α (also referred to herein as "TNF"). Other members of the superfamily include TNF-b, TNF-g, Fas ligand, CD27 ligand, CD30 ligand, CD137 ligand, CD253, CD354, APO-3L, CD256, CD257, CD258, TL1, TITR ligand, and ED1-A1.
[0266] When the oligomer core is a hetero-oligomer core, each monomer of the oligomer core may be derived from the same protein. For example, each monomer of the oligomer core may be derived from one of the proteins described above. The oligomer core may be heteromeric because the binding sites attached to the monomer subunits are different, even if the multimerization domains of each monomer subunit are identical. The oligomer core may be heteromeric because the multimerization domains of the monomer subunits are different, even if all the monomer subunits are derived from the same protein.
[0267] Those skilled in the art will appreciate that the monomeric subunits of the oligomeric core of the multivalent protein scaffolds provided herein can be further modified to provide additional functionality or beneficial properties.
[0268] For example, the monomer may be a fragment, derivative, or variant of the monomer or multimerization structural element described herein.As will be understood by those skilled in the art, a fragment of an amino acid sequence includes a deletion variant of such a sequence, in which one or more amino acids, for example, at least 1, 2, 5, 10, 20, 50, or 100, are deleted.The deletion may occur at the C-terminus or N-terminus of the native sequence, or may occur within the native sequence.Typically, the deletion of one or more amino acids does not affect the residues immediately surrounding the multimerization structural element of the subunit monomer.
[0269] Derivatives of amino acid sequences include post-translationally modified sequences, including sequences that are modified in vivo or ex vivo. Many different protein modifications are known to those skilled in the art, including modifications to introduce new functionality into amino acid residues, modifications to protect reactive amino acid residues, or modifications to couple amino acid residues to chemical moieties such as linkers for attachment to such amino acid residues or reactive functional groups on substrates (surfaces).
[0270] Derivatives of amino acid sequences also include addition variants of such sequences in which one or more, for example, at least 1, 2, 5, 10, 20, 50, or 100 amino acids have been added or introduced into the native sequence. The addition may occur at the C-terminus or N-terminus of the native sequence, or may occur internally in the native sequence. Typically, the addition of one or more amino acids does not affect residues immediately surrounding the multimerization structural elements of the subunit monomer.
[0271] Variants of amino acid sequences include sequences in which one or more amino acids of a native sequence, for example, at least 1, 2, 5, 10, 20, 50, or 100 amino acid residues, are replaced with one or more non-natural residues. Thus, such variants may contain point mutations or more significant mutations, for example, native chemical ligation can be used to splice a non-natural amino acid sequence into a partial native sequence to produce a variant of a native enzyme. Variants of amino acid sequences include sequences with naturally occurring amino acids and / or non-natural amino acids.
[0272] Variants, derivatives, and functional fragments of the above-mentioned amino acid sequences typically retain the oligomerization ability of the wild-type sequence. Preferably, variants, derivatives, and functional fragments of the above-mentioned sequences exhibit improved properties compared to the wild-type or native sequence, such as increased stability, reduced toxicity, or additional functionality, including binding sites.
[0273] binding site The multivalent protein scaffold comprises at least one first binding site and at least one second binding site. In some embodiments, typically in a modular system useful for identifying useful combinations of effector molecules in drug discovery, at least one first binding site is orthogonal to at least one second binding site. In other words, the chemical reaction by which the first binding site binds to its target (first target) is orthogonal to the chemical reaction by which the second binding site binds to its target (second target). Thus, the first target will bind to the first binding site but not to the second binding site, and the second target will bind to the second binding site but not to the first binding site. Thus, the term orthogonal is given its usual meaning in the field of protein-protein interactions, where the first binding interaction (i.e., the first binding site and the first ligand) is independent of the second binding interaction (i.e., the second binding site and the second ligand).
[0274] The binding sites of the multivalent protein scaffold allow the scaffold to be used as a modular system for binding effector moieties, with the first and second binding sites binding to their cognate targets.
[0275] The first and second binding sites can be incorporated into the multivalent protein scaffolds provided herein in any suitable manner. In one embodiment, the first and second binding sites are provided as tandem fusions attached to the or each monomer of the oligomeric core as described herein to form the multivalent protein scaffold. For comparison, SEQ ID NO: 22 is an example of two binding sites (described herein) provided as a fusion linked by an αH linker.
[0276] The binding sites are described in more detail below.
[0277] The interaction between the binding site and its target may be a non-covalent interaction. Preferably, the or each binding site can form a covalent bond with its respective target. The reactive functional group may be naturally present in the subunit monomer or effector moiety, or may be introduced, for example, by genetic engineering or chemical modification of the monomer. The reactive group may be derived from an unnatural amino acid that is incorporated into the monomer during its synthesis or expression, for example, during cell-free expression, for example, via in vitro transcription / translation.
[0278] The binding site of the multivalent protein scaffold can bind to its target via a reactive group. Any suitable reactive group can be used. For example, the reactive group can be an amine reactive group; a carboxyl reactive group; a sulfhydryl reactive group, or a carbonyl reactive group. The reactive group can include a cysteine reactive group. The reactive group can include a maleimide, an azide, a thiol, an alkyne, an NHS ester, or a haloacetamide.
[0279] The reactive group may be a group capable of reacting with an unnatural amino acid, such as 4-azido-L-phenylalanine (Faz) and any one of amino acids 1-71 in Figure 1 of Liu CC and Schultz PG, Annu. Rev. Biochem., 2010, 79, 413-444. Such a group is particularly useful when the corresponding unnatural amino acid is included in the binding site and the cognate target.
[0280] The reactive group may be a click chemistry group. Click chemistry is a term first introduced by Kolb et al. in 2001 to describe an extensive set of powerful, selective, and modular building blocks that function reliably in both small- and large-scale applications (Kolb HC, Finn, MG, Sharpless KB, Click chemistry: diverse chemical function from a few good reactions, Angew. Chem. Int. Ed. 40 (2001) 2004-2021). Kolb et al. define a set of strict criteria for click chemistry: "The reaction must be modular, broad in scope, give very high yields, produce only innocuous by-products that can be removed by non-chromatographic methods, and be stereospecific (but not necessarily enantioselective). Required process features include simple reaction conditions (ideally, the process should be unaffected by oxygen and water), readily available starting materials and reagents, no solvent or the use of a mild or easily removable solvent (such as water), and convenient product isolation. Purification, if necessary, should be by non-chromatographic methods such as crystallization or distillation, and the product should be stable under physiological conditions."
[0281] For example, the first and second binding sites may comprise orthogonal click chemistry reagents. Suitable examples of click chemistry include, but are not limited to, the following: (a) A copper-free variant of the 1,3 dipolar cycloaddition reaction in which an azide reacts with an alkyne under strain, for example, on a cyclooctane ring; (b) reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other linker; (c) Staudinger ligation, in which the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with an azide to give an amide bond; (d) Nitrone dipolar cycloaddition, (e) norbornene cycloaddition, (f) oxanorbornadiene cycloaddition, (g) tetrazine ligation, (h) [4+1] cycloaddition, (i) tetrazole photoclick chemistry, and (j) Quadricyclane ligation.
[0282] The reactive group may be a haloacetamide, such as iodoacetamide, bromoacetamide, or chloroacetamide.
[0283] The reactive groups can be selected from vinyl groups, TCOs, tetrazines, and strained alkynes; DBCOs; activated acids, such as acid chlorides; and piperazines, and reactive amines.
[0284] Host-guest chemistry can also be used to provide a reaction between a binding site and its target. For example, a binding site can include a ligand for binding to a metal complex, and the target can include a metal complex, or vice versa. Thus, a binding site can include a metal complex that can interact non-covalently through chelation or supramolecular association with its target, which includes a site that can act as a ligand to complex with a modifier molecule by forming a stable association, or vice versa.
[0285] The reactive group may be any of those disclosed in Sakamoto and Hamachi, "Recent progress in chemical modification of proteins," Anal. Sci 2019 (35) 5-27; or McKay and Finn, "Click chemistry in complex mixtures: bioorthogonal bioconjugation," Chem. Biol. 2014, 21(9) 1075-1101, both of which are incorporated herein by reference in their entirety.
[0286] The binding sites of the multivalent protein scaffold preferably comprise polypeptides such as protein domains. More preferably, a first binding site comprises a first protein domain and said second binding site comprises a second protein domain.
[0287] When the first binding site comprises a first protein domain and the second binding site comprises a second protein domain, the first binding site and / or the second binding site are preferably genetically fused to the subunit monomers to which they are attached, forming a single polypeptide chain. Typically, the first binding site and / or the second binding site are expressed from a recombinant nucleic acid molecule as a single polypeptide chain with the subunit monomers to which they are attached, e.g., as a fusion protein. This can be advantageous in that the multivalent protein scaffold can be expressed in a state ready to be conjugated to an effector moiety without the need for further chemical modifications, e.g., those required for attaching click chemistry reagents. Attachment of protein binding sites to the proteins to which they are attached, e.g., monomeric subunits of the oligomeric core of the multivalent protein scaffold, is described below.
[0288] The first binding site may comprise a first protein domain capable of forming a non-covalent bond with a first polypeptide target, and the second binding site may comprise a second protein domain capable of forming a non-covalent bond with a second polypeptide target. More preferably, the first binding site comprises a first protein domain capable of forming a covalent bond with a first polypeptide target, and the second binding site comprises a second protein domain capable of forming a covalent bond with a second polypeptide target. Any suitable covalent bond can be formed, including the above examples. Preferably, the first protein domain is capable of forming an isopeptide bond with the first polypeptide target, and the second protein domain is capable of forming an isopeptide bond with the second binding target. An isopeptide bond is, for example, an amide bond that can be formed between the carboxyl group of one amino acid and the amino group of another amino acid. At least one of these linking groups is typically part of the side chain of one of these amino acids.
[0289] Preferably, the first and second binding sites each comprise a different split protein domain, such as a split ligand-binding protein domain. As used herein, a ligand-binding protein domain refers to a domain of a protein-binding ligand. While any suitable protein can be used, proteins that are naturally stabilized by intrachain covalent bonds, such as isopeptide bonds, are particularly useful. In such cases, the portion of the protein containing the isopeptide bond donor residue is separated from the portion of the peptide containing the isopeptide bond acceptor residue. The two protein fragments can be attached to additional polypeptides, such as the oligomer core monomers and / or polypeptide targets described herein, for example, by genetic fusion. Contacting the two separate fragments typically results in the creation of an isopeptide bond that irreversibly attaches the two fragments. Thus, split-protein approaches to generating binding sites and complementary tags are particularly useful. Such binding site / tag pairs are typically orthogonal, since one protein fragment will preferentially or exclusively bind to its natural partner (i.e., the complementary portion of the protein from which it was derived) over any other potential partners. These principles are discussed, for example, in Reddington & Howarth, Curr. Op. Chem. Biol. 29, 94-99 (2015) and Keeble et al, PNAS 2019 116(52) 26523.
[0290] Preferably, one of the first binding site and the second binding site comprises a split Streptococcus pyogenes fibronectin binding protein domain, and the other of the first binding site and the second binding site comprises a split Streptococcus pneumoniae adhesion domain.
[0291] Preferably, the first protein domain and first polypeptide target, and the second protein domain and second polypeptide target may each comprise a peptide linker pair such as those described in WO 2016 / 193746, WO 2018 / 197854, WO 2018 / 189517, Keeble et al. (PNAS 116(52), 2019: 26523-26533), Fierer et al. (PNAS 111(13), 2014: E1176-E1181).
[0292] Preferably, the first binding site and the second binding site each independently have at least 50% amino acid identity with any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18.
[0293] Preferably, the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are each independently selected from: (i) a combination of any one of SEQ ID NO: 4, 6, or 8 with any one of SEQ ID NO: 5, 7, or 9; (ii) a combination of SEQ ID NO: 12 with SEQ ID NO: 13 or 15; (iii) a combination of SEQ ID NO: 5 with SEQ ID NO: 11; (iv) a combination of SEQ ID NO: 15 with SEQ ID NO: 16; (v) a combination of SEQ ID NO: 17 with SEQ ID NO: 18, or (vi) a combination of SEQ ID NO: 23 with SEQ ID NO: 16.
[0294] More preferably, the first protein domain and first polypeptide target, and the second protein domain and second polypeptide target are each independently selected from the following pairs:
[0295] [Table 3-1]
[0296] [Table 3-2]
[0297] The protein domain and targeting domain may have at least 50% amino acid identity, for example at least 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% amino acid identity, with the sequences set forth above, while retaining the ability of the protein domain to specifically bind to the targeting domain.
[0298] In some embodiments, the first binding site is a protein domain, and the first target is a tag that binds to the first protein domain. The second binding site is a protein domain, and the second target is a tag that binds to the second protein domain. In some embodiments, the first binding site is a tag, and the first target is a protein domain that binds to the first tag. The second binding site is a tag, and the second target is a protein domain that binds to the second tag. In some embodiments, the first binding site is a protein domain, and the first target is a tag that binds to the first protein domain. The second binding site is a tag, and the second target is a protein domain that binds to the second tag. Preferably, both the first and second binding sites are protein domains, and the first and second targets are tags that specifically bind to the first and second protein domains, respectively.
[0299] More preferably, the first protein domain and first polypeptide target, and the second protein domain and second polypeptide target are each independently selected from the following pairs:
[0300] [Table 4]
[0301] The protein domain and targeting domain may have at least 50% amino acid identity, for example at least 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% amino acid identity, with the sequences set forth above, while retaining the ability of the protein domain to specifically bind to the targeting domain.
[0302] The above binding groups and targets can be divided into the following subgroups: Subgroup A: - SpyCatcher(SEQ ID NO:4) / SpyTag(SEQ ID NO:5); - SpyCatcher(SEQ ID NO:4) / SpyTag002(SEQ ID NO:7); - SpyCatcher(SEQ ID NO:4) / SpyTag003(SEQ ID NO:9); - SpyCatcher002 (SEQ ID NO: 6) / SpyTag (SEQ ID NO: 5); - SpyCatcher002 (SEQ ID NO: 6) / SpyTag002 (SEQ ID NO: 7); - SpyCatcher002 (SEQ ID NO: 6) / SpyTag003 (SEQ ID NO: 9); - SpyCatcher003(SEQ ID NO:8) / SpyTag(SEQ ID NO:5); - SpyCatcher003 (SEQ ID NO: 8) / SpyTag002 (SEQ ID NO: 7); - SpyCatcher003 (SEQ ID NO: 8) / SpyTag003 (SEQ ID NO: 9); - SpyTag (SEQ ID NO: 5) / K-tag (SEQ ID NO: 11) (mediated by SpyLigase (SEQ ID NO: 10)) Subgroup B: - SnoopCatcher(SEQ ID NO:12) / SnoopTag(SEQ ID NO:13); - SnoopCatcher(SEQ ID NO:12) / SnoopTagJr(SEQ ID NO:15); - SnoopTagJr (SEQ ID NO: 15) / DogTag (SEQ ID NO: 16) (mediated by SnoopLigase (SEQ ID NO: 14) / - DogCatcher (SEQ ID NO: 23) / DogTag (SEQ ID NO: 16) Subgroup C - Pilin-C (SEQ ID NO: 17) / IsopepTag (SEQ ID NO: 18)
[0303] Preferably, the first binding site / target pair is selected from subgroup A and the second binding site / target pair is selected from subgroup B and subgroup C, or the first binding site / target pair is selected from subgroup B and the second binding site / target pair is selected from subgroup A and subgroup C, or the first binding site / target pair is selected from subgroup C and the second binding site / target pair is selected from subgroup A and subgroup B.
[0304] Even more preferably, the first protein domain-polypeptide target pair and the second protein domain-polypeptide target may be selected from the group consisting of: (i) SpyCatcher / SpyTag and SnoopCatcher / SnoopTag; (ii) SpyCatcher002 / SpyTag002 and SnoopCatcher / SnoopTag; (iii) SpyCatcher003 / SpyTag003 and SnoopCatcher / SnoopTag; (iv) SpyCatcher / S pyTag and Pilin-C / Isopeptag; (v) SpyCatcher002 / SpyTag002 and Pilin-C / Isopeptag; (vi) SpyCatcher003 / SpyTag003 and Pilin-C / Isopeptag; (vii) Pilin-C / Isopeptag and SnoopCatcher / SnoopTag; (viii) SpyCatcher / SpyTag and SnoopTagJr / DogTag; (ix) SpyCatcher002 / SpyTag002 and SnoopTa gJr / DogTag; (x) SpyTag / K-tag and SnoopCatcher / SnoopTag; (xi) SpyTag / K-tag and SnoopTagJr / DogTag; (xii) SpyTag / K-tag and Pilin-C / Isopeptag; and (xiii) SnoopTagJr / DogTag and Pilin-C / Isopeptag; and (xiv) SpyCatcher003 / SpyTag003 and DogCatcher / DogTag; and (xv) SpyCatcher003 / Sp yTag002 and DogCatcher / DogTag; and (xvi) SpyCatcher003 / SpyTag and DogCatcher / DogTag; (xvii) SpyCatcher003 / SpyTag003 and SnoopCatcher / SnoopTagJr; and (xviii) SpyCatcher003 / SpyTag002 and SnoopCatcher / SnoopTagJr; and (xix) SpyCatcher003 / SpyTag and SnoopCatcher / SnoopTagJr.
[0305] Those skilled in the art will understand that when using SpyLigase / SpyTag / K-tag or SnoopLigase / SnoopTagJr / DogTag, the "ligase" catalyzes the attachment of two tags. Thus, the first binding site and the first polypeptide target, and the second binding site and the second polypeptide target are selected from these two "tags." The ligase may be added exogenously to catalyze the attachment of the two "tags," or may be non-covalently or covalently connected to the multivalent protein scaffold, such as by genetically fusing with the multivalent protein scaffold. Which tag is included in the multivalent protein scaffold is interchangeable.
[0306] Other binding site / tag pairs include SdyTag / SdyCatcher (Tan et al, PLOS One 11(1) e0165074) and Cpe0147, derived from the Clostridium perfringens cell surface adhesion protein Cpe0147. 439~563 / Cpe0147 565~587 (Young et al, Chem Comm. 53(9) 1502).
[0307] "Specifically bind," as used herein in the context of the binding between a binding moiety and its target, refers to the ability of the binding moiety to bind to its complementary binding moiety with greater affinity than it binds to an unrelated control. The unrelated control may be an unrelated control protein. For example, SnoopCatcher specifically binds to SnoopTag with greater affinity than it binds to an unrelated control protein. The binding is preferably covalent, such as by formation of an isopeptide bond. Preferably, the control protein is bovine serum albumin, and the binding moiety binds to the complementary binding moiety with at least 10-fold, at least 50-fold, at least 100-fold, at least 500-fold, or at least 1000-fold greater affinity than the control protein. Affinity can be determined by methods known in the art. For example, affinity can be determined by ELISA assays, biolayer interferometry, surface plasmon resonance, kinetic methods, or equilibrium / solution methods. Those skilled in the art will recognize which binding moiety pairs specifically bind to generate protein complexes that can be used in the methods of the present invention.
[0308] The at least one first binding site and the at least one second binding site preferably do not comprise an antibody or antibody fragment, and more preferably do not comprise an antigen-binding fragment of an antibody, such as an Fab or Fc region.
[0309] In all discussions herein of "binding sites" that bind to a "target," one of skill in the art will readily understand that the chemical linkage groups are reversible. In other words, although the above bonds are described in terms of the reaction of reactive group A at the binding site with a corresponding reactive group B on the target for that binding site, reactive group B at the binding site is capable of an equivalent chemical reaction with reactive group A on the target.
[0310] When a multivalent protein scaffold includes one or more binding sites that are protein domains, e.g., protein domains attached to monomeric subunits of an oligomeric core of the multivalent protein scaffold, the protein domains can be attached to the multivalent protein scaffold (e.g., can be attached to monomeric subunits of the oligomeric core of the multivalent protein scaffold) by any suitable means.
[0311] The binding sites may be attached to the multivalent protein scaffold by a linker (e.g., attached to a monomer subunit of the oligomer core of the multivalent protein scaffold). In one embodiment, the same linker may be used at each end of the subunit monomer of the oligomer core. In another embodiment, different linkers may be used at each end of the subunit monomer of the oligomer core.
[0312] The binding site is preferably covalently attached to the oligomer core (or subunit monomer). The covalent linkage may be, for example, a peptide bond, a disulfide bond, or a click chemistry linkage. More preferably, the covalent linkage includes at least one amino acid, i.e., a peptide linker, that forms part of the same polypeptide chain as the subunit monomer at the binding site.
[0313] The peptide linker used to attach the binding site to the monomer subunit of the oligomeric core of the multivalent protein scaffold may be genetically fused to the subunit monomer and / or binding site. A linker is genetically fused when it is expressed from a single polynucleotide coding sequence as a single construct with the subunit monomer and / or binding site. The length, flexibility, and hydrophilicity of the peptide linker are typically designed to allow the binding sites to be located on the same face of the oligomeric core or multivalent protein scaffold. The peptide linker typically allows for directional tethering of the binding site.
[0314] Suitable peptide linkers for use in connecting binding sites to monomeric subunits are typically 1-100, 1-50, 1-25, 1-20, 1-15, or 1-10 amino acids in length. Linkers may be composed of, for example, one or more of the following amino acids: lysine, serine, arginine, proline, glycine, and alanine. An example of a suitable flexible peptide linker is a stretch of 2-20, e.g., 4, 6, 8, 10, or 16, serine and / or glycine amino acids. An example of a rigid linker is a stretch of 2-30, e.g., 4, 6, 8, 16, or 24, proline amino acids. Examples of suitable linkers include, but are not limited to, the following: GGGS, PGGS, PGGG, RPPPPP, RPPPP, VGG, RPPG, PPPP, RPPG, PPPPPPPPP, PPPPPPPPPPPP, RPPG, GG, GGG, SGSG, SGSGSG, SGSGSGSG, SGSGSGSGSG, SGSGSGSGSG, and SGSGSGSGSGSGSGSG, where G is glycine, P is proline, R is arginine, S is serine, and V is valine. Other exemplary linkers include GSGS, GGGGS, GGGGSGGGGS, and GGGGSGGGGSGGGGGS. Suitable linking groups can be designed using conventional modeling techniques. Linkers typically have sufficient flexibility to allow the binding sites and monomer subunits to assemble into the corresponding protein oligomer.
[0315] The oligomeric core of the multivalent protein scaffold preferably comprises at least one first binding site at an end of a subunit monomer and at least one second binding site at an end of a subunit monomer, e.g., the first binding site may be located at a first end of a subunit monomer and the second binding site may be located at a second end of a subunit monomer.
[0316] The termini are preferably determined with reference to the termini of the subunit monomers of the oligomer core, excluding any linker or binding site. More preferably, the termini of the subunit monomers are selected from the N-terminus and / or C-terminus of the subunit monomer. When the binding site (e.g., a protein domain) forms part of the same polypeptide as the subunit monomer, the N-terminus and C-terminus preferably refer to the amino acids corresponding to the respective ends of the monomers in the absence of the binding site. Similarly, when the linker forms part of the same polypeptide as the subunit monomer, the N-terminus and C-terminus preferably refer to the amino acids corresponding to the respective ends of the monomers in the absence of the linker.
[0317] Preferably, the ends of the subunit monomers to which the binding sites are attached are on the same face of the oligomeric core or multimeric protein scaffold, as defined in more detail above.
[0318] In some cases, the oligomer core comprises only a single end of each subunit monomer on a given face. In this case, at least one first binding site and at least one second binding site are typically both attached to the same end, thereby presenting the same face of the multivalent protein scaffold. For example, the oligomer core may comprise multiple subunit monomers, each subunit monomer comprising a first binding site attached to a first end of the monomer and a second binding site attached to the first binding site. The oligomer core may comprise multiple subunit monomers, with at least one first binding site attached to a first end of a first subunit monomer and at least one second binding site attached to a second end of a second subunit monomer (e.g., a hetero-oligomer core). The oligomer core may comprise a combination of the attachment methods described above.
[0319] The subunit monomers of the oligomeric core of the multivalent protein scaffold preferably each comprise two ends on a single face of the monomer (and thus on a single face of the oligomeric core and the multivalent protein scaffold). The two ends are preferably the N- and C-termini of the monomeric polypeptide. Each monomer more preferably comprises a first binding site attached to a first end of the monomer and a second binding site attached to a second end of the monomer. The first and second binding sites may be at the N- and C-termini, respectively, or at the C- and N-termini, respectively.
[0320] In some cases, a monomer may comprise more than one binding site at each end. For example, a subunit monomer may comprise at each end: (i) a first binding site attached to the end of the monomer and a second binding site attached to the first binding site, or vice versa, (ii) a first binding site attached to the end of the monomer and at least one additional first binding site attached to the first binding site, (iii) a second binding site attached to the end of the monomer and at least one additional second binding site attached to the second binding site, or (iv) a single first or second binding site attached to the end of the monomer.
[0321] Location of binding sites in multivalent protein scaffolds As described above, in the multivalent protein scaffolds provided herein, at least one first binding site and at least one second binding site are located on the same surface of the scaffold.Similarly, in typical embodiments of multi-domain polypeptide constructs, the first binding domain and the second binding domain are located on the same surface of the polypeptide construct.
[0322] By being located on the same face of the multivalent protein scaffold (or multi-domain polypeptide construct), the at least one first binding site and the at least one second binding site are ultimately positioned such that effector moieties attached to the multivalent protein scaffold via the binding sites can interact with their respective biological targets (e.g., cell surface receptors) on a single surface or plane.
[0323] At least one first binding site and at least one second binding site are located on the same face of the multivalent protein scaffold (or multi-domain polypeptide construct). Preferably, at least one first binding site and at least one second binding site are located on the same face of the oligomer core. Preferably, at least one first binding site and at least one second binding site are located on the same face of the subunit monomer to which they are attached. As described herein, the subunit monomer is typically a structural domain of the multi-domain polypeptide construct.
[0324] The term "at least one first binding site and at least one second binding site are located on the same face of the scaffold" can be understood as follows with reference to Figures 1 and 2.
[0325] The multivalent protein scaffold (1) or oligomeric core (10) contains a conceptual axis of rotational symmetry (20) corresponding to the number of monomers in the core. For example, a homotrimeric core contains a C3 axis of symmetry. For example, a homooligomeric pentameric core contains a C5 axis of symmetry. Similarly, a heterooligomeric core contains a conceptual axis of rotation that passes through the center of the oligomeric core and extends parallel to the interfaces between each subunit. For example, the axis of rotation of a heterodimer extends through the oligomeric core parallel to the length of the interfaces between the monomers. The axis of rotation of a heterotrimer extends through the oligomeric core as parallel as possible to the length of at least two interfaces between the monomers. A plane (21) can be defined as being perpendicular or nearly perpendicular (e.g., about 80° to about 100°, e.g., about 85° to about 95°, e.g., about 88° to about 92°, e.g., about 90°) to the axis of rotation and extending through the center of the oligomeric core. At least one first binding site (11) and at least one second binding site (12) are located on the same side of the plane (21), and thus on the same face of the multivalent protein scaffold (1). This is shown schematically in Figure 1, which shows a trimeric oligomer core and, for clarity, only one first binding site and one second binding site. In contrast, Figure 2 shows the contrasting situation in which at least one first binding site (11) and at least one second binding site (12) are located on opposite sides of the plane (21), and thus on opposite faces of the multivalent protein scaffold (1).
[0326] Those skilled in the art will understand that even when the monomers of the oligomer core are linked by covalent fusion, for example as described herein, a conceptual axis of symmetry still exists.
[0327] Thus, the "same face of a protein scaffold" may be a solvent-exposed surface of a multivalent protein scaffold that is perpendicular to the highest order rotational symmetry axis of the oligomer core of the multivalent protein scaffold and on one side of a plane extending through the center of the multivalent protein scaffold. Similarly, a face of an oligomer core may be a solvent-exposable surface of an oligomer core (preferably defined with no binding sites attached thereto) that is perpendicular to the highest order rotational symmetry axis of the oligomer core and on one side of a plane extending through the center of the oligomer core.
[0328] Preferably, the face of the multivalent protein scaffold is the solvent-exposed portion of the multivalent protein scaffold that contacts a single surface, e.g., the surface of a cell or protein complex, such as a cell wall, cell membrane, etc. Thus, referring to the schematic diagram of Figure 3 (which shows, for clarity, multiple first and second binding sites (11, 12) attached to the oligomeric core (10) of the multivalent protein scaffold (1)), at least one first binding site (11) and at least one second binding site (12) are preferably located on the multivalent protein scaffold (1) such that they can both contact the surface (30). This would not be possible if the at least one first binding site (11) and at least one second binding site (12) were located on opposite faces of the multivalent protein scaffold (1), as shown schematically in Figure 4. For example, a first binding site may be able to contact the surface (30), while a second binding site may not be able to contact the surface (30).
[0329] The at least one first binding site and the at least one second binding site are preferably arranged in a front-to-front orientation or a side-to-front orientation, or anywhere therebetween. As used herein, a front-to-front orientation refers to both the first binding site and the second binding site being located on the same side of the multivalent protein scaffold, with attachment substantially parallel to the axis of rotational symmetry through the multivalent protein scaffold. This is shown schematically in FIG. 5. A side-to-front orientation refers to one of the first binding site and the second binding site being located on one side of the multivalent protein scaffold, with attachment substantially parallel to the axis of rotational symmetry through the multivalent protein scaffold, and the other of the first binding site and the second binding site being located on the same side of the multivalent protein scaffold, with attachment substantially perpendicular to the axis of rotational symmetry through the multivalent protein scaffold (i.e., substantially parallel to plane (21)). This is shown schematically in FIG. 6. Of course, any position between these extremes can be used, for example, the first binding site and / or the second binding site are on the same face of the multivalent protein scaffold and the attachment is at approximately 45° to the axis of rotational symmetry through the multivalent protein scaffold, as shown schematically in Figure 7.
[0330] Those skilled in the art will understand that the angle between one or more first binding sites (11) and the axis (20) and the angle between one or more second binding sites (12) and the axis (20) do not have to be the same. For example, one or more first binding sites (11) may be present on the "front" of the multivalent protein scaffold, and one or more second binding sites (12) may be present on the "side" of the multivalent protein scaffold, i.e., a front-to-side orientation as described above. Alternatively, one or more first binding sites (11) may be present on the "side" of the multivalent protein scaffold, and one or more second binding sites (12) may be present on the "front" of the multivalent protein scaffold, i.e., a side-to-front orientation as described above. Both one or more first binding sites (11) and one or more second binding sites (12) may be present on the "front" of the multivalent protein scaffold, i.e., a front-to-front orientation as described above.
[0331] Assuming that the first and second binding sites are both present on the same face of the multivalent protein scaffold, and that the first and second binding sites are both attached to a given subunit monomer (2), typically the angle formed between the first and second binding sites and the center of that monomer (indicated by X in FIG. 8) is at most 160°, such as at most 140°, such as at most 120°, such as at most 100°, or at most 90°. Typically, the angle formed between the first and second binding sites and the center of that monomer is at least 10°, such as at least 20°, such as at least 30°, such as at least 45°, or at least 60°.
[0332] The structure can also be visualized by placing a flat target plane at any position in a 3D coordinate system so that it does not intersect with the surface of the oligomer core determined from protein structural data (NMR, X-ray) or structural prediction. For each fusion site, the distance of the shortest path to the target plane that does not intersect with the surface of the oligomer core (other than the original fusion site) can be determined. Preferably, for a given structure, the position of the target plane can be found such that all such shortest paths are less than 50%, 45%, or 40% of the longest protein cross-section perpendicular to the target plane.
[0333] In some embodiments, the maximum shortest path lengths to contact the same plane are less than 100 nm, such as less than 50 nm, for example less than 20 nm, for example less than 10 nm, for example less than 5 nm, for example less than 2 nm. In preferred situations, all shortest path lengths to the target plane end within a circular region on the target plane having a radius of less than 50 nm, for example less than 25 nm, for example less than 10 nm, for example less than 5 nm.
[0334] In some embodiments, the cis orientation of the protein fused to scaffold core can be determined by structural prediction.Alphafold can take into account not only distance but also the geometric shape of linker and the interaction between fusion binding domain and scaffold core protein.Preferably, scaffold core is predicted to retain its oligomerization property even after being fused to preferred binding site via linker, preferably via short linker, for example via GSGS, for example via GGGGS, for example via GGGGSGGGGS, for example via GGGSGGGGSGGGGGS, and it is predicted that binding site is approximately represented in cis geometric configuration.
[0335] Domain Insertion The multivalent protein scaffold comprises an oligomeric core comprising multiple subunit monomers and at least one first binding site that is orthogonal to at least one second binding site, preferably at least one first binding site and at least one second binding site located on the same face of the multivalent protein scaffold. The multivalent protein scaffold may further comprise a domain insert. The domain insert is a protein domain. The domain insert may be located on the same face of the multivalent protein scaffold as the binding site, or on a different face.
[0336] In some cases, when the oligomer core and / or subunit monomers comprise at least one free end that is not attached to a binding site, the multivalent protein scaffold comprises a domain insert at the free end. The domain insert may also be located within a loop region of the oligomer core, and thus within a loop region of a subunit monomer. Preferably, the multimeric protein scaffold comprises at least one domain insert on the opposite face of the oligomer core and / or multimeric protein scaffold to the face on which the binding site is located, e.g., within 90° of the opposite end of the axis.
[0337] As used herein, a domain insert is a polypeptide sequence that encodes a protein domain, i.e., an autonomously folding functional unit of a protein. The domain insert does not interfere with the structure and folding of the oligomeric core or binding site.
[0338] The domain insert preferably has an effector function. The domain insert may comprise an antibody, antibody fragment, or antigen-binding fragment, such as an antigen-binding fragment capable of binding to CD3 or CD16. For example, the domain insert can bind to an immunomodulatory protein such as a cytokine, or a chemotherapeutic agent, or a cancer immunotherapy agent (i.e., a treatment that uses a subject's immune system to treat cancer). The domain insert can form a protein that induces cell death upon contact with a biological system. The domain insert may induce apoptosis, increase anti-tumor responses, or exhibit other beneficial activities. The domain insert may have complement inhibitory or complement stimulatory activity.
[0339] For the avoidance of doubt, domain insertions are typically made within the structural domains of the multi-domain polypeptide constructs described herein.
[0340] protein complexes Also provided herein is a protein complex comprising a multivalent protein scaffold, as described in more detail herein, attached to at least one first effector moiety and at least one second effector moiety, wherein each first effector moiety is attached to a first target bound to a first binding site of the multivalent protein scaffold, and each second effector moiety is attached to a second target bound to a second binding site of the multivalent protein scaffold.
[0341] The target is preferably a polypeptide target, more preferably a partner of a peptide linker pair as described above.
[0342] Each effector moiety is attached to a target, thereby binding each effector moiety to a multivalent protein scaffold. The first and second effector moieties can be the same or different, preferably different. Any of the attachment routes described above in the context of the binding site can be used. Each effector moiety can be directly attached to a target using conventional organic chemistry routes available to those skilled in the art, thereby binding each effector moiety to a multivalent protein scaffold. Suitable chemical methods are described in textbooks such as March's Advanced Organic Chemistry (Wiley 2020).
[0343] Preferably, each effector moiety is covalently linked to the target. More preferably, the target is a polypeptide target and may be genetically fused to the effector moiety. In other words, preferably, the effector moiety or each effector moiety is genetically fused to the polypeptide target by encoding them in the same polynucleotide so that they are expressed as a single polypeptide chain. The effector may be genetically fused to a first polypeptide target, a cleavage site, and a second polypeptide target, where the first polypeptide target is orthogonal to the second polypeptide target. The cleavage site may be a TEV cleavage site. When both the first and second polypeptide targets are present in the effector moiety, only the terminal polypeptide target is functional (i.e., can bind to its cognate binding site on the multivalent protein scaffold). The cleavage site can be used to separate the terminal polypeptide targets so that only a single target is present. After conjugation of the effector moieties is complete, the terminal polypeptide target can be specifically unfolded.
[0344] The effector moiety is preferably a protein domain. The protein domain is preferably a soluble protein domain. The protein domain preferably comprises a domain of a secreted protein or an extracellular domain of a transmembrane protein. The protein domain more preferably comprises an extracellular domain of a cell surface receptor, such as a human cell surface receptor, or a ligand of a cell surface receptor.
[0345] The effector moiety is preferably a moiety that exerts a therapeutic effect upon contact with a biological system. For example, the effector moiety may be an immunomodulatory protein such as a cytokine, or may be a chemotherapeutic agent or a cancer immunotherapeutic agent (i.e., a treatment that uses a subject's immune system to treat cancer). The effector moiety may induce cell death upon contact with a biological system. The effector moiety may induce apoptosis, increase an anti-tumor response, or exhibit other beneficial activity. The effector moiety may have complement inhibitory activity or complement stimulatory activity. The effector moiety may result in altered gene expression, receptor internalization, cytokine release, cell death, or sensitivity to a therapeutic molecule.
[0346] In one embodiment, the effector moiety can be a synthetic organic molecule or a synthetic inorganic molecule.Suitable molecules can be chemotherapeutic agents.Suitable molecules can be toxic agents, for example, agents with an EC50 of less than about 100 μM, for example, less than about 10 μM, for example, less than about 1 μM or less than about 100 nM, where EC50 is the concentration of the agent required to cause 50% cytotoxicity as assessed by a suitable cell assay.Suitable cell assays can be, for example, sulforhodamine B (SRB) assays.
[0347] Suitable synthetic molecules may be enzyme activators or enzyme inhibitors. Suitable molecules may be inhibitors of one or more of serine / threonine / tyrosine kinases, matrix metalloproteinases (MMPs), heat shock proteins (HSPs), and proteasomes. Suitable molecules can act as alkylating agents (e.g., nitrogen mustards, nitrosoureas, tetrazines, aziridines, cisplatin, and their derivatives); antimetabolites (e.g., antifolates, fluoropyrimidines, deoxynucleoside analogs, and thiopurines); antimicrotubule agents (e.g., vinca alkyloids or taxanes); topoisomerase inhibitors (e.g., inhibitors of topoisomerase I, irinotecan and topotecan; topoisomerase II toxins such as etoposide, doxorubicin, mitoxantrone, and teniposide, or topoisomerase II inhibitors such as novobiocin, mervalone, and aclarubicin), or cytotoxic antibiotics (e.g., anthracyclines and bleomycin). Suitable molecules may have a molecular mass of about 50 to about 5000 g / mol, such as about 100 to about 1000 g / mol, for example about 250 to about 500 g / mol.
[0348] In another embodiment, the effector moiety preferably comprises an antibody or an antigen-binding fragment thereof. The term "antibody or antigen-binding fragment thereof," as used herein with respect to an effector moiety, may refer to a whole antibody (i.e., comprising two heavy and two light chain elements interconnected by disulfide bonds) as well as antigen-binding fragments thereof. An antibody typically comprises an immunologically active portion of an immunoglobulin (Ig) molecule, i.e., a molecule that contains an antigen-binding site that specifically binds (immunoreacts with) an antigen. As used herein, the terms "specifically bind" or "immunoreact with," in the context of the interaction of an antibody or fragment thereof with an antigen, mean that the antibody reacts preferentially with one or more antigenic determinants of the desired antigen compared to other polypeptides. Each heavy chain is composed of a heavy chain variable region (abbreviated herein as HCVR or VH) and at least one heavy chain constant region. Each light chain is composed of a light chain variable region (abbreviated herein as LCVR or VL) and a light chain constant region. The heavy and light chain variable regions contain binding domains that interact with the antigen. The VH and VL regions can be further subdivided into regions of hypervariability called complementarity-determining regions (CDRs), interspersed with more conserved regions called framework regions (FRs). Antibodies can include, but are not limited to, polyclonal, monoclonal, chimeric, dAb (domain antibodies), single chain, Fab, Fab', and F(ab')2 fragments, scFv, and Fab expression libraries. The antibody can be selected from the group consisting of, for example, single-chain antibodies, single-chain variable fragments (scFvs), variable fragments (Fvs), fragment antigen-binding regions (Fabs), recombinant antibodies, monoclonal antibodies, fusion proteins or aptamers containing the antigen-binding domain of a natural antibody, single-domain antibodies (sdAbs), also known as VHH antibodies, nanobodies (single-domain antibodies derived from camelids), single-domain antibody fragments derived from shark IgNARs called VNARs, diabodies, triabodies, anticalins, aptamers (DNA or RNA), and active components or fragments thereof.
[0349] The "Fab fragment" (also called fragment antigen binding, or Fab region) contains the light chain constant domain (CL) and the first heavy chain constant domain (CH1), along with the light and heavy chain variable domains VL and VH, respectively. The variable domains contain the complementarity determining loops (CDRs, also called hypervariable regions) involved in antigen binding. Fab' fragments differ from Fab fragments by the addition of a few residues at the carboxy terminus of the heavy chain CHI domain, including one or more cysteines from the antibody hinge region.
[0350] A "single-chain Fv" or "scFv" comprises the VH and VL domains of an antibody, wherein these domains are present in a single polypeptide chain. In one embodiment, the Fv polypeptide further comprises a polypeptide linker between the VH and VL domains, which enables the scFv to form the desired structure for antigen binding. For a review of scFvs, see Pluckthun in The Pharmacology of Monoclonal Antibodies, vol. 113, Rosenburg and Moore eds., Springer-Verlag, New York, pp. 269-315 (1994). Exemplary scFv fragments include antibody scFv fragments described in WO 93 / 16185; U.S. Pat. No. 5,571,894; and U.S. Pat. No. 5,587,458.
[0351] The effector moiety may be the Fab region of a therapeutic antibody. For example, the effector moiety may be the Fab region of a monoclonal antibody such as muromomab, abciximab, rituximab, daclizumab, basiliximab, palivizumab, infliximab, trastuzumab, etanercept, gemtuzumab, alemtuzumab, ibritomomab, adalimumab, alefacept, omalizumab, tositumomab, efalizumab, cetuximab, bevacizumab, natalizumab, ranibizumab, panitumumab, eculizumab, or certolizumab.
[0352] The effector moiety can target any receptor associated with a pathological condition, e.g., a pathological condition described herein. For example, the effector moiety can target any receptor whose binding is associated with a clinical benefit, e.g., a hormone receptor.
[0353] In some embodiments, an effector moiety may have a target (e.g., a receptor) not previously known to be associated with a pathological condition, for example, in some cases targeting such a receptor has been found to exhibit therapeutic effects.
[0354] Therefore, those skilled in the art will understand that the protein complex provided herein can be used to simultaneously engage and thereby simultaneously contact two targets in a biological system.Targets can be, for example, derived from the same cell.For example, the protein complex provided herein can be used to bind to two different types of receptors on the surface of the same cell.
[0355] The protein complexes provided herein typically contain multiple first binding sites and multiple second binding sites on a multivalent protein scaffold, and thus can bind multiple first and second effector moieties. This is particularly beneficial because such "high valency" compounds can enable improved or previously unknown effector functions. It has previously been shown that multiple copies of a single effector moiety can lead to improved therapeutic responses when contacted with a biological system (e.g., Brune et al. (above); and Khairil Anuar et al., Nature Communications 10.1 (2019): 1-13). However, contacting multiple copies of multiple different effector moieties is a complex technical challenge, which is solved by the protein complexes provided herein.
[0356] In some embodiments, the effector function can arise only as a result of the interaction of the combination of effector moieties with a biological system in contact therewith. For example, a first effector moiety (e.g., an effector moiety attached to a first binding site) may be shown to have a therapeutic effect only in combination with a second effector moiety (e.g., an effector moiety attached to a second binding site) in situations where neither the first nor the second effector moiety alone exhibits therapeutic efficacy.
[0357] Those skilled in the art will appreciate that the platforms and methods identified herein allow for screening of new combinations of therapeutically useful effector moieties and identification of useful candidates.
[0358] Screening Platform Also provided herein is a screening platform. The screening platform includes a library, the library including a plurality of populations of protein complexes of the present invention, each of the populations of protein complexes including a different combination of a first effector moiety, a second effector moiety, and / or an oligomer core. Such a library is also provided herein.
[0359] For example, libraries can be used to screen for new combinations of effector moieties. Thus, libraries can contain multiple samples of different protein complexes. Each sample can be homogenous, i.e., each sample can contain only one type of protein complex. Each sample can also be different from each of the other samples. Thus, each sample can contain protein complexes that contain a different combination of first and second effector moieties compared to the combination of first and second effector moieties in the protein complexes of each of the other samples. For example, a library can contain about 1 or about 2 to about 1,000,000 samples, e.g., about 10 to about 100,000 samples, e.g., about 50 to about 50,000 samples, e.g., about 100 to about 10,000 samples, e.g., about 500 to about 1,000 samples. Each sample may contain a different type of protein complex, and the protein complex in each sample has a different combination of first effector moiety, second effector moiety, and oligomeric core compared to the protein complex in all other samples.
[0360] In some embodiments, the library may be a "1D" library. Thus, in some embodiments, all samples in the library may have the same or substantially the same oligomer core and first effector moiety, but may differ with respect to the second effector moiety. In other embodiments, all samples in the library may have the same or substantially the same oligomer core and second effector moiety, but may differ with respect to the first effector moiety. In some other embodiments, all samples in the library may have the same or substantially the same first and second effector moieties, but may differ with respect to the oligomer core. A polypeptide that is substantially the same as a given polypeptide (e.g., an oligomer core, or a first or second polypeptide binding site) may, for example, have at least 90% sequence identity with the given polypeptide, e.g., at least 95% sequence identity with the given polypeptide, e.g., at least 97%, 98%, 99%, 99.9%, or 99.99% sequence identity. A polypeptide (e.g., an oligomer core, or a first or second polypeptide binding site) that is substantially the same as a given polypeptide may differ from the given polypeptide by including, for example, one or more sequence additions, deletions, or insertions, or mutations described herein. A polypeptide (e.g., an oligomer core, or a first or second polypeptide binding site) that is substantially the same as a given polypeptide may differ from the given polypeptide, for example, with respect to post-translational modifications made to the polypeptide, such as its glycosylation or phosphorylation pattern.
[0361] In some embodiments, the library may be a "2D" library. Thus, in some embodiments, all samples in the library may have the same or substantially the same oligomer core, but may differ in terms of the combination of first and second effector moieties. In other embodiments, all samples in the library may have the same or substantially the same first effector moieties, but may differ in terms of the combination of oligomer core and second effector moiety. In some other embodiments, all samples in the library may have the same or substantially the same second effector moieties, but may differ in terms of the combination of oligomer core and first effector moiety.
[0362] In some embodiments, the library may be a "3D" library. Thus, in some embodiments, all samples in the library may differ with respect to the combination of oligomer core, first effector moiety, and second effector moiety.
[0363] The screening platform may also include other constituent moieties in addition to the library. For example, the screening platform may include any or all of the following: - a biological system for contacting the samples in the library; - a detection system for detecting changes in the biological system resulting from contact of the biological system with the samples in the library; - reagents and / or buffers, and - Optical, electrical, or spectroscopic means for detecting the change reported by the detection system.
[0364] The biological system may be a cell culture, such as a mammalian cell culture, preferably a human cell culture, more preferably an immune cell culture and / or a cancer cell line culture. The biological system may be a biological sample, such as a blood sample, a serum sample, a plasma sample, or a tissue or organ sample. Biological samples include tumor samples, cells, cell lysates, urine, amniotic fluid, and other biological fluids. The biological sample is preferably mammalian. The sample may be human or non-human.
[0365] The detection system can be any suitable detection system.The detection system can be a dye or stain, such as a cell viability stain.Suitable stains can include, for example, trypan blue, (fluorescein diacetate) green, propidium iodide, and Hoechst 33258.
[0366] Reagents include components necessary for cell survival, including cell growth medium components, and may include therapeutic molecules.
[0367] Buffers include, for example, aqueous compositions that may contain buffer salts.Preferred buffer salts that can be used include Tris; phosphate; citric acid / NaHPO; citric acid / sodium citrate; sodium acetate / acetic acid; NaHPO / NaHPO; imidazole (glyoxalin) / HCl; sodium carbonate / sodium bicarbonate; ammonium carbonate / ammonium bicarbonate; MES; Bis-Tris; ADA; aces; PIPES; MOPSO; Bis-Tris propane; BES; MOPS; TES; HEPES; DIPSO; MOBS; TAPSO; Trizma; HEPPSO; POPSO; TEA; EPPS; Tricine; Gly-Gly; Bicine; HEPBS; TAPS; AMPD; TABS; AMPSO; CHES; CAPSO; AMP; CAPS; and CABS. Selection of an appropriate buffer for a desired pH is routine for one of skill in the art, and guidelines are available, for example, at http: / / www.sigmaaldrich.com / life-science / core-bioreagents / biological-buffers / learning-center / buffer-reference-center.html. Buffer salts are preferably used at concentrations of 1 mM to 1 M in solution, preferably 10 mM to 100 mM, e.g., about 50 mM.
[0368] Means for detecting the changes reported by the detection system include electrical means such as microscopy (optical or electronic), electrophysiological (e.g., patch clamp) devices; and spectroscopic means such as instruments for UV / VIS spectroscopy, NMR spectroscopy, mass spectroscopy, IR spectroscopy, Raman spectroscopy, circular dichroism spectroscopy, etc.
[0369] method Also provided is a method for identifying a therapeutic drug analog, comprising: Providing a protein complex as described herein; contacting the protein complex with a biological system; and Determining whether a protein complex induces a desired change in a property of a biological system A method is provided which includes:
[0370] The method optionally further comprises selecting a protein complex that induces a desired change in a property of the biological system.
[0371] Also provided are methods for identifying therapeutic combinations of effector molecules (e.g., antigen-binding domains), comprising: Providing a protein complex as described herein; contacting the protein complex with a biological system; and Determining whether a protein complex induces a desired change in a property of a biological system A method is provided which includes:
[0372] The biological system may be a cell culture, such as a mammalian cell culture, preferably a human cell culture, more preferably an immune cell culture and / or a cancer cell line culture.
[0373] The biological system may be a biological sample, such as a blood sample, a serum sample, a plasma sample, or a tissue or organ sample. Biological samples include tumor samples, cells, cell lysates, urine, amniotic fluid, and other biological fluids. The biological sample is preferably mammalian. The sample may be human or non-human.
[0374] The change in the property of the biological system can be any change associated with the desired activity of the intended therapeutic agent. In some embodiments, the desired change is cell death. This can be particularly useful when developing cancer therapeutic agents.
[0375] Other changes include changes in effector function, and thus the method may include determining whether the protein complex induces an effector function in a biological system.
[0376] The change in effector function can include altered gene expression, protein modification, such as altered phosphorylation, receptor internalization, cytokine release, cell death, sensitivity to therapeutic molecules, etc. The effector function can be high affinity binding to a biological sample, which can be measured by a series of techniques, such as ELISA. High affinity binding to a target biological system can enable the effector domain discussed above to specifically affect the target biological system, which may be a specific cell type, such as a cancer cell.
[0377] Effector function can be assessed relative to a control, which can be a protein complex lacking the effector moiety.
[0378] The control can be a protein complex with only a single type of effector moiety attached (i.e., only one type of effector moiety is attached to the multivalent protein scaffold). In this case, the method can be used to identify effector moieties that have "synergistic function" or "synergistic biological function," which refers to an effector function or level of effector function not observed with the individual fusion protein components until use of a bispecific multivalent protein complex, or a higher or lower activity compared to the activity observed when the first and second effector moieties of the protein complex are used individually, i.e., an activity that is observed only when both effector moieties are used together in the complex.
[0379] The method may further comprise identifying a molecule in the biological system to which the effector moieties of the protein complex bind. The method may preferably comprise selecting as the selected protein complex a combination of effector moieties that specifically bind to the same molecule in the biological system, such as the effector moieties themselves.
[0380] The method may further include synthesizing a therapeutic drug candidate or drug comprising a selected combination of effector moieties or their analogs. The therapeutic drug candidate or drug may comprise an oligomeric core and the effector moiety of a therapeutic drug analog, but the binding site and targeting functionality are replaced with a covalent linkage, such as a gene fusion, as described in more detail herein. The therapeutic drug or drug candidate may comprise the same oligomeric core as the therapeutic drug analog identified by the methods of the present disclosure. Alternatively, the therapeutic drug candidate may comprise a different oligomeric core from the therapeutic drug analog identified by the methods of the present disclosure. The therapeutic drug candidate may have an oligomeric core selected or designed to confer additional therapeutic benefits, for example, additional effector functions.
[0381] Also provided are therapeutic drug candidates obtainable according to the methods of the present disclosure.
[0382] Therapeutic drug candidates, therapeutic drugs Further provided is a therapeutic drug candidate comprising an oligomer core comprising a plurality of subunit monomers attached to one or more first effector moieties and one or more second effector moieties, wherein the one or more first effector moieties and the one or more second effector moieties are located on the same face of the oligomer core, and wherein (i) the one or more first effector moieties comprise two or more first effector moieties and the one or more second effector moieties comprise two or more second effector moieties; and / or (ii) the oligomer core does not comprise an antibody or antibody fragment.
[0383] Therapeutic drugs having the same characteristics are also provided.
[0384] Typically, the oligomer core is an oligomer core as described in more detail herein. Typically, the subunit monomers are as described in more detail herein. Typically, the first effector moiety and the second effector moiety are as described in more detail herein. The first and second effector moieties can be attached to the subunit monomers of the oligomer core in any suitable manner, including by any of the attachment means described herein. In some embodiments, the attachment of the first and second effector moieties comprises the first and second binding moieties and the first and second polypeptide targets described herein. However, in other embodiments, the attachment of the first and second effector moieties does not comprise the first and second binding moieties and the first and second polypeptide targets described herein, but instead may comprise simple covalent attachment, such as gene fusion and / or click chemistry linkage, as described herein.
[0385] Specific Embodiments In a first preferred aspect, the following is provided herein: - a multivalent protein scaffold comprising an oligomeric core comprising a plurality (e.g., 3-6, preferably 3) of monomers each comprising an amino acid sequence having at least 30% or at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 1, wherein each monomer comprises a first binding site and a second binding site, and the first binding site is linked to the second binding site. and a multivalent protein scaffold, wherein the first binding site and the second binding site each independently have at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) to any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18, and each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first and second binding sites has at least 50% amino acid identity to SEQ ID NO: 4, 6, or 8, and the other has at least 50% amino acid identity to SEQ ID NO: 12. A preferred multivalent protein scaffold of this embodiment comprises a monomer of SEQ ID NO: 21, or a fragment thereof (e.g., comprising residues 14-348 thereof). a protein complex comprising the multivalent protein scaffold of the first aspect, wherein a first binding site binds to a first polypeptide target attached to a first effector moiety and a second binding site binds to a first polypeptide target attached to a second effector moiety, wherein the first and second effector moieties may be the same or different, preferably different, and wherein the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are each independently selected from: (i) a combination of any one of SEQ ID NOs: 4, 6 or 8 with any one of SEQ ID NOs: 5, 7 or 9; (ii) a combination of SEQ ID NO: 12 with SEQ ID NO: 13 or 15; (iii) a combination of SEQ ID NO: 5 with SEQ ID NO: 11; (iv) a combination of SEQ ID NO: 15 with SEQ ID NO: 16; (v) a combination of SEQ ID NO: 17 with SEQ ID NO: 18; or (vi) a combination of SEQ ID NO: 23 with SEQ ID NO: 16. - a screening platform comprising a library comprising multiple populations of protein complexes of the first aspect, each population comprising a different combination of first and second effector moieties. a therapeutic drug candidate comprising an oligomer core comprising a plurality (e.g., 3 to 6, preferably 3) monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 1, wherein each monomer is directly attached (e.g., as a genetic fusion) to a first effector moiety and a second effector moiety, which may be the same or different, preferably different, and preferably each monomer is directly attached via a polypeptide linker to the first and second effector moieties attached thereto.
[0386] In a second preferred embodiment, inter alia, the following is provided herein: - a multivalent protein scaffold comprising an oligomeric core comprising a plurality (e.g., 3-6, preferably 3) of monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO:2, wherein each monomer comprises a first binding site and a second binding site, and the first binding site is orthogonal to the second binding site. and wherein the first binding site and the second binding site each independently have at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) to any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18, and each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first and second binding sites has at least 50% amino acid identity to SEQ ID NO: 4, 6, or 8, and the other has at least 50% amino acid identity to SEQ ID NO: 12. A preferred multivalent protein scaffold of this embodiment comprises a monomer of SEQ ID NO: 20, or a fragment thereof (e.g., comprising residues 14-380 thereof). a protein complex comprising the multivalent protein scaffold of the second aspect, wherein a first binding site binds to a first polypeptide target attached to a first effector moiety and a second binding site binds to a first polypeptide target attached to a second effector moiety, wherein the first and second effector moieties may be the same or different, preferably different, and wherein the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are each independently selected from: (i) a combination of any one of SEQ ID NOs: 4, 6, or 8 with any one of SEQ ID NOs: 5, 7, or 9; (ii) a combination of SEQ ID NO: 12 with SEQ ID NO: 13 or 15; (iii) a combination of SEQ ID NO: 5 with SEQ ID NO: 11; (iv) a combination of SEQ ID NO: 15 with SEQ ID NO: 16; (v) a combination of SEQ ID NO: 17 with SEQ ID NO: 18; or (vi) a combination of SEQ ID NO: 23 with SEQ ID NO: 16. - a screening platform comprising a library comprising multiple populations of protein complexes of the second aspect, each population comprising a different combination of first and second effector moieties. a therapeutic drug candidate comprising an oligomer core comprising a plurality (e.g., 3-6, preferably 3) monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO:2, wherein each monomer is directly attached (e.g., as a genetic fusion) to a first effector moiety and a second effector moiety, which may be the same or different, preferably different, and preferably each monomer is directly attached via a polypeptide linker to the first and second effector moieties attached thereto.
[0387] In a third preferred aspect, inter alia, the following is provided herein: - a multivalent protein scaffold comprising an oligomeric core comprising a plurality (e.g., 3-6, preferably 3) of monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 3, wherein each monomer comprises a first binding site and a second binding site, and the first binding site is orthogonal to the second binding site. and wherein the first binding site and the second binding site each independently have at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) to any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18, and each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first and second binding sites has at least 50% amino acid identity to SEQ ID NO: 4, 6, or 8, and the other has at least 50% amino acid identity to SEQ ID NO: 12. a protein complex comprising the multivalent protein scaffold of the third aspect, wherein a first binding site binds to a first polypeptide target attached to a first effector moiety and a second binding site binds to a first polypeptide target attached to a second effector moiety, wherein the first and second effector moieties may be the same or different, preferably different, and wherein the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are each independently selected from: (i) a combination of any one of SEQ ID NOs: 4, 6, or 8 with any one of SEQ ID NOs: 5, 7, or 9; (ii) a combination of SEQ ID NO: 12 with SEQ ID NO: 13 or 15; (iii) a combination of SEQ ID NO: 5 with SEQ ID NO: 11; (iv) a combination of SEQ ID NO: 15 with SEQ ID NO: 16; (v) a combination of SEQ ID NO: 17 with SEQ ID NO: 18; or (vi) a combination of SEQ ID NO: 23 with SEQ ID NO: 16. - a screening platform comprising a library comprising multiple populations of protein complexes of the third aspect, each population comprising a different combination of first and second effector moieties. - a therapeutic drug candidate comprising an oligomer core comprising a plurality (e.g., 3 to 6, preferably 3) monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 3, wherein each monomer is directly attached (e.g., as a genetic fusion) to a first effector moiety and a second effector moiety, which may be the same or different, preferably different, and preferably each monomer is directly attached via a polypeptide linker to the first and second effector moieties attached thereto.
[0388] In a fourth preferred aspect, inter alia, the following is provided herein: - a multivalent protein scaffold comprising an oligomeric core comprising a plurality (e.g., 3-6, preferably 3) of monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 19, wherein each monomer comprises a first binding site and a second binding site, and the first binding site is orthogonal to the second binding site. and wherein the first binding site and the second binding site each independently have at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) to any one of SEQ ID NOs: 4-9, 11-13, 23, or 15-18, and each first binding site and each second binding site are independently genetically fused to the monomer to which they are attached. Preferably, one of the first and second binding sites has at least 50% amino acid identity to SEQ ID NO: 4, 6, or 8, and the other has at least 50% amino acid identity to SEQ ID NO: 12. - a protein complex comprising the multivalent protein scaffold of the fourth aspect, wherein a first binding site binds to a first polypeptide target attached to a first effector moiety and a second binding site binds to a first polypeptide target attached to a second effector moiety, wherein the first and second effector moieties may be the same or different, preferably different, and wherein the first binding site / polypeptide target pair and the second binding site / polypeptide target pair are each independently selected from: (i) a combination of any one of SEQ ID NOs: 4, 6 or 8 with any one of SEQ ID NOs: 5, 7 or 9; (ii) a combination of SEQ ID NO: 12 with SEQ ID NO: 13 or 15; (iii) a combination of SEQ ID NO: 5 with SEQ ID NO: 11; (iv) a combination of SEQ ID NO: 15 with SEQ ID NO: 16; (v) a combination of SEQ ID NO: 17 with SEQ ID NO: 18; or (vi) a combination of SEQ ID NO: 23 with SEQ ID NO: 16. - a screening platform comprising a library comprising multiple populations of protein complexes of the fourth aspect, each population comprising a different combination of first and second effector moieties. - a therapeutic drug candidate comprising an oligomer core comprising a plurality (e.g., 3 to 6, preferably 3) monomers each comprising an amino acid sequence having at least 50% amino acid identity (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid identity) with the amino acid sequence of SEQ ID NO: 4, wherein each monomer is directly attached (e.g., as a genetic fusion) to a first effector moiety and a second effector moiety, which may be the same or different, preferably different, and preferably each monomer is directly attached via a polypeptide linker to the first and second effector moieties attached thereto.
[0389] In a fifth preferred embodiment, inter alia, the following is provided herein: - a polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, the first and second binding domains being separated by a structural domain, and the first antigen-binding domain and the second antigen-binding domain being capable of binding to their targets when the target molecules are expressed on a single cell or immobilized on a plate or single bead. - an oligomer of polypeptides, wherein each polypeptide of the oligomer comprises or consists of a polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, the first and second binding domains being separated by a structural domain, and the first antigen-binding domain and the second antigen-binding domain being capable of binding to their targets when the target molecules are expressed on a single cell or immobilized on a plate or single bead. A polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, the first and second binding domains being separated by a structural domain, and the first and second antigen-binding domains being capable of binding to targets when the target molecules are expressed on a single cell or immobilized on a plate or single bead. The first and second binding domains are each a catcher domain capable of forming an isopeptide bond with a cognate peptide. Such cognate peptides are often referred to as tag peptides; for example, SpyTag forms an isopeptide bond with a SpyCatcher domain, as is known in the art and discussed below. The cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain. - An oligomer of polypeptides, each polypeptide of the oligomer comprising or consisting of a polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, the first and second binding domains being separated by a structural domain, and the first and second antigen-binding domains being capable of binding to targets when the target molecules are expressed on a single cell or immobilized on a plate or single bead. The first and second binding domains are each a catcher domain capable of forming an isopeptide linkage with a cognate peptide. Such cognate peptides are often referred to as tag peptides; for example, SpyTag forms an isopeptide bond with a SpyCatcher domain, as is known in the art and discussed below. The cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain.
[0390] Additional Aspects of the Disclosure Also provided herein are polynucleotides that encode at least one monomer of the oligomeric core of the multivalent protein scaffold described in more detail herein. Also provided herein are polynucleotides that encode the multi-domain polypeptide constructs described in more detail herein, including a first binding domain, a second binding domain, and a structural domain.
[0391] Further provided are vectors comprising the polynucleotides; cells comprising the vectors; and methods for producing the monomers, the oligomer cores, and / or the multivalent protein scaffolds, the methods comprising culturing the cells in a culture medium to produce the protein scaffolds.
[0392] Further provided are vectors comprising the polynucleotides; cells comprising the vectors; and methods for producing a multi-domain polypeptide construct, the method comprising culturing the cells in a culture medium to produce the multi-domain polypeptide.
[0393] Selection of an appropriate polynucleotide sequence encoding at least one monomer; an appropriate expression vector; and an appropriate cell for expressing the monomer, oligomeric core, and / or multivalent protein scaffold is routine for one of ordinary skill in the art.
[0394] Therapeutic Efficacy The protein complexes, therapeutic drug analogs, and therapeutic drug candidates provided herein are useful for therapeutic purposes. Multi-domain polypeptide constructs are typically useful for therapeutic purposes. Such provided substances are also referred to herein as "therapeutic protein complexes."
[0395] Thus, the present invention provides therapeutic protein conjugates and constructs as described herein for use in medicine. The present invention provides therapeutic protein conjugates as described herein for use in the treatment of the human or animal body. The present invention provides therapeutic protein constructs as described herein for use in the treatment of the human or animal body.
[0396] The present invention provides methods for treating a human or animal in need of such treatment, comprising administering to the human or animal in need thereof a protein complex, multi-domain polypeptide construct (in monomeric or oligomeric form), therapeutic drug analog, therapeutic drug candidate, or drug described herein.
[0397] Also provided are pharmaceutical compositions comprising one or more therapeutic protein conjugates described herein together with a pharmaceutically acceptable carrier or diluent. Typically, the compositions comprise up to 85% by weight of a therapeutic protein conjugate of the invention. More typically, the compositions comprise up to 50% by weight of a therapeutic protein conjugate of the invention. Preferred pharmaceutical compositions are sterile and pyrogen-free.
[0398] Also provided are pharmaceutical compositions comprising one or more multi-domain polypeptide constructs described herein together with a pharmaceutically acceptable carrier or diluent. Typically, the compositions comprise up to 85% by weight of a therapeutic multi-domain polypeptide construct of the invention. More typically, the compositions comprise up to 50% by weight of a therapeutic multi-domain polypeptide construct of the invention. Preferred pharmaceutical compositions are sterile and pyrogen-free.
[0399] The compositions of the invention may be provided as kits that include instructions to enable use of the kit in the methods described herein or details as to which subjects the methods may be used for.
[0400] As explained above, the therapeutic protein complexes and constructs provided herein are useful for treating or preventing a variety of disorders. Disorders that may be treated using the provided therapeutic protein complexes include cancer, autoimmune diseases (e.g., ankolysing spondylitis), psoriasis, eye diseases such as age-related macular degeneration, multiple sclerosis, cardiovascular disorders, infectious diseases including viral and bacterial infections, Crohn's disease, rheumatoid arthritis, osteoarthritis, Alzheimer's disease, transplant and allograft rejection, and hematopoietic stem cell disorders. More broadly, the therapeutic protein complexes provided herein also find utility in treating any and all conditions that are treated using antibodies, particularly bispecific antibodies.
[0401] Cancers, e.g., acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related lymphoma, primary CNS lymphoma, anal cancer, astrocytoma, brain cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer (e.g., Ewing's sarcoma, osteosarcoma, and malignant fibrous histiocytoma), breast cancer, bronchial tumors, medulloblastoma and other CNS embryonal tumors, cervical cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative neoplasms, colorectal cancer, craniopharyngioma , endometrial cancer, ependymoma, esophageal cancer, esthesioneuroblastoma, Ewing's sarcoma, extragonadal germ cell tumor, intraocular melanoma, retinoblastoma, fallopian tube cancer, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (gist), germ cell tumor, extragonadal germ cell tumor, ovarian germ cell tumor, testicular cancer, gestational trophoblastic disease, hairy cell leukemia, hepatocellular carcinoma, histiocytosis, Langerhans cell, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, pancreatic neuroendocrine tumor Urinary tract tumors, Kaposi's sarcoma, kidney (renal cell) cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer (non-small cell, small cell, pleuropulmonary blastoma, and tracheobronchial tumors), lymphoma, malignant fibrous histiocytoma and osteosarcoma of bone, Merkel cell carcinoma, mesothelioma, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma / plasma cell neoplasm, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative neoplasm, myeloid leukemia, neuroblastoma, non-Hodgkin's lymphoma Cancers such as lymphoma, oropharyngeal cancer, osteosarcoma and undifferentiated pleomorphic sarcoma, pancreatic cancer, pancreatic neuroendocrine tumors (islet cell tumors), papillomatosis, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pituitary tumors, prostate cancer, rectal cancer, retinoblastoma, rhabdomyosarcoma, T-cell lymphoma, testicular cancer, pharyngeal cancer, thymoma and thymic carcinoma, thyroid cancer, and tracheobronchial tumors are particularly suitable for treatment with the therapeutic protein conjugates provided herein. Autoimmune disorders amenable to treatment with the therapeutic protein conjugates provided herein include rheumatoid arthritis, systemic lupus erythematosus (lupus), inflammatory bowel disease (IBD), multiple sclerosis (MS), type 1 diabetes, Guillain-Barré syndrome, chronic inflammatory demyelinating polyneuropathy, psoriasis, Graves' disease, Hashimoto's thyroiditis, myasthenia gravis, and vasculitis.
[0402] The therapeutic protein complexes provided herein can be used as stand-alone therapeutic agents. Alternatively, the therapeutic protein complexes provided herein can be used in combination with other active agents, such as chemotherapeutic agents. For example, the therapeutic protein complexes provided herein can be used in combination with EGFR inhibitors (e.g., erlotinib, gefitinib, lapatinib, or cetuximab), immunotherapy (e.g., pembrolizumab or nivolumab), tumor-agnostic therapy (e.g., larotrectinib), or chemotherapy (e.g., 5-fluorouracil, cisplatin, or docetaxel).
[0403] When used in the treatment of cancer, the therapeutic protein conjugates provided herein can be used to alleviate, improve, or prevent the worsening of cancer symptoms. Typically, cancer treatment may involve reducing the progression of cancer, e.g., increasing progression-free survival. Cancer treatment may involve preventing or inhibiting the growth of cancer-associated tumors. Cancer treatment may involve preventing cancer metastasis. Preferably, cancer treatment may involve reducing the size of cancer-associated tumors. Thus, treatment can cause cancer tumor regression. Cancer treatment may involve reducing the number of tumors or lesions present in a patient. When treatment reduces the size of cancer-associated tumors, tumor size is typically reduced by at least 10% from baseline. The baseline is the tumor size on the day treatment with the compound first begins. Tumor size is typically measured according to RECIST version 1.1 (e.g., as described in Eisenhauer et al., European Journal of Cancer 45 (2009) 228-247).
[0404] The response to the treatment with the compound can be a complete response, a partial response, or stable disease according to RECIST version 1.1.Preferably, the response is a partial response or a complete response.The treatment can achieve a progression-free survival period of at least 60 days, at least 120 days, or at least 180 days.
[0405] The reduction in tumor size may be greater than 20%, greater than 30%, or greater than 50% compared to baseline. The reduction in tumor size may be observed after 30 days of treatment or after 60 days of treatment.
[0406] The therapeutic protein conjugates provided herein may also be useful for treating infectious diseases, such as those caused by gram-positive and / or gram-negative bacteria; and viral infections. The therapeutic protein conjugates provided herein may be designed to interact with pathogens such as bacteria, fungi, and viruses.
[0407] As described herein, the therapeutic protein complexes provided herein are useful for the treatment or prevention of various disorders. Accordingly, the present invention provides a therapeutic protein complex as described herein for use in a medicament. The present invention also provides the use of a therapeutic protein complex as provided herein in the manufacture of a medicament. The present invention also provides compositions and products comprising a therapeutic protein complex as provided herein. Such compositions and products are also useful for the treatment or prevention of disorders. Accordingly, the present invention provides a composition or product as defined herein for use in a medicament. The present invention also provides the use of a composition or product of the present invention in the manufacture of a medicament. Also provided is a method of treating a subject in need of such treatment, comprising administering a therapeutic protein complex as provided herein to the subject. In some embodiments, the subject is suffering from or at risk of suffering from one of the disorders disclosed herein.
[0408] In one embodiment, the subject is mammal, particularly human.However, the subject can also be non-human.Preferred non-human animals include but are not limited to primates such as marmosets or monkeys, commercially farmed animals such as horses, cows, sheep or pigs, and pets such as dogs, cats, mice, rats, guinea pigs, ferrets, gerbils or hamsters.The subject can be any animal that can be infected with bacteria.
[0409] The subject is typically a human patient. The patient may be male or female. The patient is typically at least 18 years old, for example, 30 to 70 years old, or 40 to 60 years old. The subject may also be a child or adolescent, for example, 6 months to 11 years old, or 12 to 17 years old.
[0410] The therapeutic protein complex, polypeptide construct, or composition of the present invention can be administered to a subject to prevent the onset or recurrence of one or more symptoms of a disorder. This is prophylaxis. In this embodiment, the subject may be asymptomatic. A prophylactically effective amount of an agent or preparation is administered to such a subject. A prophylactically effective amount is an amount that prevents the onset of one or more symptoms of a disorder.
[0411] The therapeutic protein complex, polypeptide construct or composition of the present invention ...
Claims
1. A polypeptide comprising a first binding domain at its N-terminus and a second binding domain at its C-terminus, wherein the first and second binding domains are separated by a CutA1 structural domain that contains one or more substitutions or deletions relative to wild-type CutA1, and the first and second binding domains are the same or different.
2. The polypeptide of claim 1 or 2, wherein the CutA1 structural domain is human CutA1.
3. 3. The polypeptide of claim 1 or 2, wherein the substitution is of one or more cysteine residues, or the deletion is of one or more residues at the N-terminus and / or the C-terminus.
4. The polypeptide of any one of claims 1 to 3, wherein one or more cysteine residues of CutA1 are substituted with one or more alanine, valine, or serine residues.
5. 5. The polypeptide of claim 1, wherein the substitution comprises or consists of replacing two cysteines with two alanines.
6. 5. The polypeptide of claim 1, wherein the substitutions comprise or consist of one cysteine for a valine and one cysteine for a serine.
7. The polypeptide according to any one of claims 1 to 6, wherein the cysteine residues at positions 75 and 96 of wild-type human CutA1 (SEQ ID NO: 19) are substituted with different residues.
8. 8. The polypeptide of claim 7, wherein the cysteine substitution comprises or consists of (i) C75A, C96A or (ii) C75V, C96S.
9. 10. The polypeptide of any preceding claim, wherein the CutA1 structural domain comprises a cysteine residue at a position that is not a cysteine residue in the wild-type sequence.
10. The following residues of wild-type human CutA1 (SEQ ID NO: 19): V64, E78, K79, K82, E83, K91, Q102, K110, E114, F136, S139, F158, and Q166 10. The polypeptide of claim 9, wherein one or more of the following are substituted with cysteine:
11. 10. The polypeptide of any preceding claim, wherein 5 to 60 residues, and optionally 10 to 59 residues, have been deleted from the N-terminus of CutA1.
12. 3. The polypeptide of any preceding claim, wherein 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 residues have been deleted from the N-terminus of CutA1.
13. The polypeptide of any preceding claim, wherein CutA1 is a truncated human CutA1 starting at residue 33 of SEQ ID NO:19, or starting at residue 44 of SEQ ID NO:19, or starting at residue 60 of SEQ ID NO:19, or starting at residue 24 of SEQ ID NO:
19.
14. 10. The polypeptide of any preceding claim, wherein 5 to 20 residues, and optionally 6 to 12 residues, have been deleted from the C-terminus of CutA1.
15. 10. The polypeptide of any preceding claim, wherein the CutA1 is a truncated CutA1 beginning at any of residues 30-67 and ending at any of residues 165-179 of SEQ ID NO:
19.
16. 10. The polypeptide of any preceding claim, wherein the CutA1 is a truncated CutA1 consisting of residues 44 to 179 or residues 60 to 171 of SEQ ID NO:
19.
17. the first binding domain and the second binding domain are different antigen-binding domains, and optionally one or both antigen-binding domains are antigen-binding fragments of an antibody, optionally an scFv or Fab, or a single domain antibody (sdAb), or other antibody mimetic or scaffold selected to bind to a particular target, or other protein or peptide capable of specifically binding to a biological molecule; and / or one or both of the first binding domain and the second binding domain are agonists of a TNF receptor superfamily member, and optionally the TNF receptor superfamily member is a TRAIL receptor, e.g., death receptor 5 (DR5); A polypeptide according to any preceding claim.
18. the first binding domain and the second binding domain are each a catcher domain capable of forming an isopeptide linkage with a cognate peptide; Optionally, the cognate peptide of the first binding domain is different from the cognate peptide of the second binding domain; or 17. The polypeptide of any of claims 1 to 16, wherein optionally the cognate peptide of the first binding domain is the same as that of the second binding domain, and optionally selective binding or conjugation is achieved by temporal or sequential control, such as activation or inactivation of the second binding domain, and / or by competitive binding.
19. 20. The polypeptide of claim 18, wherein each cognate peptide is attached to an antigen-binding domain, and optionally one or both cognate peptides are linked to the first and / or second catcher domain by an isopeptide bond.
20. 19. The polypeptide of claim 18, wherein one or both of the cognate peptides are linked to an agonist of a TNF receptor superfamily member, and optionally the TNF receptor superfamily member is a TRAIL receptor, e.g., death receptor 5 (DR5).
21. 21. The polypeptide of any one of claims 1 to 20, wherein the first binding domain and the second binding domain are capable of binding to their targets when expressed on a single cell or immobilized on a plate or a single bead.
22. an effector molecule, optionally a drug molecule or a dye molecule, wherein optionally the drug molecule comprises: (a) an anti-cancer drug, optionally a cytotoxic drug; (b) a tubulin inhibitor, optionally a maytansinoid, an auristatin, or a taxol derivative; (c) monomethylauristatin E (MMAE) or monomethylauristatin F (MMAF), (d) compounds derived from dolastatin 10; (e) tubulysins, such as tubulysin A; (f) a DNA damaging agent; (g) duocarmycins, (h) calicheamicin, (i) pyrrolobenzodiazepines, (j) SN-38 (the active metabolite of irinotecan); (k) an immunomodulatory agent, such as a TLR agonist or a STING agonist; (l) deruxtecan, (m) mertansine, or (n) an immunotoxin, optionally Pseudomonas exotoxin A (PE) 10. A polypeptide according to any preceding claim, comprising or consisting of:
23. The polypeptide of claim 22, wherein the drug molecule or the dye molecule is chemically conjugated to the polypeptide, optionally in the CutA1 domain, optionally to a cysteine residue in the CutA1 structural domain, and optionally the conjugated cysteine residue is not present in the native CutA1 sequence.
24. An oligomer comprising two or more polypeptides according to any preceding claim.
25. A polypeptide according to any of claims 1 to 23 or an oligomer according to claim 24, comprising the features according to any of embodiments A1 to A26.
26. a first binding domain N-terminal to a CutA1 structural domain or a cytokine structural domain and / or a second binding domain C-terminal to the CutA1 structural domain or cytokine structural domain, said first and said second binding domains being separated by said CutA1 structural domain or cytokine structural domain, if both are present, optionally said CutA1 structural domain or said cytokine structural domain comprising one or more substitutions or deletions relative to the respective wild-type CutA1 or cytokine, said first and said second binding domains being the same or different; or A polypeptide comprising a first binding domain connected to a second binding domain, wherein the second binding domain is connected to a CutA1 structural domain or a cytokine structural domain, optionally wherein the CutA1 structural domain or the cytokine structural domain comprises one or more substitutions or deletions relative to the respective wild-type CutA1 or cytokine, and wherein the first and second binding domains are the same or different.
27. 27. The polypeptide of claim 26, wherein the cytokine is TNF, TL1A, OX40L, or CD40L, SEQ ID NO: 80 or a modified version thereof, SEQ ID NO: 31 or a modified version thereof, SEQ ID NO: 78 or a modified version thereof, or SEQ ID NO: 58 or a modified version thereof.
28. A CutA1 protein comprising one or more substitutions or deletions relative to wild-type CutA1, optionally with the N-terminus and / or C-terminus of the CutA1 protein connected to a different polypeptide by peptide linkage as a fusion protein.
29. 29. The CutA1 protein of claim 28, wherein at least one naturally occurring cysteine residue has been substituted with a different amino acid residue and / or at least one naturally occurring residue that is not cysteine has been substituted with a cysteine residue.
30. A single polypeptide chain comprising multiple CutA1 sequences, optionally interposed between some or all of the CutA1 sequences by a linker polypeptide of 3 to 30 amino acids in length, and optionally wherein at least one and optionally all of the CutA1 sequences are as defined in any of claims 1 to 16.
31. A single polypeptide as described in claim 30, each comprising three human CutA1 sequences separated by a linker polypeptide, optionally comprising an isopeptide bond-forming domain at the N-terminus and / or C-terminus.