PRO-Macrobody for the Promotion of Structural Research

The Pro-macrobody, a chimeric fusion polypeptide with a di-proline linker, addresses the challenges of small membrane proteins by enhancing structural rigidity and resolution in cryo-EM and X-ray crystallography, enabling high-resolution structural analysis and drug design.

JP7839094B2Active Publication Date: 2026-04-01リーデクスプロ アーゲー
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-02-16
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

The structural determination of small membrane proteins, particularly those with molecular weights less than 100 kDa, is challenging due to their conformational flexibility, surfactant requirements, and difficulty in forming crystal contacts, which limits high-resolution structure determination by methods like X-ray crystallography and cryo-electron microscopy.

Method used

A chimeric fusion polypeptide, referred to as Pro-macrobody, is developed, comprising a VHH antigen-binding domain linked to a maltose-binding protein (MBP) scaffold via a di-proline linker, enhancing structural rigidity and facilitating high-resolution structural analysis in cryo-EM and X-ray crystallography by increasing molecular weight, improving particle contrast, and optimizing orientation classification.

Benefits of technology

The Pro-macrobody enhances the resolution of membrane protein structures by increasing molecular weight, improving particle classification, and reducing denaturation, thereby facilitating structure-based drug design and diagnostic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839094000007
    Figure 0007839094000007
  • Figure 0007839094000008
    Figure 0007839094000008
  • Figure 0007839094000009
    Figure 0007839094000009
Patent Text Reader

Abstract

The present invention relates to research tools for structural biology, particularly for facilitating the determination of three-dimensional structures of biological macromolecules. More specifically, the present invention helps to improve the overall feasibility of structure determination, resulting in higher resolution and better quality of three-dimensional structures of proteins through complex formation with novel fusion polypeptides.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] The structural biology of proteins, especially membrane proteins, is a challenging task. Consequently, the determination of membrane protein structures lags behind that of soluble proteins. Therefore, technological advancements are needed not only for membrane proteins but also for the difficult and challenging targets of soluble proteins.

[0002] In recent years, advances in protein production and structural determination techniques have provided crucial and comprehensive insights into the function and pharmacology of membrane proteins. This progress improves structure-based drug design in membrane proteins, one of the primary drug targets. Nevertheless, structural determination of membrane proteins by X-ray crystallography or single-particle cryo-electron microscopy (cryoEM) remains a significant challenge, and technological advancements are needed to improve quality and cost-effectiveness and reduce the risks of structural information generation. Intrinsic properties such as conformational flexibility, surfactant requirements, and the minimum size of target proteins are sources of demand for technological improvements to increase the resolution of membrane protein structures. CryoEM is now a promising method for determining the three-dimensional structure of membrane proteins with near atomic resolution and has evolved into a major technique in structural biology over the past decade. This method is increasingly used in structure-based drug design and is expected to become a major pathway for structural determination. Furthermore, proteins that are difficult to crystallize, especially membrane proteins, are often investigated using cryoEM. Despite technological advancements, one of the most limiting factors for successful structural determination is the size (molecular weight) of the protein of interest, particularly for membrane proteins. To date, there are few examples of high-resolution structures (resolution less than 4 Å) of membrane proteins with molecular weights less than 100 kDa [Bai, X. et al., Nature 525, 212-217 (2015); Renaud, J.-P. et al., Nature Reviews Drug Discovery 17, 471-492 (2018)]. Proteins that fall into this size category include, for example, pharmacologically important G protein-coupled receptors (GPCRs) or members of solute carrier superfamilies (SLCs), and more general monomeric membrane proteins that lack inherent symmetry and have few extramembrane features (i.e., additional domains on the cytoplasmic or luminal side). One reason why determining the structure of membrane proteins is difficult is the presence of a belt of amorphous surfactants (or lipid disks and other specialized polymers) necessary to keep membrane proteins in solution for single-particle imaging.Since most of the mass of many membrane proteins is concentrated within this transmembrane region, most of the protein is hidden and generally does not provide sufficient features in the micrograph for sub-class averaging. This factor becomes increasingly important as the molecular weight of the target protein decreases. The surfactant belt provides contrast to localize the particles, but it is more difficult to assign the respective orientations of the particles in three-dimensional space within the amorphous envelope of the surfactant. Recent successes in obtaining the structures of small soluble proteins at resolutions below 3.5 Å, such as streptavidin (Fan, X. et al., Nature Communications 10, 2386 (2019)), emphasize that the surfactant belt is a major consideration for the structure determination of small membrane proteins.

[0003] The surfactant belt also reduces the possibility of crystal contact formation and requires relatively large voids within the crystal lattice accommodated in the crystal, thus having an adverse effect on the crystallization of membrane proteins. These problems have been partially overcome by methods such as lipid cubic phase crystallization (LCP), but vapor diffusion remains an important technique for high-resolution structure determination, and various methodologies have been developed to increase the success rate of crystallization and diffraction properties. Therefore, molecular chaperones have also been proven to be very efficient for X-ray crystallography in vapor diffusion as well as lipid cubic phase crystallization [Zhou, Y., Morais-Cabral, J. H., Kaufman, A., & MacKinnon, R. Nature 414, 43 - 48 (2001); Dutzler, R., Campbell, E. B., & MacKinnon, R. Science 300, 108 (2003); Brunner, J. D. et al., bioRxiv 480863 (2018) doi:10.1101 / 480863] to bring about crystal contacts and good lattice formation.

[0004] To overcome the above challenges, complex formation with the molecular chaperone of the target protein can be performed (in X-ray crystallography) to provide further possibilities for crystal contact formation, or (in the case of cryo-EM) to enlarge particles and obtain further information about the orientation of particles in space. The latter parameters are essential for correctly classifying particles, as misclassification negatively impacts the final resolution. To date, Fab (fragment antigen binding) of monoclonal antibodies has been commonly applied for this purpose [Taylor, NMI et al., Nature 546, 504-509 (2017); Butterwick, JA et al., Nature 560, 447-452 (2018)]. These are relatively large (50 kDa) domains of antibodies with different shapes, but also possess inherent variable (clone and Ig subclass dependent) flexibility between variable (VHH of heavy and light chains) and a first constant region. Fab, due to its size and distinct donut shape, is generally well-visible in EM microscopy and thus meets expectations for single-particle cryo-EM in terms of several key parameters (increased complex size, improved subclass averaging, and randomization of particle orientation in space). The main drawback of Fab for structural purposes is its time-consuming and expensive production. Further challenges include the relatively small number of clones, expensive cell culture methods, and hybridoma culture for Ig production. In addition, while IgG (and consequently Fab) tends to bind to sequential epitopes, such as the flexible ends of proteins, structural studies require a binder that recognizes three-dimensional discontinuous epitopes.

[0005] To overcome these shortcomings, molecular chaperones consisting of single-chain antibodies have been introduced into scientific research [Steyaert, J. & Kobilka, BK Current Opinion in Structural Biology 21, 567-572 (2011)]. The unique and compact structure of the single-chain antibody VHH domain (called a nanobody) offers several significant advantages compared to normal IgG. Firstly, since camelid or cartilaginous fish VHHs are constructed from only one polypeptide chain and not assembled from two gene products like classical antibodies, they can be easily cloned from small blood samples, thereby enabling rapid generation of libraries for in vitro screening and simple design of synthetic libraries. Secondly, this simpler structure (only one disulfide bond / VHH) allows for efficient microbial overexpression / production. Thirdly, the smaller paratope footprint of single-chain antibodies (three complementarity-determining loops (CDRs) compared to the six of classical IgG (2×3 CDRs derived from the heavy and light chains, respectively) makes it easier for single-chain antibodies to bind to the latent epitopes of cretions and clevises. Importantly, single-chain antibodies exhibit similar affinity to those of classical antibodies for their epitopes. This is mainly due to the relatively elongated CDR3, which often results in many interactions with the epitopes. The wedge-shaped structure and pointed paratopes of single-chain antibodies, as well as their generally small size of only 15 kDa, underlie the fact that many single-chain antibodies bind to functionally important regions of membrane proteins and selectively recognize different conformations. This reduces motility and conformational heterogeneity in the target protein and often functionally blocks the target protein [Schenck, S. et al., Biochemistry 56, 3962~3971 (2017)]. This particular characteristic is what makes single-chain antibodies so interesting for structurally characterizing proteins, especially membrane proteins.

[0006] A small proteinoid binding agent from a synthetic library, the designed ankyrin repeat protein (DARPin) [Binz, HK et al., Nature Biotechnology 22, 575-582 (2004)], is fused to a larger protein (β-lactamase from Escherichia coli (E. coli)) used as a crystallization chaperone [Batyuk, A., Wu, Y., Honegger, A., Heberling, MM & Pluckthun, A., Journal of Molecular Biology 428, 1574-1588 (2016); Wu, Y. et al., Scientific Reports 7, 11217 (2017)]. To the best of our knowledge, the rigidity and applicability of this approach in cryo-EM have not been evaluated. The fusion polypeptide of DARPin with the scaffold protein was achieved by linking two alpha helices. For this purpose, the N-terminal or C-terminal alpha-helix of the scaffold protein was seamlessly grafted onto the terminal helix of DARPin. None of the scaffold proteins linked to DARPin, which has a beta-sheet structure at the interface, and no proline linkers were used.

[0007] Multiple DARPin molecules are introduced onto larger scaffolds with high symmetry, serving as platforms for smaller protein molecules in cryo-EM. These highly complex, manipulated scaffolds resemble viral particles or large protein complexes, presenting DARPin (which may be included regardless of their target specificity) in up to 12 symmetrically organized copies. These protein complexes are robust, interconnected with high specificity and affinity, and consist of multiple subunits, requiring purification and assembly of multiple components [Liu, Y., Gonen, S., Gonen, T. & Yeates, TOProc Natl Acad Sci USA 115, 3362 (2018)]. A similar approach has utilized the enzyme aldolase for DARPin binding site multimerization [Yao, Q., Weaver, SJ, Mock, J.-Y. & Jensen, GJ Structure 27, 1148-1155.e3 (2019)]. Here, the ligation of DARPin to aldolase was clearly more flexible than the modular cage approach [Liu, Y., Gonen, S., Gonen, T. & Yeates, TOProc Natl Acad Sci USA 115, 3362 (2018)]. In both cases, the binding agent used is DARPin, a highly soluble and well-expressed protein, but which has not been successfully implemented in membrane protein structural biology (e.g., conformation-specific DARPin for GPCRs). Another drawback of these two platforms is the steric constraints arising from the scaffold itself. Collisions with this scaffold are prone to occur depending on the shape of the epitope and antigen. The ligation of DARPin to the scaffold was performed via the alpha-helix.

[0008] The small size of just 12–15 kDa and the compactness of single-chain antibodies (more specifically, the VHH domain of single-chain antibodies in camelids) are among its greatest advantages. However, this also represents a limitation in crystal contact formation, particularly in the expansion of target proteins in cryo-EM.

[0009] Traditionally, the VHH domain has been expanded by loop extension using fusion proteins. This construct has been named "megabody" [Laverty, D. et al., Nature 565, 516~520 (2019); Uchanski, T. et al., bioRxiv 812230 (2019) doi:10.1101 / 812230], a chimeric protein in which the protein hopQ (Uniprot entry B5Z8H1, PDB entry 5LP2) or other scaffold proteins are grafted onto nanobodies to increase their size for cryo-EM purposes. Megabodies are hetero and homo GABA AIt has been used to structurally characterize the receptor [Laverty, D. et al., Nature 565, 516-520 (2019); Masiulis, S. et al., Nature 565, 454-459 (2019)]. Here, the adhesion protein HopQ secreted from Helicobacter pylori was inserted into the first loop after the first beta chain fold of the nanobody. Importantly, the megabody has two linking peptide chains between the two fusion proteins ("internal fusion protein") and potentially lacks interdomain rigidity [Uchanski, T. et al., bioRxiv 812230 (2019) doi:10.1101 / 812230]. This is shown in the PDB entries for the α1-β3-γ2-hetero-GABA receptor in complex with megabodies 38 (PDB entries 6HUG, 6HUJ, 6HUK, 6HUO, 6HUP, and 6I53) in which a model of the grafted hopQ protein was not constructed. We conclude that the grafted protein was not sufficiently degraded by cryoEM and that this chimeric construct did not provide sufficient rigidity to be clearly visible in micrographs or subclass averages. In the published subclass averages [Laverty, D. et al., Nature 565, 516~520 (2019); Uchanski, T. et al., bioRxiv 812230 (2019) doi:10.1101 / 812230], the density of the hopQ protein is blurred, suggesting high flexibility between the nanobodies and the hopQ moiety. However, the megabody had a positive effect on particle distribution, resulting in micrographs with many different orientations, which are particularly demanding for Cys loop receptors such as the GABAA receptor.

[0010] Recently, the expansion of VHHs in camelids was achieved by C-terminal fusion to maltose-binding protein (MBP), providing a novel crystal contact for the crystallization of ion channels [Brunner, JD et al. (2018) doi:10.1101 / 480863] (PDB entry 6HD8). This chimeric construct was named “macrobody”. The use of this approach successfully enhanced the diffraction properties of the crystals and provided evidence of the potential for expanding single-chain antibodies to design novel chaperones. The fusion polypeptide was derived from a recombinant VHH antigen-binding domain (also called a nanobody, Figure 1A, Figure 1B [Steyaert, J. & Kobilka, BK Current Opinion in Structural Biology 21, 567-572 (2011)]) derived from alpaca immunoassay and a C-terminal maltose-binding protein (MBP, malE from Escherichia coli K12 strain, Uniprot accession P0AEX9) (Figure 1C) (Protein Databank (PDB) entry 6HD8) (Brunner JD et al., BioRxiv 480863, 2018, https: / / doi.org / 10.1101 / 480863).

[0011] The chimeric fusion polypeptide of entry 6HD8 is shown in Figure 2, along with a Fab fragment of a monoclonal antibody (another commonly used molecular chaperone) for size comparison. The reported fusion polypeptide is linked by a valine linker amino acid between a cleaved, conserved C-terminus of the VHH antigen-binding domain (Figure 3A) and an N-terminal cleaved MBP starting at amino acid 6 (lysine) by a recombinant means. The numbering refers to the processed MBP protein lacking the signal peptide. This amino acid numbering excludes the signal peptide (signal peptide for periplasmic targeting) at amino acid positions 1-26 of the complete open reading frame, which is removed by the E. coli (E. coli) cellular machinery during secretion into the periplasmic space (Figure 3B). The secondary structure of the N-terminus, particularly at amino acid positions 1-64, is shown in Figure 4. This structural motif is shown for clarity and contains the most important building blocks to which the VHH antigen-binding domain is linked.

[0012] The structural conservation of the VHH domain in camelids is particularly high at the C-terminus (Figure 5A). Macrobodies were generated, as shown in Figures 5B and 6 (Brunner JD et al., BioRxiv 480863, 2018, https: / / doi.org / 10.1101 / 480863). While the amino acid valine in the linker is not involved in stable secondary structural elements via hydrogen bonding or hydrophobic interactions (Figure 7), the preceding amino acids in the VHH domain and the amino acids at the N-terminus of the MBP moiety are involved in hydrogen bonding of structural beta-sheet elements. The chimeric fusion polypeptide is linked only by a single polypeptide chain between the VHH antigen-binding domain and the MBP moiety. The linkage is not rigid and can rotate or bend around all three axes (Figure 7), i.e., the parts can rotate relative to each other. Therefore, unfortunately, the macrobody design has an inherent drawback: the linker of the single polypeptide chain is very flexible, which consequently reduces the opportunity to form crystal contacts that would result in good diffraction properties and applicability in cryo-EM, due to a lack of visibility in micrographs or subclass averages. [Overview of the project]

[0013] The present invention provides a description of a chimeric fusion polypeptide comprising a first polypeptide which is an antigen-binding domain and a second polypeptide which is a polypeptide scaffold linked to the C-terminus of the antigen-binding domain, beginning with a beta-sheet structure consisting of parallel or antiparallel beta chains, wherein the antigen-binding domain and the fusion polypeptide scaffold are linked by a single peptide linker, the peptide linker comprising only proline residues. In particular, the present invention includes a fusion protein having an engineered robust linker of two proline residues between the C-terminus of a VHH antigen-binding domain and the cleaved N-terminus of a second domain beginning with a beta-sheet structure, such as bacterial periplasmic maltose-binding protein (MBP). The present invention is referred to as "Pro-macrobody". Using a computational method, the sequence between the C-terminus of the VHH antigen-binding domain and the scaffold protein is...VTV PPFusion proteins with a linker containing two robust prolines, manipulated according to LVI..., were designed and initially evaluated in silico (italics: antigen-binding domain residue, bold: linker, standard letters: expanded scaffold fusion partner). Antigen-binding domains include all VHH domains of heavy chain antibodies from camelids or cartilaginous fish or synthetic libraries derived therefrom, or VHH domains of classical double-chain antibodies (such as IgG), or synthetic antigen-binding domains with immunoglobulin structures such as monobodies. The scaffold fusion partner polypeptide at the C-terminus of the antigen-binding domain is, more precisely, the cleaved N-terminus of bacterial periplasmic maltose-binding proteins (MBPs), and more broadly, all periplasmic binding proteins, but specifically the Escherichia coli protein malE (Uniprot entry P0AEX9). The resulting fusion protein is called a "Pro-macrobody," with the commonly used acronym "PMb" (e.g., monoclonal antibodies are called "MAb" or nanobodies "Nb"). The rigidity and linear linkage between the two proteins are achieved by two precisely positioned proline residues (hence the prefix "Pro-") introduced between the two parts. The VHH antigen-binding domain can be replaced with VHH of any given specificity; the resulting fusion protein functions as a molecular chaperone for expanding small proteins, particularly membrane proteins. Remarkably, the modification with the diproline linker results in a major conformational and structural rearrangement of the domains relative to each other, compared to the structure of a macrobody lacking the two proline linker. Most notably, this new protein is highly rigid with respect to the relative motion between the VHH antigen-binding domain and the MBP portion. The significantly higher rigidity of the Pro-macrobody compared to a macrobody without the proline linker provides evidence for its usefulness in structural biology.

[0014] Pro-macrobodies can be used in any structural study, particularly in single-particle cryo-electron microscopy (cryoEM) and X-ray crystallography. This method provides a means to combine the favorable properties of the VHH domain as a protein chaperone with a fusion partner protein module (extending the VHH) to facilitate structural determination in cryoEM or X-ray crystallography. In the case of cryoEM, Pro-macrobodies help maximize the achievable resolution of the target protein by (i) increasing the molecular weight of the complex, (ii) increasing the contrast of the particles in the micrograph, (iii) better single-particle classification by adding discernible structural features, (iv) interference with the preferred orientation of protein particles on the electron microscope cryogrid, and (v) reducing the denaturation of the target protein at the air-water interface. Thus, Pro-macrobodies enable high-resolution structural analysis of proteins, with or without bound potential drug compounds, which is a key method of structure-based drug design. Key advantages of Pro-macrobodies include the robust linkage achieved by linker manipulation, the preferred linear extension of the molecule which reduces the likelihood of collision with the target protein, and finally, the versatility of being able to incorporate any desired antigen-binding domain into the scaffold, as demonstrated by cryo-EM. Furthermore, the antigen-binding properties of the VHH moiety are conserved (the VHH antigen-binding domain and the corresponding Pro-macrobodycetes have the same affinity for their respective antigens). Thus, Pro-macrobodies result in particle recognition features optimized for alignment while retaining the favorable properties of the VHH domain (conformational specificity, bacterial production, molecular biology, and small footprint paratopes). Evidence of the rigidity of a newly manipulated Pro-macrobodycete containing two tandem prolines in the linker is described using chimeric fusion polypeptide design / primary sequence, evaluation of possible linkers by whole-atom molecular dynamics simulations, and single-particle electron microscopy. In addition, the crystal structure of the Pro-macrobodycete at a resolution of 2.0 Å is shown, which confirms the predicted fold and provides a high-resolution structure of the designed linker.

[0015] Accordingly, in a first aspect, the present invention provides a fusion polypeptide comprising a first polypeptide which is an antigen-binding domain and a second polypeptide which is a polypeptide scaffold, wherein the polypeptide scaffold comprises parallel or antiparallel beta chains, the antigen-binding domain is linked at its C-terminus to the N-terminus of the polypeptide scaffold by a peptide linker, and the peptide linker consists of one or more proline residues.

[0016] In another aspect, the present invention relates to an amino acid sequence encoding a fusion polypeptide according to a first aspect of the present invention, optionally including SEQ ID NO: 001 or SEQ ID NO: 002.

[0017] In another embodiment, the present invention is i) A fusion polypeptide according to a first aspect of the present invention, and ii) Target protein specifically bound to the fusion polypeptide via an antigen-binding domain Regarding complexes that include this.

[0018] In another aspect, the present invention relates to the use of a fusion polypeptide, amino acid sequence, nucleic acid molecule, and complex of the present invention according to a first aspect of the present invention for structural analysis of a target protein.

[0019] Another aspect relates to the use of the fusion polypeptide according to the first aspect of the present invention as a pharmaceutical.

[0020] Another aspect relates to the use of a fusion polypeptide according to the first aspect of the present invention for diagnostic purposes.

[0021] These and further embodiments, as well as preferred embodiments thereof, are also further defined below in the detailed description and claims.

[0022] These methods and uses will be widely used in academic laboratories, pharmaceutical companies, genomics companies, agricultural companies, chemical companies, and the biotechnology industry. [Brief explanation of the drawing]

[0023] [Figure 1A] It is the secondary structure of a VHH domain with a label of the beta chain according to the IMGT global reference nomenclature (Lefranc, M-P., Frontiers in Immunology. 2014, Volume 5(22): 1-22). The complementarity-determining regions (CDRs) that bind to the epitope are highlighted in black. The N-terminus and C-terminus are shown. [Figure 1B] It is the crystal structure of the VHH domain of a camelid (from PDB entry 6HD8). The N-terminus and C-terminus are shown. [Figure 1C] It is the crystal structure of the malE maltose-binding protein (MBP) of Escherichia coli (uniprot number P0AEX9; PDB entry 1ANF). The N-terminus and C-terminus are shown. [Figure 2] It is the crystal structure of a chimeric fusion polypeptide ("macrobody") of the VHH domain derived from PDB 6HD8 with MBP, compared with the structure of the Fab fragment. Dimensions are shown in angstroms. [Figure 3A] It is a multiple sequence alignment of the amino acid sequences of six different alpaca VHH domains. The CDRs are surrounded by gray. Identical residues in the consensus sequence are indicated by asterisks. The conserved C-terminus of the camelid VHH domain is underlined (sequence: VTVSS). The sequence of the VHH that is part of PDB entry 6HD8 is shown in italics. [Figure 3B] It is the amino acid sequence of the maltose-binding protein malE of Escherichia coli K12 (Uniprot number P0AEX9). The signal peptide is shown in the upper row and underlined. The lysine 1 position (as present in the periplasmic space) and lysine 6 position of the processed MBP are numbered. [Figure 4A]It is a conserved structural element in the N-terminal portion of MBP containing amino acids 1 to 64. The conserved beta strands A, B, and C and the two helices I and II are marked. [Figure 4B] It is a topological representation of the N-terminal portion (amino acids 1 to 64) of MBP. [Figure 4C] It is an enlarged region of the structural element surrounded in Figure 4B, having the hydrogen bonds shown within the beta strand. [Figure 4D] It is an enlarged view of amino acids 1 to 64 shown in Figure 4A. This structural building block is conserved across the entire protein family of periplasmic binding proteins and is fused to VHH at the C-terminus. [Figure 5A] It is a superposition of the VHH domains of five different camelid species from the PDB entry. The CDR and N-terminal portions are not well aligned, but the scaffold and C-terminal are highly conserved at the structural level. The C-terminal is shown. [Figure 5B] VHH and MBP that are approximately oriented with respect to each other for fusion with each other. The lower panel shows the construct design and the amino acids of the linker and the boundary of MBP used in the macrobody (PDB entry 6HD8). [Figure 6] It is a schematic diagram of the exact boundary of the macrobody in PDB entry 6HD8. [Figure 7] It is the structure of the linker region of the macrobody in 6HD8. The hydrogen bonds of the beta sheet element of VHH and the N-terminal portion (amino acids 1 to 64) of MBP are shown at distances represented in angstroms. The linker (Val-Lys) is not involved in hydrogen bonding and can thus rotate and bend. This is shown by the arrow. [Figure 8] It is a schematic diagram of the design of a new chimeric fusion polypeptide (Pro-macrobody) of VHH and MBP using a di-proline linker at the interface. The boundary for fusing the two polypeptides is shown (indicated by scissors), representing the exact arrangement of the linker of two proline residues. [Figure 9](A): Flexibility of the chimeric fusion polypeptide (macrobody and Pro-macrobody) designed for LptD in MD simulation. Contour space investigated during a 500 ns MD trajectory, starting from the elucidated X-ray structure (left) with a double proline PP-linker, and from a simulation model (right) with the linker computationally mutated to valine and lysine (VK-linker). Both trajectories were aligned to the VHH structure (residues 1-120). Snapshots were taken every 100 ns and overlaid. A translucent white surface is used to show the extent of the computationally investigated molecular surface during the simulation. (B) and (C): Magnified views of the two linkers showing greater conformational flexibility of the VK-linker by overlaying several snapshots obtained from the simulation. [Figure 10] This is an essential dynamics analysis of MD trajectories. The correlated motion of alpha carbons was extracted from the trajectories (PP (Pro-macrobody) structure) after they were aligned to the VHH portion (residues 1-120). Small arrows represent the overall motion resulting from the combination of the first three principal components of the covariance matrix. Large arrows indicate the direction in which VHH binds: (A): PP-linker, side view showing bending motion; (B): VK linker for comparison, showing much larger bending motion; (C): PP-linker, rear view showing rotational motion with bending; (D): VK linker showing large rotation. [Figure 11A] This is an analysis of interdomain (VHH to MBP) angles of macrobodies in the MD trajectory. The timeline of interdomain angles in MD starts from the coordinate set (X-ray) with the formed PP-linker and is after equilibrium (200 ns). Interdomain angles are defined as the angles between the geometric centers of residues 1-120 (Nb), residues 121-122 (linker), and residues 123-486 (MBP). [Figure 11B]This is an analysis of interdomain (VHH to MBP) angles of macrobodies in MD trajectories. The distribution of interdomain angles observed in four different MD simulations started from two different X-ray structures elucidated by either the PP-linker (star shape) or the VL-linker (white circle, 6HD8.pdb). In the simulation model, the PP-linker is shown as a continuous line, and the dashed line corresponds to the VK-linker. [Figure 12A] This is a simulated structure of the diproline linker (derived from the 6HD8 PDB entry). The amino acids in the linker and one preceding valine residue are shown as balls and sticks. The enclosed panel compares the Val-Pro-Pro peptide, in which two consecutive prolines are in trans configuration. The proline in the linker in the simulation is also present in trans configuration. [Figure 12B] This is a stretch of the poly-proline II helix for comparison (from PDB entry 1AWI). All prolines are in the trans configuration. Note the similar folding as in the simulated linker. In the poly-proline I helix, all prolines are in the cis configuration, although this occurs very rarely in nature. [Figure 13] This is the original macrobody (from PDB entry 6HD8) compared to the Pro-macrobody (MD simulation). (A): Macrobody from PDB entry 6HD8 (without antigen). (B): Pro-macrobody (one frame from a 500 ns whole-atom MD simulation). (C): Superposition of (A) and (B) manually aligned with the VHH portion. The MBP portion exhibits a very different conformation. (D): Rear view aligned with (C). One helix is ​​framed in green to show the conformational difference between the original macrobody (left) and the modified version with a diproline linker (right). The MBP portion is rotated approximately 170 degrees counterclockwise along its long axis when viewed from the C-terminus. [Figure 14] This is a structural statistic of the X-ray structure of Pro-Macrobody 21. [Figure 15A] This is a side view of the X-ray structure of Pro-Macrobody 21. The linkers are indicated by balls and sticks. [Figure 15B] This is a simulated macrobody (6HD8) with diprolinker linkers. The linkers are shown as balls and sticks. Note the very similar overall conformation between the X-ray structure and the simulation. [Figure 16] (A)(B): A superposition of C-alpha traces and simulations of the crystallized Pro-macrobody 21 with the linker proline shown by the ball and stick. (C): A magnified view of the linker enclosed in (B). [Figure 17A] This is the X-ray structure of Pro-Macrobody 21. [Figure 17B] The enclosed region in Figure 17A is shown as an electron density map with 2.5 sigma contour lines. It represents the proline of the linker. [Figure 17C] The two figures show enlarged views of the linker of Pro-macrobody 21, accompanied by electron density maps contoured at 2.5 sigma. The linker proline is shown. [Figure 18] This is a magnified view of the linker interface of the Pro-Macrobody 21 crystal structure. Linkers are indicated by ball and stick symbols. Hydrogen bonds are shown by dashed lines, with corresponding distances in angstroms. Note that regions stabilized by hydrogen bonds are interrupted only by linkers, which are themselves rigid elements. [Figure 19] This is a putty representation of the temperature factor (B factor) in the X-ray structure of Pro-macrobody 21. High B factors are indicated by thick cartoon regions, and low B factors (low atomic mobility) are indicated by thin cartoon drawings. Note that the lowest B factors in the crystal are observed in the linker region supporting very strong linkages. The highest B factors, which are the regions with the highest atomic mobility, are found (in the absence of antigen) in the C-terminal half of MBP and the CDR of the VHH antigen-binding domain. [Figure 20]This is a surface representation of the two parts of the Pro-macrobody 21. VHH is shown in light gray and MBP in dark gray. The linker is shown as a sphere. (A): Top view of the structure. The linker is not exposed to the solvent. (B): Bottom view of the structure with a visible linker and partial maltose in the MBP bonding cleft. (C): Side view of the Pro-macrobody 21. The interface regions of the two fused parts fit together very well. [Figure 21] This image shows the binding kinetics of isolated VHH21 and Pro-macrobody 21 to an immobilized antigen (NgLptDE) using waveguide interferometry (Creoptix). (A): Binding characteristics of VHH21. (B): Binding characteristics of VHH21-PP-MBP(Pro-macrobody 21). (C): Equilibrium binding characteristics of VHH21. (D): Equilibrium binding characteristics of VHH21-PP-MBP(Pro-macrobody 21). Note that the dynamical parameters remain largely unchanged. The off-rate (kd) is slightly higher for the Pro-macrobody. [Figure 22] This shows the binding kinetics of isolated VHH51 and Pro-macrobody 51 to an immobilized antigen (NgLptDE) using waveguide interferometry (Creoptix). (A): Binding characteristics of VHH51. (B): Binding characteristics of VHH51-PP-MBP(Pro-macrobody 51). (C): Equilibrium binding characteristics of VHH51. (D): Equilibrium binding characteristics of VHH51-PP-MBP(Pro-macrobody 51). Note that the dynamical parameters remain largely unchanged. [Figure 23]This is a size exclusion chromatogram of the Pro-macrobody 21-NgLptDE complex (dashed line) compared to NgLptDE (solid line) on a Superdex 200 5 / 150 column. The NgLptDE-Pro-macrobody 21 complex elutes earlier from the column and exhibits a higher molecular weight. For clarity, the shift in elution volume is shown in the panel below. Black circles and asterisks represent uncomplexed NgLptDE and the Pro-macrobody 21 / NgLptDE complex, respectively. Arrows indicate excess Pro-macrobody 21 that elutes much later from the column. [Figure 24] Preparation of ternary complexes of Pro-macrobodies 21 and 51 with NgLptDE for structural determination using cryo-EM. (A): Size exclusion chromatogram of the ternary complex injected into a Superdex 200 10 / 300 column. The fraction used for cryo-EM is highlighted with a light gray bar. Excess Pro-macrobodies are well separated from the complex. (B): SDS-PAGE analysis of the eluted fraction from (A). Unbound Pro-macrobodies (added to NgLptDE in a 3x molar excess) are well separated from the ternary complex. Pro-macrobodies 21 and 51 have nearly identical molecular weights and cannot be separated by SDS-PAGE. Residual free MBP (impurities) are also separated. Further impurities found in the gel are present at approximately 40 and 30 kDa, but did not adversely affect structural determination. [Figure 25] This is the cryo-EM subclass average (2D projection) of the NgLptDE-Pro-macrobody 21 / 51 ternary complex. The complex can be viewed in different orientations, and the bound Pro-macrobodies are clearly visible (arrows). [Figure 26]This is a magnified view of the two subclass averages from Figure 25. The subdomains (VHH-PP-MBP) within the Pro-macrobody can be clearly distinguished. The asterisk indicates the VHH antigen-binding domain, the arrow indicates the N-terminal half of MBP, and the arrowhead indicates the C-terminal half of MBP. Maltose binds to the cleft between the N-terminal and C-terminal halves of MBP. A surface model of the crystal structure of Pro-macrobody L21, scaled down to approximately 8 angstroms, is shown on the right (maltose-bound form). The subdomains are indicated by the same symbols as in the left panel. [Figure 27] This is a three-dimensional cryoEM map of the labeled Pro-macrobody subdomains and the ternary complex of NgLptDE / Pro-macrobody 21 / Pro-macrobody 51. NgLptDE is shown in light gray. [Figure 28] This is a comparison of the simulated Pro macrobody structure, the X-ray structure of Pro-macrobody 21, and the EM structures of Pro-macrobody 21 and 51. (A): A single frame of the simulated Pro-macrobody (from the 6HD8 structure). (B): X-ray structure of Pro-macrobody 21 at 2 Å resolution (left) and a simulated lower resolution (right panel). (C): Cryo-EM structure of Pro-macrobody 21. The right panel is rotated 180°. (D): Cryo-EM structure of Pro-macrobody 51. The right panel is rotated 180°. All structures are shown in the same orientation except for the right panels of (C) and (D). Note the high conformational similarity. The EM structures are in the absence of bound maltose and explain the open conformation of MBP (see also Figure 29). Pro-macrobody 21 and 51 have very similar shapes and provide evidence of a universal feature of the Pro-macrobody design. [Figure 29]This is the X-ray structure of Pro-macrobody 21 fitted to the EM map of the ternary complex NgLptDE / Pro-macrobody 21 / Pro-macrobody 51. The X-ray structure is in the maltose-bound form and is therefore more closed, while the cryo-EM structure is in the apo form, with the N-terminal and C-terminal lobes of MBP forming a larger gap between them. Pro-macrobody 51 is not shown, but its very similar conformation is evident from the comparison in Figure 28. [Figure 30] This shows the wide-range resolution of NgLptDE and NgLptDE / Pro-macrobody ternary complexes from cryo-EM. Standard Fourier shell correlation (GSFSC) is shown with an FSC threshold of 0.143. Using Pro-macrobodies, the map resolution was improved by 1.2 Å compared to structures obtained with a similar number of particles. Importantly, the resolution improvement is at a level critically important for the visualization of side chains. [Figure 31] This is a comparison of cryo-EM maps of NgLptDE obtained from uncomplexed NgLptDE (first dataset) and NgLptDE / Pro-macrobody 21 / Pro-macrobody 51 (second dataset). The second dataset shows much higher resolution in the central region of the transmembrane region less than 3 Å, while uncomplexed NgLptDE has a maximum resolution of 4–4.5 Å. Similar numbers of particles were used for both maps. [Figure 32] (A): This is a complex of NgLptDE and a macrobody (the same VHH domain (VHH21 / 51) as in Figure 31), but it is a subclass average of complexes that have a VK linker instead of a PP linker like the Pro-macrobody. The MBP portion is indicated by an arrow. (B): The lower panel shows a three-dimensional reconstruction from this dataset, and the overall resolution does not exceed that of the uncomplexed NgLptDE dataset and is compared to the dataset with the Pro-macrobody (C). The Pro-macrobody is indicated by an arrowhead. [Modes for carrying out the invention]

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art to which the invention pertains. Embodiments, preferred embodiments and very preferred embodiments described and disclosed herein should apply to all aspects and other embodiments, preferred embodiments and very preferred embodiments, whether or not they are specifically referred to again, i.e., repetition has been avoided for the sake of brevity. The articles “a” and “an,” as used herein, refer to one or more (i.e., at least one) grammatical objects of the article. The term “or,” as used herein, should be understood to mean “and / or,” unless the context clearly indicates otherwise.

[0025] In one embodiment, the present invention provides a fusion polypeptide comprising a first polypeptide which is an antigen-binding domain and a second polypeptide which is a polypeptide scaffold, wherein the polypeptide scaffold comprises parallel or antiparallel beta chains, and the antigen-binding domain is linked at its C-terminus to the N-terminus of the polypeptide scaffold by a peptide linker, the peptide linker comprising one or more proline residues.

[0026] As used herein, the “antigen-binding domain” is limited solely to binding to a target antigen. Any domain structure can be used as the antigen-binding domain, as long as it binds to a target antigen. One suitable example of the antigen-binding domain of the present invention is a single-domain antibody (sdAb).

[0027] The term "polypeptide" refers to any sequence of two or more amino acids, regardless of length, post-translational modification, or function. Polypeptides may include both native and non-native amino acids. In one embodiment, the fusion protein includes adjacent chains of a beta-sheet structure, which may be parallel or antiparallel to each other.

[0028] As used herein, the term “linker” refers to a link between two elements, such as protein domains. A linker may be a covalent bond or a spacer. The term “bond” refers to any type of bond resulting from a chemical bond, such as an amide bond or a disulfide bond, or from a chemical reaction, such as a chemical conjugation. The term “spacer” refers to a portion (e.g., a polyethylene glycol (PEG) polymer) or amino acid sequence (e.g., a 3-200 amino acid, 3-150 amino acid, or 3-100 amino acid sequence) present between two polypeptides or polypeptide domains, providing space and / or flexibility between them.

[0029] The terms "C-terminus" or "carboxy-terminus" are used interchangeably herein and are defined herein as they commonly are in the art. An amino acid structure contains a carbon atom known as the carbon to which an amine group, a carboxylic acid group, and a side chain are attached. The N-terminus (also known as the amino-terminus, NH2-terminus, N-terminus, or amine-terminus) is the start of a protein or polypeptide, pointing to a free amine group (-NH2) located at the end of the polypeptide.

[0030] In one embodiment, the present invention provides a fusion polypeptide in which the polypeptide scaffold is a maltose-binding protein.

[0031] In another embodiment, the present invention provides a fusion polypeptide in which the polypeptide scaffold is Escherichia coli maltose-binding protein, Uniprot entry P0AEX9.

[0032] In another embodiment, the present invention provides a fusion polypeptide in which the peptide linker consists of one, two, three, or four proline residues.

[0033] In another embodiment, the present invention provides a fusion polypeptide in which the peptide linker comprises two proline residues. As shown, such a robust linker comprises two proline residues between an antigen-binding domain, preferably the C-terminus of a VHH antigen-binding domain, and a second domain beginning with a beta-sheet structure, such as the cleaved N-terminus of a polypeptide scaffold, preferably a bacterial periplasmic maltose-binding protein (MBP).

[0034] In another embodiment, the present invention provides a fusion polypeptide in which the C-terminus of the antigen-binding domain consists of an amino acid sequence VTV. Such a C-terminus sequence is typically found in all VHH domains of heavy chain antibodies derived from camelids or cartilaginous fish or synthetic libraries derived therefrom, or in the VHH domain of a classical double-chain antibody (e.g., IgG), or in a synthetic antigen-binding domain having an immunoglobulin structure such as a monobody. In another embodiment, the present invention provides a fusion polypeptide in which the C-terminus of the antigen-binding domain consists of an amino acid sequence VTV, and the antigen-binding domain is a VHH antigen-binding domain.

[0035] In another embodiment, the present invention provides a fusion polypeptide in which the N-terminus of the polypeptide scaffold consists of the amino acid sequence LVI. Preferably, the polypeptide scaffold, the amino acid sequence LVI, includes a beta-sheet structure such as the cleaved N-terminus of a bacterial periplasmic maltose-binding protein (MBP), and typically begins with such a beta-sheet structure. The scaffold fusion partner polypeptide at the C-terminus of the antigen-binding domain is, more precisely, the cleaved N-terminus of a bacterial periplasmic maltose-binding protein (MBP), and more broadly, all periplasmic binding proteins, but particularly the Escherichia coli protein malE (Uniprot entry P0AEX9). In another embodiment, the present invention provides a fusion polypeptide in which the N-terminus of the polypeptide scaffold consists of the sequence LVI, the polypeptide scaffold is the Escherichia coli maltose-binding protein, Uniprot entry P0AEX9, and preferably the polypeptide scaffold is the cleaved N-terminus of a bacterial periplasmic maltose-binding protein (MBP).

[0036] In another embodiment, the present invention provides a fusion polypeptide in which the C-terminus of the antigen-binding domain is linked to the N-terminus of a polypeptide scaffold by a peptide linker consisting of two prolines, and the C-terminus of the antigen-binding domain linked to the N-terminus of the polypeptide scaffold by the peptide linker comprises the amino acid sequence VTVPPLVI (SEQ ID NO: 10).

[0037] In another embodiment, the present invention provides a fusion polypeptide comprising the amino acid sequence VTVPPLVI (SEQ ID NO: 10).

[0038] In another embodiment, the present invention provides a fusion polypeptide in which the C-terminus of the antigen-binding domain, which is a VHH antigen-binding domain, is linked to the N-terminus of the polypeptide scaffold, which is Escherichia coli maltose-binding protein (Uniprot entry P0AEX9), by a peptide linker consisting of two prolines (i.e., the peptide linker is a di-amino acid PP-linker), preferably the polypeptide scaffold is the cleaved N-terminus of bacterial periplasmic maltose-binding protein (MBP), and the C-terminus of the antigen-binding domain linked to the N-terminus of the polypeptide scaffold by the peptide linker contains the amino acid sequence VTVPPLVI (SEQ ID NO: 10). Rigidity and linear linkage between the two proteins are achieved by the two precisely positioned proline residues. The VHH antigen-binding domain can be replaced with VHH of any given specificity, and the resulting fusion protein functions as a molecular chaperone for expanding small proteins, particularly membrane proteins. Surprisingly, modification with the diproline linker results in major conformational and structural rearrangements of the domains relative to each other, compared to the macrobody structure lacking the two proline-containing linker. Most notably, the linkage of this new, preferred fusion protein is highly rigid with respect to the relative motion between the VHH antigen-binding domain and the polypeptide scaffold, preferably the MBP portion.

[0039] In another embodiment, the present invention provides a fusion polypeptide in which a peptide linker links the C-terminal portion of an antigen-binding domain to the N-terminal portion of a polypeptide scaffold.

[0040] In another embodiment, the present invention provides a fusion polypeptide in which a peptide linker links the conserved cleaved carboxy terminus of a camelid single-chain antibody VHH domain following the conserved sequence valine-threonine-valine (VTV) at the end of beta chain G to the cleaved amino terminus of an Escherichia coli maltose-binding protein that begins with leucine at amino acid position 7, which is the first amino acid of beta chain A of the maltose-binding protein.

[0041] As used herein, the term "amino acid position" refers to the position number of an amino acid in a protein or protein domain.

[0042] In another embodiment, the present invention provides a fusion polypeptide in which the C-terminus of the antigen-binding domain VHH comprises three amino acids valine-threonine-valine, and the N-terminus of the maltose-binding protein comprises three amino acids leucine-valine-isoleucine.

[0043] In another embodiment, the present invention provides a fusion polypeptide in which the carboxyl terminus of a peptide linker is fused to the first amino acid of the most amino-terminal beta chain of the polypeptide scaffold.

[0044] In another embodiment, the present invention provides a fusion polypeptide comprising a polypeptide of the periplasm-binding protein superfamily of Interpro entry IPR025997, or a polypeptide of the periplasm-binding protein-like I-integral superfamily of Interpro entry IPR028082, or the periplasm-binding protein domain Pfam domain Peripla_BP_4 of PF13407.

[0045] In another embodiment, the present invention provides a fusion polypeptide of IPR025997 and IPR028082 of the Interpro superfamily or the Pfam domain PF13407, in which at least one amino acid substitution of the original database entry has occurred within this scaffold.

[0046] In another embodiment, the present invention provides a fusion polypeptide comprising additional fusion polypeptides that are incorporated at any position internally or fused to the carboxyl terminus of the polypeptide scaffold.

[0047] In another embodiment, the present invention provides a fusion polypeptide in which a fusion polypeptide belonging to the Interpro superfamily IPR025997 and IPR028082, or containing the Pfam domain PF13407, is cleaved before its native carboxyl terminus.

[0048] In another embodiment, the present invention provides a fusion polypeptide comprising, preferably, the peptide linker and the polypeptide scaffold comprising the amino acid sequence of SEQ ID NO: 11.

[0049] In another embodiment, the present invention provides a fusion polypeptide comprising the amino acid sequence of SEQ ID NO: 11.

number

[0050] In another embodiment, the present invention provides a fusion polypeptide comprising the antigen-binding domain, peptide linker, and polypeptide scaffold containing the amino acid sequence of SEQ ID NO: 12.

[0051] In another embodiment, the present invention provides a fusion polypeptide comprising the amino acid sequence of SEQ ID NO: 12.

number

[0052] In another embodiment, the present invention provides a fusion polypeptide in which the antigen-binding domain comprises an immunoglobulin-like fold.

[0053] An immunoglobulin (Ig)-like domain is a protein domain that has an amino acid sequence and structure similar to that of an immunoglobulin Ig domain (Edelman, 1987; Williams & Barclay, 1988). Structurally, the Ig domain has a unique immunoglobulin fold consisting of 70 to 110 amino acids.

[0054] In another embodiment, the present invention provides a fusion polypeptide in which the antigen-binding domain comprises an immunoglobulin (Ig) domain.

[0055] The Ig domain is 70-110 amino acids long and has a different overall structure called the Ig fold. This fold consists of two β sheets, each composed of a short, antiparallel β chain. In most Ig domains, the two β sheets are linked by a disulfide bond.

[0056] In another embodiment, the present invention provides a fusion polypeptide in which the antigen-binding domain is a VHH antigen-binding domain.

[0057] In another embodiment, the present invention provides a fusion polypeptide in which the antigen-binding domain comprises a VHH domain derived from a camelid VHH or cartilaginous fish VHH, i.e., shark VHH, or ray VHH, or sawfish VHH, or a mammalian antibody or the heavy or light chain of a monobody.

[0058] In another embodiment, the present invention provides a fusion polypeptide in which the VHH domain is derived from a single-chain antibody isolated from a member of the camelid family (i.e., Llama, Vicugna, or Camelus species).

[0059] In another embodiment, the present invention provides a fusion polypeptide in which the VHH-antigen-binding domain is the VHH domain of the heavy or light chain of a classical vertebrate Ig antibody.

[0060] In another embodiment, the present invention provides a fusion polypeptide in which the VHH antigen-binding domain is replaced by a monobody (a synthetic binder constructed on a fibronectin type III domain) or an immunoglobulin (Ig).

[0061] In another embodiment, the present invention provides a fusion polypeptide in which the VHH antigen-binding domain is selected from a synthetic library generated by recombinant technology or PCR, or presented in recombinant yeast (Saccharomyces cerevisiae), or presented in recombinant phage.

[0062] In another embodiment, the present invention provides an amino acid sequence encoding the fusion polypeptide of the present invention. In one embodiment, the amino acid sequence includes SEQ ID NO: 001 or SEQ ID NO: 002. Amino Acid Sequence ID No. 001:

number

number

[0063] In a further embodiment, the present invention provides a chemically modified and covalently labeled fusion polypeptide.

[0064] In another embodiment, the present invention is i) The fusion polypeptide of the present invention, and ii) Target protein specifically bound to the fusion polypeptide A composite is provided that includes the following.

[0065] In one embodiment, in the complex of the present invention, the target protein is bound to the antigen-binding domain of the fusion polypeptide.

[0066] In another embodiment, the target protein in the complex of the present invention has a molecular weight of less than 100 kDa. In yet another embodiment, the target protein in the complex of the present invention has a molecular weight of 90 kDa, or 80 kDa, or 70 kDa, or 60 kDa, or 50 kDa, or 40 kDa, or 30 kDa, or less than 10 kDa.

[0067] In another embodiment, the present invention provides the use of the fusion polypeptide, amino acid sequence, and complex of the present invention for structural analysis of a target protein. In one embodiment, the structural analysis includes single-particle cryo-electron microscopy (cryo-EM) or X-ray crystallography. In another embodiment, the structural analysis includes single-particle cryo-EM or X-ray crystallography or negative-stained TEM or electron diffraction or NMR or any other structural determination technique.

[0068] In another aspect, the present invention provides the use of the fusion polypeptide according to the present invention as a pharmaceutical.

[0069] In another aspect, the present invention provides the use of the fusion polypeptide according to the present invention for diagnostic purposes.

[0070] In another embodiment, the present invention provides a fusion polypeptide comprising the following Interpro entries: i) Periplasm-binding protein / LacI glycosylation domain (IPR001761) ii) Receptor and ligand binding region (IPR001828) iii) Leucine-binding protein domain (IPR028081) iv) Arabinose metabolic transcriptional repressor, ligand-binding domain (IPR033532).

[0071] Example I Evaluation of chimeric fusion polypeptides (macrobodies) and design of novel chimeric fusion polypeptides (Pro-macrobodies) using computer simulations to enhance rigidity. This document describes a workflow for analyzing the motion of chimeric fusion polypeptides using computer simulations (all-atom molecular dynamics simulations) and simulations of in silico amino acid exchange (mutants) in the linker region. Chimeric fusion polypeptides with amino acid exchange in the linker region were subjected to all-atom molecular dynamics (MD) simulations to quantify their motion. All MD simulations were fully atomic (i.e., each atom was simulated, not a group of atoms / molecules) and were performed in Desmond and OPLS3 force fields using solvents specified at room temperature (300 K, or 26.85 °C). For all simulations, a 300 ns pre-equilibrium was performed before 500 ns of production.

[0072] Simulation model We constructed simulated models of four different VHH-MBP fusion polypeptides, starting from two different coordinate sets. The first coordinate set was constructed from the PDB database code 6HD8.pdb structure (macrobody, here referred to as the VK structure) simulated using Val121 and Lys122 linkers (hereafter referred to as VK-linkers; numbering refers to Protein Databank (PDB) entry 6HD8). The second coordinate set was constructed from a newly simulated structure (described below) of a VHH-MBP fusion polypeptide (Pro-macrobody, here referred to as the PP structure) whose antigen is the Neisseria gonnorrheae protein LptD (Uniprot accession A0A1D3INQ1) and Pro122 Pro123 linkers (hereafter referred to as PP-linkers). The VK structure was simulated either in its original state or after changing its linkers to PP-linkers. Similarly, the PP structure was simulated either in its original state or after its linker was changed to a VK-linker.

[0073] The precise design of the chimeric VHH-MBP fusion polypeptide containing a diproline linker with all important boundaries is shown in Figure 8. This design was used for in silico studies as well as for overexpression and crystallization in Escherichia coli (E. coli).

[0074] result Starting with both VK and PP structures, MD simulations using both VK and PP linkers demonstrated a significant influence of linker residue type on the dynamics of the VHH-MBP fusion polypeptide. Initial results were obtained for the VK structure, suggesting that the PP-linker could be a significantly more robust chaperone. Subsequently, to test this hypothesis, the VHH-MBP fusion polypeptide was generated in silico using the PP-linker, producing a novel VHH-MBP fusion polypeptide in E. coli (E. coli) in which the VHH antigen-binding domain was bound to the aforementioned LptD protein. The X-ray structure of this designed Pro-macrobody was elucidated (Section III) and shown to be consistent with the structure predicted from MD simulations (RMSD < 2 Å). This coordinate set (PP structure) was then seeded, and additional MD simulations were performed, including a simulation model in which the VK linker was reintroduced for comparison. Clear agreement was obtained in both cases, demonstrating that the linker can modify the dynamics regardless of the initial set of coordinates selected to run the simulation.

[0075] Figure 9A shows the stiffness of both the VK-linker and the designed PP-linker using a model constructed from the VK structure. Very large bending is observed around the linker, which can also be observed by zooming in on the linker (Figure 9B).

[0076] Figure 10 shows that both bending and rotational motion occur between VHH and MBP, with a significant decrease in flexibility in the case of the PP linker. The visualization is estimated from the normal mode analysis of the simulation.

[0077] Figure 11 shows that the bending of the inter-domain angle (the angle between domains – a measure of the flexibility of the domains relative to each other) at the start of the simulation is much greater for the original macrobody construct with VK linkers (up to 40 degrees), but hardly any bending occurs for the new linkers made of proline.

[0078] In Figure 11A, the simulation time series shows that the interdomain angle for the VK-linker oscillates with a much larger amplitude than that for the PP-linker. Figure 11B shows histograms of four simulations using different linker-structure combinations, showing that the interdomain angle in the PP-linker simulations converges towards 167.6 and 169.5 degrees, respectively, when starting from the VK and PP structures. The results indicate that the initial coordinates did not significantly affect the equilibrium interdomain angle, which is found to converge to the same value. This angle is in good agreement with the angle observed in the X-ray structure of the PP-linker fusion polypeptide at 172.6 degrees. The standard deviation for the interdomain angle in these simulations was 5.6° for the VK→PP linker simulation and 3.6° for the simulations starting from the PP structure. As expected, the lower standard deviation indicates a much more robust structure for the PP-linker compared to the simulations performed using the VK-linker. Simulations using a VK linker, starting from either a PP or VK structure, have been shown to exhibit a much wider range of values, although they do not converge to a stable angular distribution during the simulation.

[0079] Evaluation of Results Analysis revealed that the two-fusion polypeptide with a VK linker possessed a large conformational degree of freedom, allowing rotation around the polypeptide backbone with the valine-lysine linker even on a 500 ns timescale. Both Val121 and lysine 6 (Lys122 in the chimeric fusion polypeptide) in MBP were not stabilized, or at least insufficiently stabilized, by hydrogen bonding interactions within the beta sheet to prevent rotation. In solution, outside of protein crystals, the chimeric fusion polypeptide derived from structure 6HD8 (Figure 2) rotated and, on average, existed in several different conformational states (Figures 9–11). When the VHH antigen-binding domain was attached to its respective antigen, the MBP portion was not rigidly linked but instead wobbled with considerable flexibility. Therefore, the chimeric fusion polypeptide would not meet the requirements to function as a molecular chaperone for protein complexes in solution (e.g., in single-particle cryomicroscopy). To increase the rigidity between the VHH antigen-binding domain and the MBP portion, a linker consisting of proline was tested (Figure 8).

[0080] Proline's rotation is restricted due to the unique linkage of its side chain to the backbone. Proline is an amino acid with exceptional conformational rigidity. The α-amino group is directly linked to the main chain, making the α-carbon a direct substituent on the side chain. Therefore, L-proline lacks the rotational freedom of all other native amino acids. From the standpoint of its rigidity, L-proline would be a suitable amino acid to test. However, L-proline in polypeptides introduces a kink, where its nitrogen atom cannot contribute as a hydrogen bond donor but only as an acceptor. This results in the destabilization or termination of the alpha-helix or beta-sheet secondary structure. The effect of the L-proline residue on the stability of the linkage between the VHH antigen-binding domain and the MBP domain was unclear, particularly due to its effect on the secondary structure elements essential to the fusion polypeptide, and due to the introduced kink(s) which would affect the desired elongated shape of the fusion polypeptide.

[0081] The linker array is shown in Figure 6...VTV V KLI... from... VTV PP When LI... is changed (Figure 8), the protein adopts a new conformation in molecular dynamics simulations (italics: C-terminal residue from the VHH domain; bold and underlined: newly introduced linker amino acid; standard letters: N-terminal residue from MBP). The two proline residues exhibit a trans configuration, similar to most proteins where two subsequent prolines appear (Figure 12A), resembling a short stretch of the polyproline II helix (Figure 12B). The turn introduced into the linker by the two consecutive proline residues rotates the MBP portion by approximately 170 degrees along its long axis (counterclockwise when viewed from the end opposite the VHH antigen-binding domain) compared to the original crystal structure of 6HD8 (Figures 13A, D) (Figure 13B). The protein shape remains elongated, and the enlarged MBP portion is on the opposite side of the paratope and therefore as far away from the epitope as possible, making it an optimal configuration. Figure 13C shows the superposition of Figures 13A and B, respectively.

[0082] To the inventors' surprise, the modification with the diproline linker resulted in significant conformational and structural changes compared to the structure of the fusion polypeptide in structure 6HD8. Most notably, this new protein is remarkably stable with respect to the relative motion between the VHH antigen-binding domain and the MBP moiety. This is evident from a comparison of the trajectories in molecular dynamics simulations of the original chimeric fusion polypeptide with the VK linker and the new fusion polypeptide with the diproline (PP) linker (Figures 9-11). Comparing the bending of the two constructs in Figure 11, we show that the diproline linker substantially reduces the bending between the two fusion moieties. Only bending of ±3.6° (within a 95% confidence interval) may be observed, while bending motion of up to 40° may be observed in 6HD8. Furthermore, torsional motion is also significantly reduced, as shown in Figure 10 (normal mode analysis). In summary, only slight vibrational motion may be observed. The next section describes the X-ray structure of a chimeric fusion polypeptide consisting of a VHH antigen-binding domain linked by a diproline linker and an MBP moiety, confirming the findings in this section.

[0083] Example II Structural determination of chimeric fusion polypeptides with diproline linkers between fusion regions using X-ray crystallography. To confirm the predicted structure, a novel fusion polypeptide (Figure 8) (the clone of the VHH antigen-binding domain is L21, sequence number 1) was recombinantly produced in Escherichia coli and homogeneously purified by chromatography (Brunner JD et al., BioRxiv 480863, 2018, https: / / doi.org / 10.1101 / 480863) in the same manner as the original fusion polypeptide with a valine linker. The purified protein with a PP linker was concentrated to 10 mg / ml and subjected to crystallization. The crystals were grown, cryoprotected, rapidly frozen in liquid nitrogen, and subjected to X-ray diffraction at the Swiss Light Source (beamline X10SA) at the Paul Scherrer Institute. The crystals were diffracted to a resolution of 2.0 Å and belonged to space group I2. The structure was elucidated by molecular substitution and refined with excellent statistics (Figure 14). All residues could be clearly assigned to the electron density map. The structure is remarkably similar to the structure predicted from the aforementioned computer simulations (Figures 15 and 16). A high-resolution electron density map around the linker region is shown in Figure 17. A more detailed map of the linker region reveals the conservation of hydrogen bonds at the C-terminus of the VHH domain and the N-terminal domain of the fused MBP (Figure 18). In the MBP portion, the first amino acid is leucine (leucine 7), which is stabilized in the beta chain. In the VHH domain, the last amino acid preceding the diproline linker is valine, which also interacts with the hydrogen bonds of the main chain in the beta chain. Thus, the diproline motif functions as a rigid bridge between the two rigid parts, resulting in a rigid fusion polypeptide as a whole. This is further reflected by the very low relative B factor (temperature factor) of the linker region in the crystal structure (Figure 19). The distribution of B factors reveals that the motility in the protein is lowest near the linker region and highest at the C-terminus of the MBP. This distribution is remarkably similar to the results from the normal-mode analysis shown in Figure 10. Furthermore, there is almost no space remaining between the VHH antigen-binding domain and MBP.The two surfaces of the protein fit together very well (Figure 20), and not only are there no observable collisions, but there is also no extra space to allow for linker shortening or to provide further advantages. Polyproline helices (PPII helices) can be found in protein structures, and manipulated PPII helices have been used as molecular rulers for studying protein interactions and interdomain FRET signaling [Bonger et al. 2010; Adzhubei et al., 2013; Dobitz et al. 2017]. However, there are no reports of fusion polypeptides with diproline linkers that perform molecular chaperone functions for structural biology, or related functions, or fusion polypeptides constructed for the purpose of robust linking between fusion partners.

[0084] Example III Chimeric fusion polypeptides containing a diproline linker retain the antigen-binding properties of the coupled VHH antigen-binding domain. The binding of the VHH antigen-binding domain to each antigen was measured using waveguide interferometry, a biophysical method for investigating direct ligand binding to immobilized molecular target proteins, which allows for the evaluation of binding affinity, stoichiometry, and kinetics. The binding kinetics of unfused VHH antigen-binding domains were measured against antigens coupled to sensor chips by biotin-neutraavidin interactions, according to the manufacturer's instructions (Creoptix, Wadenswil, Switzerland). The binding kinetics of the corresponding VHH antigen-binding domains in chimeric fusion proteins containing MBP and linked by a diproline linker were also measured and compared (Figures 21 and 22). All dynamic parameters (ka, kd, and KD) are very similar, even for equilibrium kinetics. In Figure 21, VHH21-PP-MBP is compared to VHH21, and in Figure 22, VHH51-PP-MBP is compared to VHH51. Both VHH antigen-binding domains were directed towards NgLptDE and used in cryo-EM as described below. The data show that chimeric fusion of the VHH antigen-binding domain with MBP is an essential parameter for transforming the VHH antigen-binding domain into a diproline-linked chimeric fusion polypeptide without altering binding to each antigen.

[0085] Example of applying PRO macrobody Example IV We purified complexes of small membrane proteins and chimeric fusion polypeptides with diproline linkers and analyzed them by single-particle cryo-EM. We prepared complexes of chimeric fusion polypeptides with diproline linkers and corresponding antigens (bacterial transporters for lipopolysaccharide (LPS)). The bacterial LPS transporter complex LptDE (composed of proteins LptD(Uniprot) and LptE(Uniprot)) from Neisseria gonnorrhoeae was selected as an example (NgLptDE). The protein is a relatively small membrane protein target for cryo-EM, asymmetric, and almost exclusively constructed from β-chains, with a density of only 110 kDa. These characteristics make NgLptDE a very suitable example for investigating the positive influence of Pro-macrobodies on structural determination. Low-resolution structures obtained beforehand without Pro-macrobodies served as controls for direct comparison. VHH antigen-binding domains that bind to NgLptDE were constructed and pre-characterized. Two specific VHH antigen-binding domains were selected and transformed into Pro-macrobodies. These are clones L21 and L51.

[0086] The complex of NgLptDE and Pro-macrobodies was purified by size exclusion chromatography (Figure 23) and eluate fractionation after mixing the two components. For cryoEM, several complexes were prepared using one or both of the two Pro-macrobodies (ternary complex). Peak fractions containing the transport protein and the chimeric fusion polypeptide complex were analyzed by SDS-PAGE (sodium dodecyl sulfate polyacrylamide gel electrophoresis) (Figure 24). Fractions containing the complex protein at a concentration of 1 mg / ml were applied to a cryoEM grid for vitrification in liquid ethane (ternary complex). The grid quality was inspected with a screening electron microscope, then a Titan 300kV class microscope operating at cryogenic temperatures and equipped with a K2 direct electron detector camera (Gatan / ThermoFisher, USA). The Krios electron cryomicroscope (C-CINA / Basel, Switzerland (FEI / ThermoFisher)) was used to image micrographs. Particles from the micrographs were collected using an automated method with CryoSparc software, classified into subclasses (different views on the three-dimensional particle depending on their orientation in vitrified ice), and then averages (called 2D classes, as these are projections of the three-dimensional object) were automatically generated by summing the individual particles. From the 2D classes, the three-dimensional volume of the object (in this case, a protein complex) can be calculated. The 2D classes (subclass averages) are calculated according to their orientation. Several details of the object have already appeared (Figure 25), and also provide information about the quality of the sample and the flexibility of the object. In the case of flexible subdomains in the object, the details cancel each other out on an overall average because the flexible portion was frozen in many different conformations. When the subdomain is rigidly attached to the object, it is clearly identifiable. Importantly, rigidly attached subdomains of the object also provide information about the relative orientation of the object in space, increasing the contrast of the object in the electron micrograph. This makes the classification process more accurate, particles are correctly assigned to subclasses, and the quality of the calculated volume of the object is improved.The present invention aims precisely to provide such additional subdomains, which are particles rigidly linked to objects lacking other features or being too small, for accurate classification. In the subclass average of the complex (consisting of the outer membrane NgLptDE transporter and a chimeric fusion polypeptide containing both diproline linkers), the chimeric fusion polypeptide is clearly visible. This is further a relatively well-defined structure, with the two domains of the MBP portion (N-terminal and C-terminal domains) already identifiable (Figure 26). This is due to the strong support of a rigid molecule, as otherwise it would be "canceled out" in the image by motion. Using the 2D class average, we were able to reconstruct the three-dimensional volume and generate a map showing the isolated chimeric fusion polypeptide attached to the target protein (Figure 27). The structure of the chimeric fusion polypeptide with diproline linkers is also nearly identical to the structure elucidated by X-ray crystallography and structural predictions from molecular dynamics simulations (Figure 28). This consistently demonstrates that a novelly designed chimeric fusion polypeptide of a VHH domain linked to the Escherichia coli (E. coli) maltose-binding protein malE by two proline peptide linkers folds as predicted and is a robust entity in solution. The chimeric fusion polypeptide model from the X-ray structure (Figure 15A) can be readily adapted to the density of the electron microscopy map, as shown in Figure 29. Furthermore, although binding to different epitopes, both Pro-macrobodies L21 and L51 share the same fold and degrade equally well in three-dimensional reconstruction. This indicates that the Pro-macrobody format can be applied to different VHH antigen-binding domains and possesses universal characteristics.

[0087] This invention enables higher-resolution EM mapping of target proteins. A difference of 1.2 angstroms was observed in the resolution of density maps of bacterial transporters in the absence or presence of a specific chimeric fusion polypeptide (i.e., uncomplexed compared to a complexed transporter with two Pro-macrobodies). Applying representative Fourier Shell correlation (GSFSC), a standard measure for obtaining the overall resolution of a given density map from electron microscope volume reconstruction (FSC cutoff of 0.143), the resolution was 4.6 Å for the uncomplexed sample and 3.4 Å for the transporter complexed with the chimeric fusion polypeptide (Figure 30). Figure 31 compares the improved resolution of the sample complexed with the Pro-macrobodies to that of the sample without the chimeric fusion polypeptide. This data provides proof of concept of the usefulness of target proteins and the effective increase in resolution by applying a chimeric fusion polypeptide of an antigen-binding domain linked to MBP with a diproline linker (hereinafter referred to as "Pro-macrobodies").

[0088] Therefore, fusion proteins having diproline linkers, as described in Examples I-IV, are unique in that they allow for the exchange of one domain (VHH domain) without sacrificing the rigidity of the fusion protein. A single polypeptide linker can always be assumed to be flexible, failing to stabilize unless the two linked domains exhibit specific interaction at the interface. Using fusion proteins having diproline linkers, we have generated two rigidly linked, non-interacting proteins, and the identified double proline linker serves the objectives of (i) rigid linkage and (ii) keeping the VHH domain freely interchangeable.

[0089] It is important to note that chaperone rigidity becomes increasingly important as the size of protein targets decreases, requiring extremely rigid chaperones. The chimeric fusion polypeptides having diproline linkers described herein represent a significant technological improvement in molecular chaperones for structural biology.

[0090] The complete amino acid sequence of the chimeric fusion polypeptide (Pro-macrobody 21) of the VHH antigen-binding domain (clone L21) and MBP with the diproline linker is sequence number 001 below:

number

[0091] The complete amino acid sequence of the chimeric fusion polypeptide (Pro-macrobody 51) of the VHH antigen-binding domain (clone L51) and MBP with the diproline linker is sequence number 002 below:

number

[0092] The VHH domain (antigen-binding) is shown in italics and can be replaced with any other VHH due to its conserved structure. The linker is shown in bold and underlined. The remaining sequence is the expanded domain (malE of *E. coli*). The VHH domain shown here is antigen-binding domain clone 21, which was crystallized and used for electron microscopy.

[0093] In the final experiment, Pro-macrobodies were compared to macrobodies (including a VK linker, as in PDB entry 6HD8). For this purpose, NgLptDE was complexed with macrobodies 21 and 51 (each with the same VHH domain as the Pro-macrobodies), purified by chromatography in the same manner, and subjected to analysis by cryo-EM. As shown in Figure 32, it was found that the macrobodies could not significantly improve resolution compared to uncomplexed NgLptDE. The MBP portion is not sufficiently degraded due to its high intrinsic motion in the linker when not linked with a PP linker. In three-dimensional reconstruction from this dataset, a clear density of the MBP portion was not evident. As a result, the overall resolution was not significantly improved. This experiment suggests that linking using two prolines, as in the Pro-macrobodies, offers a greater advantage over the previous macrobodies. [Sequence Listing Free Text]

[0094] Sequence Listing 1 <223> A chimeric fusion polypeptide (Pro-macrobody 21) of the VHH antigen-binding domain (clone L21) and MBP with a diproline linker. Sequence Listing 2 <223> A chimeric fusion polypeptide (Pro-macrobody 51) of the VHH antigen-binding domain (clone L51) and MBP with a diproline linker. Sequence Listing 3 <223> E02_12652 Sequence Listing 4 <223> D05_12648 Sequence Listing 5 <223> A11_12648 Sequence Listing 6 <223> F10_12652 Sequence Listing 7 <223> A09_12649 Sequence Listing 8 <223> H01_12651 Sequence Listing 9 <223> P0AEX9|MALE_ECOLI Sequence Listing 10 <223> ABD's N-terminus - Peptide linker - PPS's C-terminus Sequence Listing 11 <223> Diprolyn Linker and MBP Sequence Listing 12 <223> N-terminal ABD-PP-MBP

Claims

1. A fusion polypeptide comprising a first polypeptide which is an antigen-binding domain and a second polypeptide which is a polypeptide scaffold, The antigen-binding domain is a VHH antigen-binding domain, and the polypeptide scaffold is a periplasm-binding protein. A fusion polypeptide comprising a polypeptide scaffold containing parallel or antiparallel beta chains, wherein the C-terminus of the antigen-binding domain is linked to the N-terminus of the polypeptide scaffold by a peptide linker, the peptide linker comprising one, two, three, or four proline residues.

2. The fusion polypeptide according to claim 1, wherein the polypeptide scaffold is a maltose-binding protein.

3. The fusion polypeptide according to claim 1 or 2, wherein the polypeptide scaffold is an Escherichia coli maltose-binding protein.

4. The fusion polypeptide according to any one of claims 1 to 3, wherein the peptide linker consists of two proline residues.

5. The fusion polypeptide according to any one of claims 1 to 4, wherein the antigen-binding domain comprises an immunoglobulin-like fold.

6. The fusion polypeptide according to any one of claims 1 to 5, wherein the C-terminus of the antigen-binding domain consists of the amino acid sequence of VTV.

7. The fusion polypeptide according to any one of claims 1 to 6, wherein the antigen-binding domain comprises a VHH domain derived from a camelid animal, a cartilaginous fish, or a mammalian antibody or the heavy or light chain of a monobody.

8. The fusion polypeptide according to claim 7, wherein the antigen-binding domain comprises shark VHH, ray VHH, stingray VHH, or sawfish VHH.

9. A fusion polypeptide according to any one of claims 1 to 8, comprising the amino acid sequence of SEQ ID NO: 11 or SEQ ID NO:

12.

10. A fusion polypeptide according to any one of claims 1 to 9, comprising the amino acid sequence of SEQ ID NO: 001 or SEQ ID NO:

002.

11. i) A fusion polypeptide according to any one of claims 1 to 10, and ii) Target protein specifically bound to the fusion polypeptide A complex that includes this.

12. Use of a fusion polypeptide according to any one of claims 1 to 10, or a complex according to claim 11, for structural analysis of a target protein.

13. The fusion polypeptide according to any one of claims 1 to 10, wherein the polypeptide scaffold comprises an amino acid sequence consisting of amino acids located at positions 32 to 396 of the sequence shown in Sequence ID No.

9.

14. A fusion polypeptide according to any one of claims 1 to 10, comprising the amino acid sequence shown in SEQ ID NO:

11.

15. A fusion polypeptide according to any one of claims 1 to 10, comprising the amino acid sequence shown in SEQ ID NO: 12.

Citation Information

Patent Citations

  • Stabilized reverse transcriptase fusion protein

    JP2012519489A