De novo designed miniprotein agonists and antagonists targeting g protein-coupled receptors

WO2026192791A1PCT designated stage Publication Date: 2026-09-17UNIV OF WASHINGTON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/017383
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-14
Filing Date
2026-03-03
Publication Date
2026-09-17

Smart Images

  • Figure US2026017383_17092026_PF_FP_ABST
    Figure US2026017383_17092026_PF_FP_ABST
Patent Text Reader

Abstract

Polypeptides are provided that have an amino acid sequence at least 50% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-217, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to a GPCR target, fusion proteins thereof, and their use for treating limiting development of migraine headaches, pain, itching, obesity, and / or diabetes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] De novo designed miniprotein agonists and antagonists targeting G protein-coupled receptors

[0002] Federal Funding Statement

[0003] This invention was made with government support under Grant No. W81XWH-21-1-0891, awarded by the Department of Defense and Grant No. DE031436, awarded by the National Institutes of Health and Grant No. 2143160, awarded by the National Science Foundation. The government has certain rights in the invention.

[0004] Sequence Listing Statement

[0005] A computer readable form of the Sequence Listing is filed with this application by electronic submission and is incorporated into this application by reference in its entirety. The Sequence Listing is contained in the file created on February 27, 2026 having the file name “25-0402-WO.xml” and is 192,323 bytes in size.

[0006] Background

[0007] G protein-coupled receptors (GPCRs) are the largest and most diverse family of membrane receptors in the human genome and play critical roles in many physiological processes. GPCRs are implicated in a wide array of diseases including cancer, cardiovascular and metabolic diseases, and neurological disorders, and hence are at the forefront of drug discovery and development. Over the past decades, biologies including antibodies, nanobodies, and peptides have gained momentum as GPCR therapeutics and tools. However, the design of biologies modulating GPCR signaling remains an outstanding challenge, often requiring a combination of strategies such as the insertion of peptide fragments from native proteins or screening of random libraries. It has been particularly difficult to generate GPCR agonists, which has necessitated considerable antibody and receptor engineering efforts.

[0008] Summary

[0009] In a first aspect, the disclosure provides polypeptides comprising an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-217, not including any amino acidinsertions at identified loop positions, wherein the polypeptide binds to a GPCR target. In one embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to calcitonin gene-related peptide receptor (CGRPR). In some embodiments, the polypeptide comprises an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 11, 12, and 23 (i.e., dC2_049, dC2_050, and mC2_022.).

[0010] In another embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:26-32 and 213-215, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to the Mas-related GRP family member XI (MRGPRX1) receptor.

[0011] In a further embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:44-117, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to the glucagon receptor (GCGR).

[0012] In one embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 118-212, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to gastric inhibitory polypeptide receptor (GIPR).

[0013] In another embodiment, the polypeptides comprise an amino acid sequence at least 60% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions. In a further embodiment, the polypeptides comprise amino acid sequence at least 70% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions. In one embodiment, the polypeptides comprise an amino acid sequence at least 80% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and44-215, not including any amino acid insertions at identified loop positions. In a further embodiment, the polypeptides comprise an amino acid sequence at least 90% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215.

[0014] In various embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or all identified interface residues (i.e., residues at the binding interface with the relevant target) are identical (not substituted), or conservatively substituted, relative to the reference sequence. In one embodiment, at least 5 identified interface residues are identical (not substituted), or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 10 identified interface residues are identical, or conservatively substituted, relative to the reference sequence. In another embodiment, all identified interface residues are identical, or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or all identified interface residues are identical (not substituted), relative to the reference sequence

[0015] In one embodiment, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all identified core residues (i.e., residues present in the protein core) are identical, or conservatively substituted, relative to the reference sequence. In one embodiment, at least 5 identified core residues are identical (not substituted), or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 10 identified core residues are identical (not substituted), or conservatively substituted, relative to the reference sequence. In another embodiment, all identified core residues are identical, or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all identified core residues are identical, relative to the reference sequence.

[0016] In one embodiment, all interface and all core residues are identical relative to the reference sequence.

[0017] In another embodiment that can be combined with any other embodiment herein, the polypeptide comprises an amino acid insertion at one or more identified loop positions relative to the reference sequence. In one embodiment, the polypeptide does not comprise an amino acid insertion at any identified loop positions relative to the reference sequence. In some cases, the insertion may result in deletion of one or more loop residue. In another embodiment, the insertion may not lead to deletion of one or more loop residue.

[0018] In one embodiment, the disclosure provides fusion proteins, comprising:(a) the polypeptide of any embodiment or combination of embodiments herein; and

[0019] (b) one or more functional domains at the N-terminus and / or at the C-terminus of the polypeptide. In all embodiments, the polypeptide and one or more functional domains may be directly fused or may be separated by an amino acid linker of any length and amino acid composition as appropriate for an intended use.

[0020] In another aspect the disclosure provides nucleic acids encoding the polypeptide or fusion protein of any embodiment or combination of embodiments of the disclosure. In a further aspect, the disclosure provides expression vectors comprising the nucleic acid of any aspect of the disclosure operatively linked to a suitable control sequence. In another aspect, the disclosure provides host cells that comprise the polypeptide, fusion protein nucleic acid or expression vector (i.e.: episomal or chromosomally integrated) disclosed herein, wherein the host cells can be either prokaryotic or eukaryotic.

[0021] In another aspect, the disclosure provides pharmaceutical compositions, comprising: (a) the polypeptide, the fusion protein, the nucleic acid, the expression vector, and / or the host cell of any embodiment or combination of embodiments herein; and

[0022] (b) a pharmaceutically acceptable carrier.

[0023] In another aspect, the disclosure provides methods for treating or limiting development of migraine headaches and / or pain, comprising administering to a subject in need thereof an amount effective to treat the migraine and / or pain of a CGRPR antagonist or negative allosteric modulator polypeptide or fusion protein of the disclosure, a nucleic acid encoding a CGRPR antagonist or negative allosteric modulator polypeptide or fusion protein, or an expression vector, cell, or pharmaceutical composition thereof.

[0024] In one aspect, the disclosure provides methods for treating or limiting development of itching or pain, comprising administering to a subject in need thereof an amount effective to treat the itching or pain of an MRGPRX1 agonist polypeptide or fusion protein of the disclosure, a nucleic acid encoding an MRGPRX1 agonist polypeptide or fusion protein, or an expression vector, cell, or pharmaceutical composition thereof.

[0025] In a further aspect, the disclosure provides method for treating or limiting development of obesity or diabetes (including but not limited to Type 2 diabetes), comprising administering to a subject in need thereof an amount effective to treat the obesity or diabetes of an GCGR or GIPR antagonist or negative allosteric modulator polypeptide or fusion protein of the disclosure, a nucleic acid encoding a GCGR or GIPR antagonist or negativeallosteric modulator polypeptide or fusion protein, or an expression vector, cell, or pharmaceutical composition thereof.

[0026] Figure Legends

[0027] Figure 1. Characterization of MetaGen™-designed MRGPRX1 agonists, a 13,000 miniprotein binders were designed as agonists with MetaGen™ and tested using OPS-RD. The colocalization (binding signal) induced in cells with the same binder was compared to the colocalization distribution across all imaged cells (>2.5 million). P-values were computed using the Kolmogorov- Smirnov (K-S) test and adjusted using the Benjamini-Hochberg procedure, b Concentration-response curves of three agonist hits and the positive control agonist BAM (8-22) measured in a calcium flux assay (n=3). c Computational models of three agonist hits. Receptor structures are truncated for clarity, d Concentration-response curves of two optimized MRGPRX1 agonists in comparison to the initial hit mMl_034 and BAM (8-22) in calcium flux experiments in CHO-K1 cells overexpressing MRGPRX1. Data points represent the mean ± SEM of four to five independent experiments.

[0028] Figure 2. Characterization of CGRPR antagonists, a Computational design models of MetaGen™ mCl_023, mC2_022 and RFdiffusion dC2_049 generated antagonists bound to the CGRPR (PDB ID: 6E3Y). Receptor structures are truncated for clarity, b-d Concentration response curves of antagonists were generated in the presence of CGRP and their functional estimates (mean ± SEM) are mC1_023 (pA2of 8.3 ± 0.1 (5 nM) and slope 0.96 ± 0.04, n=4), mC2_022 (pA2= 7.9 ± 0.2 (13 nM), slope 0.61 ± 0.04, n=4) and dC2_049 (pA2= 8.4 ± 0.1 (3.9 nM), slope = 0.82 ± 0.05, n=4).

[0029] Figure 3. Distribution of secondary structure content of metaproteome-derived miniprotein scaffold library. The library is composed of diverse scaffolds comprising helices, loops and strands.

[0030] Figure 4. Calcium mobilization assay of binders at the MRGPRX1. a-d The ability of binders designed by MetaGen approach to activate or inactivate the MRGPRX1 was examined in a calcium mobilization assay in agonist and antagonist mode. The native BAM 8-22 peptide served as agonist in both agonist and antagonist mode. Vehicle (buffer) served as agonist negative control and antagonist positive control. Binders that display both agonistic and antagonistic activity are false positive antagonists, due to desensitization of calcium channels in the assay. Data are shown as mean ± SD of technical replicates from a single experiment.Figure 5. Biophysical characterization of MRGPRX1 binders, a Size-exclusion chromatography (SEC) traces, b circular dichroism (CD) spectra and c melting curves of MetaGen™ mMl_034, mM1_060 and mMl_064 binders targeting MRGPRX1.

[0031] Figure 6. Agonist screen of optimized MRGPRX1 hits. Binders were optimized via structure-guided site saturation mutagenesis (SSM) and their ability to activate MRGPRX1 was evaluated in a calcium mobilization assay in CH0-K1 cells overexpressing MRGPRX1. Data were normalized to 100% = 1 pM BAM 8-22, 0% = buffer controls. Data are shown as mean from two screening experiments performed in technical duplicates.

[0032] Figure 7. Yeast display of GIPR binders using soluble ECD. Representative flow cytometry plots of binders are shown. X-axis represents binder display and Y-axis target GIPR binding, a Designed library of binders without the target GIPR extracellular domain (ECD) represents a negative control. Yeast cells expressing binders on their surface were incubated with b 100 nM c 10 nM or d 1 nM of GIPR ECD. Data are shown from the third round of sorting.

[0033] Figure 8. Yeast display of GCGR binders using soluble ECD. Representative flow cytometry plots of binders are shown. X-axis represents binder display and Y-axis target GCGR binding, a Designed library of binders without the target GCGR extracellular domain (ECD) represents a negative control. Yeast cells expressing binders on their surface were incubated with b 1 nM c 10 nM or d 100 nM of GCGR ECD. Data are shown from the third round of sorting.

[0034] Figure 9. Binding and biophysical characterization of GIPR binders, a 96 GIPR designs from yeast library screening were purified. Binding kinetics of binders were measured using SPR. The resulting equilibrium dissociation constants (KD(M)) are plotted, b Computational design models, c SEC traces and d sensorgrams of the three most potent binders dGPl_015 (left), dGPl_035 (middle) and dGPl_040 (right), respectively, identified in a functional cAMP assay.

[0035] Figure 10. Binding and biophysical characterization of GCGR binders, a 96 GCGR designs from yeast library screening were purified. Binding kinetics of binders were measured using SPR. The resulting equilibrium dissociation constants (KD(M)) are plotted, b Computational design models, c SEC traces and d sensorgrams of representative binders dGCl_012 (left), dGCl_015 (middle) and dGCl_053 (right), respectively, showing strong binding affinity.Figure 11. One-point cAMP assay of RFdiffusion-designed miniprotein binders at the CGRPR. a-c Binders were tested at a concentration of 1-10 pM. Rimegepant (RG) served as a positive control (10 pM). Data are shown as mean ± SD (n=2).

[0036] Figure 12. cAMP assay of dC1_021 at the CGRPR. The concentration-response curve of dC1_021 was generated in the presence of an EC80value of αCGRP and increasing concentrations of the binder. The IC50value of the binder was 440 ± 40 nM. Data are shown as mean ± SEM (n=3).

[0037] Figure 13. One-point luciferase assay of miniprotein binders at the CGRPR designed by MetaGen™ approach, a-c Binders were tested at a concentration of 1-10 pM. Data are shown as mean ± SEM (n=3).

[0038] Figure 14. Concentration-response curves of miniprotein binders to the CGRPR designed by MetaGen™. Concentration-response curves of mC1_023, mC1_044 and mC2_022 were generated by measuring luciferase activity in the presence of an EC80concentration of αCGRP following pre-incubation with increasing concentrations of miniproteins in CHO-K1 / Cre-Luc / CGRPR cells. The IC50values for mC1_023 and mC2_022 are 37 ± 2 nM and 420 ± 60 nM whereas mC1_044 had an IC50value in the low micromolar range. Data are shown as mean ± SEM from at least three independent experiments.

[0039] Figure 15. One-point cAMP assay of miniprotein binders at the CGRPR following partial diffusion, a-c Binders were tested at a concentration of 1-10 pM.

[0040] Rimegepant (RG) served as a positive control (10 pM). Data are shown as mean ± SD (n=2).

[0041] Figure 16. Concentration-response curves of miniprotein binders at the CGRPR following partial diffusion, a-d Antagonism by miniprotein binders was determined by coincubation with an EC80of αCGRP and measuring cAMP accumulation. Data are shown as mean ± SEM (n=3-4).

[0042] Figure 17. Biophysical characterization of mCl_023, mC2_022 and dC2_049 binders at the CGRPR. a Size-exclusion chromatography (SEC) traces, b circular dichroism (CD) spectra and c melting curves of MetaGen™ mCl_023 and mC2_022 and RFdiffusion dC2_049 binders at the CGRPR.

[0043] Figure 18. Biophysical characterization and pharmacology of the dC2_050 binder at the CGRPR. a Computational design model of dC2_050 bound to the CGRPR receptor (PDB ID: 6E3Y). b Concentration response curves of αCGRP in the presence of increasing concentrations of dC2_050 binder. The calculated pA2value is 8.1 ± 0.1 (7.9 nM) and the slope is 0.90 ± 0.05 with R2 = 0.98. Data are shown as mean ± SEM (n=4). c Size-exclusion chromatography (SEC) traces, d circular-dichroism (CD) spectra and e melting curve of dC2_050 binder.

[0044] Detailed Description / Claims

[0045] All references cited are herein incorporated by reference in their entirety. Within this application, unless otherwise stated, the techniques utilized may be found in any of several well-known references such as: Molecular Cloning: A Laboratory Manual (Sambrook, et al., 1989, Cold Spring Harbor Laboratory Press), Gene Expression Technology (Methods in Enzymology, Vol. 185, edited by D. Goeddel, 1991. Academic Press, San Diego, CA), “Guide to Protein Purification” in Methods in Enzymology (M. P. Deutshcer, ed., (1990) Academic Press, Inc.); PCR Protocols: A Guide to Methods and Applications (Innis, et al.

[0046] 1990. Academic Press, San Diego, CA), Culture of Animal Cells: A Manual of Basic Technique, 2ndEd. (R. I. Freshney. 1987. Liss, Inc. New York, NY), Gene Transfer and Expression Protocols, pp. 109-128, ed. E. J. Murray, The Humana Press Inc., Clifton, N. J.), Dang, B. et al. SNAC-tag for sequence-specific chemical protein cleavage. Nat. Methods 16, 319-322 (2019), and the Ambion 1998 Catalog (Ambion, Austin, TX).

[0047] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise.

[0048] As used herein, the amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).

[0049] Any N-terminal methionine residue in any polypeptide of the disclosure may be present or may be deleted. In all embodiments of the polypeptides disclosed herein, 1, 2, 3, 4, or 5 residues may be deleted from the N-terminus and / or the C-terminus of the polypeptide while retaining activity.

[0050] All embodiments of any aspect of the disclosure can be used in combination, unless the context clearly dictates otherwise.

[0051] Unless the context clearly requires otherwise, throughout the description and the claims, the words ‘comprise’, ‘comprising’, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”. Words using the singular or plural number also include theplural and singular number, respectively. Additionally, the words “herein,” “above,” and “below” and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of the application.

[0052] In a first aspect, the disclosure provides polypeptides comprising an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-217, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to a GPCR target.

[0053] The polypeptides of the disclosure are de novo-designed antagonists, agonists or allosteric modulators of specific GPCR targets, as described in more detail below, have binding affinity and potency in the nanomolar or picomolar range, and can be used, for example, in therapeutic treatments described below.

[0054] The amino acid sequences of SEQ ID NO: 1-23, 26-32, and 44-217 are provided in Tables 1-4 below, together with an annotation immediately below the sequence of the position of helices (H), strands (S), and loop residues (L). The sequences are also annotated to show the position of interface residues (bold font) and core residues (underlined).

[0055] Table 1

[0056] dCl_021 (SEQ ID NO: 1) GTDIDLGVEYYFLALQALVFGDFESALKYATKAKEYFSKSPDKETAAKYLKLVDKVIDLATAQLAAKAA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL dC2_001 (SEQ ID NO: 2 )

[0057] DENIELATEYLMLARQALLHNSPEGAI RYAN KAI ETASKSSDKEKAQKI IDRAKKVI EQAKELAELLKN LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHLL dC2_007 (SEQ ID NO: 3 )

[0058] DLTLELAVEYEMLARQALLHNSPEGVKKYAKKAI ETAKKAKDKEKAKKVI ERAKKL I EEAEKLAELLAS LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL dC2_011 (SEQ ID NO: 4)

[0059] NTNLELGTEYFNQGLLALVHGSYEEAI KYFEKAI EYFAKI PDKEVSDKNI ALAKKYI EKAKELLAKAKA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL dC2_019 (SEQ ID NO: 5) NPNDELAVELYMEGLQAYVHGDFETAIEYFKKAIEKAKQGTNEKVKKAyTTNSEKYIAEAEKLLAERAA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL dC2_026 (SEQ ID NO: 6)

[0060] D INVEEASEYLMQGLQAIVFGDYEAAI EYFTKAI ELLDKS SDKETASKLKKTAQKYLALAKELAKKAKE LHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL dC2_030 (SEQ ID NO: 7)

[0061]

[0062] DTNLDLGVETLLQGIQALIHGDYELAI KKATEAI KYLEKSKDKEKAKKWIAFAEKVI KTAKALLAEKAA LHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0063] dC2_039 (SEQ ID NO: 8) DTNFELGVEYIMEGLLALAHGDYETAEKYFTKALEYLNKSPDKEKAQPWIDLAKKYLALAKEKIKEKKA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0064] dC2_042 (SEQ ID NO: 9)

[0065] DPNIELATEYLMQARLALLHNSPEAAI KYANKAI ETAKKSKDKEKAQKI IDQAKKVI EEAKKLAELLKN LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0066] dC2_045 (SEQ ID NO: 10)

[0067] DTNLELGVQEYMDGLQAYVHGSYELAI EKFKKAI EYFSKSKDKEKAEKYI KLANKYI EESKKILKEKEE LHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0068] dC2_049 (SEQ ID NO: 11) DTNFELGVEYFMLGLQALVHGDYDNAIKYFNKAIEYFKKSSDKEKAAKYIALAQKYIDEAKKLKAEKEA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0069] dC2_050 (SEQ ID NO: 12)

[0070] NDNDELAVQYYMDGLLAYVHGDYEGAI KYFNKAI EYAKKGTNEKVRTSVI SNSKKYI EEAKKLLAEKEA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0071] dC2_052 (SEQ ID NO: 13)

[0072] DENLEEGVLYLMQAALALVHGDAEGAI KYAKKAI EKI SKSKDKEKATKWLKFAEKVLEEAEKLLKEKKE LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0073] dC2_053 (SEQ ID NO: 14)

[0074] DPNVEEASEYLMLGLQAIVHGDYESAI KYFTKAI ELLKKSKDKETAAKLIKTAEKYLALAKELAEKSKK LHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHLL

[0075] dC2_055 (SEQ ID NO: 15) DINIELATEYYMLGLQALVHGDWESTVKYFTKAADTAAKSSDKEKAAKI ITLSNKKIAEAKKQLAAKAA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0076] dC2_057 (SEQ ID NO: 16) SELLEEMTEYYMLGLLALTHGDYETAIKYFEKALELAKQLPDPELREKRTTRITKYIEEAKKLLAEAEA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0077] dC2_058 (SEQ ID NO: 17)

[0078] DLTLDEAVEYLMLARQALLFNDPAAAKRYAQLAI ETAKQAKDKEKAQKVI ERAQQV I EQADKLAEELAN LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0079] dC2_063 (SEQ ID NO: 18)

[0080] DINIELATEYLMLARQGLLHNSPEAAI KYATKAI EIAKKSSDKEKAKKI I ERAKEVI KQAEELAELLKN LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0081] dC2_065 (SEQ ID NO: 19)

[0082] DINIELATEYLMLARQGLLHNSPEAAI KYAKKAI EVAKKSKDKEKADKI I KQAEEVI KQAEELAKLLAE LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0083] dC2_066 (SEQ ID NO: 20)

[0084] DLTLEEASEYFMLARQALVFNRPDLTI KYAQLAI ETAKQATDKEQAQKII TLSQQLI QQAQQLAQQLAS LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0085]

[0086] dC2_067 (SEQ ID NO: 21)

[0087] DENIELAVQYYMDGLLALAHEDYKEAI KYFDKAI EAAKASSDKEKGDKI I KLANKYKELAKKKLAEAEA LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0088] mCl_023 (SEQ ID NO: 22) QELEVRENARFyYQTLHYLGPLPLEKLKKILGLTDEQLEAALEYLKKLGRIKIEETPEKKVVSLV LHHHHHHHHHHHHHHHHHLLLS SHHHHHHHLLLLHHHHHHHHHHHHHLLLS SS S SLLLLS S S S SL mC2_022 (SEQ ID NO: 23) EEEEAMEELLAAKGDCEKMAEALARVLEVGDVGTQRLAYLYVHYTHPECSAKADEWAKHY LHHHHHHHHHHHLLLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLHHHHHHHHHHHHHHL

[0089]

[0090] Table 2

[0091] mMl_011 (SEQ ID NO: 26) MRERLYVSGLRNPFLALKVAEAVLKVPGVTHVRAFPLTQLVIVEHDGADLADIRAAIEAQGVTVDG LSSSSSSSLLLLHHHHHHHHHHHHLLLLSSSSSSSHHHLSSSSSSLLLLHHHHHHHHHHHLLSSLL mMl_034 (SEQ ID NO: 27) MPIEELVGRIRFAERAAWFAGLDPVEYTKEYIKEEFSEEEREKLLKAQKEGDPRMTPAQKEALKEL LLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHLLLLLLLHHHHHHHHHL mMl_059 (SEQ ID NO: 28) MKKLVIHVEGIHTYAEQWKLVDALMTLPGVASVHPWPRSGILTVHYDPSKVSEEEI INKLKEIGAKW LSSSSSSSLLLLLHHHHHHHHHHHHHLLLSSSSSSLHHHLSSSSSSLLLLLLHHHHHHHHHHHLLLSL

[0092] mM1_060 (SEQ ID NO: 29) MNEAFERALEEAVRAGMPRERAEYWARKLMLTDPFIKYEDLVKELKKLA LHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHL

[0093] mMl_063 (SEQ ID NO: 30) MDILQLKVSGLTSAADAAAVAQALLKVPGVIRVNVDLATQTASIVYGGKLDEAAVKAAVAAAGFTASP LSSSSSSSSLLLLHHHHHHHHHHHHHLLLSSSSSSSLLLLSSSSSSLLLLLHHHHHHHHHHLLLSSSL

[0094] mM1_068 (SEQ ID NO: 31) MSFEELAEEALEALKAGDKEKAKKLVEKLWEIAEKDKRHMKYFLFLYPFFELGLK LLHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHL

[0095] mMl_086 (SEQ ID NO: 32) MEHTTLHVGGIRNFVDMDKVAEALLALPGVTHVFVDLLTQTVSVWYDPAQVSEEELRAAVEAVGFEVA LSSSSSSSSLLLLHHHHHHHHHHHHLLLLSSSSSSSLLLLSSSSSSLLLLLLHHHHHHHHHHHLLSSL

[0096] mM1_034_F12W_Y27F (SEQ ID NO: 213 ) MPIEELVGRIRWAERAAWFAGLDPVEFTKEYIKEEFSEEEREKLLKAQKEGDPRMTPAQKEALKEL LLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHLLLLLLLHHHHHHHHHL mM1_034_F12W_A58M (SEQ ID NO: 214)

[0097] MPIEELVGRIRWAERAAWFAGLDPVEYTKEYIKEEFSEEEREKLLKAQKEGDPRMTPMQKEALKEL

[0098]

[0099] LLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHLLLLLLLHHHHHHHHHLTable 3

[0100] dGCl_001 (SEQ ID NO: 44) KVYRIRYSELQRVAEQIDAGESEELLATLDRVAEELATWSEEERERLTARAALVGLDAETYLAARRILIARQE A LLLLLLHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHH

[0101] L

[0102] dGCl_003 (SEQ ID NO: 45) LEELKEKAKELLRQIVFYAAMGRVDTPEYQEAMAEYDEVLKEIAEKTGKSVEEVEEEIIEEILEES LHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHL

[0103] dGCl_004 (SEQ ID NO: 46) MEKLKKEAKELLRQIVFFAVEGRVDEPRTQEAEARYDEVLKEIAEKTGKSVEEVEKEIIEEILEES LHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHL

[0104] dGCl_006 (SEQ ID NO: 47) LEELFKQAQETLRQLVALHIFGRVESEKTEEVAAKYDELVDE^SKKTGKSREE^EEEWQRILEEM LHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHL

[0105] dGCl_007 (SEQ ID NO: 48) VEELLEKAVDLLRTIAFFAVQGRVNDPRAEDAEVEYDEVLEEIAKKTGKSLEEVEKEIIELALKRS LHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHLL

[0106] dGCl_008 (SEQ ID NO: 49) ANQRAEEI_RKKMEEGKATVEELEELVDYYYFGAAVFPAKREELEEMAAEAERKIRELLEREREE LHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0107] dGCl_009 (SEQ ID NO: 50) ANQRAEEI_REKFEKGEATIEELEELIDYYYFGAAVFPAKREEYEEEAAKFEEAAKKKKEEEEKK LHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0108] dGCl_010 (SEQ ID NO: 51) VEQEAER^EKKMKEGTATIEELETLIDYYYFGAALFPKRREELEKKAAEAEEKKKELERKEKEA LHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0109] dGCl_011 (SEQ ID NO: 52) MNEELERLMKAVEEGTJETREEYETLEDILSFIQALFPKHREQAEEMIARAREVLARNEAEREAE LHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0110] dGCl_012 (SEQ ID NO: 53) MQQKLEEYEKKLEEGTATIEELELLEDYYYFGAAVFPAKREELEEKAAEAREAIEKKKKEEKEA LHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0111] dGCl_013 (SEQ ID NO: 54) SKEKLELFEEALRLLAGAAITGRPVDEIVDKLLEAGRKELAETLLRVAQSSWGEIKEELQKLKEE LHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHHL

[0112] dGCl_014 (SEQ ID NO: 55) SKEKLELFEEALKLLAGAAITGRPIEEVLDKLEEAGKLELAETLLEVAQMSWSEIKEKLHKLKEE LHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHHL

[0113] dGCl_015 (SEQ ID NO: 56) SQEELKKFEDVLKQIAFAAVTGRSIDDLIDKLEEDGDLEGAELLTEVSSMSWEEIKKKLEELKEK LHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHHL

[0114] dGCl_020 (SEQ ID NO: 57) SPEVERFREEVRRRLEEARKSGELEKVLEELKEKADRYYFAGAILGNELADDYAAA^EEAVAEAEREREEEER LHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHHL

[0115] dGCl_021 (SEQ ID NO: 58) SEEVEKFREEVKKEVEEAKKKGTIEETLKKLEEEADRYYFAGAVLGVEKADDYAAVIEEVIEEAKKELEEEKK LHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHHL

[0116] dGCl_022 (SEQ ID NO: 59) MEDVKKLNELIDSGGPKEEIVEAAERVEFTYAWLIASGRKELPKDLADALVRYDELVATDEELRETVERRREE RQ LHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHLHHHHHHHHHHHHH HL

[0117] dGCl_023 (SEQ ID NO: 60) SHEEEKLKKELKEKEEKIDRLYFQLAMNRALVNEERWLEIEEILRPEIIKHLEEAKELAKKLGDEEKAAELEE EIERLKS LHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHH HHHHHHL

[0118] dGCl_024 (SEQ ID NO: 61)

[0119]

[0120] SAERERLERELEELDEEIERLDFQMAMVTFMSSPETAAKKNDILLDQKMEALERRAEVCEELGREEECRELRRELEELRA LHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHH HHHHHHL

[0121] dGCl_025 (SEQ ID NO: 62 ) DELEKRIFGLVLMSYFADEARRERLDTAAYMYEELLELAKETGKEEEMLELADELNDTLQKSDAEFDAKYDEV KAKLEE^RKE LHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLLLLHHHHHHHHHHH HHHHHHHHHL

[0122] dGCl_026 (SEQ ID NO: 63 ) MTRLEKTEKLIEEMNKKVEEELKDPNKDKKEVLKKYLEEFEKYLVFFNITGREEEAELFLKQAEELEKKLEEL LLHHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHL

[0123] dGCl_027 (SEQ ID NO: 64) MSRLEKTKKFIEEQLKEVEEAKKDPNKDVKKVLEEKIEEFERNLVFFNIMGQEEEAELFLEAAEILEEELKKI LLHHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHL

[0124] dGCl_032 (SEQ ID NO: 65) NIAELEKIVEENKDKLSEEQKKKLEEVEKVILFENMVRGVLSDRAADYVEEKLKE^VEELKKKEE LHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHLL

[0125] dGCl_033 (SEQ ID NO: 66) SWEELEKEIKEI^KEEGDEEEKKELEEMEKAVLFDLMMRGQLSDRVIDMIEETIEKLKEKREKRKE LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHLL

[0126] dGCl_034 (SEQ ID NO: 67) NIEELEKEIEENKEYLSEEDKEKLDEIDKIYVFELMMRNQLSEKAEDYIEKKLKEIVKRyKERKK LHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHHL

[0127] dGCl_035 (SEQ ID NO: 68) NIEELEKEIEKNKEYLSEEDKKKLDEIDKIYVFELMIRGQLSEKAENYIEEKLKEVVENI_EKRKK LHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHHL

[0128] dGCl_036 (SEQ ID NO: 69) DEVKKELEELSKYLPEDLQKKIKEIPDEAYEEVKEAVKEYLKQEVFFAIMGRPLDSDEIVEAREKLEELIEKI LHHHHHHHHHHHHLLHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHLLLLLLLHHHHHHHHHHHHHHHHL

[0129] dGCl_037 (SEQ ID NO: 70) SEIREETKRLARLFDESLVKKWSLSEEEQKKVLELLKECLKLRLFLTMMGHPEDHPDLVECENKVLEAIESA LHHHHHHHHHHHHHLHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHL

[0130] dGCl_038 (SEQ ID NO: 71) KKKELLKEILHTYVWAAITGSDERSYDRIVELAEKLAELLKEQGKEEEAERVEKAIASGDITVIVELAEEMLE ELEEE LHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHLLHHHHHHHHHHHHH HHHLL

[0131] dGCl_039 (SEQ ID NO: 72 ) SVTKQLYELCRKGDYDGFEELFEKAKEKGELSEEyiKLYEEYLKKLMFFNMTGNQEGADKAWDELCKKLEAMV LHHHHHHHHHHLLLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHL

[0132] dGCl_041 (SEQ ID NO: 73 ) SPVDELAKRREELEKRVKEEJKDEKLHEKLDYLDFQVAMQLAAGNLEGAEKLEDERWELIEEAVKKLEEEKKK LHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0133] dGCl_042 (SEQ ID NO: 74) SARRAELERLERIEDYFYKNAKDQKKLDEAERLEFQAAMLLAMGNEEGADKLWEQRVKLMVEEVEKMEAEKAK LHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0134] dGCl_043 (SEQ ID NO: 75) SPVAEAAARRDRVEEEFERLAEDEELVERVERLDFQAAMLLAMGQEEKAEQVEEERFRLMEEFVERKRREREE LHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHHHL

[0135] dGCl_044 (SEQ ID NO: 76) MVTWEEYEEARKKLLELLKEEGDEEGVKEVEEDQEKLDRLYVWAEMTGSEKARKALDELMQKQYEKYKKKIEE LEKKREEKEK LLLHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHH HHHHHHHHHL

[0136] dGCl_045 (SEQ ID NO: 77) SAKRREIYDWFEERLKECLSEEVVDRLYFQVALSQAMGRERTLPEALRDHIEELDEESAKCAEELIEEYERRM AAV LHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHH HHL

[0137] dGCl_046 (SEQ ID NO: 78)

[0138]

[0139] SPEVRKIYDEFMEKAKEEFSQELLDKAYFRQAMLHAMGRDVSFEDVMLDVCEEEGEEKLKKCKELVEEYEKRLKEV LHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHH HHL

[0140] dGCl_047 (SEQ ID NO: 79) MTFEEVEKKLEENWDLLTEEEKEKIDRCYFELALLEATLGAGPDMSPEAREKLDQKIDQCIELMKEI_LEKI_EE KKK LLHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLLLLHHHHHHHHHHHHHHHHHHHHHHHHHHH HHL

[0141] dGCl_050 (SEQ ID NO: 80) VEVELPKADASDEELEEIARRQPTEEAAERAYFQCAMLLAMGREEEAEVACELADRIEEILREKEEEE LLLLLLLLLLLHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHL dGCl_051 (SEQ ID NO: 81) SEIVLPKPEASDEELERLAREIATEENIERVEFQAAMLLAMGREEEAEVALYLADLMRRYLEERKAAE LLLLLLLLLLLHHHHHHHHHHHLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHL dGCl_052 (SEQ ID NO: 82 ) SEKEKELREELDREVFKIVLRSTISGKPLDEQQDEYLDILEKKDPELAEKLLGYLRCETEEECEAYVKEKVEE E LHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLHHHHHHHHHHHHLLLHHHHHHHHHHHHHH

[0142] L

[0143] dGCl_053 (SEQ ID NO: 83 ) MKQREKELAEELERLELWAVLTGRELEELLERLPREKAIEMLEAAIKLYEEKKKKGENPDTMEEHIEYYRRLL ERLKE LLHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHLLHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHH HHHHL

[0144] dGCl_056 (SEQ ID NO: 84) SWEELEELLKKLKEIGSEEKVKEAEEYIDRTYAFAALRGRLTDRDLDYI IEKLKKLIEEEK LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHL

[0145] dGCl_057 (SEQ ID NO: 85) SWAELEELLAKLEEVADKETVEKVKEEIDRTYAFAALRGKLTDRDLDYLIELVKKKLEEXK LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHL

[0146] dGCl_058 (SEQ ID NO: 86) SVDEMMADLLEEYLEAPEEERPALIEKAERLLVGLIILGHPQEEIELATRLLDLMEAGDDEGVRAFLEELK LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL

[0147] dGCl_059 (SEQ ID NO: 87) SIDEEALRLARRALAAPPEERPA^IEE^ERKIVGLIITGHPMEDIGLFWLLDRyEAGDDEGLRAFAAELE LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL

[0148] dGCl_060 (SEQ ID NO: 88) SFDEYVAELLERYLEAPEEERKKLLEEAERSAVGLIWGAPEEKVALLNLVAELI_EEGDEEGI_RRLLEELK LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLHHHHHHHHHHHL

[0149] dGCl_061 (SEQ ID NO: 89) SDREFLAEMAERALEAEEEERPKyCEELDRQTVGAIIIGKPQEFVDELAFLADRCEEGDDEGyRRFLEELK LHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL

[0150] dGCl_062 (SEQ ID NO: 90) SYNEYLANLAEEALEAPESKRKEMAEKLDRDTVGCIIIGKPQEVCDELAYLAELyEEGDREGLEKFLEELK LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL

[0151] dGCl_063 (SEQ ID NO: 91) KKEEQIKLLEEAIKKFESDPNNIVEYIEEFEKKDKEVGQILKQAMMAEILWGKSIVDLLEKNIEYLKK LHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHL dGCl_064 (SEQ ID NO: 92 ) MKEKQIKLLKEYEKKIAEEPDWLEVIEEAEKEDKEVADILKSALMAEALLGRDFSDVIDKNIAALEK LHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHL dGCl_065 (SEQ ID NO: 93 )

[0152] SKLLEEYI KEAVEAAEKGDEEALEEI I KKMKEALEKGEFSEEDKKOEEyi ELIDRLFFQVALVTFQSNNEEK IDKLWKDAAEELKKM LHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHHHHHHHLLLHHH HHHHHHHHHHHHHLL

[0153] dGCl_066 (SEQ ID NO: 94) SSHYEKELEEVTKELREAIKKGDEELVLTIDRLIETHSPSFLEEAIKKLSEEEQKKVKEyLKQLVFLAIQGKS NPAQELIDRLK LLHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHLLHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHLLLL

[0154]

[0155] LHHHHHHHHHLdGCl_067 (SEQ ID NO: 95) MVSREELAAAVEDLLARLDELPLELARQEVAGLYEELLLLEGVDEDRASFQATMLYTNLRTREEjERRTLERML AHLKKE LLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHLLLLLHHHHHHHHHHHH HHHHHL

[0156] dGCl_068 (SEQ ID NO: 96) EEELEELIKE^EEDIERFDELSLEECKTLVAYHLRKLLLLKGVPEDEADFRAAMLVATTRTKEECLERLKKLL KEAKSL LHHHHHHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHH HHHHHL

[0157] dGCl_069 (SEQ ID NO: 97) KEELIEELEERSDRLYLHAAIVRPGDEALDALAWTYEEMAELLREGNMEGLREWEEALEYFKTYDEETARRY QSDIREALEKIEEA LHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLHHHHHHH HHHHHHHHHHHHHL

[0158] dGCl_070 (SEQ ID NO: 98) MSDIERAELLLRQDIEDGVFSPELTECLRRLLERGISPEFAEEYLRLHFQVAMTLAQDPSKTERAEDYMRERL LELCREELER LLHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHH HHHHHHHHHL

[0159] dGCl_071 (SEQ ID NO: 99) YSDIERAELILRQDIEDGVFDPETTELIKKALEKG^SEELAKK^LELHFKVAMTLAQDPSKVPEQEEYARKEI HELLKEELEK LLHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHH HHHHHHHHHL

[0160] dGCl_072 (SEQ ID NO: 100) SREEELRAARKVTKTVLELLEENDLETTCERLEFLAAMQIAMGRDVKAEHLLHFADLCREAIKEGKSKEEVKE RVREWLERLEEE LHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHLLLLHHHHHH HHHHHHHHHHHL

[0161] dGCl_073 (SEQ ID NO: 101) SMEEEKETIEKVLKTVEELLKEYSQEEARDRLYFLSAVQTAMGRKLKAEYLGFAADKLEEAIKKGLSKEEALK LIKEEFEKLLKE LHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHLLLLHHHHHH HHHHHHHHHHHL

[0162] dGCl_074 (SEQ ID NO: 102 ) SKAEEERAIRKVLKTTRELLEELSPEETAERLEFLAAVQIAMGRETLAEHLLAAADLVREAIAEGASREEILR RLEEWAAEEIAA LHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHLLLLHHHHHH HHHHHHHHHHHL

[0163] dGCl_075 (SEQ ID NO: 103 ) GEEEERRRRIKALETACELLEENSIEEAIDRLYFQAAMQEAMGRTDMAEELREVADWLEEQLKEGKSKEEAIE ECRKRIEELRRE LHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHLLLLHHHHHH HHHHHHHHHHHL

[0164] dGCl_078 (SEQ ID NO: 104) SVEELLRELRELAERIDRLEFQRAMQIAMGSPNVELVEDELLEAQRRFVEVAEELPEEALTEEQRERLEEARE ALARAA LHHHHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHLLHHHLLHHHHHHHHHHHH HHHHHL

[0165] dGCl_079 (SEQ ID NO: 105) SNEKLKEEFEKVKEEIERAEFKLAMLEALGSPRVELAEDEIYELRKKAVEIAEKLPEELLTEEEKEFIEEAKK EIEEYK LHHHHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHLLHHHLLHHHHHHHHHHHH HHHHHL

[0166] dGCl_080 (SEQ ID NO: 106) DEKKEE^LKE^KKYCEEAGKTLVFFNIVGRQEEAGLWGMVYDECEE^LKYMEENYDELTVEEAEEMLEKVKKL IEEAEKEIEE LHHHHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHLHHHLLHHHHHHHHHHHHHH HHHHHHHHHL

[0167]

[0168] dGCl 081 (SEQ ID NO: 107)DECLKEYLEELKKLAEEAGKTLVFFNIVGRTDEAGLAASVYDELEEMVKYMEENFDEySCEEAKEMLEKAKKL IEEMMKELKE LHHHHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHH HHHHHHHHHL

[0169] dGCl_083 (SEQ ID NO: 108) DEERKQWELEQCQQEAERTGSELLRLECELLEKDYDAYERIQFQAAMLEAMGVPQDEVEDRTIELjEEEKLKE LHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHL

[0170] dGCl_084 (SEQ ID NO: 109) TEEVEKEVRELLERRDRLIMWQSITGVSEKAEEEIDRLLRRAIELAKENPDLSRELLERLAEEAELFGYKEEA EELRKMLE LHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHLLLHHHH HHHHHHHL

[0171] dGCl_085 (SEQ ID NO: 110) TEEVEKELRELLEERDRLLMWLAIHPESQKALDQIDELLRRAVELAKENPDVSEELLRELLEEAELFGYKEEA EELRKMLE LHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHLLLHHHH HHHHHHHL

[0172] dGCl_086 (SEQ ID NO: 111) MTLEELAEDALELLDEIGSLEEVRDRLEFRAAMFIAQGREEEAEWYLTFADILAEHIESGDLEAAKKAIEELV EE LLHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHH HL

[0173] dGCl_089 (SEQ ID NO: 112 ) SRCQELLERIDKVDFQLALLDAHRNPNKEEKEQKLVDEFLELVDELEACAAEEEISPEQQAKLDEYRAEKAVY EAAKA LHHHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHH HHHHL

[0174] dGCl_090 (SEQ ID NO: 113 ) SPEEARRREVEELAEEMVDWYMRCTAFGTMVGADYDDRCRDVFRERXKKLPAETRRALLEVLRERAEEDKSFR ADFYREMAAMLEEAL LHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHH HHHHHHHHHHHHHHL

[0175] dGCl_091 (SEQ ID NO: 114) SHEEEKLKRAEERAEEMIDYLMRCSAFASIMGPEIEDRCLDVJRKE^AELPEEERKLLLEVLKRKAKEDSSVR SDYARELAE I I EE VL LHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLLLHH HHHHHHHHHHHHHHL

[0176] dGCl_093 (SEQ ID NO: 115) GSEKKRKEELEKLVEEKVEYLMRCYAFSTLVGDDVLDRCLDVLREELKKLDKETREEMLKILKEKAKKNKSVE QDAYEELAKLIEEVL LHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHH HHHHHHHHHHHHHHL

[0177] dGCl_095 (SEQ ID NO: 116) SAEEAFVKEyAERILEEVAKEQGAPVELVSLTESLEELKEILEAyERGEERTLEVSIFNPFTLKDYYVEyPAE RIRAIAEEVRRERA LHHHHHHHHHHHHHHHHHHHHHLLLHHHSSSLLLHHHHHHHHHHHHLLLLLLSSSSSSSLLLLSSSSSSSLHH HHHHHHHHHHHHHL _

[0178] dGCl_018 (SEQ ID NO: 117)

[0179] SEWRKKELEEFEEYVEKKLKEGKVEEAEEEAKERVDRLYFGAALFGSEELADLAVDIEEILEKKLKELEEKKK

[0180]

[0181] LHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHHHL

[0182] Table 4

[0183] dGPl_005 (SEQ ID NO: 118) VPERRRRLATAEIYLEAERAGVEISVEELEYLVDIATRTGLAFDRSEEELLAAWPRVEARLAAQRAA LLHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHLLLLLLLHHHHHHLHHHHHHHHHHHHHL

[0184] dGPl_006 (SEQ ID NO: 119)

[0185] GAEEQLVRDSAELVLELYEAGKEPAPHPWTLEDVAAAMRVIELLAARDPRALELPEDFVLNVVVEVRKKY

[0186]

[0187] LHHHHHHHHHHHHHHHHHHLLLLLLLLLLLHHHHHHHHHHHHHHHHHLHHHHHLLHHHHHHHHHHHHHHLdGPl_007 (SEQ ID NO: 120) VTPEEVEAARDDPEALEAVATRGLAHVIADPGTTFEDFVLALQAVDVAREVLGPEAVERILARADELA LLHHHHHHLLLLHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHL dGPl_008 (SEQ ID NO: 121) KKEKLEKVLKLLEIAIEKINGNFYDIAESLIAMEKIKELSPESFELIATTPLEELKKNKKEILEEAKKIYE LHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLHHHHHHHHLLLHHHHLLLHHHHHHHHHHHHL

[0188] dGPl_009 (SEQ ID NO: 122 ) AAAKQAKIEAEKAKAVELLAKALVGTASFEETVEGILHVVAIKEAGAPEAVVQAALDAADAKAKELA LHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHL

[0189] dGPl_010 (SEQ ID NO: 123 ) AARYEALRKLAEEAILASRRGTLEEYVTAIQALFLLEAGDFEGAEAKLLELAKKRGTYEEMKKKIEAA LHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLHHHHHHHHHHL dGPl_011 (SEQ ID NO: 124) AAEREALLALAEEFILEAQRYTFEDFVTAIQARFLLEAGDVEGARALLLELARRRGIYEEVRARIEAA LHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHLLHHHHHHHHHHL dGPl_012 (SEQ ID NO: 125) ASAARIAAEKASVRAAVRELGSSAEDVLEFVIGLRNYLTPEQLQAYIAYLEEEDPELAAKARAVV LLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLHHHHHHHHHHL

[0190] dGPl_013 (SEQ ID NO: 126)

[0191] MSEEI IRRVLELLREGDEEGALRLVRESGASLEDVDRILLLAAESGDFLVFARALYLYGRYRAELEAAA LLHHHHHHHHHHHHLLLHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHL dGPl_014 (SEQ ID NO: 127) GLEEALAEARALLARVKALAEAGELTLESALELFKEWVDFYWSLKEFTPELEAVLRAALEHIEAALKRLA LHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHL dGPl_015 (SEQ ID NO: 128) GIRAEARSFVRGLLSLPADEAAAALRERILAADEELAVLLMIEALKAMLYLSPEAQAKVLAAIREAATARS LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHL

[0192] dGPl_016 (SEQ ID NO: 129) LREHYRERAREALGLSPEERRAAIDELVAEAVQHPDGAELILAMVEVFWDAGDEAGLRAVWEALRRHLAA LHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHL dGPl_017 (SEQ ID NO: 130) AAAAARARTAERVGGLLRELGYEDAAAVFEAALAAGLPAVEIMVRLIRHLAAQGVPLEQAEEVLRLAAEELDR A LHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHH

[0193] L

[0194] dGPl_018 (SEQ ID NO: 131) ALPLVEEARRLAELSPEERAALAERLIAAVAADPGLNGARALVAFLRTLSYLPPETRAALLARLEEVMATLA LHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHL

[0195] dGPl_019 (SEQ ID NO: 132 ) LRAAVAAEIPDPEFAAAVYDTTATLAAAGVPPTEIMVAVLRVLAELVDRGLPLEAAEAAFRVAWAYLEA LHHHHHHHLLLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHL dGPl_020 (SEQ ID NO: 133 ) RLAEVVERLRGLSPEERRELLVREMVEWAREGVPADEVMVLAIRVLAELEVPTLEEVEETLRAWEALEAL LHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHL

[0196] dGPl_021 (SEQ ID NO: 134) RLQEWIERLQPLSPEEATAALVAEMVALVEEGVPPDEAWYAVRVLAELEVPTLELAEAALRAWEALEAA LHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHL

[0197] dGPl_022 (SEQ ID NO: 135) GLYERLHEAVLELARYAGDEKTVARLEALKSKPPEEQAVEGFVLALRVAARDGTPEARALLERVLALARELGE AQ LHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHH HL

[0198] dGPl_023 (SEQ ID NO: 136) GALEELERLREEQTALVREAADRDPALAPYAMVQAVLAFSEFAEESGLYEEAEAALDEVLREGAAYAS LHHHHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHL dGPl_024 (SEQ ID NO: 137) AAEARETLERSLAAIAEIDPEIAEEARVIAYGDDEEAVRALVRIFRRLSRLTPEEREAVLAELRRIIEAKERL LHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHL

[0199] dGPl_025 (SEQ ID NO: 138)

[0200] REAVIAALRAAAAASGPAEATLLRVAVEALEAGRSPVEAGVAAIRTAAELSGLTNEESLRRVLALIARAEEE

[0201]

[0202] LHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLSLLSLHHHHHHHHHHHHHHHHLdGPl_026 (SEQ ID NO: 139) AAEELLAAALRRLAAIEEAHPATAHLTVRAGVRLILLVSELREAGASPAEIEAAVDAWAATEAAAAAALAA LHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHHHHL

[0203] dGPl_027 (SEQ ID NO: 140) AATLARp’ETFIATARALAEGRVTLEHVRAFIEAAEAYRELVGTPDAEEEVRALIAEAEEEVAAEL LHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHL

[0204] dGPl_028 (SEQ ID NO: 141) AETEARARETFRRTARALALGRVTLADVRAFIEAAETYRRLVGTPDAEEAVRALIAEAEEEVRAEL LHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHL

[0205] dGPl_029 (SEQ ID NO: 142 ) SAELEELEERLREAARKFIEGKVTTEDVLEFIEAAERLEELLGSEEETRRRIQE^EAEVRAEVAAEKA LHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHL dGPl_030 (SEQ ID NO: 143 ) SAELEELKEELREAIRHFVLGDVKTEDVLRFLRAAARIEEVTGSEEEARRIVEETEKEVREEIEKEEK LHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHHL dGPl_031 (SEQ ID NO: 144) EAKEIIAKANKEGALAMIEIVRAHPEVSDPEAVYLVGLAAGLSVSEAIYLSETISLEEVERWLRE LHHHHHHHLLHHHHHHHHHHHHLLLLLLHHHHHHHHHHHLLLLHHHHHHHHHLLLHHHHHHHHHL

[0206] dGPl_032 (SEQ ID NO: 145)

[0207] RAEEI IARADRAGAEAMLAIVRADPSVSDREAVFLTALAAGLSFSDAIYLSETVTLAEVEALLAA LHHHHHHHLLHHHHHHHHHHHHLLLLLLHHHHHHHHHHHLLLLHHHHHHHHHHLLHHHHHHHHHL

[0208] dGPl_033 (SEQ ID NO: 146) RAEELISSADPEGVTAMIELVRADPSVSEEEGVFLAALAGGLSFSDAIYLSETVTLDEVERWLAE LHHHHHHHLLHHHHHHHHHHHHLLLLLLHHHHHHHHHHHLLLLHHHHHHHHHLLLHHHHHHHHHL

[0209] dGPl_034 (SEQ ID NO: 147) EEEELVKKVVDAMKAKGDTETADKVEKLLTEGNSFEKVIGIVLATKAAKELGLDVLPDIQKLLDLLVG LHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHL dGPl_035 (SEQ ID NO: 148) GEKEKLRKEALEKAKELGLDVESFYKAAVALGTIEAVLGAIIYYKLIKDPSISPEELKILLKVLEAIEKRLKA Q LHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHLLHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHH

[0210] L

[0211] dGPl_036 (SEQ ID NO: 149) AEREELIARAEELVATITEEQLTELALTARTLEEFVGAILLYKLKTDKSITAAEALELNPEALTAGVELAERT LG LHHHHHHHHHHHHHLLLLHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLLLHHHHHHHLHHHHHHHHHHHHHH HL

[0212] dGPl_037 (SEQ ID NO: 150) AEREELLARAEALVARVTEEQLTELALTARTLEEYVGAILLYRARTDASVTVAELFEENPEALRAGVELAERT LG LHHHHHHHHHHHHHLLLLHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLLLHHHHHHHLHHHHHHHHHHHHHH HL

[0213] dGPl_038 (SEQ ID NO: 151) AERAEVLARAAALVATVSEEQLTELALRARTLEEYVGALLLYRARHDPSVTVEELYDENPEALRAGVALAERT LG LHHHHHHHHHHHHHLLLLHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLLLHHHHHHHLHHHHHHHHHHHHHH HL

[0214] dGPl_039 (SEQ ID NO: 152 ) AEREELVRAFREVVETVSDEDLTRLALSARTLEEYVGALLIYRARHDPSVTVEELVEENPVAARAGVELARTT LG LHHHHHHHHHHHHHHHLLHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLLLHHHHHHHLHHHHHHHHHHHHHH HL

[0215] dGPl_040 (SEQ ID NO: 153 ) AEREEILARAEEVVASVTEEQLTELALNARTLEEYVGALLIYRAKFDPSVTVAELVDENPEAVAAGVELAEKT LG LHHHHHHHHHHHHHHHLLHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLLLLHHHHHHHLHHHHHHHHHHHHHH HL

[0216] dGPl_041 (SEQ ID NO: 154) SAAREAEIAAIRAEMEARIARAKEELKGLRGLQNTYAAIAYWAAALRFEELGEEELAAEYRAKAEEALA LHHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHL

[0217]

[0218] dGPl 042 (SEQ ID NO: 155)GEKEELAE KAREL VKKALEAGKEGSLDWFYIIEAVGYVRKLAELDPELAAEVMKEVEEAYQKAVG LHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHLL

[0219] dGPl_043 (SEQ ID NO: 156) SDAEVAAEAEALIKEAVAEHVAAGVPRSFAILLAIIDAIRRAKEEGKEAVAARLEAMSLEEAGAIARAAYEEI SA LHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHLLHHHHHHHHHHHHHHH HL

[0220] dGPl_044 (SEQ ID NO: 157) MEEEVDREAEELIREAVRELVEEGVPRGIAEELAIILAIERARREGREDVAARLEAMSLEEARRIARRAYEEI TA LHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHLLHHHHHHHHHHHHHHH HL

[0221] dGPl_045 (SEQ ID NO: 158) AAAEVDARARALIREAVKEHVAAGVPRSFAVLLAIIEAIRRAEEAGDEAVAARLRAMSLEEAGAIAREAYEAV SA LHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHLLHHHHHHHHHHHHHHH HL

[0222] dGPl_046 (SEQ ID NO: 159) LRAEAEALLAKIREADPEAYAEIWRLFSTGSSVDLAIASWAWRYALEAKAAGHPIAPVYEQLAEVLLQLAIA LHHHHHHHHHHHHHHLHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHL

[0223] dGPl_047 (SEQ ID NO: 160) EEEERALKEYFEKRGYYGLYVLVSTAAKLDGISYAEAVALAIEVLQSAPEEVLEALREARELLEKV LHHHHHHHHHHHHLLLHHHHHHHHHHHHHLLLLHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHL

[0224] dGPl_048 (SEQ ID NO: 161) GEEEEEELLERGFRLAYERAARLSEEERALLVAGHLLHYWGYRSPWAQEVRRAGVEVYKELVAREEAEAA LHHHHHHHHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHHHHHHL dGPl_049 (SEQ ID NO: 162) REAVIAYLRARREELPPEVAELVDVAIAELEAGKSPVEAAVAVIRRAAELGKDRETVDTVLRLLAEAEA LHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHL dGPl_050 (SEQ ID NO: 163) SEKAKKLLEEFKEAIENGDYETAKKHLFELYKGRNNFDGLVGFLIALDYLKEAGNQKLIDYFFETLAKEEAE LHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHL

[0225] dGPl_051 (SEQ ID NO: 164)

[0226] MEEVREARERILEMIESGEIDPEEFRRLVEI IARERSTFEGLAAYLEVLLHLKEAGHEEEFELALEIADEV LHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHL

[0227] dGPl_052 (SEQ ID NO: 165) MEEVRRARDAILAALASGDIDPEEFERNVAVVARARHTFEGLVAFLEILRAAREAGAEEEFARALEIADEV LHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHL

[0228] dGPl_053 (SEQ ID NO: 166) EAEALLAETRAALEAARDDLAAAARVLARSLARAVALGNFEAISETLYLAGEVAAHFGLSPEEFAALLDAAA LHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHL

[0229] dGPl_054 (SEQ ID NO: 167) QSARARALVEGLRAATDTEERVALVRELLALPESAFPGFLELAEVLQTLAEAAEEDPAVREAIEREIA LLHHHHHHHHHHHHLLLHHHHHHHHHHHHLLLHHHLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHL dGPl_055 (SEQ ID NO: 168) LSPEERLYEAALRALDRSDDEYVALVRELARVAFGSGSFRGIALLMDLLARAKERDPALAERLLAAIDEVLAE LA LLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHH HL

[0230] dGPl_056 (SEQ ID NO: 169) AAAEAAARAARKEAAVAEGEALARALRYAGVYGTFEEFVAHVLAVERLLAAAEAEGLEEAVAA^RAALDASSW AE LHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHLLL LL

[0231] dGPl_057 (SEQ ID NO: 170) EAKELLAQLEVHKDDPAKRLEILIQLLVLAVSTGSLNFEELSRVLGLLAEVGAELGLSPEQVAALADAAQAA LHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHL

[0232] dGPl_058 (SEQ ID NO: 171) GAEEAEREAEEARRRAKELAARLVANSSNFLALAAALAEIEEYYRKLSPEGQEAFLAVLDATPLP LHHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHLLLL

[0233]

[0234] dGPl 059 (SEQ ID NO: 172)SPLEKKVKELEKEMLKNAENEEKLVAILAELVKLALKDFEGFALAMEALEKLSKVLSKEKVLRLLELAEK LHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHL dGPl_060 (SEQ ID NO: 173 )

[0235] MTVEEWTGI IELVRDPATPLEELVALVRAAGEVALRASFEETVAVLQAIDEAMKIDPEKLRAALEAAE LLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHLL dGPl_061 (SEQ ID NO: 174) MTVEEAVEGIVALVADPATPLEELVALVAAAGEVAVRASFEEYARILQAFDEARKIDPEKLDAALEAAR LLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHLL dGPl_062 (SEQ ID NO: 175) MTVEEWAGIVALVADPATPLAELVALVREAGEVALRASFEEYVAVLQAFDEARRIDPAKLDAALEAAR LLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLHHHHHHHHHHLL dGPl_063 (SEQ ID NO: 176) AKLAELVAAFRAETDREAQIALLVELLKEALKSKKFEDVAAYLNAVAELREEAGLSSEDVIAAANAAD LHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHL dGPl_064 (SEQ ID NO: 177) IAEKAREVAEKWATGTATEEELKALGEALASFLGVAAVLQVLEEYKPDPEKFETLLKALDEAIALR LHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHL

[0236] dGPl_065 (SEQ ID NO: 178) DLEKYAEEKIREVFEKFKDQELSPEALQLLVDLLKRTSEIKEEAIKKGDYEEGKKKIEELAEEAIKEIEKIL LHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHL

[0237] dGPl_066 (SEQ ID NO: 179) ELDEFAEEQIRKVFEKFAGKTLSPEGQQLLVDLLKETSEIRNEAIRSGKLEEGKKEIEKIAEEYIKKLSEIL LHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHL

[0238] dGPl_067 (SEQ ID NO: 180) ALEAHAEAEIRRVFSRFAGQELTPEGQKLLVDLIKETSRIVEAAKRSGDLEGGKAEISAIAEEYIARLEEIL LHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHL

[0239] dGPl_068 (SEQ ID NO: 181) VEEAREAIREAFRRGASPEEIVALVRELAARDPSRAHEIAVEWRVASLEYPGGLEEKEALIDAAFRAIIEVE EA LHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHH HL

[0240] dGPl_069 (SEQ ID NO: 182 ) SRAAQLAAIEA^IRDTLTKHADNADPRVTTLAAQMFLAVGREPDLADAVAAVTALAAEVDALAAAL LHHHHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHL

[0241] dGPl_070 (SEQ ID NO: 183 ) AEKEKKFKEYRKAAEDALTEALQKRPLSLEETLFAIKSLEKLKKAEEAGDLKEAEKISKEAVDFAKSLL LHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHL dGPl_071 (SEQ ID NO: 184) GELARLIEEAREVAARIEQGDAPLPANLPFRALVHGLRIYRLMKEEGLDLDAATAVVAAE^EREVAA LHHHHHHHHHHHHHHHHHHLLLLLLLLLLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHL

[0242] dGPl_072 (SEQ ID NO: 185) AAAAAAAAEAERAKREAAREAYRALARGLTGTTRDVLAGLALAERVRKEHGLTPEELLALLREEEAE LHHHHHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHL

[0243] dGPl_073 (SEQ ID NO: 186) SAAALQAEYERYLALAEEELRRAREEGASRLHLVYATIFYYEAAKLAEKLGADPAFIAALRARAEAAWAA LHHHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHL dGPl_074 (SEQ ID NO: 187)

[0244] AEGLAVALAAGDRVIADPGATFEDVLAWEALRLAREEGVS PEELVERMAPHAKLAAAI REVLEA LHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHLLLLHHHHHHHLLLLHHHHHHHHHHHHL

[0245] dGPl_075 (SEQ ID NO: 188) EEAKKLAEELLNSLSESDTLLKEELTKFIETGSLYALVYAVIFAYRYGTPEQHELADKLYEVAWSS LHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHL

[0246] dGPl_076 (SEQ ID NO: 189) GEAKELANELLSSLSEDDTLIKEELTKFIETGSLYALVYAWFAYLYGTPEQHELADRLYEVAWSS LHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHL

[0247] dGPl_077 (SEQ ID NO: 190) KEEYIERLEKELEELKKKVEAGTATLLDLVYGTIFSLELAKIYEEEGNEEKAKEYRELAEEFFKLYQER LHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHL dGPl_078 (SEQ ID NO: 191)

[0248] MTPEDWAEVRARFPEEAAEAVEVARELILAGKSIEEAMVEVIRRLAELGFPEEWDAVLRLFVELEA

[0249]

[0250] LLHHHHHHHHHHHLLHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLdGPl_079 (SEQ ID NO: 192)

[0251] MTLEE I LERQARI QAVI DEAVARGLADPRI S PVEAAVLALRLVAERGDTSLTSEEMRTIILTTAARVAA LLHHHHHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHLLLLLLHHHHHHHHHHHHHHHHL dGPl_080 (SEQ ID NO: 193) ALPELRALLESLREVGDPALADRALALVEEGRYVEAAVLVIRAASEIGDLERAEEVARRAFELLAELE LHHHHHHHHHHHHHHLLHHHHHHHHHHHHLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHL dGPl_081 (SEQ ID NO: 194) EAREKLRELIESNFPPELQKILLEALEKADTPENIVRLGVLFIRLSSLLPLELAEKWREAFELLIELE LHHHHHHHHHHHHLLHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHL dGPl_082 (SEQ ID NO: 195) GQEAFIRAMFEADPAENKAAIARDPEGFRLAMNALAYFMPDPEKAAEIANALIDEMLAEVEEEKRRA LHHHHHHHHHHHLHHHHHHHHHHLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHHHHHHL

[0252] dGPl_083 (SEQ ID NO: 196) ALEEARRRLEELAERARAALEDPETPEEERVRLLLELIRQASELTYLDPSLRPVAEGWREVGELAVEAS LHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHHHHHL dGPl_084 (SEQ ID NO: 197) AEAKRAAARARLEA^LAATEEQAKKLSPEEAQKLYLQAALD^ILGAADDGYLTEEDVDAALQRWDAEKEYSA LHHHHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHHHHL

[0253] dGPl_086 (SEQ ID NO: 198) LEEE^ISFFEERRERIENGDEEEAVEALLEGFLQAGYFNDRATPEQRKI IEALLEELGELARKKS LHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHHHHHL

[0254] dGPl_087 (SEQ ID NO: 199) MEEEFEEAVERIEEIAARLLASGRALTPEEDAVLYEGLLILIRYVGKAQTPEVREKVEAAFAAFGRLVAAQA LHHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHHL

[0255] dGPl_088 (SEQ ID NO: 200) SSEALYAALRAAAEAARAEGLEAFKAyVYELFKVGLKHLEGVWyFLEAVSSLSPEEQAEVVAYADARQAAEA LHHHHHHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHLLHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHL

[0256] dGPl_089 (SEQ ID NO: 201) GTLQALRDIAARAPELSQEELVAAIDQLSVEPLGYQGLLHLLAAARLARENPDPAARPEIARHLLAAAADAEA K LHHHHHHHHHHHHHHLLHHHHHHHHHHLLLLLLLHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHL

[0257] L

[0258] dGPl_090 (SEQ ID NO: 202) MSPATRAALRSAFARLDALLDADPLRFSFEDWALLRARELIADPATTDAEAAAQVAALRAAVARLE LLHHHHHHHHHHHHHHHHHHHHLHHHLLHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHL

[0259] dGPl_091 (SEQ ID NO: 203) LTAEEIRARLERLIADPAASPEEVAAAIVELIETDKSFEGLVTALQAAVAAEERFGVSPEELARLLDAAE LLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL dGPl_092 (SEQ ID NO: 204)

[0260] LTPEE I LARLEEL VEDPAS S PEEVAAAIVAL IKTDKSFEGLVAALEAARRAEEKFGVSPEE IARLLDEAE LLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL dGPl_093 (SEQ ID NO: 205) LTPEEIKARLEELVADPAASPEEVAKAIVELIKTDKSFEGLVAALKAAERARERFGVSPEEIARLLDRAE LLHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL dGPl_094 (SEQ ID NO: 206) SREELLARTRAALREATRIYVTAQLAGDGPAQQQGIDALLRTAYLAREVLGLSEEEVLA^IAAAEAEVRAA LHHHHHHHHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHL

[0261] dGPl_095 (SEQ ID NO: 207) MAEAEAILEAMRNDQRPITPAEAEAAVRRAPE^YERVLAEHPELDPETAERVAFAIALREAVEEAA LHHHHHHHHHHHLLLLLLLHHHHHHHHHHHHHHHHHHHHHLLLLLHHHHHHHHHHHHHHHHHHHHL

[0262] mGPl_001 (SEQ ID NO: 208) KAEEKRKEAEKAIEEGDYEKAIKIAAEMIETGEDEVNALKLAAKAYEAMGDTSVAALLRTMAEEAAKKA LHHHHHHHHHHHHHHLLHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHL mGPl_002 (SEQ ID NO: 209) EAEELRERLERALEEGDYRRALELAVEYDEASGDTEGALRAAIRAYEALGDTSVADVLRQVLEGLREAR LHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHLLHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHL mGPl_003 (SEQ ID NO: 210) MLDEVAKRIAERLTEEQFVKVLKAAAEVLQRLGITDIDSEEAAEVLEKAGEVMEETGGDAEEVAEKVYKEF LHHHHHHHHHHHLLHHHHHHHHHHHHHHHHHLLLLLLLLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHL

[0263] mGPl_004 (SEQ ID NO: 211)

[0264]

[0265] MRERGVGKSLVAVMKVIYELAKEAGTATPEMEERLEDMIEAAKAAGVTDEEIEAEAEKIKAQLLLLLLLHHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHHHHHLLLLHHHHHHHHHHHHLL mGPl_096 (SEQ ID NO: 212 )

[0266] MPEEI IKRAEEAMTGSDPNEALRLLSELLEYGEERPRALEMAAKLYEKMGHTEVAALLRTVLEEERRRA

[0267]

[0268] LLHHHHHHHHHHHHHHLHHHHHHHHHHHHHLHHHHHHHHHHHHHHHHHLLLHHHHHHHHHHHHHHHHHL In one embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to calcitonin gene-related peptide receptor (CGRPR). In some embodiments, the polypeptide comprises an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 11, 12, and 23 (i.e., dC2_049, dC2_050, and mC2_022.). The amino acid sequences of SEQ ID NO: 1-23 and annotations are shown in Table 1.

[0269] CGRPR is a heterodimer consisting of calcitonin-receptor like receptor (CLR) and receptor activity modifying protein 1 (RAMP1) and is an established target for developing migraine and / or pain therapeutics. The polypeptides of this embodiment are antagonists or negative allosteric modulators of CGRPR, and thus can be used, for example, to treat migraines and / or pain.

[0270] In another embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:26-32 and 213-215, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to the Mas-related GRP family member XI (MRGPRX1) receptor. The amino acid sequences of SEQ ID NO:26-32 and 213-215 and annotations are shown in Table 2.

[0271] Human MRGPRX1 is a G protein-coupled receptor predominantly expressed in small-diameter primary sensory neurons of the dorsal root ganglia (DRG) and peripheral nerves, where it regulates nociceptive and pruritic signaling. Agonist activation of MRGPRX1 induces pruritus, while chronic activation in nociceptors has been associated with analgesic effects, making it a key target for therapeutic intervention in itch and pain. The disclosed polypeptides function as agonists (full or partial agonists) of MRGPRX1 and may be utilized in methods of treating, modulating, or preventing pruritus and pain. Agonists, such as full and partial agonists, offer a unique therapeutic approach by selectively activatingMRGPRX1 to mitigate excessive pruritogenic signaling while harnessing its analgesic properties, thereby achieving symptom relief without fully engaging peripheral itch pathways.

[0272] In a further embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:44-117, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to the glucagon receptor (GCGR). The amino acid sequences of SEQ ID NO:44-117 and annotations are shown in Table 3.

[0273] The GCGR is a G-protein-coupled receptor that plays a crucial role in regulating blood glucose levels and glucose homeostasis by mediating the effects of glucagon. GCGR is primarily found in the liver and islet cells, but is also present in other tissues like the kidney, heart, and brain. When glucagon binds to GCGR, it promotes liver glycogen breakdown (glycogenolysis) and increases blood glucose levels, ultimately stimulating insulin release. It also plays a role in gluconeogenesis (the production of glucose from noncarbohydrate sources). The polypeptides of this embodiment are antagonists or negative allosteric modulators of GCGR, and can thus be used, for example, for treating or limiting obesity and diabetes (such as type 2 diabetes).

[0274] In one embodiment, the polypeptides comprise an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 118-212, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to gastric inhibitory polypeptide receptor (GIPR). The amino acid sequences of SEQ ID NO: 118-212 and annotations are shown in Table 4.

[0275] GIPR (also known as the glucose-dependent insulinotropic polypeptide receptor), is a G-protein-coupled receptor that plays a crucial role in energy metabolism and insulin secretion. GIPR is a protein encoded by the GIPR gene, acting as a receptor for the incretin hormone gastric inhibitory polypeptide (GIP). GIPR is involved in stimulating insulin secretion in response to glucose ingestion and plays a role in regulating energy metabolism. The polypeptides of this embodiment are antagonists or negative allosteric modulators of GIPR and can thus be used, for example, for treating or limiting obesity and diabetes (such as type 2 diabetes).

[0276] In another embodiment, the polypeptides comprise an amino acid sequence at least 60% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions. In a further embodiment, the polypeptides comprise amino acid sequence at least 70% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions. In one embodiment, the polypeptides comprise an amino acid sequence at least 80% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions. In a further embodiment, the polypeptides comprise an amino acid sequence at least 90% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215.

[0277] In various embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or all identified interface residues (i.e., residues at the binding interface with the relevant target) are identical (not substituted), or conservatively substituted, relative to the reference sequence. The interface residues with the relevant GPCR target are shown in bold font in the sequences provided in Tables 1-4. In one embodiment, at least 5 identified interface residues are identical (not substituted), or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 10 identified interface residues are identical, or conservatively substituted, relative to the reference sequence. In another embodiment, all identified interface residues are identical, or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or all identified interface residues are identical (not substituted), relative to the reference sequence

[0278] As used herein, conservative amino acid substitutions involve replacing a residue by a residue having similar physiochemical characteristics, e.g., substituting one aliphatic residue for another (such as He, Vai, Leu, or Ala for one another), or substitution of one polar residue for another (such as between Lys and Arg; Glu and Asp; or Gin and Asn). Other such conservative substitutions, e.g., substitutions of entire regions having similar hydrophobicity characteristics, are known. Amino acids can be grouped according to similarities in the properties of their side chains (in A. L. Lehninger, in Biochemistry, second ed., pp. 73-75, Worth Publishers, New York (1975)): (1) non-polar: Ala (A), Vai (V), Leu (L), He (I), Pro (P), Phe (F), Trp (W), Met (M); (2) uncharged polar: Gly (G), Ser (S), Thr (T), Cys (C), Tyr (Y), Asn (N), Gin (Q); (3) acidic: Asp (D), Glu (E); (4) basic: Lys (K), Arg (R), His (H). Alternatively, naturally occurring residues can be divided into groups based on common side-chain properties: (1) hydrophobic: Norleucine, Met, Ala, Vai, Leu, He; (2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gin; (3) acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues that influence chain orientation: Gly, Pro; (6) aromatic: Trp, Tyr, Phe.

[0279] In one embodiment, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all identified core residues (i.e., residues present in the protein core) are identical, or conservatively substituted, relative to the reference sequence. The polypeptide core residues are shown in underlined font in the sequences provided in Tables 1-4. In one embodiment, at least 5 identified core residues are identical (not substituted), or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 10 identified core residues are identical (not substituted), or conservatively substituted, relative to the reference sequence. In another embodiment, all identified core residues are identical, or conservatively substituted, relative to the reference sequence. In a further embodiment, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all identified core residues are identical, relative to the reference sequence.

[0280] In one embodiment, all interface and all core residues are identical relative to the reference sequence.

[0281] In another embodiment that can be combined with any other embodiment herein, the polypeptide comprises an amino acid insertion at one or more identified loop positions relative to the reference sequence. Loop residues are indicated by an “L” under the residue of the polypeptides shown above. In certain embodiments, amino acids or amino acid domains (such as a functional domain) may be inserted in the loop region. The polypeptides of the disclosure may include any such insertion, and in these embodiments the polypeptide would still comprise the reference amino acid sequence, with an interruption at the site of insertion. In one embodiment, the polypeptide does not comprise an amino acid insertion at any identified loop positions relative to the reference sequence. In some cases, the insertion may result in deletion of one or more loop residue. In another embodiment, the insertion may not lead to deletion of one or more loop residue.

[0282] In one embodiment, the disclosure provides fusion proteins, comprising:

[0283] (a) the polypeptide of any embodiment or combination of embodiments herein; and

[0284] (b) one or more functional domains at the N-terminus and / or at the C-terminus of the polypeptide. In this embodiment, any functional domain may be inserted (as an insertion at a loop region, and / or at one or both termini of the fusion protein). In various non-limiting embodiments, the functional domain may comprise, for example, a targeting domain, a detectable domain, a scaffold domain, a secretion signal, an Fc domain, or a furthertherapeutic peptide domain. In all embodiments, the polypeptide and one or more functional domains may be directly fused or may be separated by an amino acid linker of any length and amino acid composition as appropriate for an intended use.

[0285] In another aspect the disclosure provides nucleic acids encoding the polypeptide or fusion protein of any embodiment or combination of embodiments of the disclosure. The nucleic acid sequence may comprise single stranded or double stranded RNA or DNA in genomic or cDNA form, or DNA-RNA hybrids, each of which may include chemically or biochemically modified, non-natural, or derivatized nucleotide bases. Such nucleic acid sequences may comprise additional sequences useful for promoting expression and / or purification of the encoded peptide or chimeric molecular construct, including but not limited to polyA sequences, modified Kozak sequences, and sequences encoding epitope tags, export signals, and secretory signals, nuclear localization signals, and plasma membrane localization signals. It will be apparent to those of skill in the art, based on the teachings herein, what nucleic acid sequences will encode the polypeptide or fusion protein of the disclosure.

[0286] In a further aspect, the disclosure provides expression vectors comprising the nucleic acid of any aspect of the disclosure operatively linked to a suitable control sequence.

[0287] “Expression vector” includes vectors that operatively link a nucleic acid coding region or gene to any control sequences capable of effecting expression of the gene product. “Control sequences” operably linked to the nucleic acid sequences of the disclosure are nucleic acid sequences capable of effecting the expression of the nucleic acid molecules. The control sequences need not be contiguous with the nucleic acid sequences, so long as they function to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between a promoter sequence and the nucleic acid sequences and the promoter sequence can still be considered “operably linked” to the coding sequence. Other such control sequences include, but are not limited to, polyadenylation signals, termination signals, and ribosome binding sites. Such expression vectors can be of any type, including but not limited plasmid and viral-based expression vectors. The control sequence used to drive expression of the disclosed nucleic acid sequences in a mammalian system may be constitutive (driven by any of a variety of promoters, including but not limited to, CMV, SV40, RSV, actin, EF) or inducible (driven by any of a number of inducible promoters including, but not limited to, tetracycline, ecdysone, steroid-responsive). The expression vector must be replicable in the host organisms either as an episome or by integration into host chromosomal DNA. In various embodiments, the expression vector may comprise a plasmid, viral-based vector, or any other suitable expression vector.In another aspect, the disclosure provides host cells that comprise the polypeptide, fusion protein nucleic acid or expression vector (i.e.: episomal or chromosomally integrated) disclosed herein, wherein the host cells can be either prokaryotic or eukaryotic. The cells can be transiently or stably engineered to incorporate the expression vector of the disclosure, using techniques including but not limited to bacterial transformations, calcium phosphate coprecipitation, electroporation, or liposome mediated-, DEAE dextran mediated-, polycationic mediated-, or viral mediated transfection.

[0288] In another aspect, the disclosure provides pharmaceutical compositions, comprising: (a) the polypeptide, the fusion protein, the nucleic acid, the expression vector, and / or the host cell of any embodiment or combination of embodiments herein; and

[0289] (b) a pharmaceutically acceptable carrier. The compositions may be used, for example, in the methods of the disclosure. The compositions may further comprise (a) a lyoprotectant; (b) a surfactant; (c) a bulking agent; (d) atonicity adjusting agent; (e) a stabilizer; (f) a preservative and / or (g) a buffer. In some embodiments, the buffer in the pharmaceutical composition is a Tris buffer, a histidine buffer, a phosphate buffer, a citrate buffer or an acetate buffer. The composition may also include a lyoprotectant, e.g. sucrose, sorbitol or trehalose. In certain embodiments, the composition includes a preservative e.g. benzalkonium chloride, benzethonium, chlorohexidine, phenol, m-cresol, benzyl alcohol, methylparaben, propylparaben, chlorobutanol, o-cresol, p-cresol, chlorocresol, phenylmercuric nitrate, thimerosal, benzoic acid, and various mixtures thereof. In other embodiments, the composition includes a bulking agent, like glycine. In yet other embodiments, the composition includes a surfactant e.g., polysorbate-20, polysorbate-40, polysorbate- 60, polysorbate-65, polysorbate-80 polysorbate-85, poloxamer-188, sorbitan monolaurate, sorbitan monopalmitate, sorbitan monostearate, sorbitan monooleate, sorbitan trilaurate, sorbitan tristearate, sorbitan trioleaste, or a combination thereof. The composition may also include a tonicity adjusting agent, e.g., a compound that renders the formulation substantially isotonic or isoosmotic with human blood. Exemplary tonicity adjusting agents include sucrose, sorbitol, glycine, methionine, mannitol, dextrose, inositol, sodium chloride, arginine and arginine hydrochloride. In other embodiments, the composition additionally includes a stabilizer, e.g., a molecule that substantially prevents or reduces chemical and / or physical instability of the nanostructure, in lyophilized or liquid form. Exemplary stabilizers include sucrose, sorbitol, glycine, inositol, sodium chloride, methionine, arginine, and arginine hydrochloride.The polypeptide, fusion protein, nucleic acid, expression vector, and / or host cell may be the sole active agent in the composition, or the composition may further comprise one or more other agents suitable for an intended use.

[0290] In another aspect, the disclosure provides methods for treating or limiting development of migraine headaches and / or pain, comprising administering to a subject in need thereof an amount effective to treat the migraine and / or pain of a CGRPR antagonist or negative allosteric modulator polypeptide or fusion protein of the disclosure, a nucleic acid encoding a CGRPR antagonist or negative allosteric modulator polypeptide or fusion protein, or an expression vector, cell, or pharmaceutical composition thereof. CGRPR is a heterodimer consisting of calcitonin-receptor like receptor (CLR) and receptor activity modifying protein 1 (RAMP1) and is an established target for developing migraine and / or pain therapeutics. The polypeptides of this embodiment are antagonists or negative allosteric modulators of CGRPR, and thus can be used, for example, to treat migraines and / or pain.

[0291] In one aspect, the disclosure provides methods for treating or limiting development of itching or pain, comprising administering to a subject in need thereof an amount effective to treat the itching or pain of an MRGPRX1 agonist polypeptide or fusion protein of the disclosure, a nucleic acid encoding an MRGPRX1 agonist polypeptide or fusion protein, or an expression vector, cell, or pharmaceutical composition thereof. Human MRGPRX1 is a G protein-coupled receptor predominantly expressed in small -diameter primary sensory neurons of the dorsal root ganglia (DRG) and peripheral nerves, where it regulates nociceptive and pruritic signaling. Agonist activation of MRGPRX1 induces pruritus, while chronic activation in nociceptors has been associated with analgesic effects, making it a key target for therapeutic intervention in itch and pain. The disclosed polypeptides function as agonists or partial agonists of MRGPRX1 and may be utilized in methods of treating, modulating, or preventing pruritus and pain. Partial agonists offer a unique therapeutic approach by selectively activating MRGPRX1 to mitigate excessive pruritogenic signaling while harnessing its analgesic properties, thereby achieving symptom relief without fully engaging peripheral itch pathways.

[0292] In one aspect, the disclosure provides method for treating or limiting development of obesity or diabetes (including but not limited to Type 2 diabetes), comprising administering to a subject in need thereof an amount effective to treat the obesity or diabetes of an GCGR or GIPR antagonist or negative allosteric modulator polypeptide or fusion protein of the disclosure, a nucleic acid encoding a GCGR or GIPR antagonist or negative allosteric modulator polypeptide or fusion protein, or an expression vector, cell, or pharmaceuticalcomposition thereof. The GCGR is a G-protein-coupled receptor that plays a crucial role in regulating blood glucose levels and glucose homeostasis by mediating the effects of glucagon. GCGR is primarily found in the liver and islet cells but is also present in other tissues like the kidney, heart, and brain. When glucagon binds to GCGR, it promotes liver glycogen breakdown (glycogenolysis) and increases blood glucose levels, ultimately stimulating insulin release. It also plays a role in gluconeogenesis (the production of glucose from non-carbohydrate sources). The polypeptides of this embodiment are antagonists or negative allosteric modulators of GCGR, and can thus be used, for example, for treating or limiting obesity and diabetes (such as type 2 diabetes). Similarly, GIPR (also known as the glucose -dependent insulinotropic polypeptide receptor), is a G-protein-coupled receptor that plays a crucial role in energy metabolism and insulin secretion. GIPR is a protein encoded by the GIPR gene, acting as a receptor for the incretin hormone gastric inhibitory polypeptide (GIP). GIPR is involved in stimulating insulin secretion in response to glucose ingestion and plays a role in regulating energy metabolism. The polypeptides of this embodiment are antagonists or negative allosteric modulators of GIPR and can thus be used, for example, for treating or limiting obesity and diabetes (such as type 2 diabetes).

[0293] As used herein, "treat" or "treating" a disorder means accomplishing one or more of the following in a subject with the disorder: (a) reducing the severity of the disorder; (b) limiting or preventing development of symptoms characteristic of the disorder(s) being treated; (c) inhibiting worsening of symptoms characteristic of the disorder(s) being treated; (d) limiting or preventing recurrence of the disorder(s) in patients that have previously had the disorder(s); and (e) limiting or preventing recurrence of symptoms in patients that were previously symptomatic for the disorder(s).

[0294] As used herein, “limiting development” of a disorder means administering to a subject that does not have the disorder, but who is or may be at risk of developing the disorder, to limit or prevent development of the disorder and its symptoms.

[0295] The subject may be any subject that has a relevant disorder or may be at risk of the relevant disorder. In one embodiment, the subject is a mammal, including but not limited to humans, dogs, cats, horses, cattle, etc. In a specific embodiment, the subject is a human subject.

[0296] As used herein, an “effective” amount refers to an amount of the polypeptide, fusion protein, nucleic acid, expression vector, host cell, and / or pharmaceutical composition that is effective for treating or limiting development of the disorder. The polypeptides, fusion proteins nucleic acids, expression vectors, and / or host cells are typically formulated as apharmaceutical composition, such as those disclosed above, and can be administered via any suitable route, including but not limited to orally, by inhalation spray, ocularly, intravenously, subcutaneously, intraperitoneally, and intravesicularly in dosage unit formulations containing conventional pharmaceutically acceptable carriers, adjuvants, and vehicles.

[0297] Any suitable dosage range may be used as determined by attending medical personnel. Dosage regimens can be adjusted to provide the optimum desired response. A suitable dosage range for the polypeptides or fusion proteins may, for instance, be 0.1 ug / kg -100 mg / kg body weight; alternatively, it may be 0.5 ug / kg to 50 mg / kg; 1 ug / kg to 25 mg / kg, or 5 ug / kg to 10 mg / kg body weight. In some embodiments, the recommended dose could be lower than 0.1 mcg / kg, especially if administered locally (such as by intra-tumoral injection). In other embodiments, the recommended dose could be based on weight / m2(i.e. body surface area), and / or it could be administered at a fixed dose (e.g.,.05-100 mg). The polypeptides, fusion proteins, nucleic acids, expression vectors, and / or host cells can be delivered in a single bolus, or may be administered more than once (e.g., 2, 3, 4, 5, or more times) as determined by an attending physician.

[0298] Examples

[0299] Abstract

[0300] G protein-coupled receptors (GPCRs) play key roles in physiology and are central targets for drug discovery and development, but the design of protein agonists and antagonists has been challenging as GPCRs are integral membrane proteins and conformationally dynamic. Here we describe computational de novo design methods and a high-throughput “receptor diversion” microscopy-based screen for generating GPCR binding miniproteins with high affinity, potency and selectivity, and the use of these methods to generate MRGPRX1 agonists and GIPR, GCGR, and CGRPR antagonists. Cryo-electron microscopy data reveals atomic -level agreement between designed and experimentally determined structures for CGRPR-bound antagonists and MRGPRX1 -bound agonists, confirming precise conformational control of receptor function. Our de novo design and screening approach opens new frontiers in GPCR drug discovery and development.

[0301] Development of computational and experimental methods to target diverse GPCR epitopes

[0302] To enable targeting of the deeply recessed orthosteric binding site epitopes — critical for modulating class A GPCR function — we implemented two complementary designmethods to generate functional miniproteins. First, we developed a “motif-directed” RFdiffusion™ approach that rather than diffusing an entire binding protein, starts with just a five-residue peptide (the motif) to interact with target hot spot residues within the recessed binding pocket. The short peptide can more readily penetrate into the deep pocket, and once good solutions for a binding peptide are found, the interacting peptide is kept in a fixed position and full miniproteins are generated using the motif-scaffolding capabilities of RFdiffusion™. To increase diversity of designs for library-scale experimental screening, we developed an iterative partial diffusion approach which generates new designs in the vicinity of the most promising in silico solutions at each stage of the process. Second, we developed an approach, MetaGen™, that employs structurally diverse scaffolds from the AlphaFold™ (AF) predicted structural metaproteome12 13(Fig. 3) in Rosetta™ RifDock calculations. In contrast to traditional de novo miniprotein backbone libraries8, often composed of straight helices and short loops ill-suited for engaging deeply recessed epitopes, these scaffolds feature protruding elements, such as kinked helices and beta-hairpin loops, but are still confidently predicted from a single sequence — a key criterion for designability14. Following backbone design with either RFdiffusion™ or MetaGen™ we used ProteinMPNN™ for sequence design15and AF2 initial guess12as well as Rosetta™ metrics for filtering designs14Due to the challenging nature of class A GPCR epitopes, we reasoned that high-throughput screening (HTS) methods would be necessary to complement computational design for robust identification of functional binders. To address this, we developed Receptor Diversion (RD), a purification-free microscopy-based assay that operates directly in human cells. In this assay, both the membrane protein target and the candidate binder are expressed in a human cell line, with the binder localized within the secretory pathway using a genetic tag (e.g., KDEL; an endoplasmic reticulum (ER) retention signal). This allows the binder to interact with the extracellular face of the membrane protein target during transit through the secretory pathway. High-affinity interactions cause “diversion” of the target from its normal trafficking pattern, which can be visualized as increased binder-target subcellular colocalization, i.e. binding signal. Across 7 diverse GPCRs, we observed a robust binding signal suitable for high-throughput screening with a cross-GPCR Z' average of 0.47 when sampling 100 cells per binder. RD has the advantages that (i) the target can be expressed at near-endogenous levels in a relevant cell line and does not have to be produced as a stable soluble protein (challenging for GPCRs) as required for display methods, (ii) binders discovered through the screen must be efficiently translated into ER in human cells, be soluble and function in the molecularly crowded environment of the secretory pathway, and(iii) the binder must specifically bind the target in order to induce receptor diversion. To deploy the assay at library scale, we use optical pooled screening (OPS), where individual designs are encoded together with a DNA barcode, and optically genotyped using in situ sequencing. The OPS-RD platform enables screening of up to 100,000 designs through imaging of up to 107cells providing expression and co-localization data at the single-cell level.

[0303] As we were unsure how well OPS-RD would work in practice for library screening at the beginning of this study, we also explored the use of yeast display paired with either soluble GPCRs in nanodiscs17or GPCRs displayed on mammalian cells (biofloating)18.

[0304] Pharmacology and biophysical characterization of receptor penetrating class A MRGPRX1 agonists

[0305] To explore the potential of computational design to create GPCR agonists, we focused on the Mas-related G protein-coupled receptor XI (MRGPRX1), an emerging target for itch and pain20. Using the MetaGen™ approach, we targeted a large epitope within the orthosteric binding pocket spanning transmembranes 2-7 (TM2-TM7) of three active-state structures, reasoning that active-state stabilization alone could be sufficient to generate agonists21. We screened a library of 13,211 designs using OPS-RD, and succeeded in mapping optical binding phenotypes for >800,000 cells to their design genotypes. Averaging optical phenotypes across cells, we ranked each design and selected 64 designs (Fig. la), then generated an additional 27 designs using partial diffusion, resulting in a complete set of 91 designs selected for further characterization.

[0306] Of these, 50 were highly expressed in E. coli and subsequently screened in a calcium mobilization assay to explore their ability to stimulate intracellular signaling. Consistent with the design strategy, seven miniproteins demonstrated agonistic activity at 10 pM (Fig. 4). Next, we generated concentration-response curves for the seven hits and obtained two full agonists, relative to the endogenous agonist peptide BAM (8-22), with EC50values of 390 ± 70 nM (mean ± SEM, n=3) and 1000 ± 100 nM (mean ± SEM, n=3), respectively, while additionally discovering a partial agonist which displayed an EC50of 1400 ± 200 nM (mean ± SEM, n=3) (Fig. lb, Table 5). The three hits were structurally diverse (Fig. 1c), highly expressed, monomeric by SEC (Fig. 5a), and have CD spectra consistent with the expected molecular structure (Fig. 5b) as well as high thermal stability (Fig. 5c).

[0307] Table 5. Pharmacological data of MRGPRX1 bindersLigand Hill slope Potency ECso (M) / >(ECso) (%} mM1_011 2.7 3.1 x 10-65.5 101 mM1_034 2.7 ± 0.1 1.0 ± 0 1 x 10‘66.0 ± 0.0 99 ± 1 mM1_059 1.8 1.3 x 10-65.9 117 mM1_060 11.0 ± 2.8 1.4 ± 0.2 x 10-65.9 ± 0.1 22 ± 2 mM1_063 2.1 3.2 x 10-65.5 139 mM1_068 2.3 ± 0.1 3.9 ± 0.3 x 10-76.4 ± 0.0 96 ± 2 mM1_084 2.9 3.9 x 10’654 78 mM1_034 F21W 2.6 ± 0.6 4.2 ± 0.7 x 10-87.4 ± 0.1 85 ± 6 Y27F

[0308] mM1_034 F21W 2.5 ± 0.4 1.1 ± 02 x IO'770 ± 0 1 68 ± 3

[0309]

[0310] A58M

[0311] Data are shown as mean ± SEM from three to five independent experiments or as mean only (n=l). The EC50 of the native BAM (8-22) in the calcium mobilization assay was 3.0 ± 0.2 x 10-9(M) (pEC508.5 ± 0.1).

[0312] To enhance the potency of the MRGPRX1 agonists, we performed site saturation mutagenesis (SSM) in order identify mutations that would stabilize the complex between the top three hits and MRGPRX1. 239 variants were expressed, purified and tested for MRGPRX1 agonism at a concentration of 500 nM. 21 hits were identified (Fig. 6) and the most potent miniprotein, mM1_034_F12W_Y27F, with an ECso of 42 ± 7 nM (mean ± SEM, n=4), exhibited 12-fold increased potency, similar to the endogenous agonist BAM (8-22)) (Fig. Id). The other top hit, mM1_034_F12W_A58M with an ECso of 110 ± 20 nM (mean ± SEM, n=5) displayed a roughly four-fold improvement over the parental protein mMl_034, but had slightly decreased efficacy (68 ± 3%, mean ± SEM, n=5) compared to BAM (8-22) and mM1_034) (Fig. Id).

[0313] Cryo-electron microscopy (cryo-EM) structures of mM1_068 and mM1_060 miniprotein binders bound to hMRGPRXl were consistent with the design model (data not shown).

[0314] Pharmacology and biophysical characterization of antagonists targeting ECD of multiple class B GPCRs

[0315] To explore the design of antagonists of class B GPCRs, which include many therapeutic targets29, we applied our design methods to gastric inhibitory polypeptide receptor (GIPR), glucagon receptor (GCGR), and calcitonin gene-related peptide receptor (CGRPR). We targeted the soluble extracellular domain (ECD) of these receptors, reasoning that binding to the ECD should induce steric hindrance and prevent peptide interaction with the receptor, thereby resulting in antagonism (Fig. 2a). For GIPR and GCGR, we used yeast display and soluble ECDs of receptors to identify miniprotein binders. Similarly, probingyeast display libraries of 17,698 and 12,399 designs for GIPR (Fig. 7a-d) and GCGR, (Fig.

[0316] 8a-d), respectively, and expressing 96 designs for each target receptor yielded miniprotein binders with picomolar and nanomolar affinities for both targets (Fig. 9a-d, Fig. lOa-d). The GIPR binders were confirmed as antagonists in a functional cAMP assay and as selective for GIPR (data not shown).

[0317] CGRPR is a heterodimer consisting of calcitonin-receptor like receptor (CLR) and receptor activity-modifying protein 1 (RAMP1) and is an established target for developing migraine and / or pain therapeutics30. After closely inspecting the nature of the CGRPR epitope, we hypothesized that we could achieve a >1% hit rate, given the high epitope hydrophobicity and absence of loops, and did not attempt high-throughput screening techniques. We screened designs in a functional, one-point cAMP assay using a SK-N-MC cell line (Fig. 11). Out of 96 obtained RFdiffusion™ designs, 67 expressed, and a three -helix bundle miniprotein dCl_021 was identified with an ICso of 440 ± 40 nM (mean ± SEM, n=3) (Fig. 12, Table 6). Out of 89 MetaGen™ -derived backbones, 83 expressed and 4 had measurable antagonistic activity in a one-point luciferase assay (Fig. 13a-c). We identified a competitive antagonist mCl_023 of mixed aP-topology with an ICso of 37 ± 2 nM (mean ± SEM, n=4) (Fig. 14, Table 6) and a p i of 8.3 ± 0.1 (5 nM, mean ± SEM, n=4) (Fig. 2b).

[0318] Disulfide stapling31of the second hit mCl_044 with potency in the micromolar range (Fig.

[0319] 14) yielded antagonist mC2_022 with a

[0320]

[0321] of 7.9 ± 0.2 (13 nM, mean ± SEM, n=4) (Fig.

[0322] 2c) and ICso of 420 ± 60 nM (mean ± SEM, n=4) (Fig. 14, Table 6). To increase the potency of the RFdiffusion™ hit dC 1 021, we performed partial diffusion6. Out of 78 designs that expressed, 36 had measurable antagonistic activity against CGRPR in a one-point cAMP assay (Fig. 15a-c, Table 6). Concentration-response curves of the 20 most promising binders identified the competitive antagonist dC2_049 with an ICso value of 4.5 ± 0.9 nM (mean ± SEM, n=4) (Fig. 16a-d, Table 6) and a p i of 8.4 ± 0.1 (3.9 nM, mean ± SEM, n=4) (Fig.

[0323] 2d). All three antagonists migrated as single peaks in size exclusion chromatography (SEC), with mC l_023 eluting as a monomer, whereas mC2_022 and dC2_049 eluted as dimers, had CD spectra consistent with their structures, and high thermal stability Fig. 17a-c). dC2_049 was shown to exhibit high selectivity for the CGRPR with little or no significant crossreactivity at related receptors (data not shown). A sequence-similar (sequence identity 59%) RFdiffusion™ binder, dC2_050, derived by partial diffusion starting from the same parent structure as dC2_049, had similar pharmacological properties and biophysical characteristics (Fig. 18a-e, Table 6). dC2_049 and dC2_050 were confirmed as CGRPR antagonists in COS-7 cells (data not shown). Cryo-EM structures of dC2_049 and dC2_050 miniproteinbinders bound to CGRPR were determined and in good agreement with computational design models (data not shown).

[0324] Table 6. Pharmacological data of CGRPR binder antagonists.

[0325] Ligand Potency IC50(M)

[0326] mC1_023 3.7 ± 0.2 x 10-87.4 ± 0.1

[0327] mCi __044 > 1 x 10“ < 6

[0328] mC2_022 4.2 ± 0.6 x 10-76.2 ± 0.1

[0329] dC1_021 4.4 ± 0.4 x 10-76.4 ± 0.1

[0330] dC2_001 4.4 ± 0.6 x 10-76.4 ± 0.1

[0331] dC2_007 1.9 ± 0.1 x 10-76.7 ± 0.1

[0332] dC2_ Oi l 1.5 ± 0.6 x 10'“ 5.9 ± 0.2

[0333] dC2_019 5.9 ± 1.3 x 10-76.3 ± 0.1

[0334] dC2_026 3.0 ± 0.7 x 10-76.6 ± 0.1

[0335] dC2_030 3.7 ± 0.8 x 10-87.3 ± 0.1

[0336] dC2_039 4.0 ± 1.0 x 10-87.4 ± 0.1

[0337] dC2_042 1.2 ± 0.4 x 10-77.0 ± 0.1

[0338] dC2_045 1.2 ± 0.3 x 10-76.9 ± 0.1

[0339] dC2 049 4.5 ± 0 9 x 10‘98.5 ± 0 1

[0340] dC2_050 1.3 ± 0.1 x 10-87.9 ± 0.1

[0341] dC2_052 2.7 ± 1.0 x 10-76.7 ± 0.2

[0342] dC2 053 3.3 ± 0 6 x IO’76.5 ± 0 1

[0343] dC2__055 2.1 ± 0.5 x 10‘76.7 ± 0.1

[0344] dC2_057 2.4 ± 0.7 x 10-76.7 ± 0.2

[0345] dC2_058 1.8 ± 0.5 x 10-76.8 ± 0.1

[0346] dC2_063 3.4 ± 1.3 x 10-76.6 ± 0.2

[0347] dC2_065 6.9 ± 1.2 x 10-76.2 ± 0.1

[0348] dC2_066 4.2 ± 1.0 x 10-76.4 ± 0.1

[0349]

[0350] dC2_067 1.7 ± 0.6 x 10-76.8 ± 0.1

[0351] Data are from at least three independent experiments. Values represent mean ± SEM. The EC50 of the native CGRP was 1.3 ± 0.1 x 10-10(M) (pEC509.9 ± 0.1 in SK-N-MC and 3.9 ± 0.9 x 10-11(M) (pEC5010.4 ± 0.1 in CHO-K1 / Cre-Luc / CGRP and its EC80was used to generate concentration response curves of antagonists to derive IC50 values (mean ± SEM, n=4).

[0352] Discussion

[0353] GPCRs have been longstanding challenges for drug discovery and development of protein-based ligands owing to their structural complexity, recessed binding pockets, and dynamic character. We show that de novo design can address these challenges by generatingminiprotein binders targeting MRGPRX1, GIPR, GCGR, and CGRPR with diverse affinity, potency, and selectivity profiles. Agonists have been particularly challenging to obtain due to the need for conformational selectivity, requiring discrimination between subtle structural differences in the orthosteric binding site that distinguish active from inactive states. Here, we demonstrate the de novo design of two atomically accurate binders for MRGPRX1 (within 0.7 A), capable of inducing full and partial agonism. These findings establish de novo design as a viable strategy for engineering GPCR-targeting miniprotein antagonists and agonists.. Complementing our computational design approaches, our in-cell OPS-RD platform enables high-throughput screening for difficult GPCR targets by circumventing the need for engineering of soluble receptor preparations in artificial nanodiscs, liposomes or mutant receptor proxies, which can potentially alter sampling of receptor conformations and functional properties35. While the maximal capacity (<100,000 designs) of the OPS-RD platform is more limited than display technologies, and the false positive rate appears to be higher, the ability to screen against the native receptor in the membrane environment and bypass solubilization provides a major benefit.

[0354] The therapeutic potential of de novo designed GPCR antagonists and agonists is considerable given the central roles GPCRs play in cellular function and disease. The ability to computationally design conformationally selective binders is a step change in methodology for obtaining functional biologies targeting integral membrane receptors. Combined with their smaller size for improved tissue penetration, high stability, rapid design and optimization, and the potential for modifications to enhance metabolic stability, miniproteins represent an attractive modality for therapeutic applications. Beyond therapeutics, designed GPCR binders could have considerable utility as tools for drug discovery, probing pharmacology and receptor function, and stabilizing receptor conformations for structural studies. We anticipate de novo designed GPCR agonists and antagonists will be widely useful.

[0355] Methods

[0356] Binder design using RFdiffusion™ and metaproteomic scaffolds

[0357] The cryo-EM structures of GIPR (PDB ID: 7FIN, 7FIY), GCGR (PDB ID: 5XEZ), and CGRPR (PDB ID: 6E3Y) were used as targets for designing binders with RFdiffusion. Additionally, cryo-EM structures of CGRPR (PDB IDs: 3N7S, 7KNU, 6E3Y), MRGPRX1 (PDB IDs: 8DWC, 8DWG, 8DWH), GIPR (PDB IDs: 2QKH, 4HJ0, 7FIN), and GCGR(PDB IDs: 6WPW, 8JIT, 8JIU) served as targets for binder design using metaproteomic scaffolds. All target structures were truncated to the region containing the binding epitope.

[0358] Backbone generation using motif-scaffolded RFdiffusion™ targeting GIPR, GCGR or free RFdiffusion™ against CGRPR was performed as previously described7. For the GIPR and GCGR, 50,000-100,000 backbones were created using following hot spot residues chosen within the ECD of the receptor, GIPR M32 and GCGR F33, W36 and W87. For the CGRPR, three hydrophobic hotspot residues (L33, W72, F92) were chosen within the ECD of the receptor and approximately 50,000 backbones were generated. Sequences were designed using ProteinMPNN™ (10 sequences per backbone and sampling temperature of <0.1)15, followed by FastRelax™ and AF2 initial guess12. In silico cutoffs goals were defined a priori from our previous benchmarking of AF2 and Rosetta™ metrics across multiple targets7,14. Where feasible, we targeted pLDDT_binder > 90, pAE_interaction < 8, binder_RMSD < 2, and low SAP (< 35) to reduce aggregation risk while maintaining confident local structure. We treated pAE interaction and pLDDT binder as primary acceptance filters, then set Rosetta™ ddG and SAP to filter remaining candidates. When very few designs met these strict criteria, thresholds were relaxed within the target specific ranges reported above, guided by interface chemistry. Targets that can be engaged via larger hydrophobic contact patches generally support more stringent Rosetta™ ddG cutoffs.

[0359] Designs generated by RFdiffusion™ were selected based on pAE_interaction < 8, pLDDT binder > 85, Rosetta™ ddG < -45, pAE interaction < 6, pLDDT binder > 90, Rosetta™ ddG < -45 and sap < 60 for GIPR, pAE interaction < 8, pLDDT binder > 90, Rosetta™ ddG < -50 and sap < 45 for GCGR, pAE interaction < 4, pLDDT binder > 90, and pAE interaction < 8, pLDDT binder > 90 and Rosetta™ ddG < -45 for CGRPR.

[0360] Scaffolds for MetaGen™ were selected from AF generated metaproteome12,13containing 200 million protein structures to as follows: longest_loop <= 7 & longest_strand <= 10 & longest_helix <= 20 & fraction_loop <= 0.35 & length <= 90. Following this, we designed all the scaffolds using ProteinMPNN™, and computed structural metrics with Rosetta™ and AF2. We then filtered for globularity (percent_core_SCN > 20%) & surface aggregation propensity (sap <= 40) & energy (score _pcr_rcs <= -2.3) & pLDDT_scaffold > 90. Finally, we clustered all scaffolds by sequence with mmseqs36to 80% sequence identity and picked the best scaffold within each cluster by pLDDT scaffold. This resulted in 8920 scaffolds.

[0361] Metaproteome-derived designs targeting CGRPR, MRGPRX1, GIPR, and GCGR were generated using the RIFdock, motif extraction, and recycling strategy outlined in Cao etal.8. Following sequence design and prediction. Selection criteria varied by target: CGRPR designs were chosen based on pAE interaction < 8, binder RMSD < 2, and scaffold_pLDDT > 90; MRGPRX1 designs met pAE_interaction < 10 or (sap < 40 & contact_molecular_surface > 600 & membrane_insertion_energy > 4 & Rosetta™ ddG < -51); GIPR designs were selected based on pAE interaction < 6, binder RMSD < 2, ddG < -40, scaffold_pLDDT > 90, membrane insertion energy > 4, and sap < 35; and GCGR designs met pAE interaction < 12, binder RMSD < 2, ddG < -40, scaffold_pLDDT > 85, sap < 40, and membrane_insertion_energy > 4.

[0362] Partial diffusion was performed on the AF2 model of the most promising CGRPR hit (dCl_022). Roughly 3,000 backbones were designed by applying 10, 15, and 20 noising timesteps out of a total of 50 timesteps in the noising schedule followed by denoising steps (diffuser. partial T input values of 10, 15 and 20). The resulting backbone libraries after free and partial RFdiffusion™ were subjected to sequence design using ProteinMPNN™ (10 sequences per backbone)15, followed by FastRelax™ and AF2 initial guess12. The resulting libraries were filtered based on AF2 pAE interaction < 4, pLDDT binder > 90, and Rosetta™ ddG < -45.

[0363] Cloning, expression and purification of protein binders

[0364] Protein binder designs were obtained as synthetic genes (eBlocks™, Integrated DNA Technologies) with compatible Bsal overhangs to the target cloning vector, LM0627 for Golden Gate assembly40. Subcloning into LM0627 resulted in the following product: MSG-[protein]-GSGSHHWGSTHHHHHH (SEQ ID NO:24), with the C-terminal SNAC cleavage tag and 6xHis affinity tag. Briefly, Golden Gate subcloning reactions of designs were performed in 96-well PCR plates in 1 pL volume. Reaction mixtures were then transformed into a chemically competent expression strain (BL21(DE3)) and 10 mL of these split directly into four 96-deep well plates containing 990 uL of auto-induction media (autoclaved TB-II media supplemented with kanamycin, 2 mM MgSO4. IX 5052). Designs generated using the MetaGen™ pipeline were plated to single colonies and sequence verified before inoculating expression media. Post overnight incubation at 37°C (20-24 hours), cells were harvested, lysed, and clarified lysates applied to a 75 pL bed of Ni-NTA agarose resin in a 96-well fritted plate equilibrated with a Tris wash buffer. After sample application, the resin was washed, and samples were eluted in 200 pL of a Tris elution buffer containing 300 mM imidazole. Proteins were then purified via SEC using an AKTA FPLC equipped with an ALIAS autosampler capable of running samples from two 96-well source plates. ASuperdex75 Increase 5 / 150 GL column was used (Cytiva 29148722).. To verify the identity of MetaGen™ designed proteins, intact mass spectra were obtained via reverse-phase LC / MS on an Agilent G6230B TOF on an AdvanceBio RP-Desalting column and subsequently deconvoluted by way of Bioconfirm using a total entropy algorithm. RFdiffusion™ designed binders identified as hits in screens were confirmed by sequencing.

[0365] Circular dichroism

[0366] For circular dichroism (CD) measurements, diffusion-derived designs were diluted to 0.4 mg / ml in 20 mM Tris (pH 8.0) and 100 mM NaCl, while metaproteome-derived designs were analyzed at 50 pM in PBS (pH 7.4). Spectra were acquired on a JASCO™ J- 1500 CD Spectrophotometer. Thermal melt analyses were performed between 25°C and 95 °C, measuring CD at 222 nm. All reported measurements were acquired within the linear range of the instrument.

[0367] Cell culture

[0368] CHO-Kl / CRE-Luc / CGRPR (Genscript M00350) cells were cultured in Ham's F-12K (Kaighn's) Medium (Gibco) containing 10% FBS. PathHunter® CHO- and MRGPRX1 -arrestin cells (DiscoverX™ #93-0919C2) overexpressing the respective receptor and a split-P-galactosidase were cultured and plated according to the manufacturer’s instructions.

[0369] HEK293T and SK-N-MC cells were cultured in Dulbecco’s modified Eagle’s medium (DMEM) medium (Thermo Fisher) containing 10% FBS and penicillin-streptomycin (500 U / mL). COS-7 cells were cultured in DMEM containing 10% FBS only. LentiX™ 293T cells (Takara #632180) and HeLa cells, the latter optimized for optical pooled screening and kindly gifted by Iain Cheeseman, were cultured in D10 media (DMEM with GlutaMAX™, 10% (v / v) FBS, and 100 U / mL penicillin-streptomycin). All cells were grown at 37°C with a humidified atmosphere and 5% CO2.

[0370] DNA library preparation for yeast display

[0371] The DNA library was prepared as previously described8. All protein sequences were padded to a uniform length by adding a (GGGS)n linker at the C terminal of the designs, to avoid the biased amplification of short DNA fragments during PCR reactions. The protein sequences were reversed translated and optimized using DNAworks2.0 with the. S', cerevisiae codon frequency table. Homologous to the pETCON plasmid, oligo libraries encoding the designs were obtained from Twist Bioscience. Combinatorial libraries were obtained as IDT (Integrated DNA Technologies) ultramers with the final DNA diversity ranging from 1×106to 1 x 107. All libraries were amplified using Kapa HiFi™ Polymerase (Kapa Biosystems) with a qPCR machine (BioRAD CFX96). In detail, the libraries were firstly amplified in a 25 pL reaction, and PCR reaction was terminated when the reaction reached half the maximum yield to avoid over-amplification. The PCR product was loaded to a DNA agarose gel. The band with the expected size was cut out and DNA fragments were extracted using QIAquick™ kits (Qiagen, Inc.). Then, the DNA product was re-amplified as before to generate enough DNA for yeast transformation. The final PCR product was cleaned up with a QIAquick™ Clean up kit (Qiagen, Inc.). For the yeast transformation, 2-3 pg of digested modified pETcon™ vector (pETcon3) and 6 pg of insert were transformed into EBY 100 yeast strain using the protocol as described before. DNA libraries for deep sequencing were prepared using the same PCR protocol, except the first step started from yeast plasmid prepared from 5 x 107to 1 x 108cells by Zymoprep™ (Zymo Research). Illumina adapters and 6-bp pool-specific barcodes were added in the second qPCR step. Gel extraction was used to get the final DNA product for sequencing. All libraries include the native library and different sorting pools were sequenced using Illumina NextSeq™ / MiSeq™ sequencing.

[0372] Yeast display

[0373] General yeast display methodologies were carried out with EBY- 100 yeast cells, as previously described16,42. Yeast clones for biofloating assay were grown in SD-CAA medium at 30°C while shaking at 200 rpm. Yeast cultures were induced in SG-CAA medium at 20°C while shaking at 200 rpm at an initial optical density (OD) of 1.0 (1 x 107cells / mL). For soluble receptor-based approach, yeast EBY- 100 strain cultures were grown in C-Trp-Ura media and induced in SG-CAA. Cells were washed with PBSF (PBS with 1% BSA) and incubated with biotinylated (SinoBiological 13944-H49H-B, GIPR (SinoBiological, 18774-H49H-B), or GCGR (Aero Biosystems, GCR-H82E3), respectively. For the first round of sorting, cells were incubated with biotinylated ECDs of, GIPR or GCGR, or and labelled with corresponding antibodies simultaneously for 20 minutes whereas for the sorting rounds thereafter, cells were first pre-incubated with the target for 20 minutes and then labelled with corresponding antibodies for additional 20 minutes. Anti-c-Myc fluorescein isothiocyanate (FITC, Miltenyi Biotech) antibody was used for labeling cells and anti-streptavidin phycoerythrin (SAPE, Thermo Fisher). The concentration of FITC was used at 1 / 4 concentration of the Flag-tagged or biotinylated target. For the first round of sorting 1 pM concentration of the receptor target was used. The remaining subsequent sorts were performed with varying concentrations (10 pm - 1 pM) of the target. The final sorting poolsof the library were sequenced using Illumina NextSeq™ / MiSeq™ sequencing. All FACS data were analyzed in FlowJo™.

[0374] Yeast SC50 estimation from FACS and Next Generation Sequencing

[0375] In order to quantify the level of binding for each of the designs, the SC50 estimation, previously described8, was used. The NGS data was used to quantify the composition of each sorting pool, noting what proportion of the entire pool each design represented. Using the ratio of the proportions between a child and parent pool along with the fraction of the cells collected on the FACS machine, the fraction of cells containing each design that survived each sort was calculated. For each design in each sort, the probability that the data could be generated from a binder with SC50 X was assessed for all X from fM to mM. These probabilities were multiplied together for all sorts and the interval where the probability of the SC50 was greater than 0.002 was reported to give a value resembling a 99.8th percent confidence interval. We also created the SCso-RelativeEnrichment (SCsoRE). While the work from Cao and colleagues opted to eliminate designs carried by passenger plasmids with the use of minimum cell cutoffs, this method was proven to be not completely effective8. In fact, methods using enrichment (proportion of fmal _pool / proportion_of_expression_pool) do not suffer from passenger plasmids. We thus developed a method to include the power of enrichment to identify designs that were likely carried by passenger plasmids.

[0376] The SCsoRE represents the enrichment of a given design relative to the most enriched design with SC50 equal-to or worse-than the given design. In this way, if two designs claim to have the same SC50, but one of them is lOOx more enriched than the other, it is likely that the lesser-enriched design is from a passenger plasmid. In order to calculate the SCsoRE, one must choose which sorts to compare enrichment values from. We chose to use all non-avidity sorts where the sorting concentration is at least 10-fold greater than the predicted SCso. These sorts were chosen because real binders should be saturating under these conditions and have robust enrichment (versus sorts with 100-folder concentration lower than the predicted SC50 where the predicted collection fraction is < 1% and the enrichment largely based on collection noise). For each design and sort that fit the previous criteria, the most-enriched design with SC50 equal-to or worse-than the current design was identified, and the ratios of their enrichment calculated (current_design_enrichment / most_enriched_enrichment). The lowest value obtained in any single sort was assigned as the current design’s SCso-RelativeEnrichment.

[0377] Mathematically, we can write the SC50-RelativeEnrichment for design i asSC50REI = min Ei p / E™axfor p such that sort_concentrationp> SC50 ix 10 } where for each pool p, Ei p= propp i / propexprt is the enrichment compared to the expression pool and E’’iax= maXj Ep j} is the largest enrichment of any design in pool p.

[0378] Designs with SCsoRE lower than 1 / 50 were suspected to be from passenger plasmids. Such a design would have an Enrichment 50x lower than another design with similar or worse SCso.

[0379] Optical screen

[0380] Plasmids:

[0381] For MRGPRX1 GFP reporter vectors were generated by cloning full-length human MRGPRX1 (UniprotKB: Q96LB2) into a lentiviral entry vector encoding an N-terminal mIGK signal peptide, FLAG tag, and GFP, as well as a C-terminal BFP fused to a c-myc tag followed by a 2A peptide and blasticidin selection marker (pLenti / mlGK-FLAG-eGFP-BsmBI-BFP-myc-P2A-blast) using NEBridge™ Golden Gate Assembly Kit (BsmBI-v2) (New England Biolabs #E1602L). The lentiviral entry vector for binder library cloning (pLenti / puro-T2A-RUSH-C5 -m Cherry ™-BsmBI) was prepared by replacing the U6 promoter in lentiguide-BC-plasmid (Addgene #127168) with an EFla promoter, puromycin resistance marker, T2A peptide, RUSH secretion tag (Addgene #65294), C5 oligomerization domain (PDB 2B98) and mCherry™, followed by a BsmBI entry site for cloning of barcoded binders. Barcoded binders were synthesized as Twist oligo pools containing the designed binder, a C-terminal KDEL endoplasmic reticulum retention tag, stop codon, and a 10-nt barcode suitable for in situ sequencing. Thus, the final binder library construct encoded puromycin resistance separated by a 2A peptide from a protein fusion comprising a secretion tag, oligomerization domain, mCherry™ tag, designed binder, and ER retention tag, followed immediately by a non-coding barcode. Barcoded binder libraries were cloned into the lentiviral entry vector as reported previously43,44. Briefly, oligo pools were amplified with KAPA HiFi HotStart Ready Mix (Roche #KK2601), IX EvaGreen™ qPCR dye (Biotium #31000), 500 nM forward and reverse primers (dialout primer FW:

[0382] TCTGAACAGGCTcgtctct (SEQ ID NO:25), dialout primer RV: CTATCGCCAAGTcgtctct) (SEQ ID NO:33) (Integrated DNA Technologies), and 80 pg / pL of template in 30 pL reactions. PCRs were conducted with the following thermal cycling protocol: 95 °C for 3 min, 14-16 cycles of (98 °C for 20 s, 65 °C for 20 s, 72 °C for 45 s), then 72 °C for 1 min. Following amplification, reactions were gel purified using Zymoclean™ Gel DNA RecoveryKit (Zymo #D4007) and quantified with Qubit™ Broad Range dsDNA Quantitation assay (Thermo Fisher Scientific #Q32853). Plasmid libraries were then constructed using NEBridge™ Golden Gate Assembly Kit (BsmBI-v2) (New England Biolabs #E1602L), with a 3: 1 molar ratio of insertvector for 0.3 kb inserts. Assembly reactions were incubated at 42 °C for 1 h, and heat inactivated at 60 °C for 5 min. Reactions were purified using DNA Clean and Concentrator-5 kit (Zymo #D4014) and electroporated in Endura™ Competent Cells (Biosearch Technologies #60242-2) using a Gene Pulser Xcell™ (Biorad 1652662) set to 1.8 kV, 600 ohms, and 10 pF, and recovering for 60 minutes at 37 °C, 250 rpm in 1 mL of Endura™ recovery media (Biosearch Technologies #60242-2). Cultures were incubated for 6-14 h at 37 °C in 50 mL of LB media with 100 pg / mL of carbenicillin. Assembly and transformation efficiency were assessed, observing around 108colony forming units per pg of transformed DNA. The resulting plasmid library was validated via Illumina MiSeq™ sequencing (500-cycle Nano v2 kit) with a target coverage of 30-100X.

[0383] Generation of reporter cell lines for OPS-RD:

[0384] Lentivirus was generated for the MRGPRX1 GFP reporters and binder libraries as described previously. Reporter cell lines overexpressing the target receptor-GFP fusions were established using lentiviral transduction, as described in Feldman et. al.45.

[0385] Isogenic reporter cell lines were generated by single-cell sorting of GFP+ cells into 96-well plates. After outgrowth of clones, replicate plates were imaged and the final clones were selected based on the expression level and subcellular localization of target receptor-GFP fusions. The binder lentiviral library was prepared as previously described43, with the exception that lentivirus were first titered, and transduction of reporter cell lines targeted an MOI of 5-10%. Libraries were transduced in three biological replicates.

[0386] In situ sequencing:

[0387] Screening was conducted as described previously, with the following modifications. Cells were plated at a density of 15 x 104 / well in 6-well glass-bottom plates (Cellvis #P06-I.5H-N) 72 h prior to in situ sequencing to promote optimal adhesion and spreading. After rolling circle amplification, but prior to the first sequencing cycle, cells were stained with DAPI and imaged in the DAPI, GFP and mCherry™ channels to measure localization of the MRGPRX1 reporter and binder. A total of 9-10 cycles of in situ sequencing were conducted.Image analysis and ranking of designs:

[0388] Localization and in situ sequencing images were analyzed using a Python-based pipeline, as previously described43. Segmentation was done with Cellpose™ v2, using the DAPI stain and non-specific background from in situ sequencing as nuclear and cytoplasmic inputs, respectively. Each cell was assigned a binding score based on pixelwise crosscorrelation of GFP-GPCR with mCherry™-binder. Cells containing a common barcode were then clustered by spatial proximity. Each binder was then scored based on the average binding score amongst its cell clusters. For each of the top 100 binders, the corresponding clusters were inspected by eye, and 60 candidates were selected based on the strength of the localization phenotype and reproducibility across biological replicates.

[0389] In vitro GPCR pharmacology

[0390] cAMP assay for CGRPR was carried out as previously described using commercially available Gs and Gi Cisbio kits46. To measure antagonism of CGRPR binders, a concentration-response curve of the endogenous CGRP was first generated using a SK-N-MC cell line (ATCC, HTB-10). The Gs-mediated cAMP accumulation was measured in a final volume of 40 uL. The stimulation buffer containing 0.5 mM IBMX (Sigma- Aldrich) was used for serial dilutions of tested ligands. Approximately 10 pL of 2,500 cells per well was used to seed cells into a white 384-well plate. The reaction mixture was incubated at 37 °C for 30 min and the reaction was terminated by adding 10 pL of cryptate-labeled cAMP and cAMP d2-labeled antibody, respectively. Following an incubation for 1 hour at room temperature, cellular cAMP levels were quantified by homogeneous time-resolved fluorescence resonance energy transfer (HTRF, ratio 665 / 620 nm) on a Neo2™ plate reader (Agilent). Screening of CGRPR antagonist binders was conducted by analogy except for preincubating binders for 30 min at 37°C followed by CGRP incubation for an additional 30 minutes under the same conditions.

[0391] For the luciferase assay, CHO-Kl / Cre-Luc / CGRPR cells (MOO 187, GenScript) were seeded at a density of 10,000 cells per well in 20 pL of growth medium in a white 384-well plate (Cat. No.: 3570, Coming). The cells were incubated overnight (approximately 16 hours) at 37°C with 5% CO2. A concentration-response curve of the agonist aCGRP was generated to determine its ECso. The antagonistic activity of the CGRPR binders was assessed in the presence of ECso of aCGRP. After the overnight incubation, 4x working solutions of the ligands were prepared by serially diluting the antagonist or agonist in growth medium.Subsequently, 10 pL of the 4x antagonist or growth medium was added to each well. After a 30-minute incubation at 37°C, 10 pL of the 4x agonist working solution was added to each well, and the cells were further incubated for 6 hours at 37°C with 5% CO2. Following treatment, 40 pL of Bio-Gio™ Luciferase Assay Detection Solution (Cat. No.: G7941, Promega) was added to each well to initiate the luminescent reaction. Luminescence was then measured using a SpectraMax™ iD5 Multimode Plate Reader (Molecular Devices).).

[0392] For calcium mobilization assays, CHO-K1 PathHunter™ MRGPRX1 P-arrestin cells were seeded in a total volume of 20 pL / well, in black, clear-bottom, Poly-D-lysine coated 384-well microplates and incubated at 37°C. Subsequently, media was replaced with 20 pL of Dye Loading Buffer, consisting of IX Dye, IX Additive A, 2.5 mM Probenecid (freshly prepared) in HBSS / 20 mM HEPES, and incubated for 30-60 minutes at 37°C. For agonism, cells were incubated with 10 pL of HBSS / 20 mM HEPES. The vehicle (prepared at 3X concentration) was included in the buffer when generating agonist concentration response curves to obtain the ECso for subsequent antagonist screening. Cells were incubated in the dark for 30 minutes at room temperature. The agonist activity of ligands was measured on a FLIPR™ Tetra (MDS) or a FlexStation™ 3 (MDS). 10 pL of the sample (prepared at 4X concentration in HBSS / 20 mM HEPES) was added to the cells 5 seconds before calcium mobilization was monitored for 2 minutes. For antagonist measurements, after dye loading, 10 pL of the sample (prepared 3X) was added and cells were incubated for 30 minutes at room temperature. 10 pL of an ECso of the agonist, prepared in HBSS / 20 mM HEPES, was added to the cells 5 seconds before calcium mobilization was monitored for 2 minutes.

[0393] Kinetic measurements for GIPR and GCGR binders

[0394] Binding studies were executed on a Biacore™ 8K (Cytiva) instrument. The experiments were conducted at 25°C. Biotinylated GIPR (SinoBiological, 18774-H49H-B) and GCGR (Aero Biosystems, GCR-H82E3) ectodomain proteins were captured by Streptavidin using Biotin CAPture Kit (Cytiva #28920234) following the manufacturer’s guidelines. The GIPR or GCGR ectodomain samples at concentration of 0.125 pg / mL were injected at a flow rate of 10 pL / min in HBS-EP+ (0.01 M HEPES pH 7.4, 0.15 M NaCl, 3 mM EDTA, 0.005% v / v Surfactant P20, Cytiva #BR100669) aiming for a capture level of -150 response units. The kinetic measurements of the best 96 designs from yeast library screening were performed by injecting them as analytes in increasing concentrations ranging from 0.0128 nM, 0.064 nM, 0.32 nM, 1.6 nM, 8 nM, 40 nM, 200 nM, 1000 nM to 5000 nM in a single cycle with 9 steps. Analytes were diluted in HBS-EP+ and injected at a flow rateof 30 pL / min to monitor association. HBS-EP+ was used as a running buffer during dissociation at a flow of 30 pL / min. Binding kinetics were determined by global fitting of curves assuming a 1: 1 Langmuir interaction using the Cytiva evaluation software.

[0395] Targeted mutagenesis

[0396] Design models of mM1_068, mM1_060, and mMl_034, together with cryo-EM structures of mM1_068 and mM1_060, were relaxed with Rosetta™ FastRelax™ and used as inputs for mutagenesis. Positions were selected for mutation if the residue at that site (i) showed a change in solvent-accessible surface area (SASA) upon binding or (ii) had SASA < 20 A2in the unbound binder. All 19 amino acid substitutions were introduced at each site and evaluated with Rosetta™ cartesian ddg to compute changes in binding free energy (AAG, REU) and binder energy (REU). Substitutions were classified as favorable if they reduced AAG by < -1.0 REU without increasing monomer energy above +5 REU or maintained AAG < 0.0 REU while lowering monomer energy below -5 REU. Favorable mutations were visually inspected and combined into single, double, and triple variants based on stabilizing biophysical interactions, yielding 152 variants for mM1_068, 78 for mMl_034, and 58 for mM1_060, for a total of 288 variants.

[0397] Pharmacological data analysis

[0398] In vitro pharmacological analysis was carried out with GraphPad Prism (GraphPad Software, San Diego). Data are presented as means ± SEM over technical sample averaged biological replicates. For the luciferase assay, Relative Light Units (RLU) values were obtained by subtracting the luminescence values of the background (media alone + Luciferase Assay Detection Solution) to the ones of each sample. RLU values from the luciferase assay, relative fluorescence units (RFU) from calcium mobilization assay and data from cAMP were fitted to three-parameter nonlinear regression curves, a slope of one and logarithmic scale. Responses were then normalized using the following equation: (signal of test sample - signal of vehicle control) / (positive control ligand - signal of vehicle control), with positive control ligands being CGRP for CGRPR (100 nM in a cAMP assay with 1 pM in a cAMP assay with SK-N-MC cells when assaying RFdiffusion™ designs, or 1 pM in a CRE-Luc assay with CHO-K1 / CGRPR cells when assaying MetaGen™ designs), and BAM 8-22 for MRGPRX1 (0.1 pM).

[0399] For CGRPR, data were normalized to 100%, i.e. the saturating concentration of CGRP in the assay (either 100 nM or 1 pM) and fitted to three-parameter nonlinearregression curves using Global Gaddum-Schild regression analysis.. Data from calcium flux were fitted to four-parameter nonlinear regression curves. To measure antagonism, percentage inhibition was calculated by normalizing the RFU of the test sample relative to the response achieved with the ECso of BAM (8-22) control.

[0400] For MRGPRX1 calcium flux data, the peak of the fluorescence trace after ligand stimulation was divided by the average baseline signal for each well and subtracted by 1 to obtain AF / Fo ratios. These ratios were normalized to 100% = highest signal from the reference ligand BAM 8-22, 0% = buffer control. In order to determine ICso values, the miniprotein concentration-response curves were fitted to a four-parameter logistic equation (4-PL) after normalizing to the bottom and top values of the curves corresponding to unstimulated (negative) and stimulated (positive) control wells, respectively. Analysis was performed in GraphPad Prism 10.4.1.

[0401] References

[0402] 1. Hauser, A. S., Attwood, M. M., Rask-Andersen, M., Schioth, H. B. & Gloriam, D. E. Trends in GPCR drug discovery: new agents, targets and indications. Nat. Rev. Drug Discov. 16, 829-842 (2017).

[0403] 2. Congreve, M., De Graaf, C., Swain, N. A. & Tate, C. G. Impact of GPCR Structures on Drug Discovery. Cell 181, 81-91 (2020).

[0404] 3. Zhang, M. et al. G protein-coupled receptors (GPCRs): advances in structures, mechanisms, and drug discovery. Signal Transduct. Target. Ther. 9, 88 (2024).

[0405] 4. Ren, H. et al. Function-based high-throughput screening for antibody antagonists and agonists against G protein-coupled receptors. Commun. Biol. 3, 146 (2020).

[0406] 5. Ma, Y. et al. Structure -guided discovery of a single-domain antibody agonist against human apelin receptor. Sci. Adv. 6, eaax7379 (2020).

[0407] 6. Fontaine, T. et al. Structure elucidation of a human melanocortin-4 receptor specific orthosteric nanobody agonist. Nat. Commun. 15, 7029 (2024).

[0408] 7. Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089-1100 (2023).

[0409] 8. Cao, L. et al. Design of protein-binding proteins from the target structure alone. Nature 605, 551-560 (2022).

[0410] 9. Roy, A. et al. De novo design of highly selective miniprotein inhibitors of integrins avP6 and avP8. Nat. Commun. 14, 5660 (2023).

[0411] 10. Berger, S. et al. Preclinical proof of principle for orally delivered Th 17 antagonistminiproteins. Cell 187, 4305-4317. el8 (2024).

[0412] 11. Case, J. B. et al. Ultrapotent miniproteins targeting the SARS-CoV-2 receptor-binding domain protect against infection and disease. Cell Host Microbe 29, 1151-1161.e5 (2021).

[0413] 12. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021).

[0414] 13. Varadi, M. et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. 50, D439-D444 (2022).

[0415] 14. Bennett, N. R. et al. Improving de novo protein binder design with deep learning. Nat. Commun. 14, 2625 (2023).

[0416] 15. Dauparas, J. et al. Robust deep learning-based protein sequence design using ProteinMPNN. Science 378, 49-56 (2022).

[0417] 16. Chao, G. et al. Isolating and engineering human antibodies using yeast surface display. Nat. Protoc. 1, 755-768 (2006).

[0418] 17. Ju, M.-S. et al. A human antibody against human endothelin receptor type A that exhibits antitumor potency. Exp. Mol. Med. 53, 1437-1448 (2021).

[0419] 18. Krohl, P. J. et al. Discovery of antibodies targeting multipass transmembrane proteins using a suspension cell-based evolutionary approach. Cell Rep. Methods 3, 100429 (2023).

[0420] 19. Fridy, P. C. et al. A robust pipeline for rapid production of versatile nanobody repertoires. Nat. Methods 11, 1253-1260 (2014).

[0421] 20. Cao, C. et al. Structure, function and pharmacology of human itch GPCRs. Nature 600, 170-175 (2021).

[0422] 21. Liu, Y. et al. Ligand recognition and allosteric modulation of the human MRGPRX1 receptor. Nat. Chem. Biol. 19, 416-422 (2023).

[0423] 22. Olsen, R. H. J. et al. TRUPATH, an open-source biosensor platform for interrogating the GPCRtransducerome. Nat. Chem. Biol. 16, 841-849 (2020).

[0424] 23. DiBerto, J. F., Olsen, R. H. J. & Roth, B. L. TRUPATH: An Open-Source Biosensor Platform for Interrogating the GPCR Transducerome. Methods Mol. Biol. Clifton NJ 2525, 185-195 (2022).

[0425] 24. Kroeze, W. K. et al. PRESTO-Tango as an open-source resource for interrogation of the druggable human GPCRome. Nat. Struct. Mol. Biol. 22, 362-369 (2015).

[0426] 25. Guo, L. et al. Ligand recognition and G protein coupling of the human itch receptor MRGPRX1. Nat. Commun. 14, 5004 (2023).

[0427] 26. Shi, Y., Riese, D. J. & Shen, J. The Role of the CXCL 12 / CX CR4 / CXCR7 ChemokineAxis in Cancer. Front. Pharmacol. 11, 574667 (2020).

[0428] 27. Saotome, K. et al. Structural insights into CXCR4 modulation and oligomerization. Nat. Struct. Mol. Biol. 32, 315-325 (2025).

[0429] 28. Jurek, B. & Neumann, I. D. The Oxytocin Receptor: From Intracellular Signaling to Behavior. Physiol. Rev. 98, 1805-1908 (2018).

[0430] 29. Liang, Y.-L. et al. Toward a Structural Understanding of Class B GPCR Peptide Binding and Activation. Mol. Cell 77, 656-668.e5 (2020).

[0431] 30. Edvinsson, L., Haanes, K. A., Warfvinge, K. & Krause, D. N. CGRP as the target of new migraine therapies — successful translation from bench to clinic. Nat. Rev. Neurol.

[0432] 14, 338-350 (2018).

[0433] 31. Yao, S. et al. De novo design and directed folding of disulfide-bridged peptide heterodimers. Nat. Commun. 13, 1539 (2022).

[0434] 32. Cary, B. P. et al. New Insights into the Structure and Function of Class B 1 GPCRs. Endocr. Rev. 44, 492-517 (2023).

[0435] 33. Booe, J. M. et al. Structural Basis for Receptor Activity-Modifying Protein-Dependent Selective Peptide Recognition by a G Protein-Coupled Receptor. Mol. Cell 58, 1040-1052 (2015).

[0436] 34. Hauser, A. S. et al. GPCR activation mechanisms across classes and macro / microscales. Nat. Struct. Mol. Biol. 28, 879-888 (2021).

[0437] 35. Lavington, S. & Watts, A. Lipid nanoparticle technologies for the study of G protein-coupled receptors in lipid environments. Biophys. Rev. 12, 1287-1302 (2020).

[0438] 36. Steinegger, M. & Soding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026-1028 (2017).

[0439] 37. Qin, L. et al. Structural biology. Crystal structure of the chemokine receptor CXCR4 in complex with a viral chemokine. Science 347, 1117-1122 (2015).

[0440] 38. Brunette, T. J. et al. Exploring the repeat protein universe through computational protein design. Nature 528, 580-584 (2015).

[0441] 39. Hsia, Y. et al. Design of multi-scale protein complexes by hierarchical building block fusion. Nat. Commun. 12, 2294 (2021).

[0442] 40. Wicky, B. I. M. et al. Hallucinating symmetric protein assemblies. Science 378, 56-61 (2022).

[0443] 41. Cheng, Z. et al. Inhibition of BET bromodomain targets genetically diverse glioblastoma. Clin. Cancer Res. Off. J. Am. Assoc. Cancer Res. 19, 1748-1759 (2013).

[0444] 42. Boder, E. T. & Wittrup, K. D. Yeast surface display for screening combinatorialpolypeptide libraries. Nat. Biotechnol. 15, 553-557 (1997).

[0445] 43. Feldman, D. et al. Pooled genetic perturbation screens with image-based phenotypes. Nat. Protoc. 17, 476-512 (2022).

[0446] 44. Feldman, D. et al. Optical Pooled Screens in Human Cells. Cell 179, 787-799. el7 (2019).

[0447] 45. Feldman, D. et al. Pooled genetic perturbation screens with image-based phenotypes. Nat. Protoc. 17, 476-512 (2022).

[0448] 46. Muratspahic, E. et al. Design and structural validation of peptide-drug conjugate ligands of the kappa-opioid receptor. Nat. Commun. 14, 8064 (2023).

[0449] 47. Kumar, B. A., Kumari, P., Sona, C. & Yadav, P. N. GloSensor assay for discovery of GPCR-selective ligands, in Methods in Cell Biology vol. 142 27–50 (Elsevier, 2017). 48. Maharana, J. et al. Structural snapshots uncover a key phosphorylation motif in GPCRs driving P-arrestin activation. Mol. Cell 83, 2091-2107.e7 (2023).

[0450] 49. Dwivedi-Agnihotri, H. et al. An intrabody sensor to monitor conformational activation of -arrestins. in Methods in Cell Biology vol. 169 267–278 (Elsevier, 2022).

[0451] 50. Inoue, A. et al. Illuminating G-Protein-Coupling Selectivity of GPCRs. Cell 177, 1933-1947.e25 (2019).

[0452] 51. Fink, E. A. et al. Structure-based discovery of nonopioid analgesics acting through the a 2A -adrenergic receptor. Science 377, eabn7065 (2022).

[0453] 52. Singh, I. et al. Structure-based discovery of conformationally selective inhibitors of the serotonin transporter. Cell 186, 2160-2175. el7 (2023).

[0454] 53. Wang, Y. et al. Structures of the entire human opioid receptor family. Cell 186, 413-427.el7 (2023).

[0455] 54. Measuring surface expression and endocytosis of GPCRs using whole-cell ELISA, in Methods in Cell Biology vol. 149 131-140 (Elsevier, 2019).

[0456] 55. Liang, Y.-L. et al. Cryo-EM structure of the active, Gs-protein complexed, human CGRP receptor. Nature 561, 492-497 (2018).

[0457] 56. Josephs, T. M. et al. Structure and dynamics of the CGRP receptor in apo and peptide-bound forms. Science 372, eabf7258 (2021).

[0458] 57. Russo, C. J. & Passmore, L. A. Ultrastable gold substrates: Properties of a support for high-resolution electron cryomicroscopy of biological specimens. J. Struct. Biol. 193, 33-44 (2016).

[0459] 58. Scheres, S. H. W. RELION: implementation of a Bayesian approach to cryo-EM structure determination. Struct. Biol. 180, 519-530 (2012).59. Zheng, S. Q. et al. MotionCor2: anisotropic correction of beam-induced motion for improved cryo-electron microscopy. Nat. Methods 14, 331-332 (2017).

[0460] 60. Rohou, A. & Grigorieff, N. CTFFIND4: Fast and accurate defocus estimation from electron micrographs. J. Struct. Biol. 192, 216-221 (2015).

[0461] 61. Kinman, L. F., Powell, B. M., Zhong, E. D., Berger, B. & Davis, J. H. Uncovering structural ensembles from single -particle cryo-EM data using cryoDRGN. Nat. Protoc. 18, 319-339 (2023).

[0462] 62. Zhong, E. D., Bepler, T., Berger, B. & Davis, J. H. CryoDRGN: reconstruction of heterogeneous cryo-EM structures using neural networks. Nat. Methods 18, 176-185 (2021).

[0463] 63. Wagner, T. et al. SPHIRE-crYOLO is a fast and accurate fully automated particle picker for cryo-EM. Commun. Biol. 2, 218 (2019).

[0464] 64. Punjani, A., Rubinstein, J. L., Fleet, D. J. & Brubaker, M. A. cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nat. Methods 14, 290-296 (2017).

[0465] 65. Croll, T. I. ISOLDE: a physically realistic environment for model building into low-resolution electron-density maps. Acta Crystallogr. Sect. Struct. Biol. 74, 519-530 (2018).

[0466] 66. Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr. D Biol. Crystallogr. 66, 486-501 (2010).

[0467] 67. Afonine, P. V. et al. Real-space refinement in PHENIX for cryo-EM and crystallography. Acta Crystallogr. Sect. Struct. Biol. 74, 531-544 (2018).

[0468] 68. Wu, B. et al. Structures of the CXCR4 Chemokine GPCR with Small-Molecule and Cyclic Peptide Antagonists. Science 330, 1066-1071 (2010).

[0469] 69. Qin, L. et al. Crystal structure of the chemokine receptor CXCR4 in complex with a viral chemokine. Science 347, 1117-1122 (2015).We claim

[0470] 1. A polypeptide comprising an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-217, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to a GPCR target.

[0471] 2. The polypeptide of claim 1, wherein the polypeptide comprises an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to calcitonin gene-related peptide receptor (CGRPR).

[0472] 3. The polypeptide of claim 1, wherein the polypeptide comprises an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:26-32 and 213-215, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to the Mas-related GRP family member XI (MRGPRX1) receptor.

[0473] 4. The polypeptide of claim 1, wherein the polypeptide comprises an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO:44-117, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to the glucagon receptor (GCGR).

[0474] 5. The polypeptide of claim 1, wherein the polypeptide comprises an amino acid sequence at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 118-212, not including any amino acid insertions at identified loop positions, wherein the polypeptide binds to gastric inhibitory polypeptide receptor (GIPR).

Claims

6. The polypeptide of any one of claims 1-5, comprising an amino acid sequence at least 60% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions.

7. The polypeptide of any one of claims 1-6, comprising an amino acid sequence at least 70% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions.

8. The polypeptide of any one of claims 1-6, comprising an amino acid sequence at least 80% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215, not including any amino acid insertions at identified loop positions.

9. The polypeptide of any one of claims 1-6, comprising an amino acid sequence at least 90% identical to the amino acid sequence selected from the group consisting of SEQ ID NO: 1-23, 26-32, and 44-215.

10. The polypeptide of any one of claims 1-9, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11. 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or all identified interface residues are identical, or conservatively substituted, relative to the reference sequence.

11. The polypeptide any one of claims 1-9, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or all identified interface residues are identical, relative to the reference sequence12. The polypeptide of any one of claims 1-11, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all identified core residues are identical, or conservatively substituted, relative to the reference sequence.

13. The polypeptide of any one of claims 1-11, wherein at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all identified core residues are identical, relative to the reference sequence.

14. The polypeptide of any one of claims 1-13, wherein all interface and all core residues are identical relative to the reference sequence.

15. The polypeptide of any one of claims 1-14, wherein the polypeptide comprises an amino acid insertion at one or more identified loop positions relative to the reference sequence.

16. The polypeptide of any one of claims 1-14, wherein the polypeptide does not comprise an amino acid insertion at any identified loop positions relative to the reference sequence.

17. The polypeptide of any one of claims 1-16, wherein substitutions relative to the reference sequence are conservative amino acid substitutions.

18. A fusion protein, comprising:(a) the polypeptide of any one of claims 1-17; and(b) one or more functional domains at the N-terminus and / or at the C-terminus of the polypeptide.

19. The polypeptide or fusion protein of any one of claims 1-18, wherein the polypeptide binds its target with nanomolar or picomolar affinity.

20. A nucleic acid encoding the polypeptide or fusion protein of any one of claims 1-19.

21. An expression vector comprising the nucleic acid of claim 20 operatively linked to a suitable control sequence.

22. A host cell comprising the polypeptide, fusion protein, nucleic acid, and / or expression vector of any one of claims 1-21.

23. A pharmaceutical composition, comprising:(a) the polypeptide, the fusion protein, the nucleic acid, the expression vector, and / or the host cell of any one of claims 1-22; and(b) a pharmaceutically acceptable carrier.

24. A method for treating or limiting development of migraine headaches and / or pain, comprising administering to a subject in need thereof an amount effective to treat the migraine and / or pain of the polypeptide of claim 2, or the polypeptide, fusion protein, nucleic acid, expression vector, cell, or pharmaceutical composition of any one of claims 6-23 when depending from claim 2.

25. A method for treating or limiting development of itching or pain, comprising administering to a subject in need thereof an amount effective to treat the itching or pain of the polypeptide of claim 3, or the polypeptide, fusion protein, nucleic acid, expression vector, cell, or pharmaceutical composition of any one of claims 6-23 when depending from claim 3.

26. A method for treating or limiting development of obesity or diabetes (including but not limited to Type 2 diabetes), comprising administering to a subject in need thereof an amount effective to treat the obesity or diabetes of the polypeptide of claim 4 or 5, or the polypeptide, fusion protein, nucleic acid, expression vector, cell, or pharmaceutical composition of any one of claims 6-23 when depending from claim 4-5.