Smart MHC: a de-novo designed platform for soluble expression of peptide-receptive class-i MHC in e. coli

The de-novo designed fusion protein stabilizes class I MHC expression in E. coli, addressing production inefficiencies by enabling stable, efficient production of soluble pMHCs for T-cell studies and peptide library screening.

WO2025221716A1PCT designated stage Publication Date: 2025-10-23UNIV OF WASHINGTON +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024665
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2025-04-15
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

The production of recombinant soluble peptide-major histocompatibility complexes (pMHCs) is hindered by the instability of MHC molecules, leading to inefficient and costly processes that limit simultaneous study of multiple peptides and require specialized facilities for refolding.

Method used

A de-novo designed fusion protein comprising a stabilizer peptide and a truncated class I MHC protein, expressed in E. coli without the need for a placeholder peptide, allowing soluble expression and stable presentation of peptides for T-cell interaction.

Benefits of technology

Enables stable, efficient production of soluble pMHCs for various biological studies, including T-cell staining and structural analysis, without the need for refolding, facilitating in-house production and large-scale peptide library screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024665_23102025_PF_FP_ABST
    Figure US2025024665_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Fusion proteins are provided that include (a) a stabilizer peptide including domains X1-X2, wherein (i) XI includes the amino acid sequence of SEQ ID NO: 1; and (ii) X2 is a first amino acid linker of between 5-20 amino acids in length; and (b) a truncated class I major histocompatibility complex (MHC) protein directly fused to the C-terminus of the stabilizer peptide, wherein the truncated MHC protein consists of residues 1-175, 1-176, 1- 177, 1-178, 1-179, 1-180, 1-181, 1-182, 1-183, 1-184, or 1-185 of a class I MHC protein, optionally wherein the truncated class I MHC protein comprises residues 6K, 81, 121, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]UW 49565.01US1 SMART MHC: A de-novo Designed Platform for Soluble Expression of Peptide- Receptive Class-I MHC in E. Coli Federal Funding Statement This invention was made with government support under Grant No. FA8750-17-C- 0219, awarded by the Defense Advanced Research Projects Agency (DARPA) and Grant No. R01AG063845, awarded by the National Institute on Aging (NIA) [NIH] and Grant No. AI103867, awarded by the National Institutes of Health (NIH). The government has certain rights in the invention. Sequence Listing Statement A computer readable form of the Sequence Listing is filed with this application by electronic submission and is incorporated into this application by reference in its entirety. The Sequence Listing is contained in the file created on March 25, 2025 having the file name “24- 0527-WO.xml” and is 115,475 bytes in size. Background Of the many tools used to study T-cell receptor (TCR) biology, one of the most ubiquitous is the recombinantly expressed soluble peptide–major histocompatibility complex (pMHC). These soluble pMHCs are useful for a wide range of experiments where engagement of a specific T-cell population is important. They can be used as a staining reagent to identify or isolate a T-cell subset that can respond to a peptide of interest. These methods are useful for tracking the response to an infection or cancer, and when paired with TCR sequencing, provide important information about the molecular determinants of TCR- pMHC interactions. In addition to their utility as a staining reagent, recombinant pMHCs can be used directly as a stimulus to study T-cell activation in a variety of contexts, or in X-ray crystallography and cryo-electron microscopy to determine the structural basis for peptide- MHC and pMHC-TCR interactions. While all of these methods provide important information, they are all limited by the difficult process by which recombinant pMHCs are produced. This process, which involves separate E. coli expression of each of the two MHC chains as insoluble inclusion bodies, solubilization in guanidinium and subsequent refolding in the presence of the desired peptide, is expensive, slow, and inefficient. Because the peptide must be present for the refolding reaction to occur, it is difficult to study many peptides simultaneously, and other reagents are often substituted if possible. Summary In one aspect, the disclosure provides fusion proteins, comprising: (a) a stabilizer peptide comprising or consisting of domains X1-X2, wherein (i) X1 comprises or consists of the amino acid sequence of SEQ ID NO:1; and (ii) X2 is a first amino acid linker of between 5-20 amino acids in length; (b) a truncated class I major histocompatibility complex (MHC) protein directly fused to the C-terminus of the stabilizer peptide, wherein the truncated MHC protein consists of residues 1-175, 1-176, 1-177, 1-178, 1-179, 1-180, 1-181, 1-182, 1-183, 1-184, or 1-185 of a class I MHC protein, optionally wherein the truncated class I MHC protein comprises residues 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A. In one embodiment, X2 is a first amino acid linker of 12-16 amino acids in length. In another embodiment, X2 has an amino acid sequence selected from the group consisting of SEQ ID NO:2-12. In a further embodiment, the stabilizer peptide comprises or consists of the amino acid sequence selected from SEQ ID NO:13-22. In one embodiment, the truncated class I MHC protein is a truncated human class I MHC. In another embodiment, the truncated class I MHC protein is a truncated HLA-A, HLA-B, or HLA-C. In a further embodiment, the truncated class I MHC protein is a truncated HLA-E, HLA-F, or HLA-G. In other embodiments, the truncated class I MHC protein consists of the amino acid sequence selected from SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V,and optionally wherein the truncated class I MHC protein comprises residue 84A relative to the reference sequence. In one embodiment, (a) X1 consists of the amino acid sequence of SEQ ID NO:1; (b) X2 consists of the amino acid sequence selected from the group consisting of SEQ ID NO:2-12, or SEQ ID NO:2-11; and (c) the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V,and optionally wherein the truncated class I MHC protein comprises residue 84A. In another embodiment, (a) the stabilizer peptide consists of the amino acid sequence selected from SEQ ID NO:13-22; and (b),the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A. In another embodiment, the fusion protein further comprises an Aga2 domain comprising or consisting of the amino acid sequence of SEQ ID NO:85 fused N-terminal to the stabilizer peptide, optionally via an amino acid linker. In one embodiment, the fusion protein further comprises a signal sequence N-terminal to the Aga2 domain. In a further embodiment, the fusion protein further comprises a peptide tag located between the Aga2 domain and the stabilizer peptide, optionally via amino acid linkers. In one embodiment, the fusion protein further comprises a binding peptide that is capable of binding to the truncated class I MHC protein, wherein the truncated class I MHC protein comprises residue 167A. In another embodiment, the binding peptide is fused C- terminal to the truncated class I MHC protein, optionally via an amino acid linker. In a further embodiment, the fusion protein comprises, in N-terminal to C-terminal order, wherein each domain may be linked via an optional amino acid linker: (i) a signal sequence; (ii) an Aga2 domain; (iii) a peptide tag; (iv) an amino acid linker; (v) a stabilizer peptide; (vi) a truncated class I MHC protein; (vii) an amino acid linker; and (viii) a binding peptide capable of binding to the truncated class I MHC protein. In one embodiment, the fusion protein comprises the amino acid of SEQ ID NO:74 or 75. In another embodiment, the fusion protein further comprises: (i) a binding peptide that is capable of binding to the truncated class I MHC protein, wherein the peptide is fused N-terminal to the stabilizer peptide via an amino acid linker; and (ii) an oligomer-forming polypeptide fused C-terminal to the truncated class I MHC protein, wherein the oligomer-forming polypeptide is fused to the truncated class I MHC protein either directly or via an amino acid linker. In one embodiment, the oligomer-forming polypeptide comprises or consists of the amino acid sequence selected from the group consisting of SEQ ID NO:39-46 and 94. In a further embodiment, the truncated class I MHC protein and the binding peptide that binds to it are selected from a binding peptide and truncated class I MHC protein listed in one row of Table 5. In one embodiment, the fusion protein further comprises a peptide cleavage tag at the N-terminal end of the fusion protein, including but not limited to a peptide cleavage tag comprising or consisting of the amino acid sequence of SEQ ID NO:67. In one embodiment, the fusion protein comprises, in N-terminal to C-terminal order: (A) a peptide cleavage tag comprising or consisting of the amino acid sequence of SEQ ID NO:67; (B) a binding peptide comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:47-66, 83, 88, 91, and 96; (C) an amino acid linker of length 10-20 amino acids; (D) a stabilizer peptide comprising or consisting of the amino acid sequence selected from SEQ ID NO:13-22; (E) an optional amino acid linker of length 1-10 amino acids; (F) a truncated class I MHC protein consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 84A, 12I, 104D, 105E, 106N, and 110V; wherein the truncated class I MHC protein and the binding peptide that binds to it are selected from a binding peptide and truncated class I MHC protein listed in one row of Table 5; (G) an amino acid linker of length 3-30 amino acids; and (H) an oligomer-forming polypeptide comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:39-46 and 94. In another embodiment, the fusion protein may further comprise a detectable peptide tag at the C-terminus of the fusion protein, fused to the oligomer-forming polypeptide either directly or via a fourth optional amino acid linker. In various other embodiments, the fusion protein comprises or consists of the amino acid sequence selected from the group consisting of SEQ ID NO:68-77 and 97-100, wherein optional residues may be present or may be deleted, or may be substituted with other residues. The disclosure also provides oligomers, comprising a plurality of fusion proteins of any embodiment comprising an oligomer-forming polypeptide, wherein the peptide is linked to the truncated class I MHC protein in each identical fusion protein. The disclosure further provides nucleic acids encoding the fusion protein of any embodiment or combination of embodiments disclosed herein; expression vectors comprising the nucleic acid operatively linked to a suitable control sequence, such as a promoter, and host cells comprising the fusion protein, oligomer, nucleic acid, or expression vector of any embodiment or combination of embodiments disclosed herein. The disclosure further provides methods for using the fusion proteins, nucleic acids, expression vectors, and host cells of any embodiment herein, for a method including but not limited to yeast surface display, studying TCR docking geometries and other signaling mechanisms, as well as X-ray crystallography and cryo-electron microscopy to determine the structural basis for peptide-MHC and pMHC-TCR interactions, as a staining reagent to identify or isolate a T-cell subset that can respond to a peptide of interest, for tracking the response to an infection or cancer, or used for T-cell sorting and sequencing experiments to identify TCR sequences that recognize a pMHC of interest. Description of the Figures Figure 1. Design of soluble, monomeric, antigen-receptive, truncated (SMART) MHCs. A) Schematic of the initial design process used to generate a library of candidate stabilizing domains. B) Schematic of the yeast display sorting strategy. Yeast transformed with the design library were sorted first for surface display of the design (top). This sorted population was cultured and re-sorted for binding to FITC-gp33 peptide (bottom). C) Plots showing the distribution of expression (top) and peptide binding (bottom) values for the yeast display library. Dashed lines show sorting gates, with cells to the right of the gate being collected. D) KDEs of the distribution of enrichment values of all designs from the yeast display experiment. Designs are split into three categories: sequence-shuffled negative controls (shuffled), designs that do not replace W60 from β2m (-Trp), and designs that do replace W60 from β2m (+Trp). The relative enrichment for each design was calculated by dividing its abundance after the indicated sort by its abundance prior to that sort. E) SEC traces for the best hit from the yeast display library (left) and the improved version after cleavage site mutations and linker redesign (right) on an S75 column. The input material for SEC was purified on Ni-NTA resin from the soluble fraction of E. coli lysate from a 50mL culture without refolding in both cases. Figure 2. SMART H-2Dbretains native interactions and structure. Fluorescence polarization (FP) data (points) and fitted binding curves (lines) for FITC-gp33 binding to SMART H-2Dbpurified the day before the measurement (4C, 1 day; black circles and solid line), or one month before and stored in at 4oC (4C, 1 month; dark gray triangles and dashed line) or flash-frozen and stored at -80oC (Freeze / Thaw; light gray diamonds and dotted line). Figure 3. CSM8 A*02:01 retains native interactions. A) SEC trace for empty CSM8 A*02:01. Elution volume for the monomeric species is roughly 12mL. B) FP data (filled circles) and fitted binding curve (line) for CSM8 A*02:01 binding to AF488-NY-ESO- 1 peptide. Figure 4. Hit6 A*02:01 improves pMHC yeast display. A) Schematic of native A*02:01 yeast display construct. The Y84A mutation is used to allow room for the linker on the C-terminus of the peptide to leave the binding pocket. B) Schematic of SMART A*02:01 display constructs with and without TAX peptide fused. The W167A mutation is used to allow room for the linker on the N-terminus of the peptide to leave the binding pocket. LS: leader sequence for surface display. HA: HA peptide tag for measuring surface display. C) Scatterplots from flow cytometry measurements of expression levels of various A*02:01 constructs on the surface of yeast. Surface expression was measured by binding of an anti- HA antibody, and by binding of an A*02:01-specific antibody. D) Scatterplots from flow cytometry measurements ofA6 TCR binding of various A*02:01 constructs on the surface of yeast. TCR binding was measured with streptavidin tetramers of the A6 TCR. Figure 5. Peptide-fused SMART MHC oligomers stain T-cells in a TCR-specific manner. A) Twelve identical subunits assemble to form a tetrahedral oligomer. A single subunit is highlighted in dark gray. B) Schematic of the SMART MHC oligomer sequence. Ulp1 protease is used to cleave the SUMO tag from the peptide, leaving a clean peptide N- terminus to bind the SMART MHC peptide binding groove. The MHC sequence bears the Y84A mutation to allow the C-terminal linker from the peptide to leave the binding groove. The peptide-fused SMART MHC is then fused to an oligomerization domain, followed by a Myc tag. C) The SMART MHC oligomer self-assembles into the same shape as the base oligomer in (A). Each subunit presents a peptide-fused SMART MHC to the T-cell and a Myc tag for secondary staining, but only one subunit is shown with these for simplicity. D) Flow cytometry measurements of T-cell staining of P14 T-cells (dark gray) or control TCR T-cells (light gray) with SMART H-2Dboligomers fused to several variants of the gp33 peptide. Numerical values under the peptide names indicate the binding strength of each peptide in the native MHC context for the P14 TCR relative to gp33 (KD,gp33 / KD,mut). Values for V3P, PF, and Y4F were previously published by Duru et al. Value for M9C was published previously by Boulter et al. Figure 6. Stabilization of multiple HLA alleles using the SMART system A) SEC traces of empty HLA alleles fused to SMART stabilizer. Expected monomer range for the S75 column on the HPLC is 2.1-2.4 minutes. B) Plot of SMART HLA expression statistics showing total area under the SEC curve versus the fraction of the sample that is monomeric. C) Bar plot of native and SMART HLA soluble expression levels with (+) and without (-) a linked peptide expressed in vitro. Soluble protein (µM) was quantified using radioactive14C- leucine incorporation. Average of three replicates (n = 3) is shown for each construct. Figure 7. Alternative linkers maintain gp33 peptide binding. FP data (points) and fitted binding curves (lines) for FITC-gp33 binding to SMART H-2Dbusing either the L8 (black circles and solid line), L11 (dark gray triangles and dashed line), or L15 (light gray diamonds and dotted line) linker variants. The final SMART MHC design uses linker L11. Figure 8. Alternative oligomerization domains can be used for T-cell staining. Histograms showing staining of WT Jurkat cells (light gray) or P14 TCR Jurkats (dark gray) using SMART H-2Dbfused to the gp33 peptide, and either the SB175 (top left), C6-79-10 (top middle), HE0521 (top right), HE0427 (middle left), HE0381 (middle middle), HE0433 (middle right), or HE0368 (bottom left) oligomer. All oligomers were tagged with a C- terminal GFP tag, and cells were stained with 500nM of the designated oligomer, and then with an anti-GFP antibody for flow cytometric analysis. Note: HE0521 and HE0433 did not oligomerize consistently – some other batches performed more poorly (data not shown). Detailed Description All references cited are herein incorporated by reference in their entirety. Within this application, unless otherwise stated, the techniques utilized may be found in any of several well-known references such as: Molecular Cloning: A Laboratory Manual (Sambrook, et al., 1989, Cold Spring Harbor Laboratory Press), Gene Expression Technology (Methods in Enzymology, Vol.185, edited by D. Goeddel, 1991. Academic Press, San Diego, CA), “Guide to Protein Purification” in Methods in Enzymology (M.P. Deutscher, ed., (1990) Academic Press, Inc.); PCR Protocols: A Guide to Methods and Applications (Innis, et al. 1990. Academic Press, San Diego, CA), Culture of Animal Cells: A Manual of Basic Technique, 2ndEd. (R.I. Freshney.1987. Liss, Inc. New York, NY), Gene Transfer and Expression Protocols, pp.109-128, ed. E.J. Murray, The Humana Press Inc., Clifton, N.J.), Dang, B. et al. SNAC-tag for sequence-specific chemical protein cleavage. Nat. Methods 16, 319–322 (2019), and the Ambion 1998 Catalog (Ambion, Austin, TX). As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. As used herein, the amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V). Any N-terminal residues are optional, and may be present, or may be deleted. In some embodiments, 1, 2, 3, 4, or 5 amino-terminal and / or carboxy-terminal residues may be deleted from fusion proteins and polypeptides of the disclosure. All embodiments of any aspect of the disclosure can be used in combination, unless the context clearly dictates otherwise. Unless the context clearly requires otherwise, throughout the description and the claims, the words ‘comprise’, ‘comprising’, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”. Words using the singular or plural number also include the plural and singular number, respectively. Additionally, the words “herein,” “above,” and “below” and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of the application. In a first aspect, the disclosure provides fusion proteins, comprising: (a) a stabilizer peptide comprising or consisting of domains X1-X2, wherein (i) X1 comprises or consists of the amino acid sequence DREDVERLLRSVEWAIKAGDPYSARILVELAREDAEKIGDERLRREVEELLRELEEL (SEQ ID NO:1); and (ii) X2 is a first amino acid linker of between 5-20 amino acids in length; (b) a truncated class I major histocompatibility complex (MHC) protein directly fused to the C-terminus of the stabilizer peptide, wherein the truncated MHC protein consists of residues 1-175, 1-176, 1-177, 1-178, 1-179, 1-180, 1-181, 1-182, 1-183, 1-184, or 1-185 of a class I MHC protein, optionally wherein the truncated class I MHC protein comprises residues 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A or 167A. Class I MHCs can be treated identically in this context because of the high degree of conservation of their structure and the portions of their sequence that maintain their structure. Herein, residues 1-175 (and so on) refer to the amino terminal amino acids C-terminal to the signal peptide of a class I MHC. In their native context on the surface of a cell, the MHCs will not include the signal sequence. The fusion proteins of the disclosure can be solubly expressed in E. coli without the need for a placeholder peptide, and thus provide a solution for the inherent instability of the MHC molecule. The fusion proteins can be used, for example, in yeast surface display, allowing screening of large peptide libraries without prior optimization of the MHC sequence for yeast display. The fusion proteins also can be used as the basis for oligomeric reagents which can stain T-cells in a TCR- and peptide-specific manner. The fusion proteins also can be used in studying TCR docking geometries and other signaling mechanisms, as well as X- ray crystallography and cryo-electron microscopy to determine the structural basis for peptide-MHC and pMHC-TCR interactions. In some embodiments, the truncated class I MHC protein comprises residues 6K, 8I, 12I, 104D, 105E, 106N, and 110V. These modifications are near potential proteolytic cleavage sites on the fusion protein that were observed when it was expressed in E. coli and can help reduce the rate of proteolysis. In other embodiments, the truncated class I MHC protein comprises residue 84A or 167A. These modifications allow the fusion of a linker and a peptide by providing space for the linker to exit the peptide binding groove.84A is used when a peptide and linker are fused to the amino terminus, while 167A is used when the linker and peptide are fused to the carboxy terminus. In other embodiments, the truncated MHC protein consists of residues 1-178, 1-179 or 1-180 of a class I MHC protein. The residues carboxy-terminal to these are omitted because they are not involved directly in contacts with a peptide or TCR. Additionally, they contain structural elements which could interfere with folding and solubility when expressed in E. coli. A range of truncation sites is possible because residues 178-185 constitute a flexible loop connecting the peptide binding groove (residues 1-177) to another domain of the protein. Any truncation within this loop is unlikely to have a large impact on the structure and solubility of the resulting fusion protein. The amino acid linker may be of any length as deemed appropriate for an intended use. In some embodiments, X2 is a first amino acid linker of 12-16 amino acids in length. In all amino acid linkers of the disclosure, the linker may be of any amino acid composition. In one embodiment, the linkers may have any sequence composed exclusively of G, S, and A residues. In other embodiments, X2 has an amino acid sequence selected from the group consisting of SEQ ID NO:2-12, as shown in Table 1. Table 1 In one specific embodiment, X2 is EEELARLPKLPPD (SEQ ID NO:6). In other embodiments, the stabilizer peptide (X1-X2) comprises or consists of the amino acid sequence selected from SEQ ID NO:13-22, as shown in Table 2. Table 2 Any truncated class I MHC protein may be used as deemed appropriate for an intended use. In one embodiment, the truncated class I MHC protein is a truncated human class I MHC. Human class I MHCs are also known as human leukocyte antigens (HLAs), and include 6 (HLA-A, -B, -C, -E, -F, -G). In some embodiments, the truncated class I MHC protein is a truncated HLA-A, HLA-B, or HLA-C. In other embodiments, the truncated class I MHC protein is a truncated HLA-E, HLA-F, or HLA-G. In other embodiments, the truncated class I MHC protein consists of the amino acid sequence selected from SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V,and optionally wherein the truncated class I MHC protein comprises residue 84A. The amino acid sequences of SEQ ID NO:23-38 are provided in Table 3. In some embodiments, the truncated class I MHC protein has the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V relative to the reference sequence. In other embodiments, the truncated class I MHC protein comprises residue 84A relative to the reference sequence. Table 3 In some embodiments of the fusion proteins: (a) X1 consists of the amino acid sequence of SEQ ID NO:1; (b) X2 consists of the amino acid sequence selected from the group consisting of SEQ ID NO:2-12, or SEQ ID NO:2-11; and (c) the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A relative to the reference sequence. In further embodiments, the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38 with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V. In other embodiments of the fusion proteins: (a) the stabilizer peptide consists of the amino acid sequence selected from SEQ ID NO:13-22; and (b) the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A relative to the reference sequence. In further embodiments, the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V. In some embodiments, the truncated class I MCH protein does not include residue 84A relative to the reference sequence. In other embodiments, the truncated class I MCH protein includes residue 84A relative to the reference sequence. Residue 84A provides a space for the peptide linker to exit the binding pocket properly and permit the oligomer to function properly in embodiments with an oligomer. In another embodiment, the fusion protein further comprises an Aga2 domain (SEQ ID NO:85) fused N-terminal to the stabilizer peptide, optionally via an amino acid linker. This embodiment is particularly useful for yeast cell surface display studies because it can be displayed at much higher levels than conventional class I MHCs, suggesting the ability to enhance detection of weak binding to TCRs. Aga2 domain: QELTTICEQIPSPTLESTPYSLSTTTILANGKAMQGVFEYYKSVTFVSNCGSHPSTTSK GSPINTQYVFKDNSSTIEGRTR (SEQ ID NO:85) In a further embodiment, the fusion protein further comprises a signal sequence N- terminal to the Aga2 domain. Any signal sequence may be used as appropriate for secretion in a host cell of interest. In one embodiment, the leader sequence comprises or consists of the amino acid sequence of MQLLRCFSIFSVIASVLA (SEQ ID NO:86). In another embodiment, the fusion protein may further comprise a peptide tag located between the Aga2 domain and the stabilizer peptide, optionally via amino acid linkers. The peptide tag is useful for measuring surface display on yeast cell surface display embodiments. Any peptide tag may be used as appropriate for an intended use. In one embodiment, the peptide tag comprises or consists of the HA peptide (YPYDVPDYA; SEQ ID NO:90). In one embodiment, the fusion protein further comprises a binding peptide that is capable of binding to the truncated class I MHC protein, wherein the truncated class I MHC protein comprises residue 167A. Residue 167A allows the peptide linker (the linker connecting the binding peptide to the construct) to properly exit the binding pocket when the peptide is fused at the C-terminus of the construct. Thus, in another embodiment, the binding peptide is fused C- terminal to the truncated class I MHC protein, optionally via an amino acid linker. In a further embodiment, the fusion protein comprises, in N-terminal to C-terminal order, wherein each domain may be linked via an optional amino acid linker: (i) a signal sequence; (ii) an Aga2 domain; (iii) a peptide tag; (iv) an amino acid linker; (v) a stabilizer peptide of any embodiment or combination of embodiments disclosed herein; (vi) a truncated class I MHC protein of any embodiment or combination of embodiments disclosed herein; (vii) an amino acid linker of any embodiment or combination of embodiments disclosed herein; and (viii) a binding peptide capable of binding to the truncated class I MHC protein. In exemplary embodiments, a fusion protein having this arrangement may comprise or consist of the amino acid of SEQ ID NO:74 or 75. In a different embodiment, the fusion protein further comprises (i) a binding peptide that is capable of binding to the truncated class I MHC protein, wherein the peptide is fused N-terminal to the stabilizer peptide via an amino acid linker; (ii) an oligomer-forming polypeptide fused C-terminal to the truncated class I MHC protein, wherein the oligomer-forming polypeptide is fused to the truncated class I MHC protein either directly or via an amino acid linker. The oligomer-forming polypeptide may be any protein that can oligomerize as part of the recited fusion protein. In non-limiting embodiments, the oligomer-forming polypeptide comprises or consists of the amino acid sequence selected from the group consisting of SEQ ID NO:39-46 and 94, shown in Tables 4 and 10. Table 4 In one specific embodiment, the oligomer-forming polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:46. In another embodiment, the truncated class I MHC protein and the binding peptide that binds to it are selected from a binding peptide and truncated class I MHC protein listed in one row of Table 5. In another embodiment, the fusion protein further comprises a peptide cleavage tag at the N-terminal end of the fusion protein. As used herein, a “cleavage tag” is a peptide that can be cleaved completely leaving no amino acids N-terminal to the cleavage tag. Any cleavage tag can be used as appropriate for an intended use. In one embodiment, the cleavage tag comprises or consists of a small ubiquitin-like modifier (SUMO) tag or variants thereof. In one embodiment, the peptide cleavage tag comprises or consists of the amino acid sequence of SEQ ID NO:67. MDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSL RFLYDGIRIQADQAPEDLDMEDNDIIEAHREQIGG (SEQ ID NO:67; SUMO) In another embodiment, the fusion protein comprises, in N-terminal to C-terminal order: (A) a peptide cleavage tag comprising or consisting of the amino acid sequence of SEQ ID NO:67; (B) a binding peptide comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:47-66, 83, 88, 91, and 96; (C) an amino acid linker of length 10-20 amino acids; (D) a stabilizer peptide comprising or consisting of the amino acid sequence selected from SEQ ID NO:13-22; (E) an optional amino acid linker of length 1-10 amino acids; (F) a truncated class I MHC protein consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 84A, 12I, 104D, 105E, 106N, and 110V; wherein the truncated class I MHC protein and the binding peptide that binds to it are selected from a binding peptide and truncated class I MHC protein listed in one row of Table 5; (G) an amino acid linker of length 3-30 amino acids; and (H) an oligomer-forming polypeptide comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:39-46 and 94. In other embodiments, the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, with the following substitutions relative to the reference sequence: 6K, 8I, 84A, 12I, 104D, 105E, 106N, and 110V. In a further embodiment, the fusion protein further comprises a detectable peptide tag at the C-terminus of the fusion protein, fused to the oligomer-forming polypeptide either directly or via an optional fourth amino acid linker. The detectable peptide tag may be any peptide sequence that can be detected by its own fluorescent or luminescent properties, or by the binding of a fluorescent or luminescent reagent to it. In a non-limiting embodiment, the detectable tag comprises or consists of the amino acid sequence of EQKLISEEDL (SEQ ID NO:89). In all embodiments disclosed herein, the fusion protein may further comprise a purification tag, to aid in purifying the expressed fusion protein. Any purification tag may be used, as deemed suitable for an intended purpose. In one embodiment, the purification tag is located at the C-terminus of the fusion protein. Table 5 In various embodiments, the fusion proteins comprise or consist of the amino acid sequence selected from the group consisting of SEQ ID NO:68-77 and 97-100, wherein residues in parentheses are optional and may be present or may be deleted, or may be substituted with other residues. The amino acid sequences of SEQ ID NO:68-77 and 97-100 are provided in Table 6 and Table 12. Table 6 In another embodiment, the disclosure provides oligomers, comprising a plurality of fusion proteins of any embodiment or combination of embodiments herein that comprise an oligomer-forming polypeptide, wherein the peptide is linked to the truncated class I MHC protein in each identical fusion protein. In some embodiments, the oligomer is a homo- oligomer. In other embodiments, the oligomer is a hetero-oligomer. In these embodiments, the oligomers are soluble pMHCs that are useful in a wide range of applications where engagement of a specific T-cell population is important. For example, they may be used as a staining reagent to identify or isolate a T-cell subset that can respond to a peptide of interest, for tracking the response to an infection or cancer, or used for T-cell sorting and sequencing experiments to identify TCR sequences that recognize a pMHC of interest. In another aspect the disclosure provides nucleic acids encoding the fusion protein of any embodiment or combination of embodiments of the disclosure. The nucleic acid sequence may comprise single stranded or double stranded RNA (such as mRNA) or DNA in genomic or cDNA form, or DNA-RNA hybrids, each of which may include chemically or biochemically modified, non-natural, or derivatized nucleotide bases. Such nucleic acid sequences may comprise additional sequences useful for promoting expression and / or purification of the encoded polypeptide, including but not limited polyA sequences, modified Kozak sequences, and sequences encoding epitope tags, export signals, and secretory signals, nuclear localization signals, and plasma membrane localization signals. It will be apparent to those of skill in the art, based on the teachings herein, what nucleic acid sequences will encode the polypeptides of the disclosure. In some embodiments, the nucleic acid comprises a nucleotide sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the nucleotide sequence selected from SEQ ID NO:78-82, The polynucleotide sequences of SEQ ID NO:78-82 are provided in Table 7. Table 7 In a further aspect, the disclosure provides expression vectors comprising the nucleic acid of any aspect of the disclosure operatively linked to a suitable control sequence, such as a promoter. “Expression vector” includes vectors that operatively link a nucleic acid coding region or gene to any control sequences (such as promoters) capable of effecting expression of the gene product. “Control sequences” operably linked to the nucleic acid sequences of the disclosure are nucleic acid sequences capable of effecting the expression of the nucleic acid molecules. The control sequences need not be contiguous with the nucleic acid sequences, so long as they function to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between a promoter sequence and the nucleic acid sequences and the control / promoter sequence can still be considered “operably linked” to the coding sequence. Other such control sequences include, but are not limited to, polyadenylation signals, termination signals, and ribosome binding sites. Such expression vectors can be of any type, including but not limited plasmid and viral- based expression vectors. The control sequence used to drive expression of the disclosed nucleic acid sequences in a mammalian system may be constitutive (driven by any of a variety of promoters, including but not limited to, CMV, SV40, RSV, actin, EF) or inducible (driven by any of a number of inducible promoters including, but not limited to, tetracycline, ecdysone, steroid-responsive). The expression vector must be replicable in the host organisms either as an episome or by integration into host chromosomal DNA. In various embodiments, the expression vector may comprise a plasmid, viral-based vector, or any other suitable expression vector. In another aspect, the disclosure provides host cells that comprise fusion proteins, oligomers, nucleic acids or expression vectors (i.e.: episomal or chromosomally integrated) disclosed herein, wherein the host cells can be either prokaryotic or eukaryotic. The cells can be transiently or stably engineered to incorporate the expression vector of the disclosure, using techniques including but not limited to bacterial transformations, calcium phosphate co- precipitation, electroporation, or liposome mediated-, DEAE dextran mediated-, polycationic mediated-, or viral mediated transfection. In one embodiment, the host cells comprise E. coli cells. In other embodiments, the host cells comprise yeast cells. In some embodiments, the fusion protein is expressed on the yeast cell surface (i.e., embodiments in which the fusion protein comprises the Aga2 domain). In some embodiments, the host cell comprises a nucleic acid or expression vector that encodes and is capable of expressing Ulp1 (SEQ ID NO:87). MGGLVPELNEKDDDQVQKALASRENTQLMNRDNIEITVRDFKTLAPRRWLNDTIIEF FMKYIEKSTPNTVAFNSFFYTNLSERGYQGVRRWMKRKKTQIDKLDKIFTPINLNQS HWALGIIDLKKKTIGYVDSLSNGPNAMSFAILTDLQKYVMEESKHTIGEDFDLIHLDC PQQPNGYDCGIYVCMNTLYGSADAPLDFDYKDAIRMRRFIAHLILTDALKG (SEQ ID NO:87) Ulp1 can cleave the SUMO tag, and thus this embodiment is particularly useful in embodiments where the fusion protein comprises the cleavage tag comprising the amino acid sequence of SEQ ID NO:67. In one embodiment, Ulp1 is encoded by a different nucleic acid or expression vector than the fusion protein is encoded by. The fusion proteins, acids, expression vectors, and host cells of any embodiment herein, can be used for methods including but not limited to yeast surface display, studying TCR docking geometries and other signaling mechanisms, as well as X-ray crystallography and cryo-electron microscopy to determine the structural basis for peptide-MHC and pMHC- TCR interactions, as a staining reagent to identify or isolate a T-cell subset that can respond to a peptide of interest, for tracking the response to an infection or cancer, or used for T-cell sorting and sequencing experiments to identify TCR sequences that recognize a pMHC of interest. Examples It is of high interest to create a more approachable recombinant MHC molecule. The limitations of current approaches stem from the inherent instability of the MHC molecule. If it were more stable, it might be possible to produce soluble pMHCs in bulk without the need for refolding. Instead of relying on specialized facilities to produce refolded pMHCs for staining experiments, immunologists could produce them in-house, dramatically improving their ability to iterate through multiple peptide variants or even MHC alleles. Similarly, a stable MHC molecule could be easily expressed on the surface of yeast, allowing screening of large peptide libraries without prior optimization of the MHC sequence for yeast display. Finally, if the MHC were more stable, it could also be more easily modified in other ways, allowing a more precise exploration of TCR docking geometries and other signaling mechanisms. To facilitate these uses for recombinantly produced MHC molecules, we developed a de-novo protein domain that, when fused to the peptide binding groove a class I MHC, allows soluble expression of that MHC in E. coli without the need for a placeholder peptide. The resulting soluble, monomeric, antigen-receptive, truncated (SMART) MHC is the subject of this patent. Design Strategy The structure of a native class I MHC is composed of two chains composed of four total domains: the α1, α2, and α3, domains constitute the MHC chain and the β2 microglobulin (β2m) domain heterodimerizes with it to form the full MHC structure19. The α1 and α2 domains consist of a pair of α-helices lying over a β-sheet that forms a groove where peptides can be loaded for presentation to T-cells19. The α3 and β2m domains are more membrane proximal and function as structural support for the α1 and α2 domains19. The α3 domain also provides a binding site for the CD8 co-receptor on T-cells20. Since all MHCs share very similar structures, we selected one representative MHC, the mouse H-2Dballele, to use as the basis for our stabilizing design process. In order to begin stabilizing the MHC fold, we first removed as much of the structure as possible without disrupting the ability of the MHC to present peptides and interact with TCRs. We removed the α3 and β2m domains completely (fig.1A, left). The truncated MHC contains a relatively hydrophobic patch on the underside (facing away from the peptide) of the β-sheet which we used as the starting point for our stabilizing domain design. Our design process generated small protein domains to cover this hydrophobic patch and support the folding of the H-2Dbα1 and α2 domains. We used computational methods to design protein binders for arbitrary target proteins23to design this stabilizing domain (fig.1A, middle). For initial screening, the designed stabilizing domain was linked to the N-terminus of the truncated MHC by a flexible GS linker (fig.1A, right). We screened roughly 104designs generated by this method using yeast surface display24(fig.1B). We first sorted for designs that were able to be displayed on the yeast surface, and within that sorted population, we further sorted for designs which were able to bind to fluorophore-labeled gp33 peptide, which is known to be bound strongly by native H- 2Db 25(fig.1C). We then determined which designs were enriched throughout this sorting process to find those that are able to be displayed on the yeast surface and are well-folded enough to bind a peptide (fig.1D). During this analysis, we found that designs which contained a Trp residue placed similarly to how W60 of β2m is placed in the native structure were much more likely to be positively enriched. This result aligns well with previous work showing that mutations at W60 can significantly destabilize the interaction between β2m and the MHC26,27. The hits selected by our enrichment analysis were then tested for soluble expression in E. coli, resulting in identification of a single stabilizing domain (referred to as hit6) that was able to preserve MHC folding in both yeast and E. coli expression systems. Although hit6 was able to allow soluble expression of the α1 and α2 domains of H-2Dbin E. coli, a large portion of the material produced was insoluble and the soluble protein was susceptible to proteolysis, resulting in very low yields (fig.1E, left). In order to improve on this design, we first identified likely cleavage sites based on the masses of the proteolytic fragments, and then redesigned the sequence near those sites. This redesign process was restricted to prevent mutations in amino acids that directly interact with the peptide, and to bias mutations towards amino acids that are commonly observed at those sites in other MHC alleles. We tested 12 variants of these cleavage site mutations (CSMs) and found that one set of mutations (CSM8) was best able to prevent cleavage. Finally, we used machine learning methods28–30to design a more structured linker to replace the flexible GS linker we had previously used. Of the 24 linker variants we tested, linker 11 (L11) had the largest improvement in soluble monomeric yield (fig.1E, right). Most of the data presented below use the fully optimized design (CSM8 mutations and L11), referred to as SMART MHC, however a few earlier results use the version with CSM8 mutations and the GS linker, referred to as CSM8 MHC. To verify that the soluble material we produced in E. coli retained the function characteristics of the MHC, we first measured the binding affinity of SMART H-2Dbfor the gp33 peptide using fluorescence polarization (FP) measurements of the same fluorophore- labeled peptide as in our yeast display experiments. Because SMART MHCs can be produced without a place-holder peptide, we can directly measure binding by titrating empty SMART MHC with a constant concentration of peptide. These measurements indicate that SMART H- 2Dbbinds the gp33 peptide with high pM affinity (fig.2), an affinity that is stronger than the reported value of 5nM for native H-2Db 31. We also found that SMART H-2Dbwas highly stable, retaining this tight binding affinity after storage for at least one month at 4oC, and after one freeze-thaw cycle (fig.2). Subsequent studies suggested that SMART H-2Dbbinds the gp33 peptide in a nearly identical manner and binds to a relevant TCR with a similar affinity to native H-2Db, (data not shown). Finally, the crystal structure of CSM8 H-2Dbin complex with the gp33 peptide was determined and showed excellent alignment with both the native structure33and the design model (data not shown). Overall, these results demonstrate that CSM8 H-2Dbis able to retain all of the necessary structural features that allow peptide and TCR binding, without the refolding requirements of the native MHC. Biochemical validation of SMART A*02:01 Next, we wanted to test the ability of our stabilizing domain to generalize to other MHC alleles. Based on the structural and sequence similarity of H-2Dbto a variety of human MHCs, we did not make any modifications to the stabilizing domain for the following experiments, instead varying only the MHC sequence. We selected HLA A*02:01 to test initially due to its abundance in the human population, and because there are many well- characterized peptides and TCRs that bind to it34,35. We found that CSM8 A*02:01 could be expressed solubly in E. coli, but that roughly half of the material produced was in a dimeric (and likely misfolded) state (fig. 3A). For all of the following characterization we used only the monomeric fraction, as purified by SEC. To test the ability of CSM8 A*02:01 to present peptides, we selected the NY-ESO-1 peptide because of its relevance in cancer36and the existence of a well-characterized TCR that binds the NY-ESO-1 / A*02:01 complex35. Using FP methods very similar to those used for SMART H-2Db, we measured the binding affinity of CSM8 A*02:01 to NY-ESO-1 to be roughly 190nM (fig.3B), similar to the 83nM affinity measured for the native A*02:01. The difference in affinity between the native and CSM8 A*02:01 could be due to partial unfolding of the CSM8 A*02:01, or to contamination of the monomeric species with dimer due to imperfect separation by SEC. Overall this difference is relatively minor, and the FP measurements indicate that peptide binding is maintained with CSM8 A*02:01. Given that CSM8 A*02:01 can bind the NY-ESO-1 peptide, it was determined whether it also retained binding to the 1G4 TCR36. The affinity of the NY-ESO-1 / CSM8 A*02:01 complex to the 1G4 TCR with two variants of the peptide was determined (data not shown), and the measured affinities were roughly 7-9 fold weaker than the same values for native A*02:01, similar to the difference in peptide binding affinity. Despite this difference in absolute binding affinity the ranking of the affinities for peptide variants was maintained, indicating that, although some of the CSM8 A*02:01 might be improperly folded, the proportion that is folded can be recognized by a relevant TCR. Overall, these data suggest that, although it is less effective, our stabilizing domain can stabilize an allele other than the one it was designed for. Yeast Display of Peptide-fused SMART A02 As previously mentioned, yeast display can be a powerful tool to screen libraries of peptide variants for TCR binding15,17. We again used A*02:01 as a test case, this time using the TAX peptide and A6 TCR that recognizes the A*02:01 / TAX complex34. The SMART A*02:01 / TAX complex was displayed on the C-terminus of Aga2 using a mutant A*02:01 (W167A) to allow the peptide linker to leave the peptide binding groove. We found that SMART A*02:01 is able to be displayed at much higher levels than native A*02:01, regardless of whether a peptide is fused to it (fig.4A). The elevated expression levels also lead to an increase in binding of an A6 TCR tetramer only when the TAX peptide was present (fig.4B). These results indicate that SMART MHCs can significantly improve yeast display off MHCs, and can be used in the future to facilitate large peptide library screens. T-cell Staining with Oligomerized SMART MHCs Another important application of pMHCs in immunological research is the staining of T-cells with pMHC tetramers1. We therefore tested whether SMART MHCs could be converted into a similar multimeric staining reagent. Rather than using streptavidin to tetramerize the SMART MHCs, as is typically done, we chose to directly fuse them to a de- novo designed oligomeric protein which assembles into a tetrahedral symmetry containing 12 subunits. This allowed us to skip the biotinylation step which is necessary for streptavidin- based tetramerization. In order to improve folding and assembly of the SMART MHC- tetrahedron fusion, we additionally fused a peptide of interest to the N-terminus of the SMART MHC along with a SUMO tag. Co-expression of the Ulp1 protease allows the N- terminus of the peptide to be cleanly cleaved, allowing it to bind properly in the peptide binding groove. Additionally, the Y84A mutation in the MHC sequence was used to allow the peptide linker to exit the peptide binding groove. Finally, on the C-terminus of the oligomer, we fused a Myc tag to allow for antibody staining (fig.5A). To a assess the ability of the SMART MHC oligomers to stain T-cells, we mixed two populations of T-cells: one expressing the P14 TCR and the other expressing an unrelated TCR along with mTagBFP to allow the two cell types to be distinguished independently of TCR staining. We then made SMART H-2Dboligomers fused to several variants of the gp33 peptide with known affinities to the P14 TCR25,32and assessed their ability to stain the T-cell mixture using an anti-Myc antibody. We found that peptides with affinities similar to unmutated gp33 (V3P and M9C) showed bright staining of the P14 T-cells relative to the controls, with V3P achieving both the brightest staining and highest affinity. In contrast, we were unable to detect binding to gp33 variants with significantly reduced affinities (V3P+Y4F and Y4F). The lack of staining we observed with V3P+Y4F suggests that this staining method is somewhat less sensitive than the SPR methods used to measure the binding affinities. Overall, the success of the SMART MHC tetrahedra in staining T-cells in both a TCR- and peptide- specific manner indicates that they can be used in place of pMHC tetramers without the need for refolding and biotinylation. The crystal structure of CSM8 A*02:01 / TAX9 complexed with the A6c134 TCR was solved to determine whether SMART A*02:01 presents peptides and interacts with TCR in a native-like way. Excellent agreement between CSM8 and native A*02:01 within the HLA and TAX9 peptide was found, and the TCR variable domains engaged CSM8 A*02:01 nearly identically to native A*02:01 (data not shown). The designed stabilizing domain solubilizes several additional common HLA allomorphs We expressed 15 human HLA allomorphs using our SMART stabilizing domain in E. coli and carried out small-scale purification using HPLC. SMART versions of many allomorphs exhibited peaks in the expected monomeric range, though many also displayed dimeric and aggregate peaks (fig.6A). Interestingly, HLA A*03:01 and HLA A*01:01 demonstrated high protein expression, and HLA B*07:02 showed the highest monomeric fraction (fig.6B). The presence of a significant dimeric or aggregated population, particularly in allomorphs with high expression, indicates that the stabilizer was not effective in simultaneously promoting high protein expression and monomeric behavior across diverse HLAs. However, we observed higher expression levels in four of the 15 allomorphs, and reduced dimerization in three, when compared to SMART A*02:01. These results suggest that the stability and expression of these allomorphs could be improved through minimal redesign of the stabilizing domain. We next compared the expression of native, full-length HLAs (native) with their SMART versions, with and without genetically linked peptides, using CFE. To determine whether SMART HLAs have improved solubility over native HLAs, we measured soluble protein using radiolabeled14C-Leucine incorporation which allows for precise quantification of protein yields (fig.6C). We found that almost all SMART HLAs had improved soluble yields compared to native HLAs, and that many SMART HLAs that did not express well in E. coli had significant CFE expression levels. The improvement in soluble yields in the CFE system could arise from the lower protein concentrations, the oxidative conditions or the presence of the DsbC chaperone. Both E. coli and CFE experiments suggest that the designed stabilizing domain enables the solubilization of different allomorphs without the need for extensive inclusion body preparation. Further design optimization could further improve expression and reduce aggregation. Methods Stabilizer library design We used computational methods to design protein binders for arbitrary target proteins23to design this stabilizing domain. The “target” supplied to this method was the α1 and α2 domains of the H-2Dbstructure (PDB: 1S7U). We ran two versions of the design protocol. The first version was a completely de-novo approach which allowed any amino acid to be placed anywhere near the bottom of the H-2Dbpeptide binding groove in order to make a favorable interaction. The second approach specifically focused on making designs which placed amino acids in locations that allow them to replace the exact interactions that β2m makes in the native structure. Designs from both of these protocols were pooled and filtered for the quality of the stabilizer and the interactions it made with the MHC, as previously described23. We also included a set of negative control designs which were made by randomly scrambling the sequence (while conserving the pattern of hydrophobic and hydrophilic residues) of a random subset of the designs. Yeast display screening of stabilizing domains Yeast display was performed as previously described23with the following changes. Rather than screening our designs for binding to the α1 and α2 domains of H-2Db, we fused the designs to those domains using a flexible poly-GS linker and screened them for surface display and binding to a FITC-labeled gp33 peptide (KAVYNFATM (SEQ ID NO: 47)), with FITC linked to the amine on the lysine sidechain). Sorted populations were subsequently cultured and plasmid DNA was extracted for sequencing. Protein expression and purification Genes encoding the designed protein sequences were synthesized and cloned into modified pET-29b(+) E. coli plasmid expression vectors with a 6xHis tag added to the N- terminus (for monomeric versions) or the C-terminus (for oligomeric / peptide-fused versions). Plasmids were transformed into chemically competent E. coli BL21 (DE3) cells (NEB). E. coli cells were grown in LB medium at 37 °C until the cell density reached 1.0 at OD600. Then, IPTG was added to a final concentration of 1mM and the cells were grown overnight at 16°C for expression. The cells were collected by spinning at 4,000g for 5min and then resuspended in lysis buffer (150 mM NaCl, 25mM Tris-HCL (pH 8.0), 25mM imidazole, 1mM PMSF, and 5% glycerol) with RNase. The cells were lysed with a Qsonica Sonicators sonicator for 7 min in total (3.5 min each time, 10 s on, 10 s off) with an amplitude of 80%. The soluble fraction was clarified by centrifugation at 14,000g for 30 min. The soluble fraction was purified by immobilized metal affinity chromatography (Qiagen) followed by FPLC SEC on a SuperdexTM7510 / 300 GL column (GE Healthcare) for monomeric versions, and a SuperoseTM610 / 300 GL column (GE Healthcare) for oligomeric versions. All protein samples were characterized by SDS–PAGE, and purity was greater than 95%. Protein concentrations were determined by absorbance at 280 nm measured with a NanoDropTMspectrophotometer (Thermo Scientific) using predicted extinction coefficients. Cleavage site mutation design Cleavage site design was used to make mutations to the MHC sequence that would reduce proteolysis without impacting peptide or TCR binding. First, a multiple sequence alignment of MHC protein sequences was collected using PSI-BLAST. Next, the MSA was converted into a position-specific score matrix (PSSM) denoting the likelihood of observingeach possible amino acid at each position in the MSA. Standard RosettaTM design protocolswere modified to allow mutations to amino acids that had likelihoods above a specified cutoff, and to only allow mutations at sequence positions near the pre-identified cleavage. Design was further restricted to prevent mutations in residues with sidechains that could interact with a bound peptide or TCR. All cleavage site mutants were designed based on the hit6 design model as a starting point, and the highest scoring designs were selected for experimental testing based on a combination of RosettaTMmetrics relating to the overall RosettaTMenergy of the design model, the complementarity of the stabilizer / MHC interface, and the number hydrogen bond donors / acceptors that were left unbonded. Linker design Linker design was used to replace the flexible GS linker that was used in the yeast display screening with a more structured and shorter linker. Using the design model of the CSM8 variant as a starting point, we used previously developed “inpainting” methods28to fill in a small segment of protein structure to bridge the gap between the C-terminus of the stabilizer and N-terminus of the MHC. Sequences for the protein backbones produced by this method were designed using ProteinMPNNTM, restricted to changing only the amino acids in the “inpainted” structure. Finally, the resulting designs were evaluated using AlphaFoldTMpredictions where the MHC structure was provided as a template, but the stabilizer and linker were not. Designs with high overall pLDDT scores and low PAE scores for residues in the MHC / stabilizer interface were selected for experimental testing. Peptide binding affinity measurements Three technical replicates of varying concentrations of SMARTTMMHC were mixed with a constant concentration of fluorophore-labeled peptide (300pM FITC-gp33 for H-2Dband 10nM AF488-NY-ESO-1 for A*02:01) and incubated overnight at room temperature to allow equilibration. Fluorescence polarization measurements of these samples were made with a SynergyTMNeo2 plate reader (BioTek instruments) with a 485 / 530 FP filter. Binding curves (using the non-simplified equilibrium binding equation37) were fitted separately to each of the triplicate measurements and averaged to determine the KD. TCR binding affinity measurements TCR binding affinities were measured as previously described32, using SMARTTMMHCs as the mobile phase instead of native MHCs. All measurements were performed on a BIAcore T200 (GE Healthcare) at 4°C in the buffer containing 10 mM HEPES pH7.4, 150 mM NaCl, 0.005% Tween-20, 3 mM EDTA. Soluble P14-his6 was noncovalently coupled to the anti-his antibody, immobilized on a CM5-chip via standard amine coupling, and around 4000 response units of anti-his antibody was coupled, immobilizing around 1000 response units of P14-his5 (0.75 uM). A control surface was generated the same way, and up to 100 μM or 200 μM of freshly produced WWDb / peptide complexes (2-fold dilutions from stock, at least 10 concentrations in duplicate) were injected over the chip surfaces at 30 μL / mL. The sample rack was cooled to 4°C during the run. Chip surfaces were regenerated using 0.1 M Glycine-HCl pH 2.5, 500 mM NaCl, Tween 0.05% at 30 μL / min after each injection. The WWDb / peptide was injected over a control surface, and the final signal was calculated by subtracting the signal obtained on the control surface from the signal on the TCR-coupling surface, to remove the contributions of the bulk effect and possible non-specific binding. The data were then analyzed using BIAevaluation 3.0 software. The KD values were obtained from steady-state fitting of equilibrium-binding curves. Peptide-fused yeast display Yeast display of peptide-fused SMARTTMA*02:01 was performed as previously described15. In brief, 50ng pCT or pYAL plasmids encoding corresponding full length or SMARTTMA02 constructs with TAX were electroporated into competent EBY100. The EBY100 was cultured in YPD medium for 1 hour at 30oC, spun down and continued to grow in SDCAA medium for 48 hours before induction in SGCAA for 48 hours. Display levels were evaluated with fluorophore-conjugated anti-HA and anti-A*02:01 (clone BB7.2) antibodies, and TCR binding was evaluated with TCR tetramers made by combining soluble biotinylated A6 TCR with fluorophore-conjugated streptavidin. Small-scale expression and SEC of HLA Alleles Linear DNA encoding native, full-length HLA alleles, with or without covalently linked peptides, were cloned into modified pET-29b(+) vector containing an N-terminal 6xHis tag via Golden Gate Assembly (BsaI-HF® v2). Plasmids were transformed into chemically competent E. coli BL21(DE3) cells (NEB) and received in LB media. Cells were diluted and grown in Terrific Broth II (TB-II) media + 50 µg / mL Kanamycin at 37°C in 96 well microplates with long drip spouts until reaching an OD600 of 1.0, followed by the addition of 1 mM IPTG. Plates were transferred to 16°C for protein expression. Plates were spun down for 5 minutes at 4000xg and cells were resuspended in a lysis buffer consisting of BugBuster® Protein Extraction Reagent, 0.1 mg / mL lysozyme, 0.01 mg / mL benzonase, 1 mM PMSF. Plates were shaken at 1000 rpm at 37°C for 15 minutes. Lysate was spun down and purified in 96 well 25 µm polyethylene fritted plates (Agilent) by immobilized metal affinity chromatography (Qiagen). Soluble, filtered samples were analyzed using high performance liquid chromatography (Agilent 1260 InfinityTMII LC System) on a Superdex™ 75 Increase 5 / 150 small-scale SEC column (Cytiva). T-cell staining Two Jurkat T-cell lines were used: one expressing the P14 TCR, and the other expressing an unrelated TCR (a3a TCR, recognizing the MAGE-A3 / A*01:01 complex) as well as mTagBFP. Both cell lines were mixed in equal numbers and resuspended to 1M cells / mL in 100uL of staining solution (500nM SMARTTMMHC oligomer, 25mM Tris-HCl pH 8.0, 150mM NaCl, and 5% glycerol) and incubated at 4oC for 30min. Cells were then washed twice with 100uL of Fc block (HBH (HBSS with 0.5% BSA, and 10mM HEPES) supplemented with 10% 2.4G2 cell culture supernatant). Cells were then resuspended in 50uL AF647-anti-Myc antibody diluted in Fc block and incubated at 4C for 20min. Cells were then washed twice with 100uL of HBH and resuspended to a final volume of 150uL of HBH for analysis on an Attune Nxt flow cytometer. Cells were gated into P14 (mTagBFP-) or a3a (mTagBFP+) populations and the brightness of the AF647 stain for each population was compared. Sequence Information for Oligomers and Stabilizing Domain Table 11. Peptides Table 12. MHCs (with CSM8 mutations) Alternative linkers We tested a total of 24 linker variants to replace the flexible GS linker with a structured linker, and tested their ability to stabilize the H-2Dballele. Of these, 11 showed improved soluble expression in E. coli, relative to the flexible linker version. Of these, 9 showed high monomeric yields and preliminary evidence of peptide binding. The 3 with the highest yields were selected for further characterization of peptide binding (shown below). All 3 of these showed high affinity peptide binding, comparable to the flexible linker version. Of these 3, we selected L11 to move forward with due to its high expression levels. Table 13. Linkers with verified gp33 peptide binding Alternative oligomerization domains We tested 63 oligomerization domains in combination with several tags and linker lengths connecting the oligomerization domain to the SMARTTMMHC. Of these 63, 42 showed high expression in E. coli, but only 10 showed the correct oligomeric state (measured by SEC). We tested 8 of these, and all 8 showed some degree of TCR specific T-cell staining. Out of these 8, we selected TetT=1-4 because it had the best combination of proper oligomeric yield and T-cell staining brightness. All tests were done using the gp33 peptide, SMART H- 2Db, and the P14 TCR. Table 14. Oligomers with verified T-cell binding MHC Alleles and peptides Table 15 lists exemplary MHCs and peptides that go with them. Expression yields estimated by SDS-PAGE gel, proportion monomer estimated from SEC chromatograms. Yield and monomer data for non-oligomerized versions with no peptide fused. Table 15 References 1. Wooldridge, L. et al. Tricks with tetramers: how to get the most from multimeric peptide–MHC. Immunology 126, 147–164 (2009). 2. Murali-Krishna, K. et al. Counting Antigen-Specific CD8 T Cells: A Reevaluation of Bystander Activation during Viral Infection. Immunity 8, 177–187 (1998). 3. Lee, P. P. et al. Characterization of circulating T cells specific for tumor-associated antigens in melanoma patients. Nat. Med.5, 677–685 (1999). 4. Linnemann, C. et al. High-throughput identification of antigen-specific TCRs by TCR gene capture. Nat. Med.19, 1534–1541 (2013). 5. Aleksic, M. et al. Dependence of T cell antigen recognition on T cell receptor-peptide MHC confinement time. Immunity 32, 163–174 (2010). 6. Lin, J. J. Y. et al. Mapping the stochastic sequence of individual ligand-receptor binding events to cellular activation: T cells act on the rare events. Sci. Signal.12, eaat8715 (2019). 7. Sibener, L. V. et al. Isolation of a Structural Mechanism for Uncoupling T Cell Receptor Signaling from Peptide-MHC Binding. Cell 174, 672-687.e27 (2018). 8. Zareie, P. et al. Canonical T cell receptor docking on peptide-MHC is essential for T cell signaling. Science 372, eabe9124 (2021). 9. Sušac, L. et al. Structure of a fully assembled tumor-specific T cell receptor ligated by pMHC. Cell 185, 3201-3213.e19 (2022). 10. Production Protocols | NIH Tetramer Core Facility. https: / / tetramer.yerkes.emory.edu / support / protocols#4. 11. Bakker, A. H. et al. Conditional MHC class I ligands and peptide exchange technology for the human MHC gene products HLA-A1, -A3, -A11, and -B7. Proc. Natl. Acad. Sci. 105, 3825–3830 (2008). 12. Luimstra, J. J. et al. A flexible MHC class I multimer loading system for large-scale detection of antigen-specific T cells. J. Exp. Med.215, 1493–1504 (2018). 13. Overall, S. A. et al. High throughput pMHC-I tetramer library production using chaperone-mediated peptide exchange. Nat. Commun.11, 1909 (2020). 14. Birnbaum, M. E., Dong, S. & Garcia, K. C. Diversity-oriented approaches for interrogating T-cell receptor repertoire, ligand recognition, and function. Immunol. Rev. 250, 82–101 (2012). 15. Adams, J. J. et al. Structural interplay between germline interactions and adaptive recognition determines the bandwidth of TCR-peptide-MHC cross-reactivity. Nat. Immunol.17, 87–94 (2016). 16. Crawford, F., Huseby, E., White, J., Marrack, P. & Kappler, J. W. Mimotopes for Alloreactive and Conventional T Cells in a Peptide–MHC Display Library. PLOS Biol.2, e90 (2004). 17. Yang, X. et al. Autoimmunity-associated T cell receptors recognize HLA-B*27-bound peptides. Nature 612, 771–777 (2022). 18. Dobson, C. S. et al. Antigen identification and high-throughput interaction mapping by reprogramming viral entry. Nat. Methods 19, 449–460 (2022). 19. Madden, D. R. The Three-Dimensional Structure of Peptide-MHC Complexes. Annu. Rev. Immunol.13, 587–622 (1995). 20. Wang, R., Natarajan, K. & Margulies, D. H. Structural Basis of the CD8αβ / MHC Class I Interaction: Focused Recognition Orients CD8β to a T Cell Proximal Position12. J. Immunol.183, 2554–2564 (2009). 21. Jones, L. L. et al. Engineering and Characterization of a Stabilized α1 / α2 Module of the Class I Major Histocompatibility Complex Product Ld*. J. Biol. Chem.281, 25734– 25744 (2006). 22. Birnbaum, M. E. et al. Deconstructing the peptide-MHC specificity of T cell recognition. Cell 157, 1073–1087 (2014). 23. Cao, L. et al. Design of protein-binding proteins from the target structure alone. Nature 605, 551–560 (2022). 24. Boder, E. T. & Wittrup, K. D. Yeast surface display for screening combinatorial polypeptide libraries. Nat. Biotechnol.15, 553–557 (1997). 25. Boulter, J. M. et al. Potent T cell agonism mediated by a very rapid TCR / pMHC interaction. Eur. J. Immunol.37, 798–806 (2007). 26. Achour, A. et al. Structural Basis of the Differential Stability and Receptor Specificity of H-2Db in Complex with Murine versus Human β2-Microglobulin. J. Mol. Biol.356, 382–396 (2006). 27. Li, Z. et al. The Mechanism of β2m Molecule-Induced Changes in the Peptide Presentation Profile in a Bony Fish. iScience 23, 101119 (2020). 28. Wang, J. et al. Scaffolding protein functional sites using deep learning. Science 377, 387– 394 (2022). 29. Dauparas, J. et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science 378, 49–56 (2022). 30. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). 31. Vita, R. et al. The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res. 47, D339–D343 (2019). 32. Duru, A. D. et al. Tuning antiviral CD8 T-cell response via proline-altered peptide ligand vaccination. PLOS Pathog.16, e1008244 (2020). 33. Velloso, L. M., Michaëlsson, J., Ljunggren, H.-G., Schneider, G. & Achour, A. Determination of Structural Principles Underlying Three Different Modes of Lymphocytic Choriomeningitis Virus Escape from CTL Recognition1. J. Immunol.172, 5504–5511 (2004). 34. Utz, U., Banks, D., Jacobson, S. & Biddison, W. E. Analysis of the T-cell receptor repertoire of human T-cell leukemia virus type 1 (HTLV-1) Tax-specific CD8+ cytotoxic T lymphocytes from patients with HTLV-1-associated disease: evidence for oligoclonal expansion. J. Virol.70, 843–851 (1996). 35. Chen, J.-L. et al. Structural and kinetic basis for heightened immunogenicity of T cell vaccines. J. Exp. Med.201, 1243–1255 (2005). 36. Chen, Y. T. et al. A testicular antigen aberrantly expressed in human cancers detected by autologous antibody screening. Proc. Natl. Acad. Sci. U. S. A.94, 1914–1918 (1997). 37. Hulme, E. C. & Trevethick, M. A. Ligand binding assays at equilibrium: validation and interpretation. Br. J. Pharmacol.161, 1219–1237 (2010). 38. Vonrhein, C. et al. Data processing and analysis with the autoPROC toolbox. Acta Crystallogr. D Biol. Crystallogr.67, 293–302 (2011). 39. McCoy, A. J. et al. Phaser crystallographic software. J. Appl. Crystallogr.40, 658–674 (2007). 40. Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr. D Biol. Crystallogr.66, 486–501 (2010). 41. Afonine, P. V. et al. Towards automated crystallographic structure refinement with phenix.refine. Acta Crystallogr. D Biol. Crystallogr.68, 352–367 (2012). 42. Berkholz, D. S., Shapovalov, M. V., Dunbrack, R. L. & Karplus, P. A. Conformation Dependence of Backbone Geometry in Proteins. Struct. Lond. Engl.199317, 1316–1325 (2009). 43. Williams, C. J. et al. MolProbity: More and better reference data for improved all-atom structure validation. Protein Sci. Publ. Protein Soc.27, 293–315 (2018).

Claims

We claim 1. A fusion protein, comprising: (a) a stabilizer peptide comprising or consisting of domains X1-X2, wherein (i) X1 comprises or consists of the amino acid sequence SEQ ID NO:1; and (ii) X2 is a first amino acid linker of between 5-20 amino acids in length; (b) a truncated class I major histocompatibility complex (MHC) protein directly fused to the C-terminus of the stabilizer peptide, wherein the truncated MHC protein consists of residues 1-175, 1-176, 1-177, 1-178, 1-179, 1-180, 1-181, 1-182, 1-183, 1-184, or 1-185 of a class I MHC protein, optionally wherein the truncated class I MHC protein comprises residues 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A. 2 The fusion protein of claim 1, wherein the truncated class I MHC protein comprises residues 6K, 8I, 12I, 104D, 105E, 106N, and 110V.

3. The fusion protein of claim 1 or 2, wherein the truncated MHC protein consists of residues 1-, 1-179 or 1-180 of a class I MHC protein.

4. The fusion protein of any one of claims 1-3, wherein X2 is a first amino acid linker of 12-16 amino acids in length.

5. The fusion protein of any one of claims 1-4, wherein X2 has an amino acid sequence selected from the group consisting of SEQ ID NO:2-12.

6. The fusion protein of any one of claims 1-5, wherein X2 is SEQ ID NO:

6.

7. The fusion protein of any one of claims 1-5, wherein the stabilizer peptide comprises or consists of the amino acid sequence selected from SEQ ID NO:13-22.

8. The fusion protein of any one of claims 1-7, wherein the truncated class I MHC protein is a truncated human class I MHC.

9. The fusion protein of claim 8, wherein the truncated class I MHC protein is a truncated HLA-A, HLA-B, or HLA-C.

10. The fusion protein of claim 8, wherein the truncated class I MHC protein is a truncated HLA-E, HLA-F, or HLA-G.

11. The fusion protein of any one of claims 1-10, wherein the truncated class I MHC protein consists of the amino acid sequence selected from SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V,and optionally wherein the truncated class I MHC protein comprises residue 84A relative to the reference sequence.

12. The fusion protein of claim 11, wherein the fusion protein has the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V.

13. The fusion protein of any one of claims 1-12, wherein: (a) X1 consists of the amino acid sequence of SEQ ID NO:1; (b) X2 consists of the amino acid sequence selected from the group consisting of SEQ ID NO:2-12, or SEQ ID NO:2-11; and (c) the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V,and optionally wherein the truncated class I MHC protein comprises residue 84A.

14. The fusion protein of claim 13, wherein the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38 with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V.

15. The fusion protein of any one of claims 1-14, wherein: (a) the stabilizer peptide consists of the amino acid sequence selected from SEQ ID NO:13-22; and (b) the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V, and optionally wherein the truncated class I MHC protein comprises residue 84A.

16. The fusion protein of claim 15, wherein the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, with the following substitutions relative to the reference sequence: 6K, 8I, 12I, 104D, 105E, 106N, and 110V.

17. The fusion protein of any one of claims 1-16, wherein the truncated class I MCH protein does not include residue 84A relative to the reference sequence.

18. The fusion protein of any one of claims 1-16, wherein the truncated class I MCH protein includes residue 84A relative to the reference sequence.

19. The fusion protein of any one of claims 1-18, further comprising an Aga2 domain comprising or consisting of the amino acid sequence of SEQ ID NO:85 fused N-terminal to the stabilizer peptide, optionally via an amino acid linker.

20. The fusion protein of claim 19, further comprising a signal sequence N-terminal to the Aga2 domain.

21. The fusion protein of claim 19 or 20, further comprising a peptide tag located between the Aga2 domain and the stabilizer peptide, optionally via amino acid linkers.

22. The fusion protein of any one of claims 1-21, further comprising a binding peptide that is capable of binding to the truncated class I MHC protein, wherein the truncated class I MHC protein comprises residue 167A.

23. The fusion protein of claim 22, wherein the binding peptide is fused C-terminal to the truncated class I MHC protein, optionally via an amino acid linker.

24. The fusion protein of claim 23, wherein the fusion protein comprises, in N-terminal to C-terminal order, wherein each domain may be linked via an optional amino acid linker: (ix) a signal sequence; (x) an Aga2 domain; (xi) a peptide tag;(xii) an amino acid linker; (xiii) a stabilizer peptide; (xiv) a truncated class I MHC protein; (xv) an amino acid linker; and (xvi) a binding peptide capable of binding to the truncated class I MHC protein.

25. The fusion protein of any one of claims 19-24, comprising the amino acid of SEQ ID NO:74 or 75.

26. The fusion protein of claim 18, wherein the fusion protein further comprises: (i) a binding peptide that is capable of binding to the truncated class I MHC protein, wherein the peptide is fused N-terminal to the stabilizer peptide via an amino acid linker; and (ii) an oligomer-forming polypeptide fused C-terminal to the truncated class I MHC protein, wherein the oligomer-forming polypeptide is fused to the truncated class I MHC protein either directly or via an amino acid linker.

27. The fusion protein of claim 26, wherein the oligomer-forming polypeptide comprises or consists of the amino acid sequence selected from the group consisting of SEQ ID NO:39- 46 and 94.

28. The fusion protein of claim 26, wherein the oligomer-forming polypeptide comprises or consists of the amino acid sequence of SEQ ID NO:

46.

29. The fusion protein of claim 27 or 28, wherein the truncated class I MHC protein and the binding peptide that binds to it are selected from a binding peptide and truncated class I MHC protein listed in one row of Table 5.

30. The fusion protein of any one of claims 26-29, further comprising a peptide cleavage tag at the N-terminal end of the fusion protein.

31. The fusion protein of claim 30, wherein the peptide cleavage tag comprises or consists of the amino acid sequence of SEQ ID NO:67.

32. The fusion protein of any one of claims 14-20, wherein the fusion protein comprises, in N-terminal to C-terminal order: (A) a peptide cleavage tag comprising or consisting of the amino acid sequence of SEQ ID NO:67; (B) a binding peptide comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:47-66, 83, 88, 91, and 96; (C) an amino acid linker of length 10-20 amino acids; (D) a stabilizer peptide comprising or consisting of the amino acid sequence selected from SEQ ID NO:13-22; (E) an optional amino acid linker of length 1-10 amino acids; (F) a truncated class I MHC protein consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, optionally with the following substitutions relative to the reference sequence: 6K, 8I, 84A, 12I, 104D, 105E, 106N, and 110V; wherein the truncated class I MHC protein and the binding peptide that binds to it are selected from a binding peptide and truncated class I MHC protein listed in one row of Table 5; (G) an amino acid linker of length 3-30 amino acids; and (H) an oligomer-forming polypeptide comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:39-46 and 94.

33. The fusion protein of claim32, wherein the truncated class I MHC protein consists of the amino acid sequence selected from the group consisting of SEQ ID NO:23-38, with the following substitutions relative to the reference sequence: 6K, 8I, 84A, 12I, 104D, 105E, 106N, and 110V.

34. The fusion protein of any one of claims 26-33, further comprising a detectable peptide tag at the C-terminus of the fusion protein, fused to the oligomer-forming polypeptide either directly or via a fourth optional amino acid linker.

35. The fusion protein of claim 21, wherein the detectable tag comprises or consists of the amino acid sequence of SEQ ID NO:

89.

36. The fusion protein of any one of claims 1-35, comprising or consisting of the amino acid sequence selected from the group consisting of SEQ ID NO:68-77 and 97-100, whereinresidues in parentheses are optional and may be present or may be deleted, or may be substituted with other residues.

37. An oligomer, comprising a plurality of fusion proteins of any one of claims 26-36, wherein the peptide is linked to the truncated class I MHC protein in each identical fusion protein.

38. The oligomer of claim 37, wherein the oligomer is a homo-oligomer.

39. The oligomer of claim 37, wherein the oligomer is a hetero-oligomer.

40. A nucleic acid encoding the fusion protein of any one of claims 1-36.

41. The nucleic acid of claim 40, wherein the nucleic acid comprises a nucleotide sequence at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the nucleotide sequence selected from SEQ ID NO:78-82 42. An expression vector comprising the nucleic acid of claim 40 or 41 operatively linked to a suitable control sequence, such as a promoter.

43. A host cell comprising the fusion protein, oligomer, nucleic acid, or expression vector of any one of claims 1-42.

44. The host cell of claim 43, wherein the cell is an E. coli cell.

45. The host cell of claim 43, wherein the cell is a yeast cell.

46. The host cell of claim 45, wherein the fusion protein further comprises the fusion protein of any one of claims 19-25, and wherein the fusion protein is expressed on the yeast cell surface.

47. The host cell of any one of claims 43-46, wherein the host cell comprises a nucleic acid or expression vector that encodes and is capable of expressing Ulp1 (SEQ ID NO:87)48. A method for using the fusion proteins, nucleic acids, expression vectors, and host cells of any embodiment herein, for a method including but not limited to yeast surface display, studying TCR docking geometries and other signaling mechanisms, as well as X-ray crystallography and cryo-electron microscopy to determine the structural basis for peptide- MHC and pMHC-TCR interactions, as a staining reagent to identify or isolate a T-cell subset that can respond to a peptide of interest, for tracking the response to an infection or cancer, or used for T-cell sorting and sequencing experiments to identify TCR sequences that recognize a pMHC of interest.

Citation Information

Patent Citations

  • A peptide-MHC-i-antibody fusion protein for therapeutic use in a patient with amplified immune response

    US20220073630A1

  • SARS-COV-2 inhibitors

    US20230250134A1

  • Ultrahigh-affinity small protein targeting PD-l1 and use

    WO2023016559A1