MHC protein production method, TCR protein production method, expression system, bacterium belonging to genus agrobacterium, plant cell, method for producing transformed plant, expression vector, and method for detecting binding between MHC protein and TCR protein

WO2026177131A1PCT designated stage Publication Date: 2026-08-27UNIV OF TSUKUBA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/005716
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-19
Filing Date
2026-02-17
Publication Date
2026-08-27

Smart Images

  • Figure JPOXMLDOC01-APPB-T000001
    Figure JPOXMLDOC01-APPB-T000001
  • Figure JPOXMLDOC01-APPB-T000002
    Figure JPOXMLDOC01-APPB-T000002
  • Figure JPOXMLDOC01-APPB-T000003
    Figure JPOXMLDOC01-APPB-T000003
Patent Text Reader

Abstract

Provided are: an MHC protein production method with which it is possible to obtain, at a low cost, a recombinant MHC protein that functions sufficiently; a TCR protein production method with which it is possible to obtain a recombinant TCR protein that functions sufficiently; an expression system for use in the production method; a bacterium belonging to the genus agrobacterium and being transformed using the expression system; a plant cell into which the expression system has been introduced; a method for producing a transformed plant by using the expression system; an expression vector for use in creating the expression system; and a method for evaluating the binding strength between an MHC protein and a TCR protein. This MHC protein production method comprises a step for expressing an MHC protein in plant cells to obtain the MHC protein.
Need to check novelty before this filing date? Find Prior Art

Description

Method for producing MHC protein, method for producing TCR protein, expression system, Agrobacterium bacterium, plant cell, method for producing transformed plant, expression vector, method for detecting binding between MHC protein and TCR protein

[0001] The present invention relates to a method for producing an MHC protein, an expression system used in the production method, an Agrobacterium bacterium transformed with the expression system, a plant cell into which the expression system has been introduced, a method for producing a transformed plant using the expression system, and an expression vector used for preparing the expression system. This application claims priority based on Japanese Patent Application No. 2025-025148 filed in Japan on February 19, 2025, the content of which is incorporated herein by reference.

[0002] Major histocompatibility complex (MHC) proteins are membrane proteins present on the cell surface of vertebrates and play an essential role in self / nonself discrimination and acquired immune responses.

[0003] MHC class I proteins and MHC class II proteins present antigen peptides to T cells (see, for example, Non-Patent Document 1). Attempts have been made to express recombinant MHC proteins in Escherichia coli, insect cells, and mammalian cultured cells.

[0004] Neefjes J et al., (2011) Towards a systems understanding of MHC class I and MHC class II antigen presentation. Nat Rev Immunol. 2011,11:823-836., https: / / doi.org / 10.1038 / nri3084

[0005] However, when recombinant MHC class II proteins are expressed in Escherichia coli, there are problems such as the tertiary structure of the resulting protein not being properly formed and endotoxin being mixed into the resulting protein. When recombinant MHC class II proteins are expressed in mammalian cultured cells, the cost increases, and pathogens may be mixed into the resulting protein.

[0006] Therefore, the present invention aims to provide a method for producing an MHC protein that can be obtained inexpensively and with sufficient functionality, a method for producing a TCR protein that can be obtained with sufficient functionality, an expression system used in the production method, Agrobacterium bacteria transformed with the expression system, plant cells into which the expression system has been introduced, a method for producing a transformed plant using the expression system, an expression vector used in the production of the expression system, and a method for detecting the binding of an MHC protein to a TCR protein.

[0007] The present invention includes the following embodiments: [1] A method for producing an MHC protein, comprising the step of expressing an MHC protein in plant cells to obtain the MHC protein. [2] The method for producing an MHC protein according to [1], wherein the MHC protein is a mammalian protein. [3] The method for producing an MHC protein according to [1], wherein the MHC protein is a protein selected from (a) to (f) below. (a) Proteins containing an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (b) Proteins containing a portion of an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (c) Proteins containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added to an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (d) Proteins containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added to a portion of an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (e) Proteins containing an amino acid sequence having 90% or more sequence identity with an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (f) Proteins containing an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus A protein comprising an amino acid sequence having 90% or more sequence identity with a portion of the amino acid sequence encoded by the microglobulin gene locus [4] The method for producing an MHC protein according to [2], wherein the MHC protein is an MHC class II protein. [5] The method for producing an MHC protein according to [4], wherein the MHC class II protein is an HLA-DR protein, an HLA-DQ protein, or an HLA-DP protein. [6] The method for producing an MHC protein according to [1], wherein the MHC protein is a fusion protein in which a zipper-like helix domain is fused.[7] A method for producing an MHC protein according to [1], wherein a protein derived from the α chain of an MHC class I protein and a protein derived from β2 microglobulin of an MHC class I protein are co-expressed to obtain a dimer of the protein derived from the α chain of an MHC class I protein and the protein derived from β2 microglobulin of an MHC class I protein, or a protein derived from the α chain of an MHC class II protein and a protein derived from the β chain of an MHC class II protein are co-expressed to obtain a dimer of the protein derived from the α chain of an MHC class II protein and the protein derived from the β chain of an MHC class II protein. [8] The step of obtaining the MHC protein comprises a step of introducing an expression system into the plant cells, the expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette of the MHC protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette of a Rep / RepA protein derived from a geminivirus, wherein the expression cassette of the MHC protein comprises, in this order, a promoter, a nucleic acid fragment encoding the MHC protein, and two or more linked terminators, the method for producing the MHC protein according to [1].

[0008] [9] An expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette for the MHC protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette for the Rep / RepA protein derived from a geminivirus, wherein the expression cassette for the MHC protein comprises, in this order, a promoter, a nucleic acid fragment encoding the MHC protein, and two or more linked terminators.

[10] An Agrobacterium bacterium transformed with the expression system described in [9].

[11] A plant cell into which the expression system described in [9] has been introduced.

[12] A method for producing a transformed plant, comprising introducing the expression system described in [9] into a plant cell.

[13] An expression vector comprising a first nucleic acid fragment comprising a geminivirus-derived LIR, a geminivirus-derived SIR, and an expression cassette of an MHC protein linked between the LIR and the SIR, wherein the expression cassette of the MHC protein comprises, in this order, a promoter, a multicloning site, and two or more linked terminators.

[0009]

[14] A method for producing a TCR protein, comprising the step of expressing a TCR protein in plant cells to obtain the TCR protein.

[15] The method for producing a TCR protein according to

[14] , wherein the TCR protein is a fusion protein into which a zipper-like helix domain is fused.

[16] The method for producing a TCR protein according to

[14] , wherein the TCR protein is a fusion protein into which an apoplast retention signal is fused.

[17] The method for producing a TCR protein according to

[14] , comprising co-expressing a protein derived from the α chain of a TCR protein and a protein derived from the β chain of a TCR protein to obtain a dimer of the protein derived from the α chain of a TCR protein and a protein derived from the β chain of a TCR protein, or co-expressing a protein derived from the γ chain of a TCR protein and a protein derived from the δ chain of a TCR protein to obtain a dimer of the protein derived from the γ chain of a TCR protein and a protein derived from the δ chain of a TCR protein.

[18] The step of obtaining the TCR protein comprises a step of introducing an expression system into the plant cells, the expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette of the TCR protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette of a Rep / RepA protein derived from a geminivirus, wherein the expression cassette of the TCR protein comprises, in this order, a promoter, a nucleic acid fragment encoding the TCR protein, and two or more linked terminators, the method for producing the TCR protein according to

[14] .

[0010]

[19] An expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette for a TCR protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette for a Rep / RepA protein derived from a geminivirus, wherein the expression cassette for the TCR protein comprises, in this order, a promoter, a nucleic acid fragment encoding the TCR protein, and two or more linked terminators.

[20] An Agrobacterium bacterium transformed with the expression system described in

[19] .

[21] A plant cell into which the expression system described in

[19] has been introduced.

[22] A method for producing a transformed plant, comprising introducing the expression system described in

[19] into a plant cell.

[23] An expression vector comprising a first nucleic acid fragment comprising a geminivirus-derived LIR, a geminivirus-derived SIR, and an expression cassette of a TCR protein linked between the LIR and the SIR, wherein the expression cassette of the TCR protein comprises, in this order, a promoter, a multicloning site, and two or more linked terminators.

[0011]

[24] A method for evaluating the binding affinity between an MHC protein fused with an antigen peptide and a TCR protein, comprising the steps of: (a1) obtaining the MHC protein by the method for producing an MHC protein described in [1]; and (b1) evaluating the binding affinity between the MHC protein obtained in step (a1) and the TCR protein.

[25] A method for evaluating the binding affinity between an MHC protein containing an antigen peptide sequence and a TCR protein, comprising the steps of: (a2) obtaining the TCR protein by the method for producing a TCR protein described in

[14] ; and (b2) evaluating the binding affinity between the MHC protein and the TCR protein obtained in step (a2).

[0012] According to the present invention, it is possible to provide a method for producing an MHC protein that can be obtained inexpensively and with sufficient functionality, a method for producing a TCR protein that can be obtained with sufficient functionality, an expression system used in the production method, Agrobacterium bacteria transformed with the expression system, plant cells into which the expression system has been introduced, a method for producing a transformed plant using the expression system, an expression vector used in the production of the expression system, and a method for detecting the binding of an MHC protein to a TCR protein.

[0013] This figure schematically shows the domain structure of the DR protein in the example. This figure schematically shows the expression construct of the DR protein in the example. This figure schematically shows the expression construct of the DQ protein in the example. This figure schematically shows the expression construct of the DP protein in the example.

[0014] This figure shows the results of the expression analysis of DR protein in Experimental Example 2. This figure shows the results of ammonium sulfate precipitation of DR protein in Experimental Example 2. This figure shows the results of purifying DR protein using cobalt resin in Experimental Example 2. This figure shows the results of purifying DR#1 protein by ion exchange chromatography in Experimental Example 2. This figure shows the results of purifying DR#2 protein by ion exchange chromatography in Experimental Example 2.

[0015] This figure shows the results of CBB staining of DR#1, DR#2 proteins, and BSA in Experimental Example 3. This figure also shows the calibration curve based on CBB-stained BSA in Experimental Example 3.

[0016] This figure shows the results of ELISA analysis for DR#1 and DR#2 proteins in Experimental Example 4. This figure shows the relative luminescence intensity of the ELISA analysis in Experimental Example 4.

[0017] This figure shows the results of the expression analysis of DQ#1 and DQ#2 proteins in Experimental Example 5. This figure shows the results of purifying DQ protein using cobalt resin in Experimental Example 5. This figure shows the results of the expression analysis of DP#1 and DP#2 proteins in Experimental Example 6. This figure shows the results of purifying DP protein using cobalt resin in Experimental Example 6.

[0018] This figure schematically shows the expression construct of the TCR#1 protein in the example. This figure schematically shows the expression construct of the TCR#2 protein in the example. This figure shows the results of the expression analysis of TCRA#1 and TCRA#2 proteins in Experimental Example 7. This figure shows the results of the expression analysis of TCRB#1 and TCRB#2 proteins in Experimental Example 7. This figure shows the results of predicting the sites of glycosylation on the TCRα chain in Experimental Example 8. This figure shows the three-dimensional structure of the complex of the extracellular domain of the TCRα chain and HLA DR1 presenting influenza HA peptide, and the sites of glycosylation. This figure shows the results of the analysis of glycans of TCRA#2 that have undergone glycosylation treatment by Western blotting in Experimental Example 9. This figure shows the results of the analysis of glycans of TCRA#2 that have undergone glycosylation treatment in Experimental Example 9. This figure shows the results of analyzing the binding of TCR#2 obtained in Experimental Example 7 to DR#1 presenting CLIP or DR#2 presenting HA in Experimental Example 10. This figure shows the results of analyzing the binding of the glycan-removed TCR#2 obtained in Experimental Example 9 to DR#1 presenting CLIP or DR#2 presenting HA in Experimental Example 10.

[0019] (Method for producing MHC protein) The method for producing MHC protein according to this embodiment includes the step of expressing MHC protein in plant cells to obtain the MHC protein.

[0020] <MHC Protein> In this specification, MHC protein refers to the proteins that constitute the Major histocompatibility complex. MHC protein is a protein expressed in vertebrate cells.

[0021] A functional MHC protein can be obtained using the method for producing MHC protein according to this embodiment. Furthermore, a dimeric MHC protein can be obtained by co-expressing multiple subunits of MHC protein. As will be described later in the examples, the obtained dimeric MHC protein can bind to TCR protein. In this embodiment, it is preferable that the MHC protein has an endoplasmic reticulum retention signal and an apoplast retention signal. In this case, the expression level of the MHC protein can be significantly increased. In this embodiment, it is preferable that the MHC protein has a zipper-like helix domain. In this case, since the subunits of the MHC protein are strongly bound by the zipper-like helix domain, they can be purified and separated in a dimerized state. Furthermore, the obtained dimeric MHC protein can bind to a dimeric TCR protein.

[0022] Examples of vertebrates from which MHC proteins originate include mammals and birds. Examples of mammals include humans and non-human mammals, and examples of non-human mammals include pets and livestock. More specifically, examples of non-human mammals include monkeys, mice, rats, rabbits, cows, pigs, horses, sheep, dogs, and cats.

[0023] MHC proteins are encoded at the MHC locus or the β2 microglobulin locus. Examples of MHC proteins include MHC class I proteins and MHC class II proteins.

[0024] MHC class I proteins are composed of an α chain and β2 microglobulin. The α chain is encoded at the MHC locus, and β2 microglobulin is encoded at the β2 microglobulin locus.

[0025] MHC class II genes consist of an α chain and a β chain. The α and β chains are encoded by the MHC locus.

[0026] In this specification, the MHC protein expressed in plant cells refers to the α chain, β chain, β2 microglobulin, and subunits of a fusion protein containing these, as described above.

[0027] Exogenous and endogenous antigens are broken down into peptide fragments, and these antigen peptides bind to grooves formed by the α-chain of MHC class I proteins and grooves formed by the α-chain and β-chain of MHC class II proteins.

[0028] In this specification, human MHC proteins may be referred to as Human Leukocyte Antigen or Histocompatibility Leukocyte Antigen (HLA).

[0029] Examples of human MHC class I include HLA-A, HLA-B, and HLA-C. Examples of human MHC class II include HLA-DR, HLA-DQ, HLA-DP, HLA-DM, and HLA-DO.

[0030] Examples of HLA-A include HLA-A1, HLA-A2, HLA-A11, HLA-A24, HLA-A26, and HLA-A33. Examples of HLA-A gene alleles include HLA-A*01:01, HLA-A*02:01, HLA-A*02:06, HLA-A*02:07, HLA-A*03:01, HLA-A*11:01, HLA-A*24:02, HLA-A*26:01, HLA-A*26:03, HLA-A*31:01, and HLA-A*33:03.

[0031] Examples of HLA-B include HLA-B7, HLA-B13, HLA-B15, HLA-B27, HLA-B35, HLA-B39, HLA-B40, HLA-B44, HLA-B61, HLA-B62, HLA-B71, and the like. Examples of HLA-B gene alleles include HLA-B*13:01, HLA-B*15:01, HLA-B*15:18, HLA-B*27:05, HLA-B*35:01, HLA-B*39:01, HLA-B*40:01, HLA-B*40:02, HLA-B*40:03, HLA-B*40:06, HLA-B*44:03, and so on.

[0032] Examples of HLA-C include HLA-Cw1, HLA-Cw3, HLA-Cw4, HLA-Cw5, HLA-Cw6, and HLA-Cw7. Examples of HLA-C gene alleles include HLA-C*01:02, HLA-C*03:03, HLA-C*03:04, HLA-C*04:01, HLA-C*05:01, HLA-C*06:02, HLA-C*07:01, and HLA-C*07:02.

[0033] Examples of HLA-DRs include HLA-DR1, HLA-DR2, HLA-DR3, HLA-DR4, HLA-DR5, HLA-DR6, HLA-DR7, HLA-DR8, HLA-DR9, HLA-DR10, HLA-DR11, HLA-DR12, HLA-DR13, HLA-DR14, HLA-DR15, HLA-DR52, HLA-DR53, etc. An allele of the gene encoding the α chain of HLA-DRs is, for example, HLA-DRA*01. An allele of the gene encoding the β chain of HLA-DRs is, for example, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, etc. Examples of alleles for genes encoding the β chain include HLA-DRB1*01, HLA-DRB1*03, HLA-DRB1*04, HLA-DRB1*07, HLA-DRB1*08, HLA-DRB1*09, HLA-DRB1*10, HLA-DRB1*11, HLA-DRB1*12, HLA-DRB1*13, HLA-DRB1*14, HLA-DRB1*15, HLA-DRB1*16, HLA-DRB3*01, HLA-DRB4*01, HLA-DRB5*01, and the like.

[0034] Examples of HLA-DQ include HLA-DQ1, HLA-DQ2, HLA-DQ3, HLA-DQ4, HLA-DQ5, HLA-DQ6, HLA-DQ7, and HLA-DQ8. Examples of alleles of the gene encoding the α chain of HLA-DQ include HLA-DQA1*01, HLA-DQA1*02, HLA-DQA1*03, HLA-DQA1*04, HLA-DQA1*05, and HLA-DQA1*06. Examples of alleles for the gene encoding the β-chain of HLA-DQ include HLA-DQB1*02, HLA-DQB1*03, HLA-DQB1*04, HLA-DQB1*05, and HLA-DQB1*06.

[0035] Examples of HLA-DP include HLA-DPw1, HLA-DPw2, HLA-DPw3, HLA-DPw4, and HLA-DPw5. Examples of alleles of the gene encoding the α chain of HLA-DP include HLA-DPA1*01, HLA-DPA1*02, HLA-DPA1*03, and HLA-DPA1*04. Examples of alleles of the gene encoding the β chain of HLA-DP include HLA-DPB1*01, HLA-DPB1*02, HLA-DPB1*04, HLA-DPB1*05, and HLA-DPB1*09.

[0036] In this embodiment, the MHC protein expressed in plant cells is a recombinant MHC protein. Hereinafter, recombinant MHC protein may simply be referred to as MHC protein.

[0037] Recombinant MHC proteins may be proteins selected from (A) to (F) below: (A) Proteins containing the amino acid sequence of a natural MHC protein; (B) Proteins containing a portion of the amino acid sequence of a natural MHC protein; (C) Proteins containing amino acid sequences in which one or more amino acids are deleted, inserted, substituted and / or added to the amino acid sequence of a natural MHC protein; (D) Proteins containing amino acid sequences in which one or more amino acids are deleted, inserted, substituted and / or added to a portion of the amino acid sequence of a natural MHC protein; (E) Proteins containing amino acid sequences having 90% or more sequence identity with the amino acid sequence of a natural MHC protein; (F) Proteins containing amino acid sequences having 90% or more sequence identity with a portion of the amino acid sequence of a natural MHC protein.

[0038] The recombinant MHC protein in (A) may be, for example, a protein containing the amino acid sequence of a natural MHC protein. Natural MHC proteins encompass polymorphisms at the locus from which the protein originates.

[0039] Protein (A) may contain amino acid sequences other than those of the natural MHC protein. Protein (A) may also be a protein in which an amino acid sequence has been added to the amino acid sequence of the natural MHC protein. If protein (A) is a protein in which an amino acid has been added to the amino acid sequence of the natural MHC protein, the addition of amino acids may occur at either the N-terminus or the C-terminus of the amino acid sequence of the natural MHC protein, or at both the N-terminus and the C-terminus.

[0040] The number of amino acids constituting the protein (A) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0041] If the protein in (A) is a protein in which one or more amino acid residues are added to the amino acid sequence that constitutes a natural MHC protein, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0042] The recombinant MHC protein of (B) may be a protein containing a portion of the amino acid sequence of a natural MHC protein (hereinafter sometimes referred to as a partial sequence). The partial sequence is not particularly limited, but examples include the amino acid sequence that constitutes the extracellular domain of a natural MHC protein. As the amino acid sequence that constitutes the extracellular domain, the amino acid sequence obtained by removing the signal sequence, transmembrane domain and intracellular domain from the full length of the amino acid sequence of a natural MHC protein is preferred.

[0043] The protein of (B) may contain an amino acid sequence other than the partial sequence. The protein of (B) may be a protein in which an amino acid sequence is added to the partial sequence. When the protein of (B) is a protein in which an amino acid is added to the partial sequence, the addition of the amino acid may be performed at either the N-terminus or the C-terminus of the partial sequence, or may be performed at both the N-terminus and the C-terminus.

[0044] The number of amino acids constituting the protein of (B) is not particularly limited, and may be, for example, 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0045] When the protein of (B) is a protein in which one or more amino acid residues are added to the partial sequence, the number of added amino acids may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0046] The protein recombinant MHC protein of (C) may be a protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added in the amino acid sequence of the natural MHC protein (hereinafter sometimes simply referred to as a mutant sequence). Alternatively, the protein of (C) may be a protein in which one or more amino acid residues are added to either or both of the N-terminus and the C-terminus of the mutant sequence.

[0047] In the present specification, "a plurality" is not particularly limited, and examples thereof include 2 to 50, 2 to 40, 2 to 30, 2 to 27, 2 to 23, 2 to 19, 2 to 16, 2 to 13, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 2 to 3, or 2.

[0048] The number of amino acids constituting the protein (C) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0049] If protein (C) is a protein in which one or more amino acid residues are added to a mutant sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0050] The recombinant MHC protein of (D) may be a protein containing an amino acid sequence (hereinafter also referred to as the "mutant subsequence") in which one or more amino acids are deleted, inserted, substituted and / or added in the subsequence described above in (B). Alternatively, the protein of (D) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant subsequence.

[0051] The number of amino acids constituting the protein (D) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0052] If protein (D) is a protein in which one or more amino acid residues are added to a mutant partial sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0053] The recombinant MHC protein of (E) may be a protein containing an amino acid sequence (hereinafter sometimes simply referred to as a homologous sequence) that has 90% or more sequence identity with the amino acid sequence of a natural MHC protein. The sequence identity of the homologous sequence may be 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more. Alternatively, the protein of (E) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the homologous sequence.

[0054] In this specification, amino acid sequence identity is a value that indicates the percentage of the target amino acid sequence that matches a reference amino acid sequence. The sequence identity of the target amino acid sequence with respect to the reference amino acid sequence can be determined, for example, as follows: First, the reference amino acid sequence and the target amino acid sequence are aligned. Here, gaps may be included in each amino acid sequence to maximize sequence identity. Next, the number of matching amino acids in the reference amino acid sequence and the target amino acid sequence is calculated, and the sequence identity can be determined according to the following formula (1): Sequence identity (%) = Number of matching amino acids / Total number of amino acids in the target amino acid sequence × 100 …(1) The value of amino acid sequence identity can be obtained by calculation based on alignment obtained by known homology search software such as BLASTP.

[0055] The number of amino acids that make up the protein (E) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0056] If protein (E) is a protein in which one or more amino acid residues are added to a mutant partial amino acid sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0057] The recombinant MHC protein of (F) may be a protein containing an amino acid sequence (hereinafter sometimes simply referred to as a homologous partial sequence) that has 90% or more sequence identity with the partial sequence described above in (B). The sequence identity of the homologous partial sequence may be 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more. Alternatively, the protein of (F) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the homologous partial sequence.

[0058] The number of amino acids that make up the protein (F) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0059] If protein (F) is a protein in which one or more amino acid residues are added to a mutant partial amino acid sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0060] The proteins (A) to (F) may be natural MHC proteins, proteins consisting of partial sequences, proteins consisting of mutant sequences, proteins consisting of mutant partial sequences, proteins consisting of homologous sequences, or fusion proteins obtained by fusing a protein consisting of a homologous partial sequence with another protein or peptide. Examples of other proteins or peptides include signal sequences; endoplasmic reticulum retention signals; peptides constituting zipper-like helix domains; antigen peptides; marker enzyme proteins such as alkaline phosphatase or partial peptides thereof; fluorescent proteins such as GFP or partial peptides thereof; tag peptides such as His tags and FLAG tags.

[0061] The natural MHC protein from which the recombinant MHC protein described above is derived may be an α-chain or a β-chain, and the allele of the gene encoding it is not particularly limited. In this embodiment, a dimer can be obtained by co-expressing a protein derived from the α-chain of an MHC class I protein and a protein derived from β2 microglobulin of an MHC class I protein. In this embodiment, a dimer can be obtained by co-expressing a protein derived from the α-chain of an MHC class II protein and a protein derived from the β-chain of an MHC class II protein. Examples of human MHC class II proteins include the proteins shown in Table 1.

[0062]

[0063] More specifically, the recombinant MHC protein may be a protein selected from (a) to (f) below. (a) Proteins containing an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (b) Proteins containing a portion of an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (c) Proteins containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added to an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (d) Proteins containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added to a portion of an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (e) Proteins containing an amino acid sequence having 90% or more sequence identity with an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (f) Proteins containing an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus A protein containing an amino acid sequence that has more than 90% sequence identity with a portion of the amino acid sequence encoded by the microglobulin gene locus.

[0064] The amino acid sequences encoded by the MHC locus and β2 microglobulin locus in protein (a) are the amino acid sequences of the native MHC protein and encompass polymorphisms of these loci.

[0065] The protein in (a) may contain an amino acid sequence other than that of a natural MHC protein. The protein in (a) may also be a protein in which an amino acid sequence has been added to the amino acid sequence of a natural MHC protein. If the protein in (a) is a protein in which an amino acid has been added to the amino acid sequence of a natural MHC protein, the addition of amino acids may occur at either the N-terminus or the C-terminus of the amino acid sequence of the natural MHC protein, or at both the N-terminus and the C-terminus.

[0066] The number of amino acids constituting the protein in (a) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0067] If the protein in (a) is a protein in which one or more amino acid residues are added to the amino acid sequence that constitutes a natural MHC protein, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0068] (b) A portion of the amino acid sequence encoded by the MHC gene locus or the β2 microglobulin gene locus is called a partial sequence of the amino acid sequence that constitutes the natural MHC protein (hereinafter sometimes simply referred to as a partial sequence). The partial sequence is not particularly limited, but for example, an amino acid sequence that constitutes the extracellular domain of the natural MHC protein is an example. As the amino acid sequence that constitutes the extracellular domain, an amino acid sequence obtained by removing the signal sequence, transmembrane domain and intracellular domain from the full length of the amino acid sequence of the natural MHC protein is preferred.

[0069] The protein in (b) may contain amino acid sequences other than the partial sequence. The protein in (b) may also be a protein in which an amino acid sequence has been added to the partial sequence. If the protein in (b) is a protein in which amino acids have been added to the partial sequence, the addition of amino acids may occur at either the N-terminus or the C-terminus of the partial sequence, or at both the N-terminus and the C-terminus.

[0070] The number of amino acids constituting the protein in (b) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0071] If the protein in (b) is a protein in which one or more amino acid residues are added to a partial sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0072] Protein (c) The protein in (c) may be a protein consisting of an amino acid sequence in which one or more amino acids have been mutated in the amino acid sequence of a natural MHC protein (hereinafter sometimes simply referred to as a mutant sequence). Alternatively, the protein in (c) may be a protein in which one or more amino acid residues have been added to either the N-terminus and / or C-terminus of the mutant sequence.

[0073] The number of mutated amino acids in protein (c) may be the same as that described above for protein (C).

[0074] The number of amino acids constituting the protein in (c) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0075] If the protein in (c) is a protein in which one or more amino acid residues are added to the mutant sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0076] Protein (d) Protein (d) may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the subsequence described above in (b) (hereinafter also referred to as the "mutant subsequence"). Alternatively, protein (d) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant subsequence.

[0077] The number of amino acids constituting the protein in (d) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0078] If the protein in (d) is a protein in which one or more amino acid residues are added to a mutant partial sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0079] Protein (e) Protein (e) may be a protein consisting of an amino acid sequence having 90% or more sequence identity with a natural MHC protein (hereinafter sometimes simply referred to as a homologous sequence). The sequence identity of the homologous sequence may be 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more. Alternatively, protein (e) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the homologous sequence.

[0080] The number of amino acids constituting the protein in (e) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0081] If the protein in (e) is a protein in which one or more amino acid residues are added to a mutant partial amino acid sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0082] Protein (f) The protein in (f) may consist of an amino acid sequence (hereinafter sometimes simply referred to as a homologous partial sequence) that has 90% or more sequence identity with the partial sequence described above in (b). The sequence identity of the homologous partial sequence may be 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more. Alternatively, the protein in (f) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the homologous partial sequence.

[0083] The number of amino acids that make up the protein of (f) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0084] If the protein in (f) is a protein in which one or more amino acid residues are added to a mutant partial amino acid sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0085] The proteins (a) to (f) may be natural MHC proteins, proteins consisting of partial sequences, proteins consisting of mutant sequences, proteins consisting of mutant partial sequences, proteins consisting of homologous sequences, or fusion proteins obtained by fusing a protein consisting of a homologous partial sequence with another protein or peptide. Examples of other proteins or peptides include those mentioned above in proteins (A) to (F).

[0086] In this embodiment, the α and β chains of the MHC protein may be co-expressed in plant cells to obtain the MHC protein. In this case, the protein folding proceeds normally, and an MHC protein with normal three-dimensional structure and function can be obtained.

[0087] Examples of combinations of α and β chains of MHC proteins to be co-expressed include any combination of the α chain of an MHC class I protein and β2 microglobulin, and any combination of the α chain of an MHC class II protein and the β chain of an MHC class II protein.

[0088] When co-expressing the α and β chains of human MHC class II proteins, the combination of α and β chains may be, for example, the α and β chains of HLA-DR, the α and β chains of HLA-DQ, the α and β chains of HLA-DP, or any combination of α and β chains.

[0089] More specific examples of MHC proteins are described in detail below.

[0090] ≪The α-chain of the DR protein≫ The α-chain of the DR protein may be a protein selected from the following (Rα1), (Rα2), and (Rα3): (Rα1) A protein containing the amino acid sequence described in SEQ ID NO: 2; (Rα2) A protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, and / or added to the amino acid sequence described in SEQ ID NO: 2; (Rα3) A protein containing an amino acid sequence having 90% or more sequence identity with the amino acid sequence of SEQ ID NO: 2;

[0091] The protein consisting of the amino acid sequence described in Sequence ID No. 2 is the extracellular domain of HLA-DRA*01:01.

[0092] The protein of (Rα1) may contain amino acid sequences other than the amino acid sequence described in SEQ ID NO: 2. The protein of (Rα1) may also be a protein in which an amino acid sequence has been added to the amino acid sequence described in SEQ ID NO: 2. If the protein of (Rα1) is a protein in which an amino acid has been added to the amino acid sequence described in SEQ ID NO: 2, the addition of amino acids may be performed at either the N-terminus or the C-terminus of the amino acid sequence described in SEQ ID NO: 2, or at both the N-terminus and the C-terminus.

[0093] The number of amino acids that make up the (Rα1) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0094] If the (Rα1) protein is a protein in which one or more amino acid residues are added to the amino acid sequence described in Sequence ID No. 2, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0095] The (Rα2) protein may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the amino acid sequence described in SEQ ID NO: 2 (hereinafter also referred to as the "mutant sequence of SEQ ID NO: 2"). Alternatively, the (Rα2) protein may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence of SEQ ID NO: 2.

[0096] The number of mutated amino acids in the amino acid sequence described in Sequence ID No. 2 may be the same as that described above for protein (c).

[0097] The number of amino acids that make up the (Rα2) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0098] If the (Rα2) protein is a protein in which one or more amino acid residues are added to the mutant sequence of Sequence ID No. 2, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0099] The protein of (Rα3) contains an amino acid sequence that has 90% or more sequence identity with the amino acid sequence described in SEQ ID NO: 2 (hereinafter, this sequence is referred to as the "homologous sequence of SEQ ID NO: 2"). The homologous sequence of SEQ ID NO: 2 in the amino acid sequence of the protein of (Rα3) may have 91% or more, or 92% or more, sequence identity with respect to the amino acid sequence of SEQ ID NO: 2.

[0100] The number of amino acids that make up the (Rα3) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0101] If the (Rα3) protein is a protein in which one or more amino acid residues are added to the homologous sequence of SEQ ID NO: 2, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0102] The protein selected from (Rα1), (Rα2), and (Rα3) may be a fusion protein formed by fusing it with another protein or peptide. For example, another protein or peptide may be bound to either or both of the N-terminus and C-terminus of a protein containing the amino acid sequence of SEQ ID NO: 2, a variant sequence of SEQ ID NO: 2, or a homologous sequence of SEQ ID NO: 2. Examples of other proteins or peptides include those mentioned above in proteins (a) to (f).

[0103] ≪Beta Chain of DR Protein≫ The beta chain of the DR protein may be a protein selected from the following (Rβ1), (Rβ2), and (Rβ3): (Rβ1) A protein containing the amino acid sequence described in SEQ ID NO: 12; (Rβ2) A protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, and / or added to the amino acid sequence described in SEQ ID NO: 12; (Rβ3) A protein containing an amino acid sequence having 90% or more sequence identity with the amino acid sequence of SEQ ID NO: 12;

[0104] The protein consisting of the amino acid sequence described in Sequence ID No. 12 is the extracellular domain of HLA-DRB1*01:01.

[0105] The (Rβ1) protein may contain amino acid sequences other than those described in SEQ ID NO: 12. The (Rβ1) protein may also be a protein in which an amino acid sequence has been added to the amino acid sequence described in SEQ ID NO: 12. If the (Rβ1) protein is a protein in which an amino acid has been added to the amino acid sequence described in SEQ ID NO: 12, the addition of amino acids may occur at either the N-terminus or the C-terminus of the amino acid sequence described in SEQ ID NO: 12, or at both the N-terminus and the C-terminus.

[0106] The number of amino acids that make up the (Rβ1) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0107] If the (Rβ1) protein is a protein in which one or more amino acid residues are added to the amino acid sequence described in Sequence ID No. 12, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0108] The (Rβ2) protein may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the amino acid sequence described in SEQ ID NO: 12 (hereinafter also referred to as the "mutant sequence of SEQ ID NO: 12"). Alternatively, the (Rβ2) protein may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence of SEQ ID NO: 12.

[0109] The number of mutated amino acids in the amino acid sequence described in Sequence ID No. 12 may be the same as that described above for the (Rα2) protein.

[0110] The number of amino acids that make up the (Rβ2) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0111] If the (Rβ2) protein is a protein in which one or more amino acid residues are added to the mutant sequence of Sequence ID No. 12, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0112] The (Rβ3) protein contains an amino acid sequence having 90% or more sequence identity with the amino acid sequence described in SEQ ID NO: 12 (hereinafter, this sequence is referred to as the "homologous sequence of SEQ ID NO: 12"). The homologous sequence of SEQ ID NO: 12 in the amino acid sequence of the (Rβ3) protein may have 91% or more, or 92% or more, sequence identity with respect to the amino acid sequence of SEQ ID NO: 12.

[0113] The number of amino acids that make up the (Rβ3) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0114] If the (Rβ3) protein is a protein in which one or more amino acid residues are added to the homologous sequence of SEQ ID NO: 12, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0115] The protein selected from (Rβ1), (Rβ2), and (Rβ3) may be a fusion protein formed by fusing it with another protein or peptide. For example, another protein or peptide may be bound to either or both of the N-terminus and C-terminus of a protein containing the amino acid sequence of SEQ ID NO: 12, a variant sequence of SEQ ID NO: 12, or a homologous sequence of SEQ ID NO: 12. Examples of other proteins or peptides include those mentioned above in the protein selected from (Rα1), (Rα2), and (Rα3).

[0116] ≪α-chain of DQ protein≫ The α-chain of the DQ protein may be a protein selected from the following (Qα1), (Qα2), and (Qα3): (Qα1) A protein containing the amino acid sequence described in SEQ ID NO: 16; (Qα2) A protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, and / or added to the amino acid sequence described in SEQ ID NO: 16; (Qα3) A protein containing an amino acid sequence having 90% or more sequence identity with the amino acid sequence of SEQ ID NO: 16;

[0117] The protein consisting of the amino acid sequence described in Sequence ID No. 16 is the extracellular domain of HLA-DQA1*01:02.

[0118] The protein of (Qα1) may contain amino acid sequences other than the amino acid sequence described in SEQ ID NO: 16. The protein of (Qα1) may also be a protein in which an amino acid sequence has been added to the amino acid sequence described in SEQ ID NO: 16. If the protein of (Qα1) is a protein in which an amino acid has been added to the amino acid sequence described in SEQ ID NO: 16, the addition of amino acids may be performed at either the N-terminus or the C-terminus of the amino acid sequence described in SEQ ID NO: 16, or at both the N-terminus and the C-terminus.

[0119] The number of amino acids constituting the (Qα1) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0120] If the protein (Qα1) is a protein in which one or more amino acid residues are added to the amino acid sequence described in Sequence ID No. 16, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0121] The (Qα2) protein may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the amino acid sequence described in SEQ ID NO: 16 (hereinafter also referred to as the "mutant sequence of SEQ ID NO: 16"). Alternatively, the (Qα2) protein may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence of SEQ ID NO: 16.

[0122] The number of mutated amino acids in the amino acid sequence described in Sequence ID No. 16 may be the same as that described above for the (Rα2) protein.

[0123] The number of amino acids that make up the (Qα2) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0124] If the (Qα2) protein is a protein in which one or more amino acid residues are added to the mutant sequence of Sequence ID No. 16, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0125] The (Qα3) protein contains an amino acid sequence having 90% or more sequence identity with the amino acid sequence described in SEQ ID NO: 16 (hereinafter, this sequence is referred to as the "homologous sequence of SEQ ID NO: 16"). The homologous sequence of SEQ ID NO: 16 in the amino acid sequence of the (Qα3) protein may have 91% or more, or 92% or more, sequence identity with respect to the amino acid sequence of SEQ ID NO: 16.

[0126] The number of amino acids that make up the (Qα3) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0127] If the (Qα3) protein is a protein in which one or more amino acid residues are added to the homologous sequence of Sequence ID No. 16, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0128] The protein selected from (Qα1), (Qα2), and (Qα3) may be a fusion protein formed by fusing it with another protein or peptide. For example, another protein or peptide may be bound to either or both of the N-terminus and C-terminus of a protein containing the amino acid sequence of SEQ ID NO: 16, a variant sequence of SEQ ID NO: 16, or a homologous sequence of SEQ ID NO: 16. Examples of other proteins or peptides include those mentioned above in the protein selected from (Rα1), (Rα2), and (Rα3).

[0129] ≪β-chain of DQ protein≫ The β-chain of the DQ protein may be a protein selected from the following (Qβ1), (Qβ2), and (Qβ3): (Qβ1) A protein containing the amino acid sequence described in SEQ ID NO: 19; (Qβ2) A protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, and / or added to the amino acid sequence described in SEQ ID NO: 19; (Qβ3) A protein containing an amino acid sequence having 90% or more sequence identity with the amino acid sequence of SEQ ID NO: 19;

[0130] The protein consisting of the amino acid sequence described in Sequence ID No. 19 is the extracellular domain of HLA-DQB1*06:02.

[0131] The protein of (Qβ1) may contain amino acid sequences other than the amino acid sequence described in SEQ ID NO: 19. The protein of (Qβ1) may also be a protein in which an amino acid sequence has been added to the amino acid sequence described in SEQ ID NO: 19. If the protein of (Qβ1) is a protein in which an amino acid has been added to the amino acid sequence described in SEQ ID NO: 19, the addition of amino acids may be performed at either the N-terminus or the C-terminus of the amino acid sequence described in SEQ ID NO: 19, or at both the N-terminus and the C-terminus.

[0132] The number of amino acids that make up the (Qβ1) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0133] If the (Qβ1) protein is a protein in which one or more amino acid residues are added to the amino acid sequence described in Sequence ID No. 19, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0134] The (Qβ2) protein may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the amino acid sequence described in SEQ ID NO: 19 (hereinafter also referred to as the "mutant sequence of SEQ ID NO: 19"). Alternatively, the (Qβ2) protein may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence of SEQ ID NO: 19.

[0135] The number of mutated amino acids in the amino acid sequence described in Sequence ID No. 19 may be the same as that described above for the (Rα2) protein.

[0136] The number of amino acids that make up the (Qβ2) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0137] If the (Qβ2) protein is a protein in which one or more amino acid residues are added to the mutant sequence of Sequence ID No. 19, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0138] The (Qβ3) protein contains an amino acid sequence having 90% or more sequence identity with the amino acid sequence described in SEQ ID NO: 19 (hereinafter, this sequence is referred to as the "homologous sequence of SEQ ID NO: 19"). The homologous sequence of SEQ ID NO: 19 in the amino acid sequence of the (Qβ3) protein may have 91% or more, or 92% or more, sequence identity with respect to the amino acid sequence of SEQ ID NO: 19.

[0139] The number of amino acids that make up the (Qβ3) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0140] If the (Qβ3) protein is a protein in which one or more amino acid residues are added to the homologous sequence of Sequence ID No. 19, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0141] The protein selected from (Qβ1), (Qβ2), and (Qβ3) may be a fusion protein formed by fusing it with another protein or peptide. For example, another protein or peptide may be bound to either or both of the N-terminus and C-terminus of a protein containing the amino acid sequence of SEQ ID NO: 19, a variant sequence of SEQ ID NO: 19, or a homologous sequence of SEQ ID NO: 19. Examples of other proteins or peptides include those mentioned above in the protein selected from (Rα1), (Rα2), and (Rα3).

[0142] ≪α-chain of DP protein≫ The α-chain of the DP protein may be a protein selected from the following (Pα1), (Pα2), and (Pα3): (Pα1) A protein containing the amino acid sequence described in SEQ ID NO: 21; (Pα2) A protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, and / or added to the amino acid sequence described in SEQ ID NO: 21; (Pα3) A protein containing an amino acid sequence having 90% or more sequence identity with the amino acid sequence of SEQ ID NO: 21.

[0143] The protein consisting of the amino acid sequence described in Sequence ID No. 21 is the extracellular domain of HLA-DPA1*02:01.

[0144] The protein of (Pα1) may contain amino acid sequences other than the amino acid sequence described in SEQ ID NO: 21. The protein of (Pα1) may also be a protein in which an amino acid sequence has been added to the amino acid sequence described in SEQ ID NO: 21. If the protein of (Pα1) is a protein in which an amino acid has been added to the amino acid sequence described in SEQ ID NO: 21, the addition of amino acids may be performed at either the N-terminus or the C-terminus of the amino acid sequence described in SEQ ID NO: 21, or at both the N-terminus and the C-terminus.

[0145] The number of amino acids constituting the (Pα1) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0146] If the protein (Pα1) is a protein in which one or more amino acid residues are added to the amino acid sequence described in SEQ ID NO: 21, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0147] The (Pα2) protein may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the amino acid sequence described in SEQ ID NO: 21 (hereinafter also referred to as the "mutant sequence of SEQ ID NO: 21"). Alternatively, the (Pα2) protein may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence of SEQ ID NO: 21.

[0148] The number of mutated amino acids in the amino acid sequence described in Sequence ID No. 21 may be the same as that described above for the (Rα2) protein.

[0149] The number of amino acids constituting the (Pα2) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0150] If the (Pα2) protein is a protein in which one or more amino acid residues are added to the mutant sequence of Sequence ID No. 21, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0151] The (Pα3) protein contains an amino acid sequence having 90% or more sequence identity with the amino acid sequence described in SEQ ID NO: 21 (hereinafter, this sequence is referred to as the "homologous sequence of SEQ ID NO: 21"). The homologous sequence of SEQ ID NO: 21 in the amino acid sequence of the (Pα3) protein may have 91% or more, or 92% or more, sequence identity with respect to the amino acid sequence of SEQ ID NO: 21.

[0152] The number of amino acids constituting the (Pα3) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0153] If the protein (Pα3) is a protein in which one or more amino acid residues are added to the homologous sequence of Sequence ID No. 21, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0154] The protein selected from (Pα1), (Pα2), and (Pα3) may be a fusion protein formed by fusing it with another protein or peptide. For example, another protein or peptide may be bound to either or both of the N-terminus and C-terminus of a protein containing the amino acid sequence of SEQ ID NO: 21, a variant sequence of SEQ ID NO: 21, or a homologous sequence of SEQ ID NO: 21. Examples of other proteins or peptides include those mentioned above in the protein selected from (Rα1), (Rα2), and (Rα3).

[0155] ≪β-chain of DP protein≫ The β-chain of the DP protein may be a protein selected from the following (Pβ1), (Pβ2), and (Pβ3): (Pβ1) A protein containing the amino acid sequence described in SEQ ID NO: 24; (Pβ2) A protein containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted, and / or added to the amino acid sequence described in SEQ ID NO: 24; (Pβ3) A protein containing an amino acid sequence having 90% or more sequence identity with the amino acid sequence of SEQ ID NO: 24.

[0156] The protein consisting of the amino acid sequence described in Sequence ID No. 4 is the extracellular domain of HLA-DPB1*01:01.

[0157] The protein of (Pβ1) may contain amino acid sequences other than the amino acid sequence described in SEQ ID NO: 24. The protein of (Pβ1) may also be a protein in which an amino acid sequence has been added to the amino acid sequence described in SEQ ID NO: 24. If the protein of (Pβ1) is a protein in which an amino acid has been added to the amino acid sequence described in SEQ ID NO: 24, the addition of amino acids may be performed at either the N-terminus or the C-terminus of the amino acid sequence described in SEQ ID NO: 24, or at both the N-terminus and the C-terminus.

[0158] The number of amino acids that make up the (Pβ1) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0159] If the (Pβ1) protein is a protein in which one or more amino acid residues are added to the amino acid sequence described in Sequence ID No. 24, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0160] The (Pβ2) protein may be a protein consisting of an amino acid sequence in which one or more amino acids are mutated in the amino acid sequence described in SEQ ID NO: 24 (hereinafter also referred to as the "mutant sequence of SEQ ID NO: 24"). Alternatively, the (Pβ2) protein may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence of SEQ ID NO: 24.

[0161] The number of mutated amino acids in the amino acid sequence described in Sequence ID No. 24 may be the same as that described above for the (Rα2) protein.

[0162] The number of amino acids that make up the (Pβ2) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0163] If the (Pβ2) protein is a protein in which one or more amino acid residues are added to the mutant sequence of Sequence ID No. 24, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0164] The (Pβ3) protein contains an amino acid sequence having 90% or more sequence identity with the amino acid sequence described in SEQ ID NO: 24 (hereinafter referred to as the "homologous sequence of SEQ ID NO: 24"). The homologous sequence of SEQ ID NO: 24 in the amino acid sequence of the (Pβ3) protein may have 91% or more, or 92% or more, sequence identity with respect to the amino acid sequence of SEQ ID NO: 24.

[0165] The number of amino acids that make up the (Pβ3) protein is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0166] If the (Pβ3) protein is a protein in which one or more amino acid residues are added to the homologous sequence of Sequence ID No. 24, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0167] The protein selected from (Pβ1), (Pβ2), and (Pβ3) may be a fusion protein formed by fusing it with another protein or peptide. For example, another protein or peptide may be bound to either or both of the N-terminus and C-terminus of a protein containing the amino acid sequence of SEQ ID NO: 24, a variant sequence of SEQ ID NO: 24, or a homologous sequence of SEQ ID NO: 24. Examples of other proteins or peptides include those mentioned above in the protein selected from (Rα1), (Rα2), and (Rα3).

[0168] <<Peptides>> As described above, recombinant MHC proteins may also be fusion proteins formed by the fusion of peptides. Examples of peptides include peptides constituting zipper-like helix domains, signal sequences, endoplasmic reticulum retention signals, apoplast retention signals, antigen peptides, and linkers. The linker may link the aforementioned natural MHC protein, a protein consisting of a partial sequence, a protein consisting of a mutant sequence, a protein consisting of a mutant partial sequence, a protein consisting of a homologous sequence, or a protein consisting of a homologous partial sequence with the aforementioned peptides.

[0169] When the MHC protein is an MHC class I protein or an MHC class II protein, it is preferable that these proteins have a peptide sequence that constitutes a zipper-like helix domain.

[0170] Proteins containing peptide sequences that constitute zipper-like helix domains promote protein-protein interactions, which facilitates dimer formation.

[0171] The zipper-like helix domain is preferably fused to the C-terminal side of the amino acid sequence derived from the MHC protein.

[0172] The peptides constituting the zipper-like helix domain are not particularly limited, and examples include the peptides shown in SEQ ID NOs. 35-36.

[0173] The peptides constituting the zipper-like helix domain may have mutations in the amino acid sequences of SEQ ID NOs. 35-36. When the amino acid sequences of the zipper-like helix domain have mutations in the amino acid sequences of SEQ ID NOs. 35-36, each amino acid sequence preferably has, for example, 70% or more sequence identity with the amino acid sequences of SEQ ID NOs. 35-36, more preferably 80% or more, even more preferably 90% or more, and particularly preferably 95% or more.

[0174] The MHC protein preferably has a signal sequence. Proteins with a signal sequence are more likely to localize to organelles. The signal sequence is preferably located at the N-terminus of the MHC protein.

[0175] The signal sequence is not particularly limited and examples include the signal sequence of extensin from tobacco benthamiana and the signal sequence of chitinase from Arabidopsis thaliana. Sequence ID 32 shows the amino acid sequence of the signal sequence of extensin from tobacco benthamiana. Sequence ID 33 shows the amino acid sequence of the signal sequence of chitinase from Arabidopsis thaliana.

[0176] The signal sequence may have mutations in the amino acid sequences of SEQ ID NOs. 32-33. If the amino acid sequences of the signal sequence have mutations in the amino acid sequences of SEQ ID NOs. 32-33, each amino acid sequence preferably has, for example, 70% or more sequence identity with the amino acid sequences of SEQ ID NOs. 32-33, more preferably 80% or more, even more preferably 90% or more, and particularly preferably 95% or more.

[0177] The MHC protein preferably has an endoplasmic reticulum (ER) retention signal. Proteins with an ER retention signal are more likely to localize to the ER and have higher expression levels. The location of the ER retention signal is preferably at the C-terminus of the MHC protein.

[0178] The endoplasmic reticulum (ER) retention signal is not particularly limited, and for example, the amino acid sequence shown in SEQ ID NO: 34 can be cited. The ER retention signal may have mutations in the amino acid sequence of SEQ ID NO: 34. If the amino acid sequence of the ER retention signal has mutations in the amino acid sequence of SEQ ID NO: 34, the number of mutated amino acids can be 1 to 2. The ER retention signal preferably has the amino acid sequence "KD".

[0179] Proteins possessing an apoplast retention signal are more likely to localize to the apoplast and have higher expression levels. The location of the apoplast retention signal is preferably at the C-terminus of the MHC protein.

[0180] The apoplast retention signal is not particularly limited, and examples include peptides consisting of the amino acid sequence of SEQ ID NO: 51 or SEQ ID NO: 59 (e.g., Jiang MC et al., Fusion of a Novel Native Signal Peptide Enhanced the Secretion and Solubility of Bioactive Human Interferon Gamma Glycoproteins in Nicotiana benthamiana Using the Bamboo Mosaic Virus-Based Expression System, Front Plant Sci. 2020 Nov 12;11:594758. doi: 10.3389 / fpls.2020.594758).

[0181] The apoplast retention signal may have mutations in the amino acid sequence of SEQ ID NO: 51. If the amino acid sequence of the apoplast retention signal has mutations in the amino acid sequence of SEQ ID NO: 51, each amino acid sequence preferably has, for example, 70% or more sequence identity with respect to the amino acid sequence of SEQ ID NO: 51, more preferably 80% or more sequence identity, even more preferably 90% or more sequence identity, and particularly preferably 95% or more sequence identity. If the amino acid sequence of the apoplast retention signal has mutations in the amino acid sequence of SEQ ID NO: 59, each amino acid sequence preferably has, for example 70% or more sequence identity with respect to the amino acid sequence of SEQ ID NO: 59, more preferably 80% or more sequence identity, even more preferably 90% or more sequence identity, and particularly preferably 95% or more sequence identity.

[0182] If the MHC protein is an MHC class I protein or an MHC class II protein, these proteins may contain the sequence of the antigen peptide. The antigen is not particularly limited and may be an endogenous antigen or an exogenous antigen.

[0183] It is preferable to fuse the antigen peptide to the N-terminal side of the amino acid sequence derived from the MHC protein.

[0184] MHC proteins preferably have a signal sequence at their N-terminus. MHC proteins preferably have a zipper-like helix domain at their C-terminus. MHC proteins preferably have endoplasmic reticulum retention signals and apoplast retention signals at their C-terminus.

[0185] The plant species that express MHC protein are not particularly limited, but include the following examples: Grasses [wheat (Triticum aestivum L.), rice (Oryza sativa), barley (Hordeum vulgare L.), maize (Zea mays L.), sorghum (Sorghum bicolor (L.) Moench), Erianthus (Erianthus spp.), Guinea grass (Panicum maximum Jacq.), Miscanthus (Miscanthus spp.), sugarcane (Saccharum officinarum L.), Napier grass (Pennisetum purpurea)] [Schumach, Pampas grass (Cortaderia argentea Stapf.), Perennial ryegrass (Lolium perenne L.), Italian ryegrass (Lolium multiflorum Lam.), Meadow fescue (Festuca pratensis Huds.), Tall fescue (Festuca arundinacea Schreb.), Orchardgrass (Dactylis glomerata L.), Timothy (Phleum pratense L.), etc.]; Legumes [Soybean (Glycine max), Adzuki bean (Vigna angularis Willd.), Green bean (Phaseolus) [Vulgaris L., broad bean (Vicia faba L.), etc.]; Malvaceae [Cotton (Gossypium spp.), kenaf (Hibiscus cannabinus), okra (Abelmoschus esculentus), etc.]; Solanaceae [Eggplant (Solanum melongena L.), tomato (Solanum lycopersicum), bell pepper (Capsicum annuum L. var. angulosum Mill.), chili pepper (Capsicum annuum L.), tobacco (Nicotiana tabacum L.), benthamiana tobacco (Nicotiana benthamiana) etc.];Brassicaceae [Arabidopsis thaliana, Brassica campestris L., Chinese cabbage (Brassica pekinensis Rupr.), cabbage (Brassica oleracea L. var. capitata L.), radish (Raphanus sativus L.), rapeseed (Brassica campestris L., B. napus L.), etc.]; Cucurbitaceae [Cucumis sativus L., melon (Cucumis melo L.), watermelon (Citrus vulgaris) [Schrad.], pumpkin (C. moschata Duch., C. maxima Duch.), etc.]; Convolvulaceae [sweet potato (Ipomoea batatas), etc.]; Liliaceae [leeks (Allium fistulosum L.), onions (Allium cepa L.), chives (Allium tuberosum Rottl.), garlic (Allium sativum L.), asparagus (Asparagus officinalis L.), etc.]; Lamiaceae [perilla (Perilla frutescens Britt. var. crispa), etc.]; Asteraceae [Chrysanthemum (Chrysanthemum morifolium), Chrysanthemum coronarium L., Lettuce (Lactuca sativa L. var. capitata L.), etc.]; Rosaceae [Rose (Rose hybrida Hort.), Strawberry (Fragaria x ananassa Duch.), etc.]; Rutaceae [Citrus (Citrus unshiu), Japanese pepper (Zanthoxylum piperitum DC.), etc.]; Myrtaceae [Eucalyptus (Eucalyptus globelus Labill), etc.]; Salicaceae [e.g., poplar (Populus nigra L. var. italica Koehne)]; Amaranthaceae [e.g., spinach (Spinacia oleracea L.), sugar beet (Beta vulgaris L.)];Gentianaceae family [Gentian (Gentiana scabra Bunge var. buergeri Maxim.), etc.]; Caryophyllaceae family [Carnation (Dianthus caryophyllus L.), etc.].

[0186] The type of plant cell expressing the MHC protein may be cells that make up plant tissues or organs. The plant tissues and organs are not particularly limited and include, for example, leaves, stems, roots, and seeds. Alternatively, the plant cells may be cultured plant cells.

[0187] In the manufacturing method according to this embodiment, the step of obtaining MHC protein may include the step of introducing the expression system described below into plant cells.

[0188] <MHC Protein Expression System> The expression system comprises a first nucleic acid fragment including a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an MHC protein expression cassette linked between the LIR and the SIR, and a second nucleic acid fragment including an expression cassette for a Rep / RepA protein derived from a geminivirus. The MHC protein expression cassette includes, in this order, a promoter, a nucleic acid fragment encoding the MHC protein, and two or more linked terminators.

[0189] By using multiple expression systems that include expression cassettes for different MHC proteins, it is possible to co-express different MHC proteins in plant cells.

[0190] When the expression system is introduced into plant cells, the Rep / RepA protein, the replication initiation protein of the geminivirus, is expressed from the second nucleic acid fragment. Then, the expression cassette of the MHC protein linked between the LIR and SIR on the first nucleic acid fragment is replicated at a high copy number by the geminivirus's rolling circle DNA replication mechanism. Subsequently, the MHC protein is expressed at a high expression level from the expression cassette of the MHC protein that has been replicated at a high copy number.

[0191] The expression system can achieve a very high level of MHC protein expression because the expression cassette of the MHC protein includes a terminator consisting of two or more linked MHC protein expression cassettes.

[0192] A terminator is a base sequence that terminates the transcription of DNA into mRNA. Terminators are not particularly limited, but examples include the terminator derived from the Arabidopsis thaliana heat shock protein 18.2 gene, the terminator of the tobacco extensin gene, the 35S terminator of cauliflower mosaic virus (CaMV), and the NOS terminator of CaMV. In the expression system described above, two or more linked terminators may each have the same base sequence or may each have different base sequences.

[0193] Sequence ID 44 shows the nucleotide sequence of the terminator derived from the Arabidopsis thaliana heat shock protein 18.2 gene. Sequence ID 45 shows the nucleotide sequence of the terminator of the tobacco extensin gene.

[0194] The base sequences of these terminators may each have mutations in the base sequences of SEQ ID NOs. 44-45, or may have deletions in parts of the base sequences of SEQ ID NOs. 44-45, as long as they have the function of terminating the transcription of DNA into mRNA.

[0195] If the base sequences of the terminators have mutations or deletions relative to the base sequences of SEQ ID NOs. 44-45, each base sequence preferably has, for example, 70% or more sequence identity with the base sequences of SEQ ID NOs. 44-45, more preferably 80% or more, even more preferably 90% or more, and particularly preferably 95% or more.

[0196] Here, the sequence identity of the target nucleotide sequence with respect to the reference nucleotide sequence can be determined, for example, as follows. First, the reference nucleotide sequence and the target nucleotide sequence are aligned. Here, gaps may be included in each nucleotide sequence to maximize sequence identity. Next, the number of matching nucleotides in the reference nucleotide sequence and the target nucleotide sequence is calculated, and the sequence identity can be determined according to the following formula (2). Sequence identity (%) = Number of matching nucleotides / Total number of nucleotides in the target nucleotide sequence × 100 (2) The value of sequence identity of a nucleotide sequence can be obtained by calculation based on alignment obtained by known homology search software such as BLASTN.

[0197] The expression system preferably includes an MHC protein expression cassette containing two terminators.

[0198] In the expression system described above, it is preferable that at least one of the terminators included in the MHC protein expression cassette is a terminator derived from the Arabidopsis thaliana heat shock protein 18.2 gene.

[0199] The expression system described above can further enhance the expression level of MHC protein by including a terminator derived from the Arabidopsis thaliana heat shock protein 18.2 gene as a terminator.

[0200] In the expression system described above, the first nucleic acid fragment comprises, in this order, a promoter, a nucleic acid fragment encoding an MHC protein, and two or more linked terminators.

[0201] The promoter can be any promoter that exhibits transcriptional activity of the DNA linked downstream in the host plant cell, and is not particularly limited. Specifically, examples include the 35S promoter of cauliflower mosaic virus (CaMV), the ubiquitin promoter, and the cassava vein mosaic virus (CsVMV) promoter.

[0202] Examples of MHC proteins include the recombinant MHC proteins mentioned above.

[0203] The terminator consists of a nucleotide sequence involved in the specific termination of RNA transcription by RNA polymerase. The terminator may be a combination of two or more ligated terminators, or two or more terminators may be used in combination. As will be described later in the examples, the expression system of this embodiment can achieve a high expression level of MHC protein by having two or more ligated terminators.

[0204] Furthermore, if at least one of the terminators is derived from the Arabidopsis thaliana heat shock protein 18.2 gene, it is possible to further increase the expression level of the MHC protein.

[0205] The first nucleic acid fragment may also have, for example, a 5'-untranslated region (UTR), a polyadenylation signal, etc., in addition to a promoter, a nucleic acid fragment encoding an MHC protein, and a terminator.

[0206] The presence of a 5'-UTR may further increase the expression efficiency of MHC proteins. Examples of 5'-UTRs include those of the tobacco mosaic virus, the Arabidopsis thaliana alcohol dehydrogenase gene, the Arabidopsis thaliana elongation factor 1α-A3 gene, and the rice alcohol dehydrogenase gene.

[0207] In the expression system described above, the second nucleic acid fragment includes an expression cassette for the Rep / RepA protein derived from geminivirus. The Rep / RepA protein expression cassette is not particularly limited as long as it can express the Rep / RepA protein in host plant cells, and may have a promoter, a nucleic acid fragment encoding the Rep / RepA protein, a terminator, a 5'-UTR, a polyadenylation signal, etc. Here, as the promoter for the Rep / RepA protein, in addition to those described above that can be used in the first nucleic acid fragment, for example, geminivirus-derived LIR also has promoter activity and may be used.

[0208] In this specification, "expression system" means a system that can express MHC protein by introducing a combination of a first nucleic acid fragment and a second nucleic acid fragment into plant cells. The expression system may consist of one nucleic acid fragment or a combination of two or more nucleic acid fragments, as long as the effects of the present invention are obtained. Here, the nucleic acid fragment may be a vector.

[0209] In other words, in the expression system, the first nucleic acid fragment and the second nucleic acid fragment may exist separately as independent nucleic acid fragments, or they may be linked together to form a single nucleic acid fragment.

[0210] Furthermore, when the first nucleic acid fragment and the second nucleic acid fragment are linked, the order of linkage is not particularly limited, and the first nucleic acid fragment may be located on the 5' side, or the second nucleic acid fragment may be located on the 5' side.

[0211] In the expression system described above, LIR, SIR, and Rep / RepA are derived from geminiviruses. Geminiviruses are not particularly limited as long as they have a rolling circle DNA replication mechanism, and examples include bean yellow dwarf virus (BeYDV), tomato golden mosaic virus (TGMV), African cassava mosaic virus (ACMV), and rose leaf curl virus (RLCV) of the Mastrevirus genus in the family Geminiviridae.

[0212] Sequence ID 46 shows the nucleotide sequence of the LIR derived from geminivirus. Sequence ID 47 shows the nucleotide sequence of the SIR derived from geminivirus. Sequence ID 48 shows the nucleotide sequence of the Rep / RepA protein derived from geminivirus (open reading frames C1 and C2 that encode the Rep / RepA protein, which is the replication initiation protein of BeYDV).

[0213] The nucleotide sequences of LIR, SIR, and Rep / RepA protein may each have mutations in the nucleotide sequences of SEQ ID NOs. 46-48, and may also have deletions in parts of the nucleotide sequences of SEQ ID NOs. 46-48, as long as the Rep / RepA protein encoded by these sequences has the function of replicating the expression cassette of the MHC protein linked between LIR and SIR at a high copy number.

[0214] If the nucleotide sequences of LIR, SIR, and Rep / RepA protein each have mutations or deletions compared to the nucleotide sequences of SEQ ID NOs. 46-48, each nucleotide sequence preferably has, for example, 70% or more sequence identity with respect to the nucleotide sequences of SEQ ID NOs. 46-48, more preferably 80% or more, even more preferably 90% or more, and particularly preferably 95% or more.

[0215] Here, the sequence identity of the target base sequence with respect to the reference base sequence can be determined, for example, according to formula (2) described above.

[0216] The expression system may further comprise a third nucleic acid fragment containing an expression cassette for a gene silencing inhibitor. This allows for a further increase in the expression level of the MHC protein. Examples of gene silencing inhibitors include gene silencing inhibitor P19 derived from tomato bushy stunt virus and gene silencing inhibitor 16K derived from tobacco rattle virus.

[0217] The third nucleic acid fragment may exist separately as a nucleic acid fragment independent of the first and second nucleic acid fragments described above, or it may be linked with the first or second nucleic acid fragment in any order, or the first, second, and third nucleic acid fragments may be linked in any order. That is, the first, second, and third nucleic acid fragments may be contained in a single vector.

[0218] In the expression system described above, the first nucleic acid fragment, the second nucleic acid fragment, and the third nucleic acid fragment are contained in a single vector, the vector further comprising the RB and LB of T-DNA, and the first nucleic acid fragment, the second nucleic acid fragment, and the third nucleic acid fragment may be located between the RB and LB of T-DNA.

[0219] T-DNA is a specific region found in Ti plasmids and Ri plasmids of pathogenic strains of Agrobacterium, the bacterium that causes crown gall, a tumor of dicotyledonous plants. When Agrobacterium containing T-DNA is introduced into the coexistence of plant cells, it transfers nucleic acid fragments located between the RB and LB into the host plant cells.

[0220] Therefore, by introducing a vector containing the first nucleic acid fragment, the second nucleic acid fragment, and the third nucleic acid fragment between RB and LB into Agrobacterium, and then introducing the Agrobacterium into a host plant, the first nucleic acid fragment, the second nucleic acid fragment, and the third nucleic acid fragment can be easily introduced into the host plant cells.

[0221] A vector in which a first nucleic acid fragment, a second nucleic acid fragment, and a third nucleic acid fragment exist between RB and LB is preferably a vector that can be used in the binary vector method.

[0222] The binary vector method is a gene transfer method for plants that utilizes a vir helper Ti plasmid, which is a Ti plasmid from which the original T-DNA has been removed, and a small shuttle vector containing artificial T-DNA. Preferably, the shuttle vector can be maintained in both E. coli and Agrobacterium.

[0223] The vir helper Ti plasmid lacks the original T-DNA and therefore cannot form crown galls in plants. However, the vir helper Ti plasmid possesses the vir region necessary for introducing T-DNA into host plant cells.

[0224] Therefore, by introducing a T-DNA containing the desired nucleic acid fragment into an Agrobacterium possessing the vir helper Ti plasmid, and then introducing the Agrobacterium into a host plant, the desired nucleic acid fragment can be easily introduced into the host plant cells.

[0225] In other words, a vector having a first nucleic acid fragment, a second nucleic acid fragment, and a third nucleic acid fragment between RB and LB may be a shuttle vector that has a replication origin for E. coli and a replication origin for Agrobacterium, and can be maintained in both E. coli and Agrobacterium.

[0226] Sequence ID 49 shows the base sequence of the RB of T-DNA. Sequence ID 50 shows the base sequence of the LB of T-DNA.

[0227] The base sequences of RB and LB may each have mutations in the base sequences of SEQ ID NOs. 49-50, or may have deletions in parts of the base sequences of SEQ ID NOs. 49-50, as long as they have the function of transferring the nucleic acid fragment present between RB and LB into the host plant cell.

[0228] If the base sequences of RB and LB each have mutations or deletions with respect to the base sequences of SEQ ID NOs. 49 to 50, then it is preferable that each base sequence has, for example, 70% or more sequence identity with respect to the base sequences of SEQ ID NOs. 49 to 50, more preferably 80% or more sequence identity, even more preferably 90% or more sequence identity, and particularly preferably 95% or more sequence identity.

[0229] Here, the sequence identity of the target base sequence with respect to the reference base sequence can be determined, for example, according to formula (2) described above.

[0230] The expression cassette of the MHC protein preferably has a sequence encoding a signal sequence, a sequence encoding a zipper-like helix domain, a sequence encoding an endoplasmic reticulum retention signal, and a sequence encoding an apoplast retention signal. The amino acid sequences of the signal sequence, zipper-like helix domain, endoplasmic reticulum retention signal, and apoplast retention signal are as described above.

[0231] Proteins containing a signal sequence are more likely to localize to organelles. The location of the sequence encoding the signal sequence is preferably at the 5' end of the open reading frame of the MHC protein.

[0232] Proteins having a zipper-like helix domain are more likely to form dimers. The location of the sequence encoding the zipper-like helix domain is preferably at the 3' end of the open reading frame of the MHC protein.

[0233] Proteins possessing an endoplasmic reticulum (ER) retention signal are more likely to localize to the ER and have higher expression levels. The location of the sequence encoding the ER retention signal is preferably at the 3' end of the open reading frame of the MHC protein.

[0234] Proteins possessing an apoplast retention signal are more likely to localize to the apoplast and have higher expression levels. The location of the sequence encoding the apoplast retention signal is preferably at the 3' end of the open reading frame of the TCR protein.

[0235] <Introduction of the MHC protein expression system> The expression system can be introduced into plant cells using Agrobacterium bacteria (hereinafter sometimes simply referred to as Agrobacterium).

[0236] Agrobacterium is not particularly limited as long as it can infect a target plant and cause transformation of its plant cells, and examples include Agrobacterium tumefaciens, Agrobacterium vitis, Agrobacterium rhizogenes, and Agrobacterium radiobacter. More specifically, examples include, but are not limited to, Agrobacterium tumefaciens strains GV3101, LBA4404, C58, EHA101, and A208, Agrobacterium vitis strains F2 / 5 and S4, and Agrobacterium rhizogenes strains A4 and LBA9402, as well as their derivatives.

[0237] The method for producing MHC protein according to this embodiment makes it possible to obtain MHC protein having a normal three-dimensional structure and sufficient function. In this production method, since recombinant MHC protein is expressed in plant cells, the cost is low and there is no risk of contamination with pathogens that can infect humans. In this embodiment, when the α-chain and β-chain of MHC protein are co-expressed in plant cells, protein folding proceeds normally, and MHC protein with normal three-dimensional structure and function can be obtained.

[0238] (MHC protein expression system) The expression system according to this embodiment includes a first nucleic acid fragment comprising a geminivirus-derived LIR, a geminivirus-derived SIR, and an MHC protein expression cassette linked between the LIR and the SIR, wherein the expression cassette comprises, in this order, a promoter, a multicloning site, and two or more linked terminators.

[0239] The expression system according to this embodiment is the same as the expression system described in the manufacturing method according to the above embodiment.

[0240] (Agrobacterium bacteria for expressing MHC protein) The Agrobacterium bacteria according to this embodiment are Agrobacterium bacteria transformed with the expression system according to the embodiment, and are suitably used in the method for producing MHC protein according to the embodiment.

[0241] The Agrobacterium bacteria according to this embodiment are not particularly limited as long as they can infect target plant cells and cause transformation of those plant cells, and examples include Agrobacterium tumefaciens, Agrobacterium vitis, Agrobacterium rhizogenes, and Agrobacterium radiobacter. More specifically, examples include, but are not limited to, Agrobacterium tumefaciens strains GV3101, LBA4404, C58, EHA101, and A208, Agrobacterium vitis strains F2 / 5 and S4, and Agrobacterium rhizogenes strains A4 and LBA9402, as well as their derivatives.

[0242] (Plant cells for expressing MHC protein) The plant cells according to this embodiment are plant cells into which the expression system according to the above embodiment has been introduced.

[0243] The plant species and type of plant cells from which the plant cells are derived are the same as those described above in the method for producing MHC proteins. The method for introducing the expression system into plant cells is the same as the method described above in the method for producing MHC proteins according to the embodiment.

[0244] (Method for producing transformed plants to express MHC protein) The method for producing transformed plants according to this embodiment includes infecting plant cells with the Agrobacterium bacterium according to the above embodiment.

[0245] (MHC Protein Expression Vector) The expression vector according to this embodiment comprises a first nucleic acid fragment including a geminivirus-derived LIR, a geminivirus-derived SIR, and an MHC protein expression cassette linked between the LIR and the SIR, wherein the expression cassette includes a promoter, a multicloning site, and two or more linked terminators in that order. The composition of the first nucleic acid fragment is as described above. The MHC protein expressed by the MHC protein expression cassette is the MHC protein described above in the above embodiment.

[0246] The expression vector of this embodiment can be suitably used to construct the expression system described above. As will be described later in the examples, the expression vector of this embodiment can achieve a very high expression level of MHC protein by introducing a gene fragment encoding an MHC protein into the multi-cloning site of the expression cassette. This is because it contains a terminator in which two or more expression cassettes are linked together.

[0247] Furthermore, in the expression vector of this embodiment, it is preferable that at least one of the terminators included in the expression cassette is a terminator derived from the Arabidopsis thaliana heat shock protein 18.2 gene.

[0248] In the expression vector of this embodiment, the LIR, SIR, promoter, and terminator are the same as those described above. That is, in the expression vector of this embodiment, the geminivirus may be bean malformation virus (BeYDV).

[0249] In this specification, a multi-cloning site is a region in which one or more nucleotide sequences recognized by restriction enzymes are arranged. That is, in the multi-cloning site of the expression vector of this embodiment, there may be one restriction enzyme site or multiple restriction enzyme sites. Since the vector of this embodiment has a multi-cloning site, nucleic acid fragments encoding MHC proteins can be easily cloned.

[0250] The vector of this embodiment may further include a second nucleic acid fragment containing an expression cassette for the Rep / RepA protein derived from geminivirus. The Rep / RepA protein is the same as described above.

[0251] The vector of this embodiment may further include a third nucleic acid fragment containing an expression cassette for a gene silencing inhibitor. The gene silencing inhibitor may be the same as described above, for example, gene silencing inhibitor P19 derived from tomato bushy stunt virus.

[0252] The vector of this embodiment further comprises RB and LB of T-DNA, and a first nucleic acid fragment, a second nucleic acid fragment, and a third nucleic acid fragment may be present between the RB and LB of T-DNA. The RB and LB are the same as those described above. Furthermore, it is preferable that the vector of this embodiment is a vector that can be used in the binary vector method.

[0253] (Method for producing TCR protein) The method for producing TCR protein according to this embodiment includes the step of expressing TCR protein in plant cells to obtain the TCR protein.

[0254] In this specification, TCR protein refers to the protein that constitutes the T cell receptor. TCR protein is a protein expressed in the cells of vertebrates. The vertebrates from which TCR protein originates include those mentioned above.

[0255] A functional TCR protein can be obtained using the method for producing TCR protein according to this embodiment. Furthermore, a dimeric TCR protein can be obtained by expressing multiple subunits of the TCR protein. As will be described later in the examples, the obtained dimeric TCR protein can bind to MHC protein. As will be described later in the examples, the TCR protein expressed in plant cells has a sugar chain attached to it, and this sugar chain attachment was necessary for the binding of the MHC protein to the TCR protein. It is known that the type of sugar chain attached differs between plant cells and vertebrates. Surprisingly, the inventors found that the TCR protein expressed in plant cells can bind to MHC protein. In this embodiment, it is preferable that the TCR protein has an apoplast retention signal. In this case, the expression level of the TCR protein can be significantly increased. In this embodiment, it is preferable that the TCR protein has a zipper-like helix domain. In this case, since the subunits of the TCR protein are strongly bound by the zipper-like helix domain, they can be purified and separated in a dimerized state. Furthermore, the resulting dimerized TCR protein can bind to the dimerized MHC protein.

[0256] Natural TCR proteins are composed of α-chains and β-chains, or γ-chains and δ-chains. In this embodiment, the recombinant TCR protein expressed in plant cells has an amino acid sequence derived from α-chains, β-chains, γ-chains, and δ-chains.

[0257] In this embodiment, the TCR protein expressed in plant cells is a recombinant TCR protein. Hereinafter, the recombinant TCR protein may simply be referred to as the TCR protein.

[0258] The recombinant TCR protein may be a protein selected from (A) to (F) below.

[0259] (A) Proteins containing the amino acid sequence of the TCR protein; (B) Proteins containing a portion of the amino acid sequence of the TCR protein; (C) Proteins containing amino acid sequences in which one or more amino acids are deleted, inserted, substituted and / or added to the amino acid sequence of the TCR protein; (D) Proteins containing amino acid sequences in which one or more amino acids are deleted, inserted, substituted and / or added to a portion of the amino acid sequence of the TCR protein; (E) Proteins containing amino acid sequences having 90% or more sequence identity with the amino acid sequence of the TCR protein; (F) Proteins containing amino acid sequences having 90% or more sequence identity with a portion of the amino acid sequence of the TCR protein

[0260] The TCR protein derived from the recombinant TCR proteins (A) to (F) may be a natural TCR protein or an artificially modified TCR protein amino acid sequence. The TCR protein derived from the recombinant TCR proteins (A) to (F) may be an α-chain, β-chain, γ-chain, or δ-chain. The TCR protein derived from the recombinant TCR proteins (A) to (F) may sometimes be simply referred to as the TCR protein.

[0261] Protein (A) may contain amino acid sequences other than the amino acid sequence of the TCR protein. Protein (A) may also be a protein in which an amino acid sequence has been added to the amino acid sequence of the TCR protein. If protein (A) is a protein in which an amino acid has been added to the amino acid sequence of the TCR protein, the addition of amino acids may occur at either the N-terminus or the C-terminus of the amino acid sequence of the TCR protein, or at both the N-terminus and the C-terminus.

[0262] The number of amino acids constituting the protein (A) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0263] If protein (A) is a protein in which one or more amino acid residues are added to the amino acid sequence that constitutes the TCR protein, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0264] The recombinant TCR protein of (B) may be a protein containing a portion of the amino acid sequence of the TCR protein (hereinafter sometimes referred to as a partial sequence). The partial sequence is not particularly limited, but examples include the amino acid sequence that constitutes the extracellular domain of the TCR protein. The amino acid sequence that constitutes the extracellular domain is preferably the amino acid sequence obtained by excluding the signal sequence, transmembrane domain and intracellular domain from the full length of the TCR protein amino acid sequence.

[0265] Protein (B) may contain amino acid sequences other than the partial sequence. Protein (B) may also be a protein in which an amino acid sequence has been added to the partial sequence. If protein (B) is a protein in which an amino acid has been added to the partial sequence, the addition of amino acids may occur at either the N-terminus or the C-terminus of the partial sequence, or at both the N-terminus and the C-terminus.

[0266] The number of amino acids that make up the protein of (B) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0267] If protein (B) is a protein in which one or more amino acid residues are added to a partial sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0268] The recombinant TCR protein in (C) may be a protein containing an amino acid sequence (hereinafter sometimes simply referred to as a mutant sequence) in which one or more amino acids are deleted, inserted, substituted and / or added in the amino acid sequence of the TCR protein. Alternatively, the protein in (C) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant sequence.

[0269] In this specification, "multiple" is not particularly limited, but examples include 2 to 50, 2 to 40, 2 to 30, 2 to 27, 2 to 23, 2 to 19, 2 to 16, 2 to 13, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 2 to 3, or 2.

[0270] The number of amino acids constituting the protein (C) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0271] If protein (C) is a protein in which one or more amino acid residues are added to a mutant sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0272] The recombinant TCR protein of (D) may be a protein containing an amino acid sequence (hereinafter also referred to as the "mutant subsequence") in which one or more amino acids are deleted, inserted, substituted and / or added in the subsequence described above in (B). Alternatively, the protein of (D) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the mutant subsequence.

[0273] The number of amino acids constituting the protein (D) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0274] If protein (D) is a protein in which one or more amino acid residues are added to a mutant partial sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0275] The recombinant TCR protein of (E) may also be a protein containing an amino acid sequence (hereinafter sometimes simply referred to as a homologous sequence) that has 90% or more sequence identity with the amino acid sequence of the TCR protein. The sequence identity of the homologous sequence may be 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more. Alternatively, the protein of (E) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the homologous sequence.

[0276] In this specification, amino acid sequence identity is a value that indicates the percentage of the target amino acid sequence that matches a reference amino acid sequence. The sequence identity of the target amino acid sequence with respect to the reference amino acid sequence can be determined, for example, as follows: First, the reference amino acid sequence and the target amino acid sequence are aligned. Here, gaps may be included in each amino acid sequence to maximize sequence identity. Next, the number of matching amino acids in the reference amino acid sequence and the target amino acid sequence is calculated, and the sequence identity can be determined according to the following formula (1): Sequence identity (%) = Number of matching amino acids / Total number of amino acids in the target amino acid sequence × 100 …(1) The value of amino acid sequence identity can be obtained by calculation based on the alignment obtained by homology search software such as BLASTP.

[0277] The number of amino acids that make up the protein (E) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0278] If protein (E) is a protein in which one or more amino acid residues are added to a mutant partial amino acid sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0279] The recombinant TCR protein of (F) may be a protein containing an amino acid sequence (hereinafter sometimes simply referred to as a homologous partial sequence) that has 90% or more sequence identity with the partial sequence described above in (B). The sequence identity of the homologous partial sequence may be 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more. Alternatively, the protein of (F) may be a protein in which one or more amino acid residues are added to either the N-terminus and / or C-terminus of the homologous partial sequence.

[0280] The number of amino acids that make up the protein (F) is not particularly limited; for example, it may be 2000 residues or less, 1000 residues or less, 800 residues or less, 600 residues or less, or 400 residues or less.

[0281] If protein (F) is a protein in which one or more amino acid residues are added to a mutant partial amino acid sequence, the number of amino acids added may be 1 to 2000 residues, 1 to 1500 residues, 1 to 1000 residues, 1 to 800 residues, 1 to 400 residues, 1 to 200 residues, or 1 to 100 residues.

[0282] In proteins (B), (D), and (F), the amino acid sequence of the TCR protein is not particularly limited as long as the effects of the present invention are achieved, and may be a sequence derived from any of the α, β, γ, or δ chains. Examples include the extracellular domain of the α chain represented by SEQ ID NO: 54 and the extracellular domain of the β chain represented by SEQ ID NO: 56.

[0283] Proteins (A) to (F) may be fusion proteins obtained by fusing a TCR protein, a protein consisting of a partial sequence, a protein consisting of a mutant sequence, a protein consisting of a mutant partial sequence, a protein consisting of a homologous sequence, or a protein consisting of a homologous partial sequence with another protein or peptide. Examples of other proteins or peptides include signal sequences; endoplasmic reticulum retention signals; peptides constituting zipper-like helix domains; antigen peptides; marker enzyme proteins such as alkaline phosphatase or partial peptides thereof; fluorescent proteins such as GFP or partial peptides thereof; tag peptides such as His tags and FLAG tags.

[0284] Examples of peptides include the peptides that constitute the zipper-like helix domain, signal sequences, endoplasmic reticulum retention signals, apoplast retention signals, antigen peptides, and linkers mentioned above.

[0285] The TCR protein preferably has a signal sequence at its N-terminus. The TCR protein preferably has a zipper-like helix domain at its C-terminus. The TCR protein preferably has an apoplast retention signal at its C-terminus.

[0286] In this embodiment, the TCR protein may be obtained by co-expressing proteins derived from the α-chain and β-chain of the TCR protein, and proteins derived from the γ-chain and δ-chain of the TCR protein in plant cells. In this case, protein folding proceeds normally, and a TCR protein with normal three-dimensional structure and function can be obtained. Co-expression can yield dimers of proteins derived from the α-chain and β-chain, and dimers of proteins derived from the γ-chain and δ-chain.

[0287] The plant species and types of plant cells that express the TCR protein are those mentioned above.

[0288] In the manufacturing method according to this embodiment, the step of obtaining the TCR protein may include the step of introducing the expression system described below into plant cells.

[0289] <TCR Protein Expression System> The expression system comprises a first nucleic acid fragment including a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette for the TCR protein linked between the LIR and the SIR, and a second nucleic acid fragment including an expression cassette for the Rep / RepA protein derived from a geminivirus. The TCR protein expression cassette includes, in this order, a promoter, a nucleic acid fragment encoding the TCR protein, and two or more linked terminators.

[0290] The components of the expression system are similar to those described above in the method for producing MHC proteins.

[0291] The TCR protein expression cassette preferably has a sequence encoding a signal sequence, a sequence encoding a zipper-like helix domain, and a sequence encoding an apoplast retention signal. The amino acid sequences of the signal sequence, zipper-like helix domain, and apoplast retention signal are as described above.

[0292] Proteins containing a signal sequence are more likely to localize to organelles. The location of the sequence encoding the signal sequence is preferably at the 5' end of the open reading frame of the TCR protein.

[0293] Proteins having a zipper-like helix domain are more likely to form dimers. The location of the sequence encoding the zipper-like helix domain is preferably at the 3' end of the open reading frame of the TCR protein.

[0294] Proteins possessing an apoplast retention signal are more likely to localize to the apoplast and have higher expression levels. The location of the sequence encoding the apoplast retention signal is preferably at the 3' end of the open reading frame of the TCR protein.

[0295] <Introduction of the TCR protein expression system> The expression system can be introduced into plant cells using Agrobacterium bacteria (hereinafter sometimes simply referred to as Agrobacterium). Examples of Agrobacterium include those mentioned above.

[0296] (TCR protein expression system) The expression system according to this embodiment includes a first nucleic acid fragment comprising a geminivirus-derived LIR, a geminivirus-derived SIR, and an expression cassette for a TCR protein linked between the LIR and the SIR, wherein the expression cassette comprises, in this order, a promoter, a multi-cloning site, and two or more linked terminators.

[0297] The expression system according to this embodiment is the same as the expression system described in the method for producing the TCR protein according to the above embodiment.

[0298] (Agrobacterium bacteria for expressing TCR protein) The Agrobacterium bacteria according to this embodiment are Agrobacterium bacteria transformed with the expression system according to the embodiment, and are suitably used in the method for producing TCR protein according to the embodiment. Examples of Agrobacterium include those described above.

[0299] (Plant cells for expressing TCR protein) The plant cells according to this embodiment are plant cells into which the expression system according to the above embodiment has been introduced.

[0300] The plant species and type of plant cells from which the plant cells are derived are the same as those described above in the method for producing MHC proteins. The method for introducing the expression system into plant cells is the same as the method described above in the method for producing MHC proteins according to the embodiment.

[0301] (Method for producing transformed plants to express TCR protein) The method for producing transformed plants according to this embodiment includes infecting plant cells with Agrobacterium bacteria according to the above embodiment.

[0302] (TCR Protein Expression Vector) The expression vector according to this embodiment comprises a first nucleic acid fragment including a geminivirus-derived LIR, a geminivirus-derived SIR, and an expression cassette for a TCR protein linked between the LIR and the SIR, wherein the expression cassette includes a promoter, a multi-cloning site, and two or more linked terminators in that order. The composition of the first nucleic acid fragment is as described above. The TCR protein expressed by the TCR protein expression cassette is the TCR protein described above in the above embodiment.

[0303] (Method for evaluating binding strength) The method for evaluating binding strength according to this embodiment is a method for evaluating the binding strength between an MHC protein fused with an antigen peptide and a TCR protein. The method according to this embodiment includes a step (a1) of obtaining an MHC protein fused with an antigen peptide by the method for producing an MHC protein according to the above embodiment, and a step (b1) of evaluating the binding strength between the MHC protein fused with the antigen peptide obtained in step (a1) and a TCR protein.

[0304] Step (a1) In step (a1), an MHC protein fused with an antigen peptide is obtained by the method for producing an MHC protein according to the embodiment described above. As described above, the MHC protein fused with an antigen peptide obtained by the method for producing an MHC protein according to the embodiment described above can bind to a TCR protein. Examples of MHC proteins fused with an antigen peptide are those described above. The obtained MHC protein is preferably a dimer.

[0305] Step (b1): In step (b1), the binding affinity between the MHC protein, to which the antigen peptide obtained in step (a1) is fused, and the TCR protein is evaluated.

[0306] The type of TCR protein is not particularly limited as long as the effects of the present invention are achieved, and may be a protein containing the amino acid sequence of a natural TCR protein, or a protein containing the amino acid sequence of an artificially modified TCR protein. The TCR protein may be an α chain, β chain, γ chain, or δ chain. The amino acid sequence constituting the TCR protein may be the amino acid sequence of the TCR protein exemplified in the method for producing the TCR protein according to the above embodiment.

[0307] The form of the TCR protein in step (b1) is not particularly limited as long as the effects of the present invention are achieved, and may be isolated from cells expressing the TCR protein or expressed in cells.

[0308] The TCR protein may be obtained by the method for producing the TCR protein according to the above embodiment. The method according to this embodiment may include a step of obtaining the TCR protein by the method for producing the TCR protein according to the above embodiment.

[0309] The MHC protein used for evaluating binding strength is preferably a dimer. The combinations of subunits constituting the dimer are as described above. The TCR protein used for evaluating binding strength is preferably a dimer. The combinations of subunits constituting the dimer are as described above.

[0310] The method for evaluating the binding affinity between MHC protein and TCR protein is not particularly limited as long as the effects of the present invention are achieved, and may, for example, be a method using an interaction detector with a Biacore® system.

[0311] (Method for evaluating binding affinity) The method for evaluating binding affinity according to this embodiment is a method for evaluating the binding affinity between an MHC protein containing the sequence of an antigen peptide and a TCR protein. The method according to this embodiment includes a step (a2) of obtaining a TCR protein by the method for producing a TCR protein according to the above embodiment, and a step (b2) of evaluating the binding affinity between an MHC protein containing the sequence of an antigen peptide and the TCR protein obtained in step (a2).

[0312] Step (a2) In step (a2), a TCR protein is obtained by the method for producing a TCR protein according to the embodiment described above. As described above, the TCR protein obtained by the method for producing a TCR protein according to the embodiment can bind to an MHC protein containing the sequence of an antigen peptide.

[0313] Step (b2): In step (b2), the binding affinity between the MHC protein containing the antigen peptide sequence and the TCR protein obtained in step (a2) is evaluated.

[0314] The type of MHC protein containing the antigen peptide sequence is not particularly limited as long as the effects of the present invention are achieved, and may be a protein containing the amino acid sequence of a natural MHC protein, or a protein containing the amino acid sequence of an artificially modified MHC protein. The amino acid sequence constituting the MHC protein may be the amino acid sequence of the MHC protein exemplified in the method for producing the TCR protein according to the above embodiment.

[0315] The form of the MHC protein in step (b1) is not particularly limited as long as the effects of the present invention are achieved, and may be isolated from cells expressing the MHC protein or expressed in cells.

[0316] The MHC protein may be obtained by the method for producing the MHC protein according to the above embodiment. The method according to this embodiment may include a step of obtaining the MHC protein by the method for producing the MHC protein according to the above embodiment.

[0317] The MHC protein used for evaluating binding strength is preferably a dimer. The combinations of subunits constituting the dimer are as described above. The TCR protein used for evaluating binding strength is preferably a dimer. The combinations of subunits constituting the dimer are as described above.

[0318] The method for evaluating the binding affinity between MHC protein and TCR protein is not particularly limited as long as the effects of the present invention are achieved, and may, for example, be a method using an interaction detector with a Biacore® system.

[0319] The present invention will now be described in more detail with reference to examples, but the present invention is not limited to the following examples.

[0320] (Materials and Methods) <Preparation of Expression Vectors> The following proteins were expressed: DR#1 to DR#9, DQ#1 to DQ#2, DP#1 to DP#2, and TCR#1 to #2. The sequences of the domains of each fusion protein are shown below.

[0321] SSExt: MGKMASLFATLLVVLVSLSLASESSA (SEQ ID NO: 32) cSP: MPPQKENHRTLNKMKTNLFLFLIFSLLLSLSSA (SEQ ID NO: 33) Endoplasmic reticulum retention signal: KDEL (SEQ ID NO: 34, sometimes referred to as the KD sequence) Apoplast retention signal: SPSPSPSPSPSPSPSPSPSP (SEQ ID NO: 51, sometimes referred to as the SP10 sequence) "SSExt" refers to the signal sequence of extensin in Citrus benthamiana. "cSP" refers to the signal sequence of chitinase in Arabidopsis thaliana. Proteins containing apoplast retention signals are secreted into the apoplast and their expression levels increase.

[0322] Zipper-like helix domains Zipper A: AQLEKELQALEKENAQLEWELQALEKELAQ (SEQ ID NO: 35) Zipper B: AQLKKKLQALKKKNAQLKWKLQALKKKLAQ (SEQ ID NO: 36) The inclusion of zipper-like helix domains facilitates the formation of dimers in the expressed protein.

[0323] TEV cleavage site: ENLYFQG (SEQ ID NO: 37) HRV3C protease cleavage site: LEVLFQGP (SEQ ID NO: 38) Thrombin cleavage site: LVPRGS (SEQ ID NO: 39)

[0324] V5 tag: GKPIPNPLLGLDST (SEQ ID NO: 40) Histidine tag: HHHHHHHHHH (SEQ ID NO: 41) Avi tag: GLNDIFEAQKIEWHE (SEQ ID NO: 42) Streptavidin-binding tag (Strep-tag II): WSHPQFEK (SEQ ID NO: 43) AGIA tag: EEAAGIARP (SEQ ID NO: 52) The AGIA tag is a peptide tag derived from the C-terminal region (404-412) of the dopamine receptor DRD1, and has high specificity and affinity with anti-AGIA antibodies.

[0325] The DR proteins, DR#1 to DR#9, all consist of a fusion protein (DRA#1) containing the α-chain water-soluble extracellular domain of the HLA-DR protein, and fusion proteins (DRB#1 to DRB#9) containing the β-chain water-soluble extracellular domain of the HLA-DR protein. The sequence numbers of the amino acid sequences of these fusion proteins are shown in Table 2. The allele names of the α-chain and β-chain from which the water-soluble extracellular domains are derived, and the sequence numbers of the amino acid sequences of the water-soluble extracellular domains in DRA#1 and DRB#1 to DRB#9 are also shown in Table 2.

[0326]

[0327] A schematic diagram of the domain structure of the DR protein is shown in Figure 1A. The expression vector for the DR protein is shown in Figure 1B. DRA#1 has the following sequences: SSExt, α-chain water-soluble extracellular domain, TEV cleavage site, Zipper A, V5 tag, HRV3C protease cleavage site, histidine tag, and endoplasmic reticulum retention signal ("KD" in the figure).

[0328] DRB#1 to DRB#9 contain the following sequences: SSExt, antigen peptide, β-chain water-soluble extracellular domain, TEV cleavage site, Zipper B, thrombin cleavage site, Avi tag, and endoplasmic reticulum retention signal ("KD" in the figure). The amino acid sequences of the antigen peptides are shown in Table 3.

[0329]

[0330] The DQ proteins, DQ#1 to DQ#2, consist of a fusion protein (DQA#1) containing the α-chain water-soluble extracellular domain of the HLA-DQ protein, and fusion proteins (DQB#1, DQB#2) containing the β-chain water-soluble extracellular domain of the HLA-DQ protein. The sequence numbers of the amino acid sequences of these fusion proteins are shown in Table 4. The allele names of the α-chain and β-chain from which the water-soluble extracellular domains are derived in DQA#1, DQB#1 to DQB#2, and the sequence numbers of the amino acid sequences of the water-soluble extracellular domains are also shown in Table 4.

[0331]

[0332] The expression vector for the DQ protein is shown in Figure 1C. DQA#1 has the following sequences: cSP, α-chain water-soluble extracellular domain, TEV cleavage site, Zipper A, thrombin cleavage site, histidine tag, and endoplasmic reticulum retention signal ("KD" in the figure).

[0333] Both DQB#1 and DQB#2 contain the following sequences: cSP, antigen peptide, β-chain water-soluble extracellular domain, TEV cleavage site, Zipper B, Avi tag, HRV3C protease cleavage site, streptavidin-binding tag (Strep-tagII), and endoplasmic reticulum retention signal ("KD" in the figure). The amino acid sequence of the antigen peptide is shown in Table 5.

[0334]

[0335] The DP proteins, DP#1 to DP#2, consist of a fusion protein (DPA) containing the α-chain water-soluble extracellular domain of the HLA-DP protein, and a fusion protein (DPB) containing the β-chain water-soluble extracellular domain of the HLA-DP protein. The sequence numbers of the amino acid sequences of these fusion proteins are shown in Table 6. The allele names of the α-chain and β-chain from which the water-soluble extracellular domains are derived, and the sequence numbers of the amino acid sequences of the water-soluble extracellular domains in DPA#1 and DPB#1 to DPB#2 are also shown in Table 6.

[0336]

[0337] The expression vector for the DP protein is shown in Figure 1D. DPA#1 has the following sequences: SSExt, α-chain water-soluble extracellular domain, TEV cleavage site, Zipper A, V5 tag, HRV3C protease cleavage site, histidine tag, and endoplasmic reticulum retention signal ("KD" in the figure).

[0338] Both DPB#1 and DPB#2 contain the following sequences: SSExt, antigen peptide, β-chain water-soluble extracellular domain, TEV cleavage site, Zipper B, thrombin cleavage site, Avi tag, and endoplasmic reticulum retention signal ("KD" in the figure). The antigen peptides are shown in Table 7.

[0339]

[0340] The TCR proteins, TCR#1 and TCR#2, consist of a fusion protein containing the extracellular domain of the HA1.7 TCRα chain (TCRα protein) and a fusion protein containing the extracellular domain of the HA1.7 TCRβ chain (TCRβ protein). Table 8 shows the sequence numbers of the amino acid sequences of these fusion proteins and the amino acid sequences of the extracellular domains. The HA1.7 TCRα chain and HA1.7 TCRβ chain are known to specifically recognize the HA antigen peptide represented by sequence number 26.

[0341]

[0342] TCRA#1 has the sequence of SSExt, the extracellular domain of the HA1.7 TCRα chain, a TEV cleavage site, Zipper A, a V5 tag, an HRV3C protease cleavage site, a histidine tag, and an endoplasmic reticulum retention signal ("KD" in the figure). TCRB#1 has the sequence of SSExt, the extracellular domain of the HA1.7 TCRβ chain, a TEV cleavage site, Zipper B, a thrombin cleavage site, an Avi tag, and an endoplasmic reticulum retention signal ("KD" in the figure).

[0343] TCRA#2 has the sequence of SSExt, the extracellular domain of the HA1.7 TCRα chain, the TEV cleavage site, Zipper A, V5 tag, apoplast retention signal, and histidine tag. TCRB#2 has the sequence of SSExt, the extracellular domain of the HA1.7 TCRβ chain, the TEV cleavage site, Zipper B, the thrombin cleavage site, apoplast retention signal, Avi tag, and AGIA tag.

[0344] To express each of the proteins mentioned above, an expression vector having the following structure was used. The structure of the expression vector is shown in Figures 1B to 1D and 7A to 7B. The names of each sequence in the expression vector refer to the following sequences.

[0345] "35S-p×2": 35S promoter of cauliflower mosaic virus (CaMV) with two enhance elements. "TMVΩ5'": 5'-UTR of tobacco mosaic virus. "HSPter": Terminator of the Arabidopsis thaliana heat shock protein 18.2 gene (SEQ ID NO: 44). "Ext3'": Terminator of the tobacco extensin gene (SEQ ID NO: 45). "LIR": Long Integrative Region of the bean malformation virus (BeYDV) genome (SEQ ID NO: 46). "SIR": Short Integrative Region of the BeYDV genome (SEQ ID NO: 47). "C1" and "C2": Open reading frames C1 and C2 encoding the Rep / RepA protein, the replication initiation protein of BeYDV (SEQ ID NO: 48). "RB": Left border sequence of T-DNA (SEQ ID NO: 49) "LB": Right border sequence of T-DNA (SEQ ID NO: 50) "Nos-p": NOS promoter "Nos-t": NOS terminator "p19": Gene encoding the gene silencing inhibitor P19 derived from the tomato bushy stunt virus

[0346] <Protein Expression in Tobacco Benthamiana> Preparation of A. tumefaciens suspension and transient expression of MHCII and TCR proteins in N. benthamiana were performed according to previously described methods (Yamamoto T et al., (2018) Improvement of the transient expression system for production of recombinant proteins in plants. Sci Rep 8:. https: / / doi.org / 10.1038 / s41598-018-23024-y).

[0347] Agrobacterium tumefaciens GV3101, containing the aforementioned binary vector, was cultured for 2 days at 28°C in L-broth medium containing antibiotics (kanamycin 100 mg / L, gentamicin 30 mg / L, rifampicin 30 mg / L).

[0348] Next, the culture medium, which had been incubated for two days, was diluted 100-fold with the same medium containing antibiotics, 10 mM MES (pH 5.6), and 20 μM acetosyringone, and incubated in a rotary shaker at 140 rpm and 28°C for 18-24 hours. After centrifugation, the Agrobacterium were immersed in infiltration buffer (10 mM MgCl). 2 Resuspend in 10 mM MES (pH 5.6), 100 μM acetosyringone, and OD 600 The ratio was adjusted to approximately 1. Equal amounts of Agrobacterium having an expression vector for a protein containing the MHCII protein α chain and Agrobacterium having an expression vector for a protein containing the MHCII protein β chain were mixed. Alternatively, equal amounts of Agrobacterium having an expression vector for a protein containing the extracellular domain of the HA1.7 TCR α chain and Agrobacterium having an expression vector for a protein containing the extracellular domain of the HA1.7 TCR β chain were mixed.

[0349] Next, tobacco leaves were immersed in an Agrobacterium suspension and subjected to vacuum osmosis (29 inHg) for 2 minutes. After osmosis, 200 mM sodium ascorbate was sprayed onto the leaf surface, and the tobacco leaves were left to stand for 3 days at 25°C with a 16-hour light / 8-hour dark light cycle. This resulted in the co-expression of proteins containing the MHCII protein α chain and proteins containing the MHCII protein β chain in the tobacco leaves. Alternatively, proteins containing the extracellular domain of the HA1.7 TCR α chain and proteins containing the extracellular domain of the TCR HA1.7 TCR β chain were co-expressed.

[0350] <Protein Extraction and Purification> Protein extraction was performed as follows, referring to the literature (Yamada Y, Kidoguchi M, Yata A, Nakamura T, Yoshida H, Kato Y, Masuko H, Hizawa N, Fujieda S, Noguchi E, Miura K (2020) High-yield production of the major birch pollen allergen Bet v 1 with allergen immunogenicity in Nicotiana benthamiana. Frontiers in Plant Science 11, 344.).

[0351] Specifically, 100 g of plant leaves were frozen in liquid nitrogen and pulverized using a Vitamix E310 blender (Vitamix Corporation). The samples were then resuspended in a 1:5 ratio with lysis buffer (50 mM Tris-HCl, pH 8.0, 120 mM NaCl, 0.2 mM sodium orthovanadate, 100 mM NaF, 10% glycerol, 0.2% Triton X-100, 5 mM DTT, 1x protease inhibitor cocktail (Nacalai Tesque Co., Ltd., Kyoto, Japan)). These suspensions were incubated on ice for 1 hour with shaking and filtered through Miracloth (Calbiochem).

[0352] Next, the mixture was centrifuged at 44,000 × g for 20 minutes at 4°C, and ammonium sulfate was added to a final concentration of 30% saturation. This solution was incubated on ice for 1 hour with shaking, and then centrifuged again. The supernatant was dialyzed through a dialysis membrane (molecular weight cutoff: 3500, AsOne) to remove the high concentration of ammonium sulfate and lysis buffer, and replaced with binding buffer (10 mM Tris, pH 7.4, 0.5 M NaCl, 20 mM imidazole, 1 mM PMSF).

[0353] Next, complexes of proteins containing the MHCII protein α chain and MHCII protein β chain, both possessing a histidine tag, were purified using cobalt resin (TALON®, Takara Bio, Inc.). The proteins containing the MHCII protein α chain were eluted into five fractions using 1×bed volume elution buffer.

[0354] The isoelectric point (pI) of the MHCII protein was calculated using the website (https: / / www.protpi.ch). The fraction was then collected from the cobalt resin and dialysis was performed using a dialysis membrane (molecular weight cutoff: 3500, AsOne, Japan), and the buffer was changed to MES buffer pH 6.5 (137 mM NaCl, 2.7 mM KCl, 24.8 mM Tris, pH 7.4).

[0355] Next, ion exchange chromatography was performed on the protein sample using an AKTA start system (Cytiva, formerly GEHealthcare) equipped with a HiTrap® QFF column (Cytiva, formerly GEHealthcare). The purified protein from the peak fraction was concentrated, and the buffer was changed to phosphate-buffered saline (PBS) pH 7.0 (137 mM NaCl, 2.7 mM KCl, 24.8 mM Tris, pH 7.4) using a 30 kDa cutoff membrane (Sartorius Vivaspin Turbo 15).

[0356] Using the same method as described above for purifying MHC proteins, a complex of proteins containing the extracellular domain of the HA1.7 TCRα chain and proteins containing the extracellular domain of the TCR HA1.7 TCRβ chain was purified using cobalt resin. After calculating the isoelectric point of the TCR protein, it was purified by ion exchange chromatography, and the purified protein was concentrated.

[0357] <CBB Staining and Protein Quantification> For protein quantitative analysis, concentrated samples of DR#1 and DR#2 proteins were obtained using a 30 kDa cutoff membrane (Sartorius Vivaspin Turbo 15). Bovine serum albumin (BSA) of different concentrations was prepared. 10 μl of protein sample was mixed with 2.5 μl of 5× sample buffer and boiled at 95°C for 5 minutes. The sample was then developed on an SDS-PAGE, stained with Coomassie Brilliant Blue (CBB), and images were detected using a WSE-6300 LuminoGraph III (ATTO).

[0358] <Western Blotting> For immunoblotting analysis, proteins developed using SDS-PAGE were transferred to a PVDF membrane (Amersham Hybrid P PVDF, GEHealthcare). As primary antibodies, anti-6×His tag antibody (rabbit polyclonal antibody, ab9108, Abcam) and anti-rabbit IgG-HRP (Nacalaitesque) were used. As secondary antibodies, anti-Avi tag antibody (mouse monoclonal antibody, Genscript), anti-mouse IgG-HRP (Jackson immuno research laboratories, inc) and anti-AGIA tag antibody (donated by Dr. Hiroyuki Takeda, Ehime University) were used. The membrane was detected using Luminata Forte Western HRP substrate (Millipore) with a WSE-6300 LuminoGraph III (ATTO).

[0359] <ELISA> The concentrations of MHCII proteins in PBS pH 7.0 were adjusted to approximately 1 μg / ml for DR#1 protein and approximately 0.7 μg / ml for DR#2 protein. Then, 10 μl of each protein was diluted with PBS containing 0.1% (v / v) Tween 20 (Wako, hereafter referred to as PBS-T) and placed in wells of a Pierce nickel-coated plate (ThermoFisher Scientific Inc, IL, USA), and incubated at room temperature for 2 hours. Lysates of mouse fibroblast cell line NIH3T3 (mock) were used as a negative control. Lysates of NIH3T3 expressing DRA*01:01 and HLA-DRB1*09:01 or HLA-DRB1*04:05 were used as a positive control.

[0360] Next, the plates were washed with PBS-T, and after washing with PBS-T, 100 μl of mouse anti-human HLA DP DQ DR monoclonal antibody (clone: ​​WR18, BIO-RAD), adjusted to a concentration of 0.1 μg / ml with PBS-T, was added to each well, and the plates were incubated at room temperature for 1 hour.

[0361] Next, the plate was washed with PBS-T, and 100 μl of goat anti-mouse IgG-HRP conjugate (SantaCruz), adjusted to a concentration of 0.1 μg / ml with PBS-T, was added to each well. The plate was incubated at room temperature for 1 hour, washed again with PBS-T, and 40 μl of ECL substrate (ECL Prime Western Blotting Detection Reagent (Cytiva)) was added to allow the reaction to proceed. After 60 seconds, luminescence was measured using Varioskan LUX (Thermo Fisher Scientific). Luminescence images were captured using Sayaca-Imager (DRC Corporation) after 60 seconds of exposure.

[0362] <Deglycosylation of TCR Protein> N-linked glycans were removed from TCRA#2 in the purified TCR protein complex by treating it with PNGase F (New England Biolabs) for 16–24 hours under non-denaturing conditions. The treated samples were analyzed using Western blotting and the Pro-Q® Emerald 300 Glycoprotein Gel and Blot Stain Kit (ThermoFisher Scientific). In the latter analysis, approximately 0.25 μg of protein was developed on a 10% SDS-PAGE gel, and glycans were analyzed using the kit. The stained gels were analyzed using a WSE-6300 LuminoGraph III (ATTO) with an excitation filter of BP300–320 nm and an emission filter of BP520–540 nm (green channel).

[0363] <Analysis of Binding Between TCR Protein and MHC Protein> The binding of TCR#2, which has an AGIA sequence, to DR#1, which presents CLIP, or DR#2, which presents HA, was analyzed as follows. The binding analysis was performed using a Biacore® T100 instrument (GE Healthcare, Chicago, IL, USA), referring to a previously described method (Mamedov A et al., Protective Allele for Multiple Sclerosis HLA-DRB1*01:01 Provides Kinetic Discrimination of Myelin and Exogenous Antigenic Peptides. Frontiers in Immunology, 10. https: / / doi.org / 10.3389 / fimmu.2019.03088). The capture method was used. The AGIA antibody was immobilized by amine coupling onto two flow cells of a CM5 sensor chip (Cytiva, Marlborough, MA, USA) at approximately 5000–6000 response units (RUs). Next, the AGIA-tagged HA1.7 TCR protein was manually injected into the flow cells at a rate of 10 μL / min over multiple 60-second cycles until 300 RUs were reached. Throughout the experiment, PBS (pH 7.4) containing 0.05% (w / v) Tween 20 was used as the running buffer. Interaction analysis was performed by injecting a diluted series of DR#1 or DR#2, prepared with the running buffer, into the flow cells at 30 μL / min. To standardize the raw signal, sensorgrams immobilized with only the AGIA antibody were subtracted. Kd was calculated by fitting equilibrium-bound data to a one-site-binding model using GraphPad PRISM.

[0364] (Experimental Example 1) Expression vectors for the DR#1-DR#9 proteins, DQ#1-DQ#2 proteins, and DP#1-DP#2 proteins, as described above in <Expression Vector Preparation>, were prepared. The domain structures of the DR#1-DR#9 proteins are shown in Figure 1A. The expression vectors are shown in Figures 1B-1D.

[0365] (Experimental Example 2) Using the methods described above in <Protein Expression in Benthamia Tobacco> and <Protein Extraction and Purification>, proteins containing an α-chain water-soluble extracellular domain and proteins containing a β-chain water-soluble extracellular domain were co-expressed in the leaves of Benthamia Tobacco for each of DR#1 to DR#9, and these were purified.

[0366] Using anti-histag antibodies, anti-Avi-tag antibodies, and secondary antibodies against them, the expression of proteins containing the MHCII protein α-chain water-soluble extracellular domain and proteins containing the MHCII protein β-chain water-soluble extracellular domain was analyzed in crude extracts obtained from tobacco leaves of Citrus benthamiana. The results are shown in Figure 2A. High expression of DRA#1 and DRB#1-DRB#9 was confirmed in Citrus benthamiana. In the figures showing the results of Western blotting, "1," "2," "3," etc., indicate samples obtained from different tobacco leaves of Citrus benthamiana. The same applies to figures of other experimental examples.

[0367] Ammonium sulfate was added to the centrifuged crude extract to precipitate insoluble proteins, and the supernatant and precipitate were analyzed by immunoblotting. The results are shown in Figure 2B. In Figure 2B, S represents the supernatant and P represents the precipitate. It was confirmed that solubilized DRA#1, DRB#1, and DRB#2 proteins were obtained.

[0368] Next, the supernatant containing the solubilized DR protein was dialyzed. Then, the post-dialysis protein was purified using cobalt resin (TALON, registered trademark). The results are shown in Figure 2C. In Figure 2C, W represents the protein before purification, Ft represents the pass-through fraction, and F1 to F5 represent the eluted fractions. The DR protein was confirmed in the eluted fraction.

[0369] Next, the DR proteins eluted from the cobalt resin were purified by ion-exchange chromatography. The results of the purification of DR#1 and DR#2 are shown in Figures 2D and 2E, respectively. In the figures, Fra indicates the fraction.

[0370] (Experimental Example 3) The expression level of DR protein expressed in Citrus benthamiana was quantified using the method described above in <CBB staining and protein quantification>.

[0371] DR#1 and DR#2 proteins were purified and concentrated from 100 g of tobacco leaves expressing DR protein, then developed on an SDS-PAGE and stained with CBB. The concentration of DR protein was quantified using a BSA calibration curve. The results are shown in Figures 3A and 3B. Figure 3A shows the results of CBB staining, and Figure 3B shows the BSA calibration curve. The concentrations of DR#1 and DR#2 proteins were 4.37 μg / gFW and 2.81 μg / gFW, respectively.

[0372] (Experimental Example 4) In <ELISA>, ELISA analysis was performed on the undenatured DR protein obtained from tobacco benthamiana using the method described above.

[0373] ELISA analysis was performed on purified and concentrated DR#1 and DR#2 proteins without denaturation, using a mouse anti-human HLA DP DQ DR monoclonal antibody (clone: ​​WR18, BIO-RAD) and a secondary antibody, goat anti-mouse IgG-HRP conjugate (SantaCruz). NIH3T3 (mock) cell extract was used as a negative control. NIH3T3 cell extract expressing HLA-DRB1*09:01 or HLA-DRB1*04:05 was used as a positive control. The results are shown in Figures 4A and 4B.

[0374] In Figure 4A, "DR#1-1" and "DR#1-2" show the results of two analyses of the DR#1 protein, and "DR#2-1" and "DR#2-2" show the results of two analyses of the DR#2 protein. In Figure 4B, "DR#1" shows the average values ​​of "DR#1-1" and "DR#1-2," and "DR#2" shows the average values ​​of "DR#2-1" and "DR#2-2."

[0375] Similar to the positive control, it was confirmed that DR#1 and DR#2 proteins can be used in ELISA analysis.

[0376] (Experimental Example 5) <Protein Expression in Benthamiana> Using the method described above, proteins containing an α-chain water-soluble extracellular domain and proteins containing a β-chain water-soluble extracellular domain were co-expressed in Benthamiana for each of the DQ#1 and DQ#2 proteins. The expression construct is shown in Figure 1C.

[0377] Recombinant protein expression was analyzed in crude extracts obtained from tobacco benthamiana using anti-histag antibodies and anti-Avi-tag antibodies. The results are shown in Figure 5A. High expression of DQA#1 and DQB#1-DQB#2 was confirmed in tobacco benthamiana.

[0378] Similar to Experimental Example 2, ammonium sulfate was added to the obtained crude extract to obtain the supernatant, which was then dialyzed. Next, the dialyzed protein was purified using cobalt resin (TALON, registered trademark). The results are shown in Figure 5B. In Figure 5B, W represents the protein before purification, Ft represents the pass-through fraction, and F1 to F5 represent the eluted fractions. Western blotting was performed using anti-His tag antibody and anti-Avi tag antibody. The DQ protein was confirmed in the eluted fraction.

[0379] (Experimental Example 6) <Protein Expression in Benthamiana> Using the method described above, DP#1 and DP#2 proteins were co-expressed in Benthamiana using proteins containing both α-chains and β-chains. The expression constructs are shown in Figure 1D.

[0380] Recombinant protein expression was analyzed in crude extracts obtained from tobacco benthamiana using anti-histag antibodies and anti-Avi-tag antibodies. The results are shown in Figure 6A. High expression of DPA#1 and DPB#1-DPB#2 was confirmed in tobacco benthamiana.

[0381] Similar to Experimental Example 2, ammonium sulfate was added to the obtained crude extract to obtain the supernatant, which was then dialyzed. Next, the dialyzed protein was purified using cobalt resin (TALON, registered trademark). The results are shown in Figure 6B. In Figure 6B, W represents the protein before purification, Ft represents the pass-through fraction, and F1 to F5 represent the eluted fractions. Western blotting was performed using anti-His tag antibody and anti-Avi tag antibody. DP protein was confirmed in the eluted fraction.

[0382] (Experimental Example 7) <Protein Expression in Benthamiana> Using the method described above, in Benthamiana, each of the TCR#1 to TCR#2 proteins was co-expressed with a protein containing the extracellular domain of the HA1.7 TCRα chain and a protein containing the extracellular domain of the TCRHA1.7 TCRβ chain. The expression constructs are shown in Figures 7A and 7B.

[0383] Recombinant protein expression was analyzed in crude extracts obtained from tobacco benthamiana using anti-histag antibodies and anti-AGIA tag antibodies. The results are shown in Figures 7C and 7D. In tobacco benthamiana, the expression levels of TCRA#1 and TCRB#1, which have endoplasmic reticulum retention signals (KD), were found to be low. In tobacco benthamiana, TCRA#2 and TCRB#2, which have apoplast retention signals (SP10), were found to be highly expressed. These results indicate that adding an apoplast retention signal (SP10) increases the expression and recovery levels of TCRs more than adding an endoplasmic reticulum signal (KD).

[0384] (Experimental Example 8) The site where N-linked glycans are attached to the HA1.7 TCRα chain was predicted. NetNGlyc 1.0 (http: / / www.cbs.dtu.dk / services / NetNGlyc / ) was used for prediction. This tool predicts the site where N-linked glycans will bind from the surrounding sequence of the Asn-Xaa-Ser / Thr motif. If the glycosylation potential score is greater than 0.5, N-linked glycan attachment was predicted. The motif location and prediction results are shown in Figure 8A and Table 9, respectively.

[0385]

[0386] The three-dimensional structure (PDB ID: 1J8H) of the complex of the extracellular domain of the HA1.7 TCRα chain and HLA DR1 presenting the influenza HA peptide was obtained from the RCSB Protein Data Bank (https: / / www.rcsb.org / 3d-view / 1J8H / 0). The position of asparagine at position 24 in this three-dimensional structure is shown in Figure 8B. Asparagine at position 24 is located at the site where the TCR and the peptide-presenting MHC interact. This suggests that glycosylation of asparagine at position 24 may be important for the binding of the TCR to the peptide-presenting MHC.

[0387] (Experimental Example 9) <Glycosyl Decongestion of TCR Protein> The glycans of TCRA#2 obtained in Experimental Example 7 were removed using the method described above, and the sample was analyzed using Western blotting and the Pro-Q® Emerald 300 Glycoprotein Gel and Blot Stain Kit (ThermoFisher Scientific).

[0388] The results of Western blotting are shown in Figure 9A. TCRA#2, which had not undergone glycan removal, was identified as multiple bands. Glycan removal shifted the band mobility of TCRA#2 towards the lower molecular weight side. The results using the Pro-Q® Emerald 300 Glycoprotein Gel and Blot Stain Kit are shown in Figure 9B. A reduction in signal was observed after glycan removal. These results indicate that glycans were attached to TCRA#2.

[0389] (Experimental Example 10) Using the method described above in <Analysis of binding between TCR protein and MHC protein>, the binding of TCR#2 obtained in Experimental Example 7 to DR#1 presenting CLIP or DR#2 presenting HA was analyzed. The binding of TCR#2 obtained in Experimental Example 9 after glycosylation was analyzed to DR#1 or DR#2.

[0390] Figure 10A shows the results of the analysis of the binding of TCR#2 obtained in Experimental Example 7 to DR#1 or DR#2. It was confirmed that TCR#2 did not bind to DR#1 but bound to DR#2. The affinity between TCR#2 and DR#2 was a dissociation constant Kd = 3.62 ± 1.59 μM.

[0391] Figure 10B shows the results of the analysis of the binding between TCR#2 after glycan removal and DR#1 or DR#2, obtained in Experimental Example 9. It was confirmed that TCR#2 after glycan removal does not bind to DR#2. It was confirmed that the glycans of TCR#2 are important for binding.

[0392] According to the present invention, a technology for producing MHC protein and TCR protein in plant cells can be provided.

Claims

1. A method for producing an MHC protein, comprising the step of expressing an MHC protein in plant cells to obtain the MHC protein.

2. The method for producing an MHC protein according to claim 1, wherein the MHC protein is a mammalian protein.

3. The method for producing an MHC protein according to claim 1, wherein the MHC protein is a protein selected from (a) to (f) below. (a) Proteins containing an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (b) Proteins containing a portion of an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (c) Proteins containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added to an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (d) Proteins containing an amino acid sequence in which one or more amino acids are deleted, inserted, substituted and / or added to a portion of an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (e) Proteins containing an amino acid sequence having 90% or more sequence identity with an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus; (f) Proteins containing an amino acid sequence encoded by the MHC locus or the β2 microglobulin locus A protein containing an amino acid sequence that has more than 90% sequence identity with a portion of the amino acid sequence encoded by the microglobulin gene locus.

4. The method for producing an MHC protein according to claim 2, wherein the MHC protein is an MHC class II protein.

5. The method for producing an MHC protein according to claim 4, wherein the MHC class II protein is an HLA-DR protein, an HLA-DQ protein, or an HLA-DP protein.

6. The method for producing an MHC protein according to claim 1, wherein the MHC protein is a fusion protein in which zipper-like helix domains are fused.

7. A method for producing an MHC protein according to claim 1, comprising co-expressing a protein derived from the α chain of an MHC class I protein and a protein derived from β2 microglobulin of an MHC class I protein to obtain a dimer of the protein derived from the α chain of an MHC class I protein and the protein derived from β2 microglobulin of an MHC class I protein, or co-expressing a protein derived from the α chain of an MHC class II protein and a protein derived from the β chain of an MHC class II protein to obtain a dimer of the protein derived from the α chain of an MHC class II protein and the protein derived from the β chain of an MHC class II protein.

8. The method for producing an MHC protein according to claim 1, wherein the step of obtaining the MHC protein includes a step of introducing an expression system into the plant cells, the expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette for the MHC protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette for the Rep / RepA protein derived from a geminivirus, the expression cassette for the MHC protein comprising, in this order, a promoter, a nucleic acid fragment encoding the MHC protein, and two or more linked terminators.

9. An expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an MHC protein expression cassette linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette for a Rep / RepA protein derived from a geminivirus, wherein the MHC protein expression cassette comprises, in this order, a promoter, a nucleic acid fragment encoding the MHC protein, and two or more linked terminators.

10. Agrobacterium bacteria transformed by the expression system described in claim 9.

11. A plant cell into which the expression system described in claim 9 has been introduced.

12. A method for producing a transformed plant, comprising introducing the expression system described in claim 9 into plant cells.

13. An expression vector comprising a first nucleic acid fragment comprising a geminivirus-derived LIR, a geminivirus-derived SIR, and an expression cassette of an MHC protein linked between the LIR and the SIR, wherein the expression cassette of the MHC protein comprises, in this order, a promoter, a multicloning site, and two or more linked terminators.

14. A method for producing a TCR protein, comprising the step of expressing a TCR protein in plant cells to obtain the TCR protein.

15. The method for producing a TCR protein according to claim 14, wherein the TCR protein is a fusion protein in which a zipper-like helix domain is fused.

16. The method for producing a TCR protein according to claim 14, wherein the TCR protein is a fusion protein into which an apoplast retention signal is fused.

17. A method for producing a TCR protein according to claim 14, comprising co-expressing a protein derived from the α chain of the TCR protein and a protein derived from the β chain of the TCR protein to obtain a dimer of the protein derived from the α chain of the TCR protein and the protein derived from the β chain of the TCR protein, or co-expressing a protein derived from the γ chain of the TCR protein and a protein derived from the δ chain of the TCR protein to obtain a dimer of the protein derived from the γ chain of the TCR protein and the protein derived from the δ chain of the TCR protein.

18. The method for producing a TCR protein according to claim 14, wherein the step of obtaining the TCR protein includes a step of introducing an expression system into the plant cells, the expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette for the TCR protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette for the Rep / RepA protein derived from a geminivirus, the expression cassette for the TCR protein comprising, in this order, a promoter, a nucleic acid fragment encoding the TCR protein, and two or more linked terminators.

19. An expression system comprising: a first nucleic acid fragment comprising a Long Integrative Region (LIR) derived from a geminivirus, a Small Integrative Region (SIR) derived from a geminivirus, and an expression cassette for a TCR protein linked between the LIR and the SIR; and a second nucleic acid fragment comprising an expression cassette for a Rep / RepA protein derived from a geminivirus, wherein the expression cassette for the TCR protein comprises, in this order, a promoter, a nucleic acid fragment encoding the TCR protein, and two or more linked terminators.

20. Agrobacterium bacteria transformed by the expression system described in claim 19.

21. A plant cell into which the expression system described in claim 19 has been introduced.

22. A method for producing a transformed plant, comprising introducing the expression system described in claim 19 into plant cells.

23. An expression vector comprising a first nucleic acid fragment comprising a geminivirus-derived LIR, a geminivirus-derived SIR, and an expression cassette of a TCR protein linked between the LIR and the SIR, wherein the expression cassette of the TCR protein comprises, in this order, a promoter, a multi-cloning site, and two or more linked terminators.

24. A method for evaluating the binding affinity between an MHC protein fused with an antigen peptide and a TCR protein, comprising: a step (a1) of obtaining the MHC protein by the method for producing the MHC protein described in claim 1; and a step (b1) of evaluating the binding affinity between the MHC protein obtained in step (a1) and the TCR protein.

25. A method for evaluating the binding affinity between an MHC protein containing an antigen peptide sequence and a TCR protein, comprising: a step (a2) of obtaining the TCR protein by the method for producing the TCR protein described in claim 14; and a step (b2) of evaluating the binding affinity between the MHC protein and the TCR protein obtained in step (a2).