Method for producing protein domains that induce immune responses and protein domains
By designing and optimizing protein domains for expression in E. coli based on predicted structure and exposure, the method addresses yield and neutralizing ability challenges, achieving efficient and cost-effective production of immunogenic viral proteins.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2026-04-01
AI Technical Summary
Existing methods for expressing viral protein domains, such as the receptor-binding domain (RBD) of the coronavirus spike protein, face challenges in achieving high expression yields and neutralizing ability, particularly for strains like Omicron and Delta, and often require animal cells, which are costly and difficult to purify.
A method for designing and producing protein domains by predicting their three-dimensional structure, physicochemical properties, and exposure levels, allowing for efficient expression in E. coli, using software like AlphaFold and DSSP, and optimizing expression and purification protocols.
The method enables the production of protein domains with high immunogenicity and neutralizing ability, reducing development time and costs, and avoiding issues associated with animal cell expression, such as purification difficulties and environmental impact.
Smart Images

Figure 2026056159000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a protein domain, which is a part of a protein, and mainly has immunogenicity against a specific virus. That is, it relates to methods for designing, selecting, and manufacturing the protein domain, and the protein domain manufactured thereby.
Background Art
[0002] Regarding viruses that cause various viral diseases, such as SARS-CoV-2 that causes COVID-19 (coronavirus disease 2019) and influenza viruses, vaccines using nucleic acids or proteins have been studied. As a vaccine using nucleic acids, for example, mRNA vaccines have conventionally been used. On the other hand, mRNA vaccines have reports of ultra-low temperature storage, instability, or side effects. For this reason, vaccines using protein domains (hereinafter also referred to as domains), which are the sequences of proteins that make up viruses or parts thereof, have also been studied.
[0003] If a protein or a part of its sequence is appropriately designed, it can be expressed by gene introduction into a microorganism, and can induce an immune response at the level of a vaccine lead candidate in terms of quality (ability to induce neutralizing antibodies) and quantity (titer). When the target is a virus or bacteria, it can be a vaccine antigen that defends against infection by them.
[0004] For example, Patent Document 1 discloses an engineered immunogen polypeptide derived from the coronavirus spike (S) protein, and a coronavirus vaccine composition using the same, comprising a modified soluble S sequence having modifications that stabilize the pre-fusion S structure compared to the wild-type soluble S protein sequence of the coronavirus, wherein the modifications include a mutation that inactivates the S1 / S2 cleavage site and a mutation in the turn region between the hepta repeat 1 (HR1) region and the central helix (CH) region that prevents HR1 and CH from forming a linear helix during fusion. This technology aims to provide a redesigned soluble coronavirus S protein-derived immunogen stabilized by specific modifications in the wild-type soluble S sequence, a nanoparticle vaccine containing the redesigned soluble S immunogen displayed on self-assembling nanoparticles, a polynucleotide sequence encoding the redesigned immunogen and nanoparticle vaccine, and a method of using the vaccine composition to prevent or treat coronavirus infection.
[0005] Non-patent document 1 discloses a technique for expressing the receptor-binding domain (RBD; aa:319-541) of the Wuhan-type coronavirus spike protein in T7 Shuffle E. coli cells using a pET15b expression vector in LB medium containing 50 μg / mL ampicillin.
[0006] Non-patent document 2 discloses a technique that uses E. coli to express the receptor-binding domain of the coronavirus spike protein and analyze the immune response. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] Special Publication No. 2023-533228 [Non-patent literature]
[0008] [Non-Patent Document 1] Brindha S, Yoshizue T, Wongnak R, Takemae H, Oba M, Mizutani T, Kuroda Y. An Escherichia coli Expressed Multi-Disulfide Bonded SARS-CoV-2 RBD Shows Native-like Biophysical Properties and Elicits Neutralizing Antisera in a Mouse Model. Int J Mol Sci. 2022 Dec 12;23(24):15744. [Non-Patent Document 2] Ke Q, Sun P, Wang T, Mi T, Xu H, Wu J, Liu B. Non-glycosylated SARS-CoV-2 RBD elicited a robust neutralizing antibody response in mice. J Immunol Methods. 2022 Jul;506:113279. doi: 10.1016 / j.jim.2022.113279. Epub 2022 May 6. PMID: 35533747; PMCID: PMC9075978. [Overview of the project] [Problems that the invention aims to solve]
[0009] The technology described in Patent Document 1 expresses the RBD domain of the S protein in yeast and animal cells. Because animal cells are used for expression, the ease of purification, cost, effort, and yield are not as good as when expression is performed using E. coli or other organisms.
[0010] Non-patent document 1 is an example of expressing the receptor-binding domain of the coronavirus spike protein in E. coli, but it focuses on the Wuhan type. Furthermore, the protein domain of the sequence described in that document has low neutralizing ability against the virus in the resulting antibodies.
[0011] Non-patent document 2 describes an example of expressing the receptor-binding domain of the coronavirus spike protein in E. coli, focusing on the Omicron and Delta strains. However, the protein domain of the sequence described in this document did not yield sufficiently during purification, and its neutralizing ability against the Omicron and Delta strains of the virus was low.
[0012] Given this background, there is a demand for technologies that offer higher predictable expression levels or immunogenicity of proteins or parts thereof than those offered by conventional methods. This invention has been made in view of the above circumstances, and its purpose is to provide a method for producing a protein domain that can induce a high immune response to viruses and maintain its natural structure, by appropriately designing the protein domain, such as by adjusting the functional site related to immunity, and to provide a protein domain produced using this method. [Means for solving the problem]
[0013] To solve the above problems, the present invention has the following aspects. [1] A method for producing a protein domain having immunogenicity against viruses, A step of estimating the immunogenic protein domain from the three-dimensional structure predicted from the amino acid sequence of the protein constituting the virus, A step of predicting the physicochemical properties, structural properties, and stability of the estimated protein domain, A step of predicting the degree of exposure of each part of the virus from the three-dimensional structure of the virus, A step of selecting a protein domain to be manufactured based on the physicochemical properties of the protein domain, the structural properties, the stability, and the exposure level of each part of the virus; A method for producing protein domains, including those mentioned above. [2] A protein domain produced by the method for producing a protein domain described in [1], The aforementioned virus is the protein domain of SARS coronavirus 2. [3] A protein domain produced by the method for producing a protein domain according to [1], wherein the virus is an influenza virus, the protein domain. [4] A protein domain produced by the method for producing a protein domain according to [1], wherein the protein domain is (I) an amino acid sequence selected from amino acid residues 323 to 538 of the spike protein of SARS coronavirus 2 - Omicron strain, an amino acid sequence in which at least one residue of cysteine in the amino acid sequence is substituted with alanine, or an amino acid sequence containing SEQ ID NO: 1, or (II) an amino acid sequence selected from amino acid residues 45 to 279 of the HA1 protein of hemagglutinin of influenza A, an amino acid sequence in which at least one residue of arginine in the amino acid sequence is substituted with alanine, or an amino acid sequence containing SEQ ID NO: 2, a protein domain having an amino acid sequence selected from the above. [Advantages of the Invention]
[0014] According to the present invention, by appropriately designing a protein domain, a method for producing a protein domain that can be simply and efficiently produced while maintaining a natural structure that induces a high immune response in terms of quality (titer) and quantity (neutralizing ability) against the virus, and a protein domain produced using the same can be provided. [Brief Description of the Drawings]
[0015] [Figure 1] It is a schematic diagram of a method for designing and estimating a viral protein domain that induces an immune response having a neutralizing ability and an explanation of each component. [Figure 2] It is a schematic diagram created with a coronavirus showing an overview of the protein domain selected in this example. [Figure 3] It is a schematic diagram (a) showing a protein domain region for the full - length S protein created with the coronavirus of this example and a schematic diagram (b) showing an overview of its purification. [Figure 4]This is a graph showing the secondary structure of COV-RBD in this example, obtained by circular dichroism spectroscopy (CD spectroscopy). The measurement temperature is shown in the figure. At room temperature (25°C), a structure close to the natural structure was observed, while at high temperature (70°C), denaturation and secondary structure content were observed. [Figure 5] This graph shows the reversibility of the thermal denaturation of COV-RBD in this example. The natural structure was observed at room temperature (25°C), denaturation occurred when the temperature rose to a high temperature (70°C), and the natural structure was observed again when the temperature dropped back down to a low temperature (25°C). This indicates that the thermal denaturation is reversible and suggests that COV-RBD forms a natural structure that maintains its activity at low temperatures. [Figure 6] This graph shows the denaturation temperature of COV-RBD in this example, monitored using the CD value at a wavelength of 222 nm. The combined denaturation curve suggests that COV-RBD forms its native structure at low temperatures. [Figure 7] This graph shows the tertiary structure of COV-RBD in this example, measured by fluorescence spectroscopy. The natural structure was observed at room temperature (25°C), denatured upon increasing temperature (70°C), and then returned to the natural structure upon decreasing temperature again. This indicates that thermal denaturation is reversible, suggesting that COV-RBD maintains its activity at low temperatures by forming a natural structure. [Figure 8] This is a schematic diagram showing the immunotherapy scheme in this embodiment. [Figure 9] This graph shows the IgG titer measurement by ELISA in RBD-immunized mice with adjuvant in this example. A strong immune response was observed in one out of three mice (M1). [Figure 10] This graph shows the IgG titer measurement by ELISA in RBD-immunized mice without adjuvant in this example. A strong immune response was observed in 1 out of 4 mice (M1), and a fairly strong immune response was observed in 3 out of 4 mice. [Figure 11]This graph shows the recognition of mammalian-expressed S1 spike protein by E. coli-expressed RBD antiserum in this example. It was shown that serum induced by E. coli-expressed COV-RBD recognizes mammalian-expressed S1 spike protein, thus suggesting the possibility of recognition of S1 on the virus. Furthermore, since the results were independent of the presence or absence of an immunostimulant, it was shown that an immunostimulant can be used in conjunction with the vaccine but is not essential. Therefore, E. coli-expressed COV-RBD is expected to be a sufficient vaccine even without an immunostimulant. [Figure 12] This graph shows the inhibition of RBD binding to hACE2 (human ACE2 receptor protein) in mice by E. coli-expressed RBD antiserum in this example (experimental results by Bio Interferometry). It was shown that serum induced by E. coli-expressed COV-RBD suppresses the binding of COV-RBD to ACE2, suggesting in vitro that the serum may suppress infection with the coronavirus. [Figure 13] This graph shows the neutralization by RBD antiserum pseudovirus in this example. It was verified in vitro that serum induced by E. coli-expressed COV-RBD completely suppressed pseudocoronavirus infection. [Figure 14] This graph shows the neutralizing titer (ID50) of the RBD antiserum pseudovirus in this example. This indicates that the serum induced by E. coli-expressed COV-RBD has the same level of neutralizing ability as the approved Novavax (the ID50 for Novavax is a literature value). [Figure 15] This is a schematic diagram illustrating the domainization of the hemagglutinin protein (HA) of the influenza virus (INFV) selected in this example. [Figure 16] This is a chart of the RP-HPLC (reverse-phase HPLC) purification of the INF-RBD protein in this example. It can be seen that the purity of the purified INF-RBD is close to 100%. [Figure 17]This graph shows the reversibility of thermal denaturation of INF-RBD in this embodiment. The natural structure was observed at room temperature (25°C), denaturation occurred when the temperature rose to a high temperature (70°C), and the natural structure was observed again when the temperature dropped back down to a low temperature (25°C). This indicates that thermal denaturation is reversible and suggests that INF-RBD maintains its activity at low temperatures by forming a natural structure. [Figure 18] This is a schematic diagram showing the immune protocol for the INF-RBD protein in this embodiment. The protocol is basically the same as that shown in Figure 7. [Figure 19] This bar graph shows the changes in antibody titers in mice immunized with the INF-RBD protein in this example. The bar graph represents the average value of 5 mice, and the error bars represent the standard deviation. [Figure 20] This graph shows the evaluation of the serum specificity of the INF-RBD protein in this example. Since Lyzoyme is not recognized at all when compared, it is suggested that serum specific to COVID-RBD was generated. [Modes for carrying out the invention]
[0016] The method for producing a protein domain and the protein domain itself according to the present invention will be described below with reference to embodiments. However, the present invention is not limited to the following embodiments.
[0017] (Method for manufacturing protein domains) The method for producing a protein domain according to this embodiment is a method for producing a protein domain having immunogenicity against a virus, and includes the steps of: estimating the immunogenic protein domain from the three-dimensional structure predicted from the amino acid sequence of the protein constituting the virus; predicting the structure and stability of the estimated protein domain; predicting the degree of exposure of each part of the virus from the three-dimensional structure of the virus; and selecting a protein domain to be produced from the information on the structure, stability, and degree of exposure of the protein domain.
[0018] In this embodiment, the term "protein domain" broadly refers to at least a portion of the amino acid sequence of a protein, but more specifically to a portion of the sequence that has a distinct structure and / or function independent of that portion.
[0019] In this embodiment, the virus includes various viruses, but is preferably one that has pathogenicity related to diseases of living organisms, and is particularly preferably a virus that has pathogenicity related to diseases of humans. Specifically, it is more preferably a coronavirus, particularly SARS coronavirus 2-Omicron strain, or influenza virus, as described later.
[0020] In this embodiment, the immunogenicity of a protein domain refers to its property of being able to induce a reaction equivalent to an immune response against the virus when administered to a living organism in place of the virus. Specifically, it refers to properties such as acting as an antigen and reacting with antibodies against the virus (antibody reactivity), or inducing the production of antibodies against the virus in living organisms (antibody-inducing ability). It also refers to properties such as competing with the virus and neutralizing its effects (replication, proliferation, etc.) (neutralizing ability). In this embodiment, any one or more of these properties are also referred to as the property of inducing an immune response. The protein domain of this embodiment may have one or more of the aforementioned properties: antibody reactivity, antibody induction ability, and neutralizing ability.
[0021] The method for producing the protein domain of this embodiment comprises the following steps. (1) A step of extracting the immunogenic protein domain from the three-dimensional structure predicted from the amino acid sequence of the protein constituting the virus. First, the three-dimensional structure of the target protein is predicted from its amino acid sequence. The amino acid sequence of the target protein can be obtained using information from the virus as appropriate, including analyzed data or information from viruses registered in databases. For example, if a known structure exists for the virus, the PDB database can be referenced. Predicting the three-dimensional structure from the amino acid sequence can be done using existing software for protein structure prediction or modeling. If it is discovered that a sequence with a known or similar structure exists based on the amino acid sequence, it is preferable to use that known structure. The prediction of the protein structure should also include means of performing modeling or simulation of the three-dimensional structure with higher accuracy using software based on the amino acid sequence. Furthermore, in the case of a known structure, the prediction should also include structural identification.
[0022] Next, the immunogenic protein domain is estimated from the three-dimensional structure. For example, a protein domain that takes on a specific structure, such as an existing protein structure (helix, sheet, combination thereof, existing structure known to have a specific action or function), is estimated, and its site and sequence are extracted. Specifically, the amino acid sequence of the protein constituting the virus is examined to determine whether there are any independent protein domains. If a protein domain is present, its amino acid sequence (part of the protein, the domain) is estimated. If there are multiple such protein domains, several are extracted as candidates, and the optimal protein domain is selected from among them in subsequent steps.
[0023] (2) Steps to predict the physicochemical properties, structural properties, and stability of the estimated protein domain. Next, the estimated protein domains (or multiple domains if multiple domains have been extracted) are verified to ensure that each site maintains a stable structure independently. The physicochemical properties, structural properties, and stability of the protein domains, as well as whether each protein domain maintains a stable structure independently, can be verified using existing protein structure prediction software. AlphaFold is one example of such existing software. Based on that verification, we will further extract candidate protein domains that can maintain a stable structure on their own.
[0024] (3) A step of predicting the degree of exposure of each part of the virus from the three-dimensional structure of the virus. In addition to the extraction of the protein domains, the degree of exposure of each part of the virus is predicted from the three-dimensional structure of the virus. That is, since antibodies are thought to bind to the parts of the virus surface that are exposed, the extent of exposure of each part of the virus is verified. Specifically, the exposed surface area is calculated from the three-dimensional structure of the virus, and the degree of exposure is estimated. Existing software can be used to identify the three-dimensional structure and exposed amino acids. Examples of existing software include DSSP.
[0025] (4) A step of selecting a protein domain to be manufactured based on the information of the structure, stability, and exposure of the protein domain. In light of the information regarding the structure, stability, and exposure of the protein domains estimated in each of the above steps, a protein domain to be manufactured is selected from among the candidate protein domains. Specifically, the protein domain to be manufactured is selected by considering advantages such as the structural advantages (for example, being suitable for expression in specific microorganisms as described later, or having a structure that is advantageous in terms of immunogenicity), stability during use, and high exposure (easily binding to antibodies).
[0026] This embodiment may include a step of further verifying the structure and function of the selected protein domain. That is, for example, it may include a step of designing the protein domain, in which the position and sequence of the protein domain are appropriately modified, taking into consideration activity and ease of expression during the selection process. For example, the process may include a step to predict the thermodynamic stability of the protein domain. To further verify the structure and function of the protein domain, existing programs can be used, such as AlphaFold (a 3D structure prediction program), gromacs.org (a molecular dynamics simulation), DROP (a protein domain boundary prediction program), and IS-DOM (a program to predict the stability of the protein domain), as appropriate.
[0027] This embodiment may further include a step of producing a selected protein domain. The step of producing the protein domain can appropriately be performed by chemical synthesis or by a gene expression system. In particular, production by a gene expression system, in which a gene encoding the selected protein domain is introduced into a microorganism and expressed in the microorganism, is preferred. Various prokaryotes, eukaryotes, etc., can be used as the microorganism, but Escherichia coli is particularly preferred. Expression using Escherichia coli allows for easier and larger-scale production than expression using other organisms.
[0028] In the process of producing protein domains, the gene sequence encoding the protein domain can be modified or modified to include sequences suitable for the expression system. For example, SEP tags and SCP tags are known as solubility tags that increase the solubility of expressed protein domains. SEP tags and SCP tags are also known as expression enhancement tags that increase the expression level. Patent No. 5273438 is known as a technique for increasing solubility, and Patent No. 6986261 is known as a method for improving expression level and yield, and these techniques can be used as appropriate.
[0029] Furthermore, when selecting the protein domain, it is preferable to select one with properties suitable for expression by E. coli. On the other hand, when creating an expression system, it is preferable to select an expression system and strain that are considered optimal based on the amino acid sequence and predicted structure, and then determine the expression and purification conditions. For example, suitable expression vectors include the pET system (Novagen) and pCold. Suitable E. coli strains include BL21, JMM109, ORIGAMI, and Tshuffle, which can be used in combination with the above expression systems as appropriate. Purification conditions such as the presence and condition of air oxidation, tag removal, and selection of the purification column may be selected based on the design of the expression system, for example, its suitability for the vector or E. coli strain.
[0030] Another embodiment may include a protein domain design method comprising steps (1) to (4) described above. The protein domain design method can also be used as a method for selecting immunogenic molecules and predicting their function. The protein domain design method can also be used to select amino acid sequence information of the protein domain and nucleic acid sequence information encoding it.
[0031] (Protein domain) The protein domain of this embodiment is a protein domain obtained by the method for producing a protein domain having immunogenicity against viruses as described above.
[0032] In this embodiment, the protein domain is preferably a coronavirus, preferably SARS coronavirus 2, and more preferably SARS coronavirus 2-Omicron strain. SARS-CoV-2 is the so-called novel coronavirus that causes COVID-19. SARS-CoV-2 Omicron variant (lineage B.1.1.529) is one of the variant strains of SARS-CoV-2. The Omicron strain is the dominant variant among the various major variants of SARS-CoV-2. Unlike other strains that mainly infect the lungs, it mainly infects cells in the upper respiratory tract, and monoclonal antibodies developed against the original strain are known not to be effective against it. The protein domain production method of this embodiment can effectively obtain a protein domain that is immunogenic against SARS coronavirus 2-omicron strain.
[0033] In this embodiment, it is preferable that the virus used to produce the protein domain is the influenza virus. The influenza virus is the cause of acute influenza virus infection (or simply influenza), and there is always a strong demand for treatment methods worldwide. The method for producing the protein domain in this embodiment can effectively obtain a protein domain that has immunogenicity against the influenza virus. Influenza viruses are known to be of types A, B, and C, but this embodiment can be used for any type of influenza virus. Of these, types A and B are known to cause epidemic acute influenza virus infections, and it is more preferable that the protein domain has immunogenicity against these types. Furthermore, the protein domain of this embodiment can also be applied to various viruses other than SARS coronavirus 2-omicron strain or influenza virus.
[0034] (Specific sequence of the protein domain) Furthermore, the protein domain of this embodiment is a protein domain produced by a protein domain production method, and the protein domain is (I) an amino acid sequence selected from amino acid residues 323-538 of the spike protein of SARS coronavirus 2-Omicron strain, an amino acid sequence obtained by substituting at least one cysteine residue of the said sequence with alanine, or an amino acid sequence containing SEQ ID NO: 1 Or, (II) An amino acid sequence selected from amino acid residues 45-279 of the hemagglutinin HA1 protein of influenza A, an amino acid sequence obtained by substituting at least one arginine residue in the said sequence with alanine, or an amino acid sequence containing SEQ ID NO: 2 It has an amino acid sequence selected from the following.
[0035] In this case, the length of the protein domain to be selected is preferably 195 residues or more and 215 residues or less when selecting from sequences (I), and preferably 214 residues or more and 234 residues or less when selecting from sequences (II).
[0036] A protein domain having an amino acid sequence selected from the sequence described in (I) above can be particularly suitably used as a protein domain having immunogenicity against SARS coronavirus 2-omicron strain. A protein domain having an amino acid sequence selected from the sequence described in (II) above can be particularly suitably used as a protein domain having immunogenicity against influenza A. The specific sequences of Sequence IDs 1 and 2 will be described later, but (I) Sequence ID 1 is obtained by substituting cysteine with alanine in amino acid residues 333-528, which are the RBD of the spike protein receptor-binding domain (RBD) of SARS coronavirus 2-Omicron strain. (II) Sequence ID 2 is obtained by substituting arginine with alanine in amino acid residue 221 of amino acid residues 55-269, which are the RBD of HA1, the hemagglutinin of influenza A.
[0037] (Effects of this embodiment) The technology of this embodiment is the first to automatically design and predict protein domains that possess vaccine candidate-level qualitative (neutralizing antibody induction ability) and quantitative (antibody titer) immunogenicity, even when expressed in non-animal insect cells such as E. coli. Furthermore, this invention is the first to experimentally demonstrate vaccine candidate-level immunogenicity (titer and neutralizing antibody induction ability) using this design and prediction technology.
[0038] In this embodiment, by selecting an antigenic fragment, expression in E. coli becomes possible. Furthermore, since the optimal protocol for expression in E. coli can be rapidly identified, development time and effort are significantly reduced. Moreover, by expressing the antigen in E. coli, a new modality (protein domain) for a subunit vaccine, a novel vaccine with vaccine candidate-level immunogenicity (potency and neutralizing ability), can be created using a simpler, safer, less environmentally impactful, lower-cost, and more versatile method than conventional techniques.
[0039] More specifically, according to this embodiment, it is possible to identify protein domains of the target virus that are in their native state or the closest to it, which are optimal for use as vaccine candidates and can be expressed in E. coli, and furthermore, to select the optimal conditions for expressing and purifying those protein domains. Furthermore, according to this embodiment, it is possible to obtain a protein domain against SARS coronavirus 2-omicron strain, and to obtain the protein domain itself and the sequence of the protein domain that have strong immunogenicity and neutralizing antibody-inducing ability against the so-called SARS coronavirus omicron strain expressed in E. coli. Furthermore, according to this embodiment, it is possible to obtain a protein domain against the influenza virus, and to obtain the protein domain itself and the sequence of the protein domain that have strong immunogenicity against the influenza virus.
[0040] Conventional technologies lacked the ability to predict viral protein domains with vaccine antigen-level antigenicity (antibody-inducing ability with titer and neutralizing power) from their amino acid sequences. Furthermore, there were no technologies to design efficient expression and purification protocols for producing these domains using E. coli. In this embodiment, expression is possible using E. coli, and antigens can be produced easily, in large quantities, in a short period of time, and with minimal environmental pollution. Furthermore, since it can be used as a protein vaccine, it has fewer problems associated with mRNA vaccines, such as the need for ultra-low temperature storage, instability, and side effects. Because it is a subunit vaccine that uses the S protein rather than the full-length protein, expression is simple, and the activity per unit mass of the compound is also high. For example, compared to the technology described in Non-Patent Document 2, the present invention yields nearly 10 times more protein domains during purification and achieves 5 to 10 times higher immunogenicity against the Omicron strain. This result is thought to be due to the optimal selection of protein domains in the present invention's technology.
[0041] This embodiment, if the protein domain expressed in E. coli is optimally designed, can induce an immune response at the level of a vaccine candidate in terms of both quality (neutralizing antibody induction ability) and quantity (potency). Furthermore, when the target is a virus or bacteria, it can serve as a vaccine antigen that protects against infection by them. Regarding the design of protein domains, there is no existing method that corresponds to this process. Until now, researchers have relied on their intuition and experience to select the appropriate domain. This embodiment systematically performs this task, resulting in a more efficient selection of antigens that can be expressed in E. coli, and furthermore, the selection of the correct and optimal expression and purification methods, thus offering high efficiency (yield, expression time, etc.). Therefore, it possesses a novel approach to selecting expression and purification methods.
[0042] In this embodiment, a structure suitable for expression in E. coli can be selected. Unlike conventional E. coli expression methods, which produce unusable proteins expressed in the precipitate fraction, this method allows for the selection of a protocol that produces proteins in a state usable as vaccine candidate candidates. Conventional methods required selecting expression and purification protocols based on intuition and experience, and determining the optimal method through trial and error, which was time-consuming and costly. This method uses the algorithm described above to automatically propose the optimal or near-optimal protocol.
[0043] In this embodiment, in particular, as shown in the examples below, it has been demonstrated that protein domains using the influenza receptor-binding domain (INF-RBD) and the COVID-19 receptor-binding domain (COV2-RBD) can be designed, and as a result, INF-RBD and COV2-RBD are effective vaccine candidates. Conventional technologies for INF-RBD and COV2-RBD include older vaccines (live, attenuated, inactivated), but this embodiment is advantageous in all respects, including cost, safety, and environmental impact during production. Furthermore, it is also advantageous in terms of cost, safety, and environmental impact during production for new vaccines (mRNA, DNA, subunits, VLPs, etc.).
[0044] According to the inventors' considerations, in the near future, the majority of the vaccine market will be occupied by novel vaccines. While mRNA vaccines have shown provisional success with COVID-19 vaccines, significant challenges remain, such as side effects and the need for cryopreservation. DNA vaccines have not received as much attention, perhaps because they have not achieved the same level of success as mRNA vaccines, and there are few reports regarding side effects. Subunits and VLPs are all protein-based vaccines, and there are examples of them being put into practical use. However, because they use animal cells or plants, their development and manufacturing periods and manufacturing costs are inferior to those of antigens expressed in E. coli. In the future, it is highly likely that multiple technologies will coexist in the medium to long term for novel vaccines. Protein domain vaccines expressed in E. coli may play an important role (market) within this context.
[0045] This embodiment has the following potential applications. (1) Vaccines against viruses and bacteria and their development The technology of this embodiment significantly improves the efficiency of vaccine development against viruses and bacteria. In particular, the expression system using E. coli is superior in terms of cost and safety. This enables a rapid response to infections caused by novel viruses and bacteria, contributing to the improvement of public health. (2) Other vaccines and development fields related to disease-causing cells and individuals The technology of this embodiment can also be applied to the development of vaccines against disease-causing cells or individuals, such as those responsible for cancer and autoimmune diseases. Because the technology of this embodiment can efficiently select and express specific antigens, it enables approaches to targets that were difficult to achieve with conventional methods. (3) Fields for the search of optimal candidate proteins and protein domains for vaccine antigens The technology of this embodiment is an innovative tool in the search for antigenic proteins and their domains for vaccine candidates. Because the technology of this embodiment enables the rapid selection of the optimal antigen, it can significantly reduce time and costs in the early stages of research and development. (4) Screening for vaccine and formulation By utilizing the technology of this embodiment, screening for vaccine candidate formulation can be performed efficiently. Using an expression system in E. coli makes it possible to evaluate multiple candidates simultaneously, more rapidly and safely than before. This significantly shortens the process of selecting the optimal vaccine formulation. The technology of this embodiment is expected to provide a new approach in these fields and improve the speed and efficiency of vaccine development.
[0046] Although embodiments of the present invention have been described above, the present invention is not limited to the above embodiments and can be modified in various ways. [Examples]
[0047] The effects of the present invention will be further clarified by the following examples and comparative examples. It should be noted that the present invention is not limited to the following examples, and can be implemented with appropriate modifications without altering its essence.
[0048] (Overview of methods for designing and estimating viral protein domains) Figure 1 is a schematic diagram of the design and prediction method for viral protein domains that induce an immune response with neutralizing ability. The following is a description of each component. In the diagram, the components of "i. AlphaFold, Molecular Dynamics, DROP, IS-DOM" are: • AlphaFold (3D structure prediction program): https: / / alphafold.ebi.ac.uk / • Molecular dynamics simulation: https: / / www.gromacs.org / • DROP (Protein Domain Boundary Prediction Program): https: / / web.tuat.ac.jp / ~domserv / DROP.html • IS-DOM (a program for predicting the stability of protein domains): https: / / pubmed.ncbi.nlm.nih.gov / 23715893 / These can be used.
[0049] In the diagram, the components of "ii. Tag, expression vector" and "iii. E. coli strain" are: • Solubility tags: SEP tags, SCP tags, Patent No. 5273438: "Method for calculating the solubility of peptide-added biomolecules, and method for designing peptide tags using the same and method for preventing inclusion body formation," Filing date: January 11, 2008 • Expression enhancement tags: SEP tag, SCP tag, Patent No. 6986261: "Method for improving the expression level and yield of recombinant polypeptides" Date: December 22, 2017 • Expression vectors: pET system (Novagen), pCold, etc. • Escherichia coli strains: BL21, JMM109, ORIGAMI, Tshuffle, etc. Estimated combinations with the above expression systems. These can be used.
[0050] In the diagram, the components of "iv. Estimation of optimal purification conditions" are: • Purification conditions • Presence and conditions of air oxidation • Tag removal • Selection of purification column These can be used.
[0051] (Test Example 1) (Design and purification of the protein domain COV2-RBD) We designed and created a protein domain (COV2-RBD) that exhibits strong immunogenicity against the novel coronavirus for expression in E. coli. We identified the epitope region and receptor-binding domain region from the nucleotide and amino acid sequence of SARS coronavirus 2-Omicron strain. We predicted and estimated the three-dimensional structure using AlphaFold (a three-dimensional structure prediction program) (https: / / alphafold.ebi.ac.uk / ) or, if the structure was known, from the PDB (Protein Data Bank). We then identified exposed amino acids using DSSP (a program that calculates secondary structure and exposed surface area from protein coordinates; Kabsch W, Sander C, Biopolymers. 1983 22 2577-2637.). By comparing these results, we estimated candidate protein domains with neutralizing antibody-inducing ability. Next, the stability of this protein domain was verified using AlphaFold, molecular dynamics simulations (https: / / www.gromacs.org / , etc.), DROP (protein domain boundary prediction program) (https: / / web.tuat.ac.jp / ~domserv / DROP.html), and IS-DOM (program for predicting protein domain stability) (https: / / pubmed.ncbi.nlm.nih.gov / 23715893 / ).
[0052] As a result, we selected the protein domain (COV-RDB) obtained by substituting cysteine with alanine in the SARS coronavirus 2-Omicron strain spike protein (S protein) receptor-binding domain RBD (amino acid residues 333-528). Figure 2 is a schematic diagram showing an overview of the protein domain selected in this example. (a) is the SARS coronavirus 2-Omicron strain, (b) is the spike protein, and (c) is the protein domain COV-RDB.
[0053] The expressed protein domain (COV-RDB) is shown in Sequence ID No. 1. Sequence ID 1: GSSHHHHHHSSGLVPRGSHMTNLCPFDEVFNAT RFASVYAWNRKRISNCVADYSVLYNFAPFFAFKCYGVSPT KLNDLCFTNVYADSFVIRGNEVSQIAPGQTGNIADYNYKL PDDFTGCVIAWNSNKLDSKVGGNYNYRYRLFRKSNLKPFE RDISTEIYQAGNKPCNGVAGVNCYFPLQSYGFRPTYGVGH QPYRVVVLSFELLHAPATVCGKK
[0054] Figure 3 is a schematic diagram showing the protein domain of this embodiment and an overview of its purification. (a) shows the sequence of the protein domain, and (b) shows the purification steps. A plasmid was prepared by inserting the nucleotide sequence encoding the aforementioned protein domain into the multi-cloning site (MCS) between the NdeI and BamHI recognition sequences in the pET-15b vector (Novagen). The 6×His tag and thrombin recognition / cleavage site are included in pET-15b, and in this example, the His tag was used as is. The prepared plasmid was transformed into E. coli by the heat shock method, and the protein was purified.
[0055] SARS-CoV-2 Omicron BA.5 RBD was expressed using an E. coli T7 Shuffle cell line designed to optimize intracellular disulfide bond formation. Protein expression and purification were controlled using a low-temperature induction method to regulate the production rate and promote correct disulfide bond pairing. Further oxidation with GuHCl pH 8.8 completed disulfide bond formation. Further purification was performed by denatured Ni-NTA chromatography and RP-HPLC.
[0056] The biophysical properties and stability of the purified protein domains were measured. All measurements were performed in 10 mM MES buffer, pH 6.5, at a protein concentration of 0.3 mg / mL.
[0057] Figure 4 is a graph of the secondary structure of COV-RBD in this embodiment, obtained by far-ultraviolet CD spectroscopy. Measurements were taken at 10°C intervals from 25°C to 70°C using far-ultraviolet CD spectroscopy (200-260 nm). Figure 5 is a graph showing the reversibility of the COV-RBD in this embodiment against thermal degradation. Figure 6 is a graph showing the melting temperature (Tm) of COV-RBD in this embodiment, monitored using a 222 nm CD spectrum. In the figure, the points represent raw data, and the results of curve fitting corresponding to each point are also shown. Figure 7 is a graph showing the tertiary structure of the COV-RBD in this embodiment, measured by fluorescence spectroscopy. The tryptophan fluorescence intensity of the RBD is shown in 10°C increments from 25°C to 70°C.
[0058] (Immunogenicity of the protein domain COV2-RBD in a mouse model) Two sets of immunization experiments for COV2-RBD (hereinafter also simply referred to as RBD) were performed using 4-week-old Jcl:ICR mice (including one control mouse). One group (4 mice in total) was immunized without adjuvant, and the other group (also 4 mice in total) was immunized in the presence of TiterMax Gold adjuvant. Figure 8 is a schematic diagram showing the immunization scheme of this example. As shown in the figure, the group without adjuvant was immunized with 30 μg at the start of the experiment (day 0), and again on days 14 and 28, and on day 131. The group with adjuvant was immunized only at the start of the experiment.
[0059] Figure 9 is a graph showing the IgG titer measurement by ELISA of adjuvanted RBD-immunized mice in this example. The IgG titer of individual mice (M1-M4) is represented by circles and lines, and the average titer is shown by bars. Figure 10 is a graph showing the IgG titer measurement by ELISA of adjuvant-free RBD-immunized mice in this example. The IgG titer of individual mice (M1-M4) is represented by circles and lines, and the average titer is shown by bars. Figure 11 is a graph showing the recognition of mammalian-expressed S1 spike protein by E. coli-expressed RBD antiserum in this example. The RBD antiserum was diluted 1:1000. The circles represent the OD (absorbance) at 492 nm for four mice in each group (adjuvant-treated group and non-adjuvant-treated group), indicating the binding of the antiserum collected 9 weeks after the initial administration. For comparison, measurements were taken using mammalian-expressed spike (S1) protein or E. coli-expressed RBD as coating antigens. The average OD values for each group are shown in the bar graph. Analysis revealed no significant difference between the S1-coated and RBD-coated mice (p>0.05). The adjuvant was used, while the other group (using a total of 4 mice) was immunized in the presence of Titremax Gold adjuvant (from Rawiwan et al, Molecules 2024).
[0060] Immunogenicity results of E. coli-expressed receptor-binding domains (RBDs) showed that even without adjuvants, E. coli-expressed RBDs elicited a robust immune response, reaching 7.3 × 10⁶ 10 days after the third dose. 5 This resulted in a high titer. In contrast, the group of mice immunized with RBD in combination with an adjuvant did not uniformly show a high immune response. This variability may be due to differences in sensitivity among individual mice, or because a single injection of RBD in combination with an adjuvant does not produce the same level of immune response.
[0061] (Virus neutralization assay) The virus neutralization assay was performed using SARS-CoV-2 Spike (Omicron BA.4&BA.5) Fluc Pseudovirus. HEK293 cells overexpressing human ACE2 (ACROBiosystems) were cultured in DMEM supplemented with 10% FBS and 1% penicillin-streptomycin. The E. coli-expressed RBD serum was diluted 1:1 and mixed with 25 μL of pseudovirus, then incubated at 37°C for 1 hour under a 5% CO2 environment. Subsequently, 4-5 × 10⁻⁶ cells were incubated. 5Cells seeded at a density of cells / mL were added to the wells and incubated under the same conditions for 48 hours. After incubation, the culture medium was discarded, leaving approximately 100 μL in each well. Then, 100 μL of Britlite Plus reporter reagent (PerkinElmer) was added and incubated at room temperature for 2 minutes. The luminescence lumens (RLU) was measured using a luminescence meter (PerkinElmer).
[0062] Figure 12 is a graph showing the inhibition of RBD binding to hACE2 in mice by E. coli-expressed RBD antiserum in this example. One mouse each (n=1) was used from the adjuvant-free group (M1, M7) and the control group. Figure 13 is a graph showing the neutralization by RBD antiserum pseudovirus in this embodiment. In the figure, the dots represent the neutralization results for each of the four mice (n=4) in the adjuvant-free group, and the bar graph shows the average. Figure 14 shows the neutralizing titer (ID) of the RBD antiserum pseudovirus in this example. 50 This is a graph showing ).
[0063] ID of RBD expressed in E. coli prepared in this example 50 The value was 901-fold dilution, which was equivalent to Novavax, an approved anti-SARS-CoV-2 subunit vaccine. The high neutralizing antibody-inducing ability of the antiserum produced by E. coli-expressed RBD is consistent with the rationale that because the protein expressed in E. coli is not post-translationally glycosylated, the antiserum can bind to unmasked epitopes. Overall, our results highlight the potential of the E. coli expression system to produce small proteins that induce the production of neutralizing antiserum, and emphasize its suitability for protein expression in vaccine development.
[0064] (Test Example 2) (Design and purification of the protein domain INF-RBD) We designed and created a protein domain (INF-RBD) that exhibits strong immunogenicity against influenza virus for expression in E. coli. Using the same method as in Test Example 1, an INF-RBD was designed from the base-amino acid sequence of influenza A virus, and the receptor-binding domain of the HA1 hemagglutinin RBD (amino acid residues 55-269) of influenza A was selected. In this example, the H3N2wt domain was used in which arginine at amino acid residue 221 was substituted with alanine. Figure 15 is a schematic diagram showing the overview of the protein domain selected in this example. (a) is a schematic diagram of influenza virus strain A, (b) is a schematic diagram of hemagglutinin, and (c) is a schematic diagram of the protein domain H3N2wt. The three-dimensional structure was predicted from the amino acid sequence using AlphaFold2.
[0065] This protein domain (H3N2wt) is shown in Sequence ID No. 2. Sequence ID 2: GSSHHHHHHSSGLVPRGSHMPHRILDGIDCTLI DALLGDPHCDVFQNETWDLFVERSKAFSNCYPYDVPDYAS LRSLVASSGTLEFITEGFTWTGVTQNGGSNACKRGPGSGF FSRLNWLTKSGSTYPVLNVTMPNNDNFDKLYIWGIHHPST NQEQTSLYVQASGRVTVSTRRSQQTIIPNIGSRPWVRGLS SRISIYWTIVKPGDVLVINSNGNLIAPAGYFKMRTGKSSI MR The molecular weight of H3N2wt was 26052.29 Da, and its isoelectric point (pI) was 9.12.
[0066] H3N2wt expression was performed using the same method as in Experimental Example 1. A plasmid was constructed by inserting the nucleotide sequence encoding H3N2wt into the multi-cloning site (MCS) between the NdeI and BamHI recognition sequences in the pET-15b vector (Novagen). The 6×His tag and thrombin recognition / cleavage site are included in pET-15b, and in this example, the His tag was used as is. The prepared plasmid was transformed into E. coli using the heat shock method, and the INF-RBD protein (protein domain) with the H3N2wt sequence was purified.
[0067] For the purification of the INF-RBD protein, E. coli BL21(DE3) was transformed with a plasmid (pET-15b with H3N2wt introduced), the entire pre-culture was added to LB / Amp liquid medium, and the culture was incubated at 37°C and 250 rpm for approximately 2 hours. When the OD value reached 0.5-0.6, 200 μL of 1 M IPTG was added to induce protein expression to a final concentration of 1 mM, and the cells were collected after 6 hours. Subsequently, the cells were centrifuged at 4°C and 8000 rpm for 20 minutes, the supernatant was removed, 1× Lysis buffer was added to the cells, and sonication was performed three times. Further, 30 mL of 1× Lysis Wash buffer was added to the precipitate and suspended, and sonication, centrifugation, and supernatant removal were performed under the same conditions. 20 mL of 10 mM Tris HCl (pH 8.7) containing 6 M GuHCl was added to the precipitate obtained by sonication, and the mixture was stirred at 25°C for approximately 48 hours (air oxidation). This was purified by Ni-NTA. Subsequently, acetic acid was added to a final concentration of 10%, and the sample was filtered through a 0.2 μm pore size filter and purified by reverse HPLC. The final sample was stored at -30°C. Finally, the purity of the purified protein was verified by reverse HPLC, and the molecular weight was verified by MALDI-TOF MS.
[0068] Figure 16 is a chart of the RP-HPLC purification of the INF-RBD protein in this example. The figure shows the RP-HPLC flow of the protein derived from the major peak. A single peak was obtained when the concentration of mobile phase B was approximately 40%, suggesting that the sample contains only the protein derived from the major peak. Figure 17 is a graph showing the reversibility of the INF-RBD in this example against thermal denaturation. Similar to Test Example 1, the secondary structure content and structural stability were shown by CD measurement.
[0069] Figure 18 is a schematic diagram showing the INF-RBD protein immunization protocol in this embodiment. Tail breeding was performed from mice approximately once a week, and INF-RBD was injected a total of three times: at the start of the study, and after 2 weeks, 4 weeks, and 20 weeks. Figure 19 is a graph showing the progression of antibody titers in mice immunized with the INF-RBD protein in this example. It shows the progression of antibody titers from the first administration until sacrifice, and the TB numbers (TB-1 to 20, etc.) indicate the number of times tail breeding was performed (i.e., approximately every week).
[0070] Figure 20 is a graph showing the evaluation of serum specificity of INF-RBD protein in this example. H3N2wt and lysozyme were used as immobilized antigens for ELISA. M1 to M5 show the results for each mouse (n=5), and H3N2wt-Mean shows the average of the results from these 5 mice. In all test groups, antibody titers were overwhelmingly higher when H3N2wt was used as the immobilized antigen.
[0071] Although embodiments of the present invention have been described above, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents. [Industrial applicability]
[0072] According to the present invention, by appropriately designing the protein domain, it is possible to provide a method for producing a protein domain that has high immunogenicity against viruses and can be produced simply and efficiently while maintaining its natural structure, and a protein domain produced using this method.
Claims
1. A method for producing a protein domain having immunogenicity against viruses, A step of estimating the immunogenic protein domain from the three-dimensional structure predicted from the amino acid sequence of the protein constituting the virus, A step of predicting the physicochemical properties, structural properties, and stability of the estimated protein domain, A step of predicting the degree of exposure of each part of the virus from the three-dimensional structure of the virus, A step of selecting a protein domain to be manufactured based on the physicochemical properties of the protein domain, the structural properties, the stability, and the exposure level of each part of the virus; A method for producing protein domains, including those mentioned above.
2. A protein domain produced by the method for producing a protein domain described in claim 1, The aforementioned virus is the protein domain of SARS coronavirus 2.
3. A protein domain produced by the method for producing a protein domain described in claim 1, The aforementioned virus is the influenza virus, and this is its protein domain.
4. A protein domain produced by the method for producing a protein domain described in claim 1, wherein the protein domain is (i) an amino acid sequence selected from amino acid residues 323 to 538 of the spike protein of SARS coronavirus 2-Omicron strain, an amino acid sequence obtained by substituting at least one cysteine residue of the above amino acid sequence with alanine, or an amino acid sequence containing SEQ ID NO: 1 Or, (II) An amino acid sequence selected from amino acid residues 45 to 279 of the HA1 protein of hemagglutinin of influenza A, an amino acid sequence obtained by substituting at least one arginine residue in the above amino acid sequence with alanine, or an amino acid sequence containing SEQ ID NO: 2 A protein domain comprising an amino acid sequence selected from [a specific source].
Citation Information
Patent Citations
Stabilized coronavirus spike (S) protein immunogens and related vaccines
JP2023533228A