Modified coronavirus S protein
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- アラミス バイオテクノロジーズ インコーポレイテッド
- Filing Date
- 2023-05-01
- Publication Date
- 2026-05-07
AI Technical Summary
The prior art is difficult to effectively improve the purity, uniformity and stability of coronavirus S proteins in hosts or host cells, which affects the production efficiency and quality of VLPs.
The structure of the coronavirus S protein is stabilized and improved by introducing specific amino acid sequence modifications, such as introducing N-glycosylation sites or deleting specific amino acid sequences, thereby improving its purity, uniformity and stability in the host or host cell.
The higher purity, uniformity and stability of S protein are achieved, the content of full-length S protein in VLPs is improved, and the immune response effect of the vaccine is enhanced.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to modified coronavirus S proteins and virus-like particles (VLPs) comprising the modified coronavirus S proteins. The present disclosure also relates to methods for increasing the purity, homogeneity, and / or stability of coronavirus S proteins produced in a host or host cell. [Background technology]
[0002] Coronaviruses (CoVs) are the largest group of viruses in the order Nidovirales, which includes the families Coronaviridae, Arteriviridae, Mesoniviridae, and Roniviridae. The Coronavirinae subfamily includes one of two subfamilies of the Coronaviridae family, the other being the Torovirinae family. The Coronavirinae subfamily is further subdivided into four genera: alpha, beta, gamma, and deltacoronaviruses. Members of the alphacoronavirus and betacoronavirus families are found exclusively in mammals. The alphacoronavirus genus includes two human virus species, HCoV-229E and HCoV-NL63. Important animal alphacoronaviruses are porcine transmissible gastroenteritis virus and feline infectious peritonitis virus.
[0003] Betacoronaviruses of clinical importance in humans include the Embecoviruses OC43 and HKU1 (which can cause the common cold), the Sarbecoviruses SARS-CoV and SARS-CoV-2, and the Merbecovirus MERS-CoV. The Sarbecovirus SARS-CoV-2, also known as 2019-nCoV and HCoV-19, first emerged in 2019 and causes coronavirus disease 2019 (COVID-19), a respiratory illness with high mortality and morbidity and significant public health impact. SARS-CoV-2 outbreaks, such as the COVID-19 pandemic that began in 2020, present challenges for healthcare organizations due to the asymptomatic incubation period and high transmissibility of the virus. Long-term management of SARS-CoV-2 outbreaks requires high-coverage vaccination coverage worldwide with an effective vaccine.
[0004] Since its initial emergence in 2019, SARS-CoV-2 has mutated into numerous additional lineages and sublineages through natural substitutions and insertion-deletion events from ancestral strains. The phylogeny of these SARS-CoV-2 lineages is typically represented by the PANGO nomenclature system (Rambaut et al. 2020). Clinically significant lineages are further noted as variants of interest (VOI) or variants of concern (VOC) based on the risk they pose to global public health and may be assigned names by the World Health Organization (WHO), such as beta lineage (corresponding to B.1.351), gamma lineage (corresponding to P.1), delta lineage (corresponding to B.1.617.2), and omicron lineage (corresponding to B.1.1.529). Mutations vary significantly among all reported SARS-CoV-2 lineages, with potential implications for the efficacy and development of vaccines containing the coronavirus spike (S) protein.
[0005] The globular-rod-shaped S protein is the most prominent structural feature of coronaviruses and is a protrusion arising from the surface of the virion. Coronavirus S protein is a glycoprotein required for host receptor recognition and fusion of the viral membrane with the host cell membrane for viral entry into cells (Belouzard et al., Viruses 2012 Jun;4(6):1011-33). As the primary glycoprotein on the surface of the viral envelope, the S protein of the Coronaviridae family is the primary target of neutralizing antibodies elicited by natural infections, including SARS-CoV-2 infection, and is an important antigen used in coronavirus vaccine formulations.
[0006] The SARS-CoV-2 S protein, like the S proteins of other coronaviruses, is first synthesized as a precursor protein. Individual precursor S proteins form homotrimers and undergo glycosylation and processing to remove the signal peptide in the Golgi compartment. S proteins require two steps of protease-mediated activation to promote membrane fusion. SARS-CoV-2 S proteins are distinguished by a polybasic RRAR furin cleavage site at the S1 / S2 junction, which is presumably processed in the Golgi compartment to yield two distinct polypeptides: the S1 polypeptide (or subunit) and the S2 polypeptide (or subunit), which remain noncovalently associated as S1 / S2 protomers within a homotrimer in a prefusion conformation (Walls et al. Cell 2020 181(2)p281-292; Li et al. eLife 2019;8:e51230). Furin cleavage at the S1 / S2 junction and further cleavage at the S2' site upstream of the fusion peptide occur during viral entry at the cell surface or endosomes and can be mediated by several proteases. Stabilization of the S protein ectodomain in the pre-fusion conformation tends to increase recombinant expression yields, likely by preventing triggering or misfolding resulting from the tendency to adopt a more stable post-fusion structure (Hsieh et al. Science 2020, 369 pp. 1501-1505).
[0007] The S1 domain is further composed of an N-terminal domain (NTD) and a receptor-binding domain (RBD). Neutralizing antibodies from individuals infected with SARS-CoV-2 have been shown to target the RBD of the S1 subunit of the S protein (Premkumar, L., 2020 Science Immunology 11 Jun 2020: Vol. 5, Issue 48). Highly protective antibodies specific to the NTD and targeting the conserved supersite have also been reported (Lok et al. 2021, Cell Host & Microbe 29).
[0008] Vaccination provides protection against disease by inducing a subject to mount an immune response against the same agent before infection. Traditionally, this has been achieved by using live, attenuated, or completely inactivated forms of infectious agents as immunogens. To avoid the drawbacks of using whole viruses (such as killed or attenuated viruses) to create vaccines, viral proteins or subunits, or their recombinant forms, have been pursued as vaccines. The main obstacle to using natural or recombinant viral proteins as vaccine agents is ensuring that the conformation of the protein mimics the antigen in its natural environment. To enhance the immune response, suitable adjuvants, and in the case of peptides, carrier proteins, can be used. In addition, viral proteins or subunits as vaccines primarily induce humoral responses and may not induce long-lasting immunity. Subunit vaccines may be ineffective against diseases for which whole inactivated viruses may be demonstrated to provide excellent protection.
[0009] Virus-like particles (VLPs) can be used in immunogenic compositions to express viral proteins in a preferred conformation that improves antigen presentation to the immune system. VLPs closely resemble mature virions, but they do not contain viral genomic material and are non-replicative, making them safe for administration as vaccines. Additionally, VLPs can be engineered to express viral glycoproteins on their surface, their natural physiological configuration. Because VLPs resemble intact virions and are multivalent particle structures, VLPs may be more effective at inducing neutralizing antibodies against glycoproteins than soluble envelope protein antigens.
[0010] VLPs self-assemble (in vivo assembly) from single or multiple viral structural proteins, such as the coronavirus S protein, in a suitable production host. Therefore, coronavirus VLPs can be produced by expressing recombinant coronavirus S protein in a host. However, the yield, homogeneity, and overall quality of the recombinant S protein can be affected by degradation of the recombinant protein within the expression host or expression host cells and / or during subsequent purification of the protein. Conventional strategies for minimizing proteolysis within a host, such as a plant, include organ-specific transgene expression, organelle-specific protein targeting, grafting stabilizing protein domains into unstable proteins, protein secretion in native fluids, and coexpression of companion protease inhibitors. While rational mutagenesis approaches may be possible for proteins for which precise information about susceptible cleavage sites is available, in most cases proteolysis occurs too rapidly to identify the initial cleavage point.
[0011] Effective scale-up and production of coronavirus VLPs in the quantities necessary to achieve widespread vaccination of the world's population requires efficient expression of coronavirus S protein of high quality, stability, and purity. Summary of the Invention
[0012] The yield, homogeneity and overall quality of the recombinant protein can be affected by degradation of the recombinant protein within the expression host or expression host cells and / or during subsequent purification of the protein.
[0013] The present disclosure provides modified coronavirus spike proteins (S proteins) that contain one or more amino acid sequence modifications when compared to the corresponding parent or unmodified amino acid sequence. The modified S proteins have improved characteristics when compared to the unmodified S protein, such as increased integrity, increased stability, increased resistance to degradation or proteolytic cleavage, increased purity and homogeneity when extracted and / or purified from a host or host cell, or a combination thereof.
[0014] In one aspect, a modified coronavirus S protein is provided, the modified coronavirus S protein comprising one or more amino acid sequence modifications when compared to a corresponding parent amino acid sequence, the one or more modifications stabilize the modified coronavirus S protein, and the one or more modifications i) the substitution of one or more amino acids to introduce an N-glycosylation site at a position corresponding to position 251, 252 or 253 of the reference sequence SEQ ID NO: 1, wherein the N-glycosylation site is an asparagine (N) within the consensus sequence NX-(S or T); or ii) A deletion of at least four consecutive amino acid residues, including at least the residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO:1.
[0015] i) The modified coronavirus S protein comprising the substitution of one or more amino acids to introduce an N-glycosylation site may further comprise the deletion of one or more amino acids.
[0016] The one or more deletions in the modified S protein may include the following deletions: i) at least amino acid residues corresponding to positions 247, 248, 249 and 250 of reference sequence SEQ ID NO: 1; ii) at least amino acid residues corresponding to positions 248, 249, 250 and 251 of reference sequence SEQ ID NO: 1; iii) at least amino acid residues corresponding to positions 249, 250, 251 and 252 of reference sequence SEQ ID NO: 1; iv) at least amino acid residues corresponding to positions 246, 247, 248, 249 and 250 of reference sequence SEQ ID NO: 1; v) at least amino acid residues corresponding to positions 247, 248, 249, 250 and 251 of the reference sequence SEQ ID NO: 1; vi) at least amino acid residues corresponding to positions 248, 249, 250, 251 and 252 of the reference sequence SEQ ID NO: 1; vii) at least amino acid residues corresponding to positions 246, 247, 248, 249, 250 and 251 of reference sequence SEQ ID NO: 1; viii) at least amino acid residues corresponding to positions 247, 248, 249, 250, 251 and 252 of the reference sequence SEQ ID NO: 1; or ix) At least the amino acid residues corresponding to positions 246, 247, 248, 249, 250, 251 and 252 of the reference sequence SEQ ID NO: 1.
[0017] The modified coronavirus S protein may include i) one or more amino acid substitutions to introduce N-glycosylation sites and may further include one or more amino acid deletions; i) an N-glycosylation site is introduced at the amino acid corresponding to position 251 of the reference sequence SEQ ID NO: 1, and the deletion includes at least the amino acid residues corresponding to positions 249 and 250 of the reference sequence SEQ ID NO: 1; ii) an N-glycosylation site is introduced at the amino acid corresponding to position 252 of reference sequence SEQ ID NO: 1, and the deletion includes at least amino acid residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1; iii) an N-glycosylation site is introduced at the amino acid corresponding to position 253 of reference sequence SEQ ID NO: 1, and the deletion includes at least amino acid residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1; iv) an N-glycosylation site is introduced at the amino acid corresponding to position 253 of the reference sequence SEQ ID NO: 1, wherein the deletion comprises at least four consecutive amino acid residues, and wherein the deletion comprises at least residues corresponding to positions 249 and 250 of the reference sequence SEQ ID NO: 1; v) an N-glycosylation site is introduced at the amino acid corresponding to position 253 of the reference sequence SEQ ID NO:1, and the deletion includes at least the amino acid residues corresponding to positions 246, 247, 248, 249, 250, 251 and 252 of the reference sequence SEQ ID NO:1.
[0018]
[0019] The modified coronavirus S protein may include a substitution of asparagine (N) at a position corresponding to position 252 or 253 of the reference sequence SEQ ID NO:1, or the amino acid sequence modification includes a substitution of asparagine (N) at a position corresponding to position 251 of the reference sequence SEQ ID NO:1 and a substitution of threonine (T) at a position corresponding to position 253.
[0020] In a further embodiment, the modified S protein may comprise 80% to 100% identity to the sequence of SEQ ID NO: 8, 10, 12, 16, 20, 39, 41, 43, 45, 47, 49, 51, 53, 55 or 57.
[0021] The modified coronavirus S protein can be a chimeric S protein, which includes a cytoplasmic tail derived from influenza hemagglutinin. In one embodiment, the coronavirus S protein is derived from a betacoronavirus. For example, the coronavirus S protein can be derived from betacoronavirus lineage A, B, C, or D. In one embodiment, the coronavirus S protein can be derived from betacoronavirus lineage B. Additionally, the modified S protein can include plant-specific N-glycans.
[0022] Further provided is a genetic construct or nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above.
[0023] In another aspect, one or more virus-like particles (VLPs) are provided that contain the modified S protein described above. The VLPs contain a greater amount of full-length S protein than VLPs assembled from S proteins that are not modified as described herein. The VLPs may further contain plant lipids.
[0024] In a further aspect, a composition is provided comprising a pharmaceutically acceptable carrier, vehicle or excipient and an effective dose of a modified S protein described herein or a VLP comprising a modified S protein described herein.
[0025] In a further aspect, a vaccine for inducing an immune response is provided. The vaccine may comprise an effective dose of the modified S protein described herein, a VLP comprising the modified S protein described herein, or a composition described herein. The vaccine may further comprise an adjuvant such as AS03.
[0026] In yet another embodiment, the vaccine may be a multivalent vaccine comprising a mixture of VLPs.
[0027] In a further aspect, there is provided a method of inducing immunity to coronavirus infection in a subject, the method comprising administering to the subject a composition or vaccine as described above.
[0028] In yet another aspect, there is provided a method of inducing an immune response in a subject, comprising administering to the subject the composition or vaccine described above. Also provided are antibodies or antibody fragments prepared using the described compositions or vaccines.
[0029] In a further aspect, there is provided a use of the above-described composition or vaccine for inducing immunity to coronavirus infection in a subject. There is also provided a use of the above-described composition or vaccine for inducing an immune response in a subject.
[0030] In another aspect, a host or host cell is provided that comprises the modified S proteins, constructs, nucleic acids and / or VLPs described herein. The host or host cell may comprise a plant, a plant part, a plant cell, a fungus, a fungal cell, an insect, an insect cell, an animal or an animal cell.
[0031] In a further aspect, a method for producing a modified S protein in a host or host cell is provided, comprising: a) introducing into the host or host cell a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above, or providing a host or host cell comprising a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above, and b) incubating the host or host cell under conditions allowing expression of the nucleic acid, thereby producing the modified S protein.
[0032] Also provided is a method of producing a modified S protein in a host or host cell, comprising: a) expressing a modified S protein described herein in the host or host cell by incubating the host or host cell under conditions that allow expression of the modified S protein, thereby producing the modified S protein. The modified S protein can be further extracted and purified from the host or host cell.
[0033] In a further aspect, a method for producing virus-like particles (VLPs) in a host or host cell is provided, comprising: a) introducing into the host or host cell a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above, or providing a host or host cell comprising a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above; and b) incubating the host or host cell under conditions allowing expression of the nucleic acid, thereby producing VLPs. The method may further comprise a step c) of harvesting the host or host cell. The VLPs may be further extracted and purified from the host or host cell. VLPs produced by the method are also provided.
[0034] In yet another aspect, a method is provided for increasing production of a full-length coronavirus S protein in a host or host cell, comprising: a) introducing into the host or host cell a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above, or providing a host or host cell comprising a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described above, and b) incubating the host or host cell under conditions allowing expression of the nucleic acid, thereby producing the modified S protein, wherein a greater amount or a greater proportion of the modified S protein is full-length modified S protein compared to unmodified S protein produced under similar conditions in the host or host cell.
[0035] In a further aspect, a method for producing a modified coronavirus S protein in a host or host cell that has increased proteolytic stability is provided, comprising: a) introducing into the host or host cell a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein, or providing a host or host cell that comprises a nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein; b) incubating the host or host cell under conditions that allow expression of the nucleic acid, thereby producing a modified S protein that has increased proteolytic stability compared to the proteolytic stability of an unmodified S protein produced in the host or host cell under similar conditions; and c) optionally extracting the modified S protein from the host or host cell. The modified S protein can be further purified from the host or host cell. Modified S proteins produced by the above methods are also provided. The modified S protein may exhibit increased proteolytic stability compared to an S protein that is not modified as described herein. One or more VLPs comprising the modified S protein are also provided.
[0036] Additionally, in a further aspect, a method is provided for producing virus-like particles (VLPs) with increased full-length S protein content in a host or host cell, the method comprising: a) introducing into a host or host cell, or providing to a host or host cell, a nucleic acid comprising a sequence encoding a modified S protein described herein; b) incubating the host or host cell under conditions allowing expression of the nucleic acid, thereby producing VLPs, wherein the VLPs have increased full-length S protein content compared to VLPs comprising an unmodified S protein produced in the host or host cell under similar conditions; and c) optionally extracting the VLPs from the host or host cell. Furthermore, the VLPs can be purified from the host or host cell. Virus-like particles produced by the method are also provided.
[0037] In another aspect, there is provided a method for increasing the proteolytic stability of a host cell or a coronavirus S protein produced in a host cell, comprising: a) modifying a parent coronavirus S protein sequence to produce a modified coronavirus S protein having a modified sequence, wherein the modified coronavirus S protein comprises one or more amino acid sequence modifications when compared to the parent coronavirus S protein, wherein the one or more modifications are: i) the substitution of one or more amino acids to introduce an N-glycosylation site at a position corresponding to position 251, 252 or 253 of the reference sequence SEQ ID NO: 1, wherein the N-glycosylation site is an asparagine (N) within the consensus sequence NX-(S or T); or ii) a deletion of at least four consecutive amino acid residues, including at least the residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1; b) expressing the modified coronavirus S protein in a host or host cell, thereby producing a modified S protein that has increased proteolytic stability compared to the proteolytic stability of the parent coronavirus S protein produced under similar conditions in the host or host cell.
[0038] 1. A method for modifying a coronavirus S protein to produce a modified coronavirus S protein having one or more amino acid sequence modifications, wherein the one or more amino acid modifications stabilize the modified coronavirus S protein, the method comprising: i) introducing one or more amino acid substitutions into the coronavirus S protein to introduce an N-glycosylation site at a position corresponding to position 251, 252 or 253 of the reference sequence SEQ ID NO: 1, wherein the N-glycosylation site is an asparagine (N) within the consensus sequence NX-(S or T); or ii) introducing a deletion of at least four consecutive amino acid residues into the coronavirus S protein, the deletion including at least the residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1, thereby modifying the coronavirus S protein.
[0039] Also provided are modified coronavirus S proteins and VLPs comprising the modified coronavirus S proteins produced by the methods described.
[0040] This summary of the invention does not necessarily describe all features of the invention. [Brief explanation of the drawings]
[0041] [Figure 1A] FIG. 1 shows a schematic diagram of the acceptor vector 8716 used to assemble vector plasmids encoding modified coronavirus S proteins. [Figure 1B]A schematic diagram of vector 9125 encoding a modified S protein from SARS-CoV-2 strain B carrying the GSAS+2P mutation is shown. [Figure 1C] A schematic diagram of vector 9801 encoding a modified S protein from SARS-CoV-2 B strain carrying the GSAS+2P+del246-252 mutations is shown. [Figure 1D] A schematic diagram of vector 9802 encoding a modified S protein from SARS-CoV-2 B strain, carrying the GSAS+2P+D253N mutations, is shown. [Figure 1E] A schematic diagram of vector 9808 encoding a modified S protein from SARS-CoV-2 B strain, carrying GSAS+2P+del246-252+D253N mutations. [Figure 1F] A schematic diagram of vector 9513 encoding a modified S protein from the SARS-CoV-2 B.1.617.2 strain, carrying the GSAS+2P mutation, is shown. [Figure 1G] A schematic diagram of vector 10090 encoding a modified S protein from SARS-CoV-2 B.1.1.529 strain with the GSAS+2P mutation is shown. [Figure 1H] A schematic diagram of vector 10346 encoding a modified S protein from SARS-CoV-2 B strain with GSAS+2P+G252N mutations is shown. [Figure 1I] A schematic diagram of vector 10351 encoding a modified S protein from SARS-CoV-2 B strain with GSAS+2P+P251N+D253T mutations is shown. [Figure 1J] A schematic diagram of vector 10502 encoding a modified S protein from SARS-CoV-2 B strain carrying GSAS+2P+del247-250 mutations is shown. [Figure 1K] A schematic diagram of vector 10503 encoding a modified S protein from SARS-CoV-2 B strain carrying the GSAS+2P+del248-251 mutations is shown. [Figure 1L]A schematic diagram of vector 10504 encoding a modified S protein from SARS-CoV-2 strain B carrying the GSAS+2P+del249-252 mutations is shown. [Figure 1M] A schematic diagram of vector 10505 encoding a modified S protein from SARS-CoV-2 B strain carrying GSAS+2P+del246-250 mutations is shown. [Figure 1N] A schematic diagram of vector 10506 encoding a modified S protein from SARS-CoV-2 B strain carrying the GSAS+2P+del247-251 mutations is shown. [Figure 1O] A schematic diagram of vector 10507 encoding a modified S protein from SARS-CoV-2 B strain carrying the GSAS+2P+del248-252 mutations is shown. [Figure 1P] A schematic diagram of vector 10508 encoding a modified S protein from SARS-CoV-2 B strain carrying GSAS+2P+del246-251 mutations is shown. [Figure 1Q] A schematic diagram of vector 10509 encoding a modified S protein from SARS-CoV-2 B strain carrying the GSAS+2P+del247-252 mutations is shown. [Figure 1R] A schematic diagram of vector 10011 encoding a modified S protein from SARS-CoV-2 B.1.617.2 strain, carrying the GSAS+2P+D253N mutations, is shown. [Figure 1S] A schematic diagram of vector 10092 encoding a modified S protein from SARS-CoV-2 B.1.1.529 strain, carrying the GSAS+2P+D253N mutations, is shown.
[0042] [Figure 2A] Electron micrograph of a virus-like particle (VLP) containing a modified S protein from SARS-CoV-2 strain B, carrying the GSAS+2P mutation (construct 9125). [Figure 2B]Electron micrograph of a virus-like particle (VLP) containing a modified S protein from SARS-CoV-2 strain B carrying the GSAS+2P+del246-252 mutation (GSAS+2P+del246-252; construct 9801) is shown. [Figure 2C] Electron micrograph of a virus-like particle (VLP) containing a modified S protein from SARS-CoV-2 strain B, carrying the GSAS+2P+D253N mutation (GSAS+2P+D253N; construct 9802). [Figure 2D] Electron micrograph of a virus-like particle (VLP) containing a modified S protein from SARS-CoV-2 strain B, carrying the GSAS+2P+del246-252+D253N mutation (GSAS+2P+del246-252+D253N; construct 9808). [Figure 2E] Electron micrograph of a virus-like particle (VLP) containing a modified S protein from the SARS-CoV-2 C.37 strain, carrying the GSAS+2P mutation (construct 9588).
[0043] [Figure 3A] Shown are the percentages of in planta intact S proteins expressing the following S proteins: parental SARS-CoV-2 B strain S protein with the GSAS+2P mutation (Wt; construct 9125), modified S protein with the GSAS+2P+del246-252 mutation (GSAS+2P+del246-252; construct 9801), modified S protein with the GSAS+2P+D253N mutation (GSAS+2P+D253N; construct 9802), modified S protein with the GSAS+2P+del246-252+D253N mutation (GSAS+2P+del246-252+D253N; construct 9808), and the S protein from the SARS-CoV-2 C.37 strain with the GSAS+2P mutation (construct 9588). [Figure 3B]Shown are the percentages of implanters expressing intact S proteins for the following S proteins: parental SARS-CoV-2 B lineage S protein with a GSAS+2P mutation (construct 9125), modified SARS-CoV-2 B lineage S protein with a GSAS+2P+D253N mutation (construct 9802), parental SARS-CoV-2 B.1.617.2 lineage S protein with a GSAS+2P mutation (construct 9513), modified SARS-CoV-2 B.1.617.2 lineage S protein with a GSAS+2P+D253N mutation (construct 10011), parental SARS-CoV-2 B.1.1.529 lineage S protein with a GSAS+2P mutation (construct 10090), and modified SARS-CoV-2 B.1.1.529 lineage S protein with a GSAS+2P+D253N mutation (construct 10092).
[0044] [Figure 4A] Purity of drug substance (DS) obtained from hosts expressing the parent SARS-CoV-2 B strain S protein with the GSAS+2P mutation (construct 9125) or the modified SARS-CoV-2 B strain S protein with the GSAS+2P+D253N mutation (construct 9802) is shown. [Figure 4B] Purity of drug substance (DS) obtained from hosts expressing the following constructs is shown: parental S protein from SARS-CoV-2 B strain with GSAS+2P mutation (Wt; construct 9125), modified S protein with GSAS+2P+del246-252 mutation (GSAS+2P+del246-252; construct 9801), modified S protein with GSAS+2P+D253N mutation (GSAS+2P+D253N; construct 9802), modified S protein with GSAS+2P+del246-252+D253N mutation (GSAS+2P+del246-252+D253N; construct 9808), and S protein from SARS-CoV-2 C.37 strain with GSAS+2P mutation (construct 9588).
[0045] [Figure 5]Figure 1 shows the increased stability of purified modified S protein (SARS-CoV-2 B strain with GSAS+2P+D253N mutations: construct 9802) compared to unmodified (parent) S protein (SARS-CoV-2 B strain with GSAS+2P mutations: construct 9125) after overnight incubation at 24°C. Modified and unmodified S proteins were extracted using enzymatic extraction at pH 5.5 (Enz. pH5.5) or pH 6.1 (Enz. pH6.1) or by mechanical extraction (Mech).
[0046] [Figure 6A] Western blot analysis of the following S proteins is shown: parental SARS-CoV-2 B lineage S protein with GSAS+2P mutations (construct 9125), modified SARS-CoV-2 B lineage S protein with GSAS+2P+D253N mutations (construct 9802), modified SARS-CoV-2 B lineage S protein with GSAS+2P+G252N mutations (construct 10346), and modified SARS-CoV-2 B lineage S protein with GSAS+2P+P251N+D253T mutations (construct 10351). [Figure 6B] Shown are the percentages of implants expressing intact S proteins for the following S proteins: parental SARS-CoV-2 B lineage S protein with GSAS+2P mutations (construct 9125), modified SARS-CoV-2 B lineage S protein with GSAS+2P+D253N mutations (construct 9802), modified SARS-CoV-2 B lineage S protein with GSAS+2P+G252N mutations (construct 10346), and modified SARS-CoV-2 B lineage S protein with GSAS+2P+P251N+D253T mutations (construct 10351). [Figure 6C]Electron micrographs of virus-like particles (VLPs) containing modified SARS-CoV-2 B lineage S protein with GSAS+2P+D253N mutations (construct 9802), modified SARS-CoV-2 B lineage S protein with GSAS+2P+G252N mutations (construct 10346), and modified SARS-CoV-2 B lineage S protein with GSAS+2P+P251N+D253T mutations (construct 10351) are shown.
[0047] [Figure 7] The percentage of implants expressing intact S proteins is shown: parental SARS-CoV-2 B lineage S protein with GSAS+2P mutation (construct 9125); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del246-252 (construct 9801); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del247-250; (construct 10502); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del248-251; (construct 10503); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del249-252; (construct 10504); modified SARS-CoV-2 with GSAS+2P+del246-250. B lineage S protein; (construct 10505); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del247-251; (construct 10506); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del248-252; (construct 10507); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del246-251; (construct 10508); modified SARS-CoV-2 B lineage S protein with GSAS+2P+del247-252; (construct 10509). DETAILED DESCRIPTION OF THE INVENTION
[0048] The following description is of a preferred embodiment.
[0049] As used herein, the terms "comprising," "having," "including," and "containing," and grammatical variations thereof, are inclusive or open-ended and do not exclude additional, unrecited elements and / or method steps. The term "consisting essentially of," when used herein in connection with a use or method, indicates that additional elements and / or method steps may be present, but that these additions do not materially affect the manner in which the recited method or use functions. The term "consisting of," when used herein in connection with a use or method, excludes the presence of additional elements and / or method steps. A use or method described herein as including particular elements and / or steps may also, in certain embodiments, consist essentially of these elements and / or steps, regardless of whether these embodiments are specifically mentioned, and in other embodiments, consist of these elements and / or steps. Additionally, the use of the singular includes the plural, and "or" means "and / or" unless otherwise stated. The term "plurality" as used herein means a plurality, for example, two or more, three or more, four or more, etc. Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. As used herein, the term "about" refers to a variation of about + / - 10% from a given value. It should be understood that such a variation is always included in any given value provided herein, regardless of whether it is specifically mentioned. The use of the word "a" or "an" when used herein in conjunction with the term "comprising" can mean "one," but is also consistent with the meaning of "one or more," "at least one," and "one or more."
[0050] Described herein are modified coronavirus spike proteins (S proteins), also referred to as "modified coronavirus S proteins" or "modified S proteins," and methods for producing the modified S proteins in a host or host cell. The modified S proteins can contain one or more modifications compared to the parent (unmodified) S protein or the wild-type S protein. Modifications, such as the substitution or deletion of specific amino acids, in coronavirus S proteins, e.g., S proteins from the B strain of SARS-CoV-2, have been observed to result in improved characteristics of the modified S proteins when compared to the parent (unmodified) S protein or the wild-type (unmodified) S protein.
[0051] "Modification," "amino acid modification," or "amino acid sequence modification" refers to the mutation, substitution, replacement, or deletion of one or more amino acid residues in a sequence compared to the original parent (unmodified) sequence. The parent sequence may be a wild-type sequence, or the parent sequence may be a sequence that already contains a modification ("parent modification") when compared to the wild-type sequence. "Amino acid substitution" or "substitution" refers to the replacement of an amino acid in the amino acid sequence of a protein with a different amino acid. The terms amino acid, amino acid residue, or residue are used interchangeably in this disclosure. One or more amino acids may be replaced with one or more amino acids different from the original amino acid at that position without changing the full length of the amino acid sequence of the protein. The substitution or replacement may be experimentally induced by changing the codon sequence in the nucleotide sequence encoding the protein to a codon sequence for a different amino acid compared to the original amino acid. Furthermore, one or more amino acids may be deleted from the amino acid sequence of the protein. The resulting protein is a modified S protein. The modified S protein does not occur in nature.
[0052] Modified S proteins include non-naturally occurring S proteins that have at least one modification compared to the parent S protein and have improved characteristics compared to the parent S protein from which the amino acid sequence of the modified S protein is derived. Modified S proteins have an amino acid sequence not found in nature that is derived by replacing one or more amino acid residues of the S protein with one or more different amino acids.
[0053] The parent S protein may also be referred to as an unmodified S protein. When the parent is referred to as "unmodified," it means that the parent sequence does not contain the substitutions and / or deletions described herein. However, the parent S protein may contain other modifications compared to the wild-type sequence. In some embodiments, the parent S protein or unmodified S protein may be wild-type HA. In other embodiments, the parent S protein or unmodified S protein may contain other modifications as described below. For example, the parent S protein may contain one or more substitutions or substitutions to stabilize the coronavirus S protein or coronavirus S protein trimer in the pre-fusion conformation. Furthermore, the parent S protein may be a chimeric S protein. For example, the ectodomain and transmembrane domain (TM) or portion of the TM of the parent S protein may be derived from a coronavirus S protein (such as SARS-CoV 2), and the cytoplasmic tail (CT) or portion of the CT may be derived from influenza HA.
[0054] "Parent S protein" refers to an S protein from which a modified S protein can be derived. For example, a parent S protein can be modified to produce a modified S protein having the modifications described herein. As described further below, the parent S protein can be derived from a first variant or strain (also referred to as an "acceptor" variant or strain), e.g., a coronavirus of the coronavirus B lineage, and one or more modifications can be derived or determined from an S protein from a second variant or strain (also referred to as a "donor" variant or strain), e.g., a coronavirus from the coronavirus C lineage.
[0055] Some of the residues identified for modification, mutation, or substitution correspond to conserved residues, while others do not. For non-conserved residues, the substitution of one or more amino acids is limited to substitutions that produce a modified S protein with an amino acid sequence that does not correspond to that found in nature. For conserved residues, such modifications, substitutions, or replacements should also not result in a naturally occurring S protein sequence.
[0056] Saved Substitutions As described herein, residues in the S protein can be identified and modified, substituted, or mutated to produce modified S proteins. Substitutions or mutations at specific positions are not limited to the amino acid substitutions described herein or shown in examples. For example, the S protein can contain conservative substitutions or conservative substitutions of the amino acid substitutions described.
[0057] As used herein, the term "conserved substitution" or "conservative substitution" and grammatical variations thereof refer to the presence of an amino acid residue within the sequence of the S protein that is different from the described substitution or described residue (i.e., a non-polar residue replacing a non-polar residue, an aromatic residue replacing an aromatic residue, a polar uncharged residue replacing a polar uncharged residue, a charged residue replacing a charged residue), but is of the same class of amino acid as the described residue. In addition, conservative substitutions may include residues having the same sign and generally similar magnitude of interface hydropathy value as the residue replacing the wild-type residue.
[0058] Conservative amino acid substitutions are likely to have the same effect on the activity of the resulting modified S protein as the original substitution or modification. Further information on conservative substitutions can be found, for example, in Ben Bassat et al. (J. Bacteriol, 169: 751-757, 1987), O'Regan et al. (Gene, 77: 237-251, 1989), Sahin-Toth et al. (Protein ScL, 3: 240-247, 1994), Hochuli et al. (Bio / Technology, 6: 1321-1325, 1988) and widely used textbooks on genetics and molecular biology.
[0059] The Blosum matrix is commonly used to determine the relatedness of polypeptide sequences. A large database of trusted alignments (the BLOCKS database) was used to generate the Blosum matrix, which counts pairwise sequence alignments related by less than some threshold percent identity (Henikoff et al., Proc. Natl. Acad. Sci. USA, 89:10915-10919, 1992). A threshold of 90% identity was used for highly conserved target frequencies in the BLOSUM90 matrix. A threshold of 65% identity was used for the BLOSUM65 matrix. A score of 0 or greater in the Blosum matrix is considered a "conservative substitution" at the selected percent identity.
[0060] Thus, the present specification relates to modified coronavirus spike proteins (S proteins) that contain one or more amino acid sequence modifications when compared to the corresponding parent or unmodified amino acid sequence. It has been discovered that naturally occurring sequence mutations or modifications specific to the S protein of a coronavirus mutant or strain can confer desirable or improved characteristics to S proteins that do not naturally possess these mutations or modifications, e.g., S proteins from different strains or mutants. It has further been discovered that non-naturally occurring sequence mutations or modifications within the N-terminal region of an S protein can also confer desirable or improved characteristics to S proteins. Thus, modified S proteins can contain mutations or modifications from S proteins from different strains or mutants, and / or can contain modifications that are not naturally occurring, i.e., not found in different strains or mutants, and the S protein can exhibit improved characteristics compared to wild-type or unmodified S proteins.
[0061] In one aspect, the modified S protein comprises one or more amino acid sequence modifications when compared to the corresponding parent amino acid sequence, wherein the one or more modifications correspond to amino acids at positions 246, 247, 248, 249, 250, 251, 252, 253 or combinations thereof of reference sequence SEQ ID NO:1.
[0062] In one embodiment, the modifications described herein can include one or more deletions in the N-terminal region of the protein. For example, the modified S protein can include one or more deletions corresponding to amino acids at positions 246, 247, 248, 249, 250, 251, 252, or a combination thereof, of the reference sequence SEQ ID NO: 1.
[0063] In another aspect, the modifications described herein may introduce one or more N-glycosylation sites into the modified S protein. Thus, the modified S protein, when compared to the parent (unmodified) S protein, may contain one or more N-glycosylation sites, where the N-glycosylation sites are asparagine (N) within the consensus sequence NXT / S (wherein X is any amino acid except proline). The N-glycosylation sites may be introduced at positions corresponding to positions 251, 252, or 253 of the reference sequence SEQ ID NO: 1.
[0064] Examples of improved characteristics of the modified S protein include, but are not limited to, increased integrity, increased stability, increased resistance to degradation or proteolytic cleavage or proteolysis, increased purity and homogeneity, or combinations thereof, of the recombinant modified S protein when expressed in a host or host cell compared to the unmodified S protein; improved integrity, stability, or both integrity and stability of the modified S protein when expressed in a host or host cell compared to the unmodified S protein; increased or improved resistance to degradation, proteolysis, cleavage, or hydrolysis (also known as "clipping") of the modified S protein when expressed in a host or host cell compared to the unmodified S protein; reduced heterogeneity and / or truncation of the modified S protein when expressed in a host or host cell compared to the unmodified S protein; and improved processing and / or folding of the modified S protein when expressed in a host or host cell compared to the unmodified S protein.
[0065] Coronavirus S proteins (e.g., parent S protein, wild-type S protein, or unmodified S protein) that can be modified as described herein to improve S protein characteristics, for example, having increased stability, integrity, purity, homogeneity, or a combination thereof, include new coronavirus S proteins that emerge over time due to naturally occurring modifications of the S protein amino acid sequence (e.g., new S protein variants as described below), or non-native S proteins that can be produced as a result of changes to the S protein (e.g., chimeric S proteins, or S proteins altered to achieve desirable properties, e.g., increased expression in a host or stabilization of the S protein in a pre-fusion conformation). Similarly, the modified S proteins described herein can be derived from wild-type S proteins, new S proteins that emerge over time due to naturally occurring modifications of the S amino acid sequence, unmodified S proteins, non-native S proteins, e.g., chimeric S proteins, or S proteins altered to achieve desirable properties, e.g., increased expression of the S protein or increased production of VLPs comprising the S protein in a host.
[0066] The modified S proteins of the present disclosure can include one or more modifications derived from an S protein from a coronavirus from a different variant or strain compared to the modified S protein. The modified S protein can be from a first variant or strain (also referred to as an "acceptor" variant or strain), e.g., a coronavirus of the coronavirus B lineage, and the one or more modifications can be derived or determined from an S protein from a second variant or strain (also referred to as a "donor" variant or strain), e.g., a coronavirus from the coronavirus C lineage.
[0067] A coronavirus variant (e.g., a SARS-CoV-2 variant) refers to a variant specimen of a coronavirus (also called a coronavirus genetic variant) that contains one or more mutations when compared to the original or ancestral virus of other viruses, e.g., SARS-CoV-2. Generally, a coronavirus variant has one or more mutations that distinguish it from other variants of the virus. A coronavirus genetic variant is genetically different from other variants, but not different enough to be called a different viral strain.
[0068] There are thousands of variants of coronaviruses, such as SARS-CoV-2, but viral subtypes can be classified into lineages or even larger groups such as subgenera or clades.
[0069] A lineage, subgenus, or clade is a genetically closely related group of viral variants derived from a common ancestor. Various nomenclatures can be used for coronavirus lineages, subgenuses, or clades, such as those identified by the Global Initiative on Sharing Avian Influenza Data (GISAID) or "Nextstrain" (Hadfield et al., Nextstrain: Real-time Tracking of Pathogen Evolution, Bioinformatics (2018)), or those defined by the PANGO nomenclature (see Rambaut et al., 2020, incorporated by reference), or by lineages defined by systems of phylogenetic nomenclature known in the art. In addition, the World Health Organization (WHO) has adopted a nomenclature limited to variants of interest (VOI) and variants of concern (VOC), in which variants are named using letters of the Greek alphabet, to facilitate discussion and communication with non-scientific audiences. For example, alpha refers to B.1.1.7 (Pango lineage), beta refers to B1.1.351 (Pango lineage), gamma refers to P.1 (Pango lineage), delta refers to B.1.617.2 (Pango lineage), lambda refers to C.37 (Pango lineage), or omicron refers to B.1.1.529 (Pango lineage) (see https: / / www.who.int / en / activities / tracking-SARS-CoV-2-variants / ).
[0070] The donor or acceptor strain may be a SARS-CoV-2 strain, which may be defined by the PANGO nomenclature (see Rambaut et al. 2020, incorporated by reference), or another system of phylogenetic nomenclature.
[0071] For example, the donor coronavirus strain can be selected from the A, B, C, Q, L, D, P, N, S, AE, AF, AZ, AY, BA, W, Y, or Z strains, as defined by the PANGO nomenclature. For example, the donor coronavirus strain can be a C strain. For example, the donor coronavirus strain can be a C.37 ("lambda") strain.
[0072] For example, the acceptor coronavirus strain can be selected from the A, B, C, Q, L, D, P, N, S, AE, AF, AZ, AY, BA, W, Y, or Z strains, as defined by the PANGO nomenclature. For example, the acceptor coronavirus strain can be the B strain. For example, the acceptor coronavirus strain can be selected from the B strain, B.1.1.7 ("alpha") strain, B.1.351 ("beta") strain, P.1 ("gamma") strain, B.1.617.2 ("delta") strain, B.1.1.529 ("omicron") strain, or another SARS-CoV-2 strain. For example, the acceptor coronavirus strain can be of any coronavirus strain, so long as the strain is different from the donor strain. For example, the acceptor coronavirus strain can be of any strain except the C strain. In one embodiment, the modification can be derived from an S protein from the C strain (donor), and the parent S protein being modified is derived from the B strain (acceptor).
[0073] In one embodiment, the modified S protein comprises modifications from a donor strain C.37 ("lambda") S protein that are introduced into an acceptor strain B coronavirus S protein and expressed in a host or host cell, such as a plant or plant cell.
[0074] For example, the C.37 ("lambda") strain contains the following modifications when compared to the S protein of the ancestral B strain (P0DTC2, SEQ ID NO: 1): G75V, T76I, del246-252, D253N, L452Q, F490S, D614G, and T859N.
[0075] It has been found that introducing modifications of amino acid residues found in the N-terminal domain of the C.37 lineage S protein into the B lineage S protein (the modified S protein) results in improved characteristics of the B lineage S protein when produced in a host, such as a plant, compared to the unmodified (parent) B lineage S protein. It has further been found that the use of a subset of modifications, as further described herein, can also improve the characteristics of the S protein. In another aspect, it has further been found that introducing one or more N-glycosylation sites into the modified S protein also improves the characteristics of the S protein.
[0076] Thus, the modified S protein may contain one or more modifications, mutations or substitutions located in the N-terminal domain of the S protein.
[0077] In one embodiment, the modified S protein may include one or more modifications, mutations or substitutions corresponding to positions 246, 247, 248, 249, 250, 251, 252, 253 or a combination thereof.
[0078] deletion The modified S protein may comprise one or more deletions corresponding to positions 246, 247, 248, 249, 250, 251, or 252 of the reference sequence SEQ ID NO: 1. For example, the modified S protein may comprise a deletion of at least four consecutive amino acid residues, with at least the residues corresponding to positions 249 and 250 of the reference sequence SEQ ID NO: 1 being deleted. For example, the modified S protein may comprise at least deletions of residues corresponding to positions 247, 248, 249, and 250 of the reference sequence SEQ ID NO: 1. In another example, the modified S protein may comprise at least deletions of residues corresponding to positions 248, 249, 250, and 251 of the reference sequence SEQ ID NO: 1. In a further example, the modified S protein may comprise at least deletions of residues corresponding to positions 249, 250, 251, and 252 of the reference sequence SEQ ID NO: 1. In another example, the modified S protein may comprise at least deletions of residues corresponding to positions 246, 247, 248, 249, and 250 of the reference sequence SEQ ID NO: 1. In a further example, the modified S protein may comprise at least a deletion of residues corresponding to positions 247, 248, 249, 250, and 251 of the reference sequence SEQ ID NO: 1. In yet another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 248, 249, 250, 251, and 252 of the reference sequence SEQ ID NO: 1. In yet another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 246, 247, 248, 249, 250, and 251 of the reference sequence SEQ ID NO: 1. In yet another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 247, 248, 249, 250, 251, and 252 of the reference sequence SEQ ID NO: 1. In a further example, the modified S protein may comprise at least a deletion of residues corresponding to positions 246, 247, 248, 249, 250, 251, and 252 of the reference sequence SEQ ID NO: 1. In a non-limiting example, the modified S protein can contain the deletions shown in Table 2.
[0079] As shown in Figure 3A, the percentage of full-length coronavirus S protein to truncated coronavirus S protein ("% full-length S protein") increased from approximately 80% full-length coronavirus S protein for the parent (B acceptor) S protein (construct 9125) to approximately 95% full-length coronavirus S protein for the modified S protein (construct 9801) in which residues corresponding to positions 246-252 of reference sequence SEQ ID NO: 1 were deleted.
[0080] As further shown in Figure 4B, clarification and purification of coronavirus S protein followed by quantification using SDS gels demonstrates an increase in percent drug substance (DS) purity. When modified S protein incorporating the del246-252 (construct 9801) modification from donor strain C.37 ("lambda") S protein was extracted and purified from plants, an increase in DS purity was observed compared to the purity of DS obtained from plants expressing the parent (unmodified) S protein (construct 9125). Notably, the increase in DS purity was also greater for the modified coronavirus S protein incorporating del246-252 (construct 9801) than for the DS obtained from the coronavirus S protein from donor strain C.37 ("lambda", construct 9588) expressed in the same system.
[0081] As further shown in Figure 7, the percentage of full-length coronavirus S protein to truncated coronavirus S protein ("% Full S Protein") was also increased for modified S proteins containing the following deletions when compared to the parental (unmodified) S protein (construct 9125): del246-252 (construct 9801), del247-250 (construct 10502), del248-251 (construct 10503), del249-252 (construct 10504), del246-250 (construct 10505), del247-251 (construct 10506), del248-252 (construct 10507), del246-251 (construct 10508), and del247-252 (construct 10509). More specifically, the ratio of full-length coronavirus S protein to truncated coronavirus S protein ("Percent Full S Protein") increased to approximately 91%-96% for modified S proteins containing the above deletions, compared to approximately 87% for the corresponding parental S protein (construct 9125).
[0082] replacement In one aspect, the modifications described herein may introduce one or more N-glycosylation sites into the modified S protein. Thus, the modified S protein, when compared to the parent (unmodified) S protein, may contain one or more N-glycosylation sites, where the N-glycosylation sites are asparagine (N) within the consensus sequence of NXT / S (where X is any amino acid except proline).
[0083] The modified S protein may contain one or more substitutions corresponding to positions 251, 252, or 253 of the reference sequence SEQ ID NO: 1. The one or more substitutions may introduce an N-glycosylation site into the modified S protein. The N-glycosylation site is an asparagine (N) within the consensus sequence N-X-(S or T), where X can be any amino acid except proline. For example, the modified S protein may contain a substitution corresponding to position 252 or position 253 of the reference sequence SEQ ID NO: 1, or the modified S protein may contain substitutions corresponding to positions 251 and 253 of the reference sequence SEQ ID NO: 1. The substitutions may introduce an N-glycosylation site at a position corresponding to position 251, 252, or 253 of the reference sequence SEQ ID NO: 1.
[0084] For example, to create an N-glycosylation site at position 253, the amino acid residue corresponding to position 253 of Reference Sequence SEQ ID NO: 1 may be modified from aspartic acid (D) to asparagine (N) (D253N). Further, to create an N-glycosylation site at position 252, the amino acid residue corresponding to position 252 of Reference Sequence SEQ ID NO: 1 may be modified from glycine (G) to asparagine (N) (G252N). Additionally, to create an N-glycosylation site at position 251, the residue corresponding to position 251 of Reference Sequence SEQ ID NO: 1 may be modified from proline (P) to asparagine (N) (P251N), and the residue corresponding to position 253 of Reference Sequence SEQ ID NO: 1 may be modified from aspartic acid (D) to threonine (T) (D253T).
[0085] As shown in Figure 3B , an even greater proportion of modified S proteins was observed as full-length S proteins (expressed as “full S protein (%)”) when the S proteins were derived from the B lineage (construct 9802), B.1.617.2 lineage (construct 10011), or B.1.1.529 lineage (construct 10092) and contained a substitution corresponding to position 253 (“modified S proteins”) compared with the corresponding unmodified S proteins (constructs 9125, 9513, and 10090, respectively).
[0086] More specifically, the percentage of full-length coronavirus S protein to truncated coronavirus S protein ("% Intact S Protein") increased to approximately 93%-96% for modified S proteins containing the D253N substitution and derived from the B lineage (construct 9802), B.1.617.2 lineage (construct 10011), or B.1.1.529 lineage (construct 10092), compared to approximately 81%-82% for the corresponding unmodified S proteins (constructs 9125, 9513, and 10090, respectively).
[0087] Furthermore, as shown in Figures 4A and 4B, clarification and purification of coronavirus S protein followed by quantification using SDS gels demonstrates an increase in percent drug substance (DS) purity. When modified S protein incorporating the D253N (construct 9802) modification from donor strain C.37 ("lambda") S protein was extracted and purified from plants, an increase in DS purity was observed compared to the purity of DS obtained from plants expressing the unmodified S protein (construct 9125). Notably, the increase in DS purity was also greater for the modified coronavirus S protein D253N (construct 9802) than for the coronavirus S protein from donor strain C.37 ("lambda", construct 9588) expressed in the same system (Figure 4B).
[0088] As further shown in Figures 6A and 6B, when a glycosylation site is introduced into the modified S protein at residues corresponding to positions 253, 252, or 251 of the reference sequence SEQ ID NO: 1, the proportion of full-length S protein versus truncated or "clipped" S protein (represented as "full S protein (%)" in Figure 6B) is significantly increased in the S protein having a glycosylation site at position 253, 252, or 251 compared to the parent S protein.
[0089] More specifically, as shown in Fig. 6B , when compared with the corresponding unmodified S protein (construct 9125), a greater proportion of modified S protein was extracted and purified from plants as full-length S protein (expressed as “full S protein (%)”) when the S protein contained substitutions corresponding to position 253 (construct 9802), position 252 (construct 10346), or positions 251 and 253 (construct 10351).
[0090] As further shown in Figure 6C, when expressed in N. benthamiana, modified S proteins with additional glycosylation sites corresponding to positions 253, 252, or 251 (constructs: 9802, 10346, and 10351) were observed to form VLPs similar to the unmodified SARS-CoV-2 S protein (Figure 2A, construct 9125).
[0091] The modified S protein (construct 9802) with the D253N modification extracted by enzymatic extraction at pH 5.5 (Enz.pH5.5) or pH 6.1 (Enz.pH6.1) or by mechanical extraction (Mech) showed increased stability compared to the unmodified S protein (construct 9125) after overnight incubation at 24°C (see Figure 5).
[0092] Deletions and substitutions In a further embodiment, the modified S protein may comprise one or more deletions corresponding to positions 246, 247, 248, 249, 250, 251 or 252 and one or more substitutions corresponding to positions 251, 252 or 253 of reference sequence SEQ ID NO:1.
[0093] For example, the modified S protein may contain one or more deletions and one or more substitutions, wherein the one or more deletions include at least four consecutive amino acid residues, and at least residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO:1 are deleted, and the one or more substitutions correspond to positions 251, 252 or 253 of reference sequence SEQ ID NO:1.
[0094] For example, the modified S protein may comprise at least a deletion of residues corresponding to positions 247, 248, 249, and 250 of the reference sequence SEQ ID NO:1, and one or more substitutions corresponding to positions 251, 252, or 253 of the reference sequence SEQ ID NO:1. In another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 248, 249, 250, and 251 of the reference sequence SEQ ID NO:1, and one or more substitutions corresponding to positions 252 or 253 of the reference sequence SEQ ID NO:1. In a further example, the modified S protein may comprise at least a deletion of residues corresponding to positions 249, 250, 251, and 252 of the reference sequence SEQ ID NO:1, and a substitution corresponding to position 253 of the reference sequence SEQ ID NO:1. In another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 246, 247, 248, 249, and 250 of the reference sequence SEQ ID NO:1, and one or more substitutions corresponding to positions 251, 252, or 253 of the reference sequence SEQ ID NO:1. In a further example, the modified S protein may comprise at least a deletion of residues corresponding to positions 247, 248, 249, 250, and 251 of the reference sequence SEQ ID NO:1, and one or more substitutions corresponding to position 252 or 253 of the reference sequence SEQ ID NO:1. In yet another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 248, 249, 250, 251, and 252 of the reference sequence SEQ ID NO:1, and a substitution corresponding to position 253 of the reference sequence SEQ ID NO:1. In yet another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 246, 247, 248, 249, 250, and 251 of the reference sequence SEQ ID NO:1, and one or more substitutions corresponding to position 252 or 253 of the reference sequence SEQ ID NO:1. In yet another example, the modified S protein may comprise at least a deletion of residues corresponding to positions 247, 248, 249, 250, 251, and 252 of the reference sequence SEQ ID NO:1, and a substitution corresponding to position 253 of the reference sequence SEQ ID NO:1. In a further example, the modified S protein may include at least a deletion of residues corresponding to positions 246, 247, 248, 249, 250, 251 and 252 of reference sequence SEQ ID NO:1 and a substitution corresponding to position 253 of reference sequence SEQ ID NO:1.
[0095] As shown in Figure 3A, the ratio of full-length coronavirus S protein ("% Full S Protein") to truncated coronavirus S protein is significantly increased for the modified S protein incorporating the deletion del246-252 and the substitution D253N ("del-246-252+D253N"; construct 9808). More specifically, the ratio of full-length coronavirus S protein to truncated coronavirus S protein ("% Full S Protein") increased from approximately 82% full-length coronavirus S protein for the B acceptor strain S protein (construct 9125) to approximately 96% full-length coronavirus S protein for the modified S protein incorporating the deletion and substitution ("del-246-252+D253N"; construct 9808).
[0096] As further shown in Figures 4A and 4B, clarification and purification of coronavirus S protein followed by quantification using SDS gels demonstrates an increase in percent drug substance (DS) purity. When modified S protein incorporating the del246-252+D253N (construct 9808) modification from donor strain C.37 ("lambda") S protein was extracted and purified from plants, an increase in DS purity was observed compared to the purity of DS obtained from plants expressing the unmodified S protein (construct 9125). Notably, the increase in DS purity was also greater for the modified coronavirus S protein incorporating del246-252+D253N (construct 9808) than for the DS obtained from the coronavirus S protein from donor strain C.37 ("lambda", construct 9588) expressed in the same system (Figure 4B).
[0097] "Corresponding to an amino acid," "corresponding to an amino acid," or "corresponding to a position," or "corresponding to a position," etc., mean that the amino acid corresponds to an amino acid in a sequence alignment with a coronavirus S protein reference sequence as described below.
[0098] Modifications from the S protein donor strain nucleic acid sequence may be introduced into the corresponding nucleic acid position of the S protein acceptor strain nucleic acid sequence. Similarly, amino acid substitutions or deletions from the S protein donor strain amino acid sequence may be introduced into the corresponding amino acid position of the S protein acceptor strain sequence. As further described herein, additional modifications not found in the S protein from the donor strain may also be introduced into the modified S protein.
[0099] Throughout this disclosure, the amino acid residue numbers or residue positions of the coronavirus S protein follow the numbering of the S protein reference sequence. For example, the S protein reference sequence is the sequence of the "ancestral" B-lineage SARS-CoV-2 S protein (UniProtKB-P0DTC2, SEQ ID NO:1). Corresponding amino acid positions can be determined by aligning the sequence of the S protein with the S protein reference sequence. For example, an amino acid sequence alignment shows that positions 246-253 of SARS-CoV-2 B (SEQ ID NO:1) correspond to positions 256-263 in the sequence of SARS-CoV-2 B.1.617.2 (SEQ ID NO:14) and positions 255-262 in the sequence of SARS-CoV-2 B.1.1.529 (SEQ ID NO:18). Corresponding positions in other sequences can be determined by similar alignments. Methods for aligning sequences for comparison are well known in the art and are described further below.
[0100] The modified coronavirus S protein may contain one or more amino acid sequence modifications when compared to the corresponding unmodified (parent) amino acid sequence, wherein the one or more modifications correspond to amino acids at positions 246, 247, 248, 249, 250, 251, 252, 253, or combinations thereof, of the reference sequence SEQ ID NO:1.
[0101] The one or more amino acid sequence modifications can include substitution of one or more amino acids or deletion of one or more amino acids. For example, the amino acid sequence modification includes a substitution of a non-glycine corresponding to position 252, e.g., a substitution of an asparagine (N) corresponding to position 252. In one embodiment, the modification is a G252N substitution, and thus the modified S protein includes a G252N substitution or modification.
[0102] In another example, the modification comprises a substitution with a non-aspartic acid corresponding to position 253, e.g., a substitution with an asparagine (N) corresponding to position 253. In one embodiment, the modification is a D253N substitution, and thus the modified S protein comprises a D253N substitution or modification.
[0103] The modification may further comprise two substitutions. For example, the modified S protein may comprise a non-proline substitution at the amino acid corresponding to position 251, e.g., an asparagine (N) substitution corresponding to position 251 and a threonine (T) substitution corresponding to position 253. In one embodiment, the modified S protein comprises a P251N+D253T substitution or modification.
[0104] Additionally, the modified S protein can comprise a deletion of one or more amino acids corresponding to positions 246, 247, 248, 249, 250, 251, 252, or a combination thereof. In one embodiment, the amino acids corresponding to positions 246-252 are deleted, and thus the modified S protein comprises a 246-252 deletion or modification (del246-252). In another embodiment, the amino acids corresponding to positions 247-250 are deleted, and thus the modified S protein comprises a 247-250 deletion or modification (del247-250). In a further embodiment, the amino acids corresponding to positions 248-251 are deleted, and thus the modified S protein comprises a 248-251 deletion or modification (del248-251). In a further embodiment, the amino acids corresponding to positions 249-252 are deleted, and thus the modified S protein comprises a 249-252 deletion or modification (del249-252). In another embodiment, the amino acids corresponding to positions 246-250 are deleted, and thus the modified S protein comprises a 246-250 deletion or modification (del246-250). In yet another embodiment, the amino acids corresponding to positions 247-251 are deleted, and thus the modified S protein comprises a 247-251 deletion or modification (del247-251). In another embodiment, the amino acids corresponding to positions 248-252 are deleted, and thus the modified S protein comprises a 248-252 deletion or modification (del248-252). In another embodiment, the amino acids corresponding to positions 246-251 are deleted, and thus the modified S protein comprises a 246-251 deletion or modification (del246-251). In another embodiment, the amino acids corresponding to positions 247-252 are deleted, and thus the modified S protein comprises a 247-252 deletion or modification (del247-252). Non-limiting examples of modified S proteins containing one or more deletions are shown in Table 2.
[0105] In one embodiment, the modified S protein may include at least one substitution and one or more deletions of amino acids.
[0106] For example, the modified S protein can include a substitution of the amino acid corresponding to position 252 with a non-glycine and a deletion of one or more amino acids corresponding to positions 246, 247, 248, 249, 250, 251, or a combination thereof.
[0107] In one embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 252 and a deletion of amino acids 246-251, such that the modified S protein comprises a G252N substitution and a 246-251 deletion ("G252N+del246-251" or "del246-252+G252N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 252 and a deletion of amino acids 247-250, such that the modified S protein comprises a G252N substitution and a 247-250 deletion ("G252N+del247-250" or "del247-250+G252N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 252 and a deletion of amino acids 248-251, such that the modified S protein comprises a G252N substitution and a 248-251 deletion ("G252N+del248-251" or "del248-251+G252N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 252 and a deletion of amino acids 249-251, such that the modified S protein comprises a G252N substitution and a 249-251 deletion ("G252N+del249-251" or "G252N+del249-251"). In a further embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 252 and a deletion of amino acids 246-250, such that the modified S protein comprises a G252N substitution and a 246-250 deletion ("G252N+del246-250" or "G252N+del246-250"). In yet another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 252 and a deletion of amino acids 247-251, such that the modified S protein comprises a G252N substitution and a 247-251 deletion ("G252N+del247-251" or "G252N+del247-251").In another embodiment, the modified S protein comprises a substitution of asparagine (N) corresponding to position 252 and a deletion of amino acids 246-251, thus the modified S protein comprises a G252N substitution and a 246-251 deletion ("G252N+del246-251" or "G252N+del246-251").
[0108] For example, the modified S protein can include a substitution of the amino acid corresponding to position 253 with a non-aspartic acid and a deletion of one or more amino acids corresponding to positions 246, 247, 248, 249, 250, 251, 252, or a combination thereof.
[0109] In one embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 246-252, such that the modified S protein comprises a D253N substitution and a 246-252 deletion ("D253N+del246-252" or "del246-252+D253N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 247-250, such that the modified S protein comprises a D253N substitution and a 247-250 deletion ("D253N+del247-250" or "del247-250+D253N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 248-251, such that the modified S protein comprises a D253N substitution and a 248-251 deletion ("D253N+del248-251" or "del248-251+D253N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 249-252, such that the modified S protein comprises a D253N substitution and a 249-252 deletion ("D253N+del249-252" or "del249-252+D253N"). In yet another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 246-250, such that the modified S protein comprises a D253N substitution and a 246-250 deletion ("D253N+del246-250" or "del246-250+D253N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 247-251, such that the modified S protein comprises a D253N substitution and a 247-251 deletion ("D253N+del247-251" or "del247-251+D253N").In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 248-252, such that the modified S protein comprises a D253N substitution and a 248-252 deletion ("D253N+del248-252" or "del248-252+D253N"). In another embodiment, the modified S protein comprises a substitution of an asparagine (N) corresponding to position 253 and a deletion of amino acids 246-251, such that the modified S protein comprises a D253N substitution and a 246-251 deletion ("D253N+del246-251" or "del246-251+D253N"). In another embodiment, the modified S protein comprises a substitution of asparagine (N) corresponding to position 253 and a deletion of amino acids 247-252, thus the modified S protein comprises a D253N substitution and a 247-252 deletion ("D253N+del247-252" or "del247-252+D253N").
[0110] For example, the modified S protein may include a substitution of the amino acid corresponding to position 251 with a non-proline, a substitution of the amino acid corresponding to position 253 with a non-aspartic acid, and a deletion of one or more amino acids corresponding to positions 246, 247, 248, 249, 250, or a combination thereof.
[0111] In one embodiment, the modified S protein comprises a substitution of asparagine (N) corresponding to position 251, a substitution of threonine (T) corresponding to position 253, and a deletion of amino acids 246-250; thus, the modified S protein comprises a P251N substitution, a D253T substitution, and a 246-250 deletion ("P251N+D253T+del246-250" or "del246-250+P251N+D253T").
[0112] In another embodiment, the modified S protein comprises an asparagine (N) substitution at position 251, a threonine (T) substitution corresponding to position 253, and a deletion of amino acids 247-250; thus, the modified S protein comprises a P251N substitution, a D253T substitution, and a 247-250 deletion ("P251N+D253T+del247-250" or "del247-250+P251N+D253T").
[0113] In another embodiment, the modified S protein comprises a substitution of asparagine (N) corresponding to position 251, a substitution of threonine (T) corresponding to position 253, and a deletion of amino acids 248-250; thus, the modified S protein comprises a P251N substitution, a D253T substitution, and a 248-250 deletion ("P251N+D253T+del248-250" or "del248-250+P251N+D253T").
[0114] For example, the modified S protein may have an amino acid sequence of SEQ ID NO: 8, 10, 12, 16, 20, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, or 57 and an amino acid sequence of about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262 9, 100%, or any amount of sequence identity or similarity therebetween, wherein the amino acids corresponding to positions 246-252, 247-250, 248-251, 249-252, 246-250, 247-251, 248-252, 246-251, or 247-252 are deleted, and / or the amino acids corresponding to positions 25 the amino acid corresponding to position 251 is asparagine (N, Asn), or the amino acids corresponding to positions 247-250, 248-251, 246-250, 247-251, or 246-251 are deleted, and / or the amino acid corresponding to position 252 is asparagine (N, Asn), or the amino acids corresponding to positions 247-250, or 246-250 are deleted, and / or the amino acid corresponding to position 251 is asparagine (N, Asn), and position 253 is threonine (T, Thr), the numbering of the positions corresponds to positions in reference sequence SEQ ID NO: 1, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0115] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 8 or 25, wherein the amino acids corresponding to positions 246-252 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0116] The modified S protein may have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100% or any amount of sequence identity or similarity therebetween with the amino acid sequence of SEQ ID NO: 10, 16, 20 or 26, wherein the amino acid corresponding to position 253 is asparagine (N, Asn), the modified S protein sequence does not occur in nature, and the S protein forms a VLP upon expression.
[0117] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 12 or 27, wherein the amino acids corresponding to positions 246-252 are deleted and the amino acid corresponding to position 253 is an asparagine (N, Asn), wherein the modified S protein sequence is not naturally occurring, and wherein the S protein forms a VLP upon expression.
[0118] The modified S protein may have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 39 or 28, wherein the amino acid corresponding to position 252 is asparagine (N, Asn), the modified S protein sequence does not occur in nature, and the S protein forms a VLP upon expression.
[0119] The modified S protein may have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 41 or 29, wherein the amino acid corresponding to position 251 is asparagine (N, Asn) and the amino acid corresponding to position 253 is threonine (T, Thr), the modified S protein sequence does not occur in nature, and the S protein forms a VLP upon expression.
[0120] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 43 or 30, wherein the amino acids corresponding to positions 247-250 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0121] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 45 or 31, wherein the amino acids corresponding to positions 248-251 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0122] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 47 or 32, wherein the amino acids corresponding to positions 249-252 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0123] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 49 or 33, wherein the amino acids corresponding to positions 246-250 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0124] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 51 or 34, wherein the amino acids corresponding to positions 247-251 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0125] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 53 or 35, wherein the amino acids corresponding to positions 248-252 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0126] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 55 or 36, wherein the amino acids corresponding to positions 246-251 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0127] The modified S protein can have an amino acid sequence having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or similarity therebetween, with the amino acid sequence of SEQ ID NO: 57 or 37, wherein the amino acids corresponding to positions 247-252 are deleted, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0128] Also provided herein is a nucleic acid comprising a nucleotide sequence encoding an S protein having a substitution corresponding to positions 251, 252, 253, a deletion corresponding to positions 246-252, 247-250, 248-251, 249-252, 246-250, 247-251, 248-252, 246-251, 247-252, or a substitution corresponding to positions 251, 252, 253 and a deletion corresponding to positions 246-252, 247-250, 248-251, 249-252, 246-250, 247-251, 248-252, 246-251, 247-252, operably linked to a regulatory region active in a host or host cell, such as a plant.
[0129] The isolation of nucleic acids encoding such S proteins is well known to those of skill in the art, as is the modification of nucleic acids to introduce amino acid sequence changes, for example, by site-directed mutagenesis.
[0130] For example, the nucleotide sequence may be any of the nucleotide sequences of SEQ ID NOs: 9, 11, 15, 19, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56 and about 50, 55, 60, 65, 70, 75, 80, 85, 87, 90, 91, 92, 93 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity therebetween, wherein the nucleotide codons encoding the amino acids corresponding to positions 246-252, 247-250, 248-251, 249-252, 246-250, 247-251, 248-252, 246-251, or 247-252 are deleted, and / or the amino acid corresponding to position 253 is asparagine (N, Asn), or the amino acid corresponding to positions 247-250, 248-251, 246-250, 247-251, or the amino acids corresponding to positions 246-251 are deleted and / or the amino acid corresponding to position 252 is asparagine (N, Asn), or the amino acids corresponding to positions 247-250, or 246-250 are deleted and / or the amino acid corresponding to position 251 is asparagine (N, Asn), and position 253 is threonine (T, Thr), the numbering of the positions corresponds to the positions in reference sequence SEQ ID NO: 1, the modified S protein sequence is not naturally occurring, and the S protein forms a VLP upon expression.
[0131] The modified coronavirus S proteins described herein can be derived from an acceptor S protein having about 70, 75, 80, 85, 87, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100%, or any amount of sequence identity or sequence similarity therebetween with the amino acid sequence of SEQ ID NO:1, 2, 4, 6, 14, or 18.
[0132] Examples of modified S proteins having enhanced or improved characteristics as described herein include, but are not limited to, the following: · B CoV S(GSAS-2P+del246-252): construct: 9801, SEQ ID NO: 8; · B CoV S(GSAS-2P+D253N): construct: 9802, SEQ ID NO: 10; · B CoV S(GSAS-2P+del246-252+D253N) construct: 9808, SEQ ID NO: 12; B.1.617.2 CoV S(GSAS-2P+D253N): construct: 10011, SEQ ID NO: 16 B.1.1.529 CoV S(GSAS-2P+D253N): construct: 10092, SEQ ID NO: 20 · B CoV S(GSAS-2P+D252N) construct: 10346, SEQ ID NO: 39; ·B CoV S (GSAS-2P+P251N+D253T) construct: 10351, SEQ ID NO: 41; ·B CoV S (GSAS-2P+del247-250) construct: 10502, SEQ ID NO: 43; ·B CoV S (GSAS-2P+del248-251) construct: 10503, SEQ ID NO: 45; · B CoV S(GSAS-2P+del249-252) construct: 10504, SEQ ID NO: 47; · B CoV S(GSAS-2P+del246-250) construct: 10505, SEQ ID NO: 49; ·B CoV S (GSAS-2P+del247-251) construct: 10506, SEQ ID NO: 51; ·B CoV S(GSAS-2P+del248-252):10507, SEQ ID NO: 53; · B CoV S(GSAS-2P+del246-251) construct: 10508, SEQ ID NO: 55; ·B CoV S (GSAS-2P+del247-252) construct: 10509, SEQ ID NO: 57.
[0133] In one embodiment, the modified coronavirus S protein can be derived from an acceptor B strain SARS-CoV-2 S protein (SEQ ID NOs: 3 and 4, construct 9125, Figure 1B) into which the del246-252 mutation from the donor C.37 ("lambda") strain has been introduced (SEQ ID NOs: 7 and 8, construct 9801, Figure 1C). In another embodiment, the modified coronavirus S protein is derived from an acceptor B strain SARS-CoV-2 S protein (SEQ ID NOs: 3 and 4) into which the D253N mutation from the donor C.37 ("lambda") strain has been introduced (SEQ ID NOs: 9 and 10, construct 9802, Figure 1D). In another embodiment, the modified coronavirus S protein is derived from an acceptor B strain SARS-CoV-2 S protein (SEQ ID NOs: 3 and 4) into which the del246-252 and D253N mutations from the donor C.37 ("lambda") strain have been introduced (SEQ ID NOs: 11 and 12, construct 9808, Figure 1E). However, as also described herein, modifications not found in the S protein from the donor strain may also be introduced into the modified S protein for improved characteristics of the modified S protein compared to the unmodified (parent) S protein.
[0134] Modified S proteins can be produced by introducing changes into the amino acid sequence of the S protein that result in improved characteristics of the modified S protein compared to the unmodified S protein, for example, increased stability against proteolysis or increased stability against degradation by proteases, increased integrity, purity and / or homogeneity of the recombinant modified S protein when expressed in a host or host cell.
[0135] Accordingly, methods are also provided for improving characteristics of a coronavirus S protein, such as increased stability against proteolysis or against degradation by proteases, increased S protein integrity, purity, and / or homogeneity, etc. The methods include: a) modifying a parent coronavirus S protein to produce a modified coronavirus S protein, wherein the modified coronavirus S protein contains one or more amino acid sequence modifications as described above when compared to the parent coronavirus S protein, wherein the one or more modifications correspond to amino acids at positions 246, 247, 248, 249, 250, 251, 252, 253, or combinations thereof, of reference sequence SEQ ID NO:1; and b) expressing the modified coronavirus S protein in a host or host cell, thereby producing a modified S protein having improved characteristics compared to the same characteristics of the parent coronavirus S protein produced under similar conditions in a host or host cell.
[0136] Also provided are methods for producing a modified coronavirus S protein described herein in a host cell or host cell, wherein the modified S protein has improved characteristics, such as increased stability against proteolysis or increased stability against degradation by proteases, increased integrity, purity, and / or homogeneity, compared to the characteristics of the unmodified S protein. The method includes: a) introducing into a host or host cell a nucleic acid encoding the modified S protein described herein, or providing a host or host cell containing a nucleic acid encoding the modified S protein described herein; and b) incubating the host or host cell under conditions that allow expression of the nucleic acid, thereby producing a modified S protein having improved characteristics compared to the same characteristics of the unmodified S protein produced in the host or host cell under similar conditions. The modified S protein can optionally be further extracted, purified, or extracted and purified from the host or host cell.
[0137] Furthermore, the modified S protein can be assembled into virus-like particles (VLPs), and the content or amount of full-length S protein within the VLPs is increased compared to the full-length S protein content of VLPs produced from unmodified S protein. By introducing the modifications described herein into coronavirus S proteins, it has been observed that the resulting modified S protein has increased resistance to N-terminal proteolysis or clipping, allowing for the production of greater amounts of full-length S protein. Thus, in another aspect, the present disclosure provides VLPs comprising the modified S proteins described herein, wherein the VLPs comprise an increased content or amount of full-length modified S protein when compared to VLPs comprising unmodified S protein.
[0138] The increase in full-length (uncleaved) S protein produced in a host, such as a plant, can be measured and expressed as an increase in the purity of the resulting product. Protein purity is commonly assessed using SDS-PAGE gels, and band intensities from the SDS-PAGE gel can be calculated by densitometry analysis. In summary, the amount of recombinant protein species (i.e., protein bands) with different molecular weights produced in the host is determined by densitometry. Protein loading correlates with the peak area of the protein species in the densitometry profile, and the peak maximum intensity is used to quantify the protein species. Purity is calculated as the sum of the relative densities of the protein bands of interest expressed as a percentage (%); i.e., the more protein of interest is produced, the higher the calculated purity.
[0139] Thus, the present disclosure further provides a drug substance (DS) comprising the above-described modified S protein as a desired product, wherein the drug substance is substantially free of product-related impurities, and the impurities are not immunoreactive. Preferred drug substances are also substantially free of process-related impurities.
[0140] Immunological activity refers to a compound such as a coronavirus S protein or S protein fragment that is recognized by an antibody specific for the RBD domain of the coronavirus S protein (anti-RBD antibody, e.g., 40592-T62 from Sino biological) and / or an antibody specific for the S2 domain of the coronavirus S protein (anti-S2 antibody, e.g., NB100-56578 from Novus Biological).
[0141] Within the context of this application, the term "drug substance" refers to i) the active ingredient of a drug or drug product, ii) the active pharmaceutical ingredient of a drug or drug product, iii) the bulk purified active ingredient of a drug or drug product, or iv) a product or active ingredient suitable for use as a bulk purified active ingredient of a drug or drug product. The drug or drug product can be a vaccine.
[0142] Therefore, a drug substance (DS) comprising a VLP containing the modified S protein described herein is further provided. The DS containing the modified S protein has a higher purity than a DS obtained from a host expressing an unmodified S protein. Therefore, a method for increasing the purity of a DS obtained from a host or host cell expressing a modified S protein is also provided, compared to the purity of a DS obtained from a host expressing an unmodified (parent) S protein.
[0143] The modified coronavirus S protein can self-assemble into virus-like particles (VLPs). Thus, a DS containing a VLP containing the modified S protein is also provided. The DS containing the VLP with the modified S protein exhibits increased purity compared to a DS produced from the parent (unmodified) coronavirus S protein.
[0144] In a further aspect, a drug product (also called a pharmaceutical formulation or pharmaceutical composition) is also provided. The drug product may be formulated as a final dosage form, for example, a solution, a capsule, or a tablet. The drug product comprises an active pharmaceutical ingredient. The drug product may further comprise other ingredients, for example, pharmaceutically acceptable carriers and / or excipients, such as a buffer system, an adjuvant, a preservative, a tonicity agent, a chelating agent, an anti-adherent agent, a vehicle, etc. Pharmaceutically acceptable carriers and excipients are well known in the art. Thus, a drug product, pharmaceutical formulation, or pharmaceutical composition comprising a pharmaceutically acceptable carrier and / or excipient and a VLP is also provided, wherein the VLP comprises a modified S protein, or the VLP comprises a viral protein, and the viral protein consists of a modified coronavirus S protein.
[0145] Other modifications The modified S protein may be a chimeric modified S protein or chimeric S protein. A "chimeric S protein" refers to a protein or polypeptide comprising amino acid sequences and / or protein domains or portions of protein domains from two or more sources fused into a single polypeptide. For example, but not limited to, the ectodomain and transmembrane domain (TM) or portion of the TM of a chimeric S protein may be derived from a coronavirus S protein (such as SARS-CoV 2), and the cytoplasmic tail (CT) or portion of the CT may be derived from influenza HA (as described in International PCT Application PCT / CA2021 / 051201, incorporated herein by reference).
[0146] The modified S protein may further comprise one or more substitutions or substitutions to stabilize the coronavirus S protein or coronavirus S protein trimer in the pre-fusion conformation. For example, the modified S protein may further comprise an alteration of the consensus RRAR furin cleavage site and two consecutive proline substitutions at or near the boundary between the HR1 domain and the central helical domain that stabilize the S ectodomain trimer in the pre-fusion conformation, e.g., as described in WO 2018 / 081318 and PCT / CA2021 / 051201, which are incorporated herein by reference.
[0147] VLP The modified coronavirus S proteins described herein can further be incorporated into virus-like particles (VLPs). The term "virus-like particle" (VLP) or "virus-like particles" or "VLPs" refers to virus-like structures that are generally morphologically and antigenically similar to virions produced during infection, but lack sufficient genetic information for replication and are therefore non-infectious. VLPs are structures that self-assemble and comprise one or more structural proteins, such as, for example, a modified coronavirus S protein. Thus, a VLP can comprise a modified coronavirus S protein. A VLP can further comprise a coronavirus protein, where the coronavirus protein consists of a modified coronavirus S protein.
[0148] As shown in Figures 2B-2D, when expressed in N. benthamiana, modified B strain S proteins (constructs: 9801, 9802, and 9808) were observed to form VLPs similar to the unmodified B strain SARS-CoV-2 S protein (Figure 2A, construct 9125). Expression of the S protein from the donor C.37 ("lambda") strain (Figure 2E, construct 9588) formed VLPs similar to the acceptor B strain SARS-CoV-2 S protein.
[0149] As further shown in Figure 6C, when expressed in N. benthamiana, the modified B strain S proteins (constructs: 9802, 10346, and 10351) were observed to form VLPs similar to the unmodified B strain SARS-CoV-2 S protein (Figure 2A, construct 9125).
[0150] VLPs can be produced in suitable hosts or host cells, including plants and plant cells. After extraction from the host or host cells, upon isolation and further purification under suitable conditions, the VLPs can be recovered as intact structures.
[0151] VLPs can be purified or extracted using any suitable method, such as chemical or biochemical extraction. VLPs are relatively sensitive to drying, heat, pH, surfactants, and detergents. Therefore, it may be useful to maximize yield, minimize contamination of the VLP fraction with cellular proteins, maintain the integrity of the protein or VLP, and, if necessary, use a method that loosens the associated lipid envelope or membrane, cell wall, to release the protein or VLP. Minimizing or eliminating the use of detergents or surfactants, such as SDS or Triton™ X-100, may be beneficial for improving the yield of VLP extraction. VLPs can then be evaluated for structure and size, for example, by electron microscopy (see Figure 4B) or size exclusion chromatography.
[0152] In the case of enveloped viruses such as coronaviruses, it may be advantageous for the lipid layer or membrane to be retained by the virus. The composition, quality, and quantity of lipids may vary depending on the system (e.g., plant-produced enveloped viruses may contain plant lipids or plant sterols within the envelope), which may contribute to improved immune responses.
[0153] Thus, VLPs produced in a host or host cell may comprise lipids derived from the plasma membrane of the host or host cell. For example, VLPs produced in plants may comprise lipids of plant origin ("plant lipids"), VLPs produced in insect cells may comprise lipids derived from the plasma membrane of insect cells (commonly referred to as "insect lipids"), and VLPs produced in mammalian cells may comprise lipids derived from the plasma membrane of mammalian cells (commonly referred to as "mammalian lipids").
[0154] The plant lipid or plant-derived lipid may be in the form of a lipid bilayer and may further comprise an envelope surrounding the VLP. The plant-derived lipid may comprise the lipid components of the plasma membrane of the plant from which the VLP is produced, including phospholipids, tri-, di-, and monoglycerides, and lipophilic sterols or sterol-containing metabolites. Examples include phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylinositol, phosphatidylserine, glycosphingolipids, plant sterols, or combinations thereof. Examples of plant sterols include campesterol, stigmasterol, ergosterol, brassicasterol, delta-7-stigmasterol, delta-7-avenasterol, daunosterol, sitosterol, 24-methylcholesterol, cholesterol, or beta-sitosterol. As those skilled in the art will understand, the lipid composition of the plasma membrane of a cell may vary depending on the culture or growth conditions of the cell or organism or species from which the cell is obtained. Generally, beta-sitosterol is the most abundant plant sterol.
[0155] Without wishing to be bound by theory, plant-made VLPs containing plant-derived lipids may induce a stronger immune response than VLPs produced in other manufacturing systems, and the immune response induced by these plant-made VLPs may be stronger than the immune response induced by live or attenuated whole virus vaccines.
[0156] Furthermore, in addition to the potential adjuvant effect of the presence of plant lipids, the ability of plant N-glycans to promote the capture of glycoprotein antigens by antigen-presenting cells may be advantageous for the production of VLPs in plants.
[0157] VLPs produced in plants can contain modified S proteins containing plant-specific N-glycans. Thus, the present disclosure also provides VLPs containing modified S proteins with plant-specific N-glycans. Additionally, VLPs are provided that contain plant lipids and modified S proteins with plant-specific N-glycans.
[0158] method Methods for producing the modified S protein, or a virus-like particle (VLP) comprising the modified S protein, in a host or host cell are also provided.
[0159] The method includes introducing into a host or host cell a nucleic acid comprising a sequence encoding a modified S protein described herein, or providing a host or host cell comprising a nucleic acid encoding a modified S protein described herein, and incubating the host or host cell under conditions that allow expression of the nucleic acid, thereby producing the modified S protein. The modified S protein can then self-assemble into VLPs comprising the modified S protein (see Figures 2B-2D and 6C).
[0160] Additionally, methods are provided for increasing the purity or homogeneity of a coronavirus S protein produced in a host or host cell, comprising introducing into the host or host cell a nucleic acid comprising a sequence encoding a modified S protein described herein, or providing a host or host cell comprising a nucleic acid encoding a modified S protein described herein, and incubating the host or host cell under conditions that permit expression of the nucleic acid, thereby producing a modified S protein, wherein the produced S protein has increased purity and / or homogeneity compared to the purity and / or homogeneity of an unmodified S protein produced under similar conditions in the host or host cell.
[0161] Further provided is a modified S protein produced by the methods described herein, wherein the modified S protein has increased purity and / or homogeneity compared to the purity and / or homogeneity of an unmodified S protein produced under similar conditions in a host or host cell.
[0162] The modified S proteins described herein can have a purity and / or homogeneity of about 80% to 98%, or any amount therebetween, for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98%, or any amount therebetween.
[0163] It is desirable that the recombinant protein sample be pure and homogeneous, ie, that the recombinant protein is present in the same consistent form and size throughout the sample.
[0164] In this application, "increased purity" or "increased homogeneity" refers to an increase in stable and uniform protein length of the recombinant S protein, e.g., a full-length expressed protein. As noted above, during expression in a host cell or subsequent purification from the host cell, the recombinant protein may be degraded at one or more residues (also referred to as "clipped," "hydrolyzed," or "truncated") by enzymatic, proteolytic, or chemical events subsequent to or coincident with expression or translation. Without wishing to be bound by theory, most protein degradation can be attributed to host cell-derived proteases.
[0165] As described herein, modifying the S protein results in a more stable and homogeneous protein sample. More specifically, it has been found that when the modified S protein is produced in a host or host cell, such as a plant or plant cell, fewer truncated or truncated forms of the S protein are observed, and the modified S protein exhibits increased stability against proteolysis compared to unmodified S protein samples.
[0166] By modifying S proteins as described herein, the proportion of truncated, truncated, or clipped S proteins is reduced and the proportion of full-length S proteins is increased. Thus, modified S proteins incorporating the modifications described herein, when produced in a host or host cell, can reduce the proportion of truncated, truncated, or clipped S proteins and increase the proportion of full-length S proteins compared to unmodified S proteins. For example, the ratio of full-length modified S proteins to truncated modified S proteins can be about 1:0.02 to 1:0.25, or any amount therebetween, such as a full-length to truncated modified S protein ratio of 1:0.05, 1:0.1, 1:0.15, 1:2, or 1:0.25.
[0167] With respect to a coronavirus S protein or a modified coronavirus S protein, a "full-length S protein" (also referred to as a complete S protein) or a "full-length modified S protein" (also referred to as a fully modified S protein) refers to an S protein or a modified S protein having a polypeptide sequence and / or polypeptide length that corresponds to the theoretical polypeptide sequence or polypeptide length of the construct in which the recombinant S protein or recombinant modified S protein is expressed. A full-length S protein or a full-length modified S protein can include a signal peptide that directs localization when expressed in a host or host cell. The signal peptide can be a native (with respect to the protein) signal sequence or leader sequence, or a heterologous signal sequence.
[0168] Thus, as described herein, modified S proteins can be produced as precursor proteins containing the modified S protein and a heterologous amino acid signal peptide sequence. For example, the modified S protein precursor can contain a signal peptide derived from protein disulfide isomerase (PDI SP; nucleotides 32-103 of Accession No. Z11499).
[0169] The full-length modified S protein can also refer to the mature modified S protein without the signal peptide, since the signal peptide is cleaved when the mature S protein is produced in a host or host cell. The full-length S protein or full-length modified S protein is in contrast to a truncated S protein, in which a portion of the N-terminus or C-terminus of the mature S protein may be proteolytically removed.
[0170] For example, a full-length S protein can correspond to the full length of the sequence of SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, or 20, which can include a signal peptide. Alternatively, a full-length S protein can correspond to the full length of the sequence of SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 39, 41, 43, 45, 47, 49, 51, 53, 55, or 57 without the signal peptide. For example, a full-length modified S protein can correspond to the full length of the sequence of SEQ ID NO: 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, or 37.
[0171] With respect to an S protein or modified S protein, "truncated S protein," "truncated S protein," "clipped S protein," or "truncated modified S protein," "truncated modified S protein," or "clipped modified S protein" refers to an expressed polypeptide of a coronavirus S protein that is cleaved or truncated at one or more residues by an enzymatic, proteolytic, or chemical event subsequent to or concurrent with expression or translation, where the cleavage or truncation is the result of unintended proteolytic processing of the protein. "truncated S protein," "truncated S protein," "clipped S protein," "truncated modified S protein," "truncated modified S protein," or "clipped modified S protein" does not include an S protein (or modified S protein) from which the signal peptide has been cleaved to produce the mature S protein or mature modified S protein.
[0172] The present disclosure also relates to a method for improving the N-terminal uniformity of coronavirus S proteins (modified S proteins), wherein 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of coronavirus S proteins having modifications have correct N-terminal amino acid sequences.
[0173] One or more modified gene constructs comprising the modified S protein of the present disclosure can be expressed in any suitable host or host cell transformed with the nucleic acid, nucleotide sequence, construct, or vector of the present disclosure. The host or host cell can be derived from any source, including plants, fungi, bacteria, insects, and animals, e.g., mammals. The host can also be a non-human host. For example, the host or host cell can be selected from a plant or plant cell, a fungus or fungal cell, a bacterium or bacterial cell, an insect or insect cell, and an animal or animal cell. The mammal or animal need not be human. In a preferred embodiment, the host or host cell is a plant, part of a plant, or plant cell.
[0174] As used herein, the terms "plant," "plant part," "plant portion," "plant matter," "plant biomass," "plant material," "plant extract," or "plant leaf" can include a whole plant, a plant cell, a tissue, a cell, or any fraction thereof, an intracellular plant component, an extracellular plant component, a liquid or solid extract of a plant, or a combination thereof, which can provide the transcriptional, translational, and post-translational machinery for expression of one or more nucleic acids described herein and / or from which expressed proteins or VLPs can be extracted and purified. Plants can include, but are not limited to, herbaceous plants. Herbaceous plants can be annual, biennial, or perennial plants. Plants include, but are not limited to, canola, Brassica spp., maize, Nicotiana spp. (tobacco), e.g., Nicotiana benthamiana, Nicotiana rustica, Nicotiana, tabacum, Nicotiana alata, Arabidopsis thaliana, alfalfa, potato, sweet potato (Ipomoea batatus), ginseng, pea, oat, rice, soybean, wheat, barley, sunflower, cotton, corn, rye (Secale cereale), and the like. Further crops may be included, including corn (Sorghum cereale), sorghum (Sorghum bicolor, Sorghum vulgare), and safflower (Carthamus tinctorius).
[0175] The term "plant part" as used herein refers to any part of a plant, including, but not limited to, leaves, stems, roots, flowers, fruits, plant cells obtained from leaves, stems, roots, flowers, fruits, plant extracts obtained from leaves, stems, roots, flowers, fruits, or combinations thereof. In one embodiment, plant parts refer to areas of a plant, such as leaves, stems, flowers, and fruits. The term "plant extract" as used herein refers to a plant-derived product obtained after treating a plant, plant part, plant cell, or combination thereof physically (e.g., by freezing followed by extraction in a suitable buffer), mechanically (e.g., by crushing or homogenizing the plant or plant part followed by extraction in a suitable buffer), enzymatically (e.g., using cell wall-degrading enzymes), chemically (e.g., using one or more chelating agents or buffers), or a combination thereof. The plant extract may be further processed to remove undesirable plant components, such as cell wall debris. A plant extract may be obtained to aid in the recovery of one or more components from a plant, plant part, or plant cell, such as proteins (including protein complexes, protein surface structures, and / or VLPs), nucleic acids, lipids, carbohydrates, or a combination thereof, from a plant, plant part, or plant cell. When a plant extract contains proteins, it may be referred to as a protein extract. The protein extract may be a crude plant extract, a partially purified plant or protein extract, or a purified product containing one or more proteins, protein complexes, such as protein trimers, protein superstructures, and / or VLPs from plant tissue. If desired, the protein extract or plant extract may be partially purified using techniques known to those skilled in the art, for example, the extract may be subjected to salt or pH precipitation, centrifugation, gradient density centrifugation, filtration, chromatography, such as size exclusion chromatography, ion exchange chromatography, affinity chromatography, or a combination thereof. The protein extract may be purified using techniques known to those skilled in the art.
[0176] The constructs of the present disclosure can be introduced into plant cells using Ti plasmids, Ri plasmids, plant viral vectors, direct DNA transformation, microinjection, electroporation, etc. For reviews of such techniques, see, for example, Weissbach and Weissbach, Methods for Plant Molecular Biology, Academy Press, New York VIII, pp. 421-463 (1988); Geison and Corey, Plant Molecular Biology, 2nd Ed. (1988); and Miki and Iyer, Fundamentals of Gene Transfer in Plants. In Plant Metabolism, 2nd Ed. D.T. Dennis, D.H. Turpin, D.D. Lefebvre, D.B. Layzell (eds), Addison Wesly, Langmans Ltd. London, pp. 561-579 (1997). Other methods include direct DNA uptake, the use of liposomes, electroporation using, for example, protoplasts, microinjection, microprojectiles or whiskers, and vacuum filtration.For example, Bilang, et al. (Gene 100:247-250(1991), Scheid et al. (Mol. Gen. Genet. 228:104-112, 1991), Guerche et al. (Plant Science 52:111-116, 1987), Neuhause et al. Genet.75:30-36,1987), Klein et al.,Nature 327:70-73(1987);Howell et al.(Science 208:1265,1980),Horsch et al.(Science 227:1229-1231,1985),DeBlock et al.,Plant Physiology 91:694-701, 1989), Methods for Plant Molecular Biology (Weissbach and See Weissbach, eds., Academic Press Inc., 1988), Methods in Plant Molecular Biology (Schuler and Zielinski, eds., Academic Press Inc., 1989), Liu and Lomonossoff (J Virol Meth, 105:343-348, 2002), U.S. Pat. Nos. 4,945,050; 5,036,006 and 5,100,792, U.S. patent application Ser. No. 08 / 438,666, filed May 10, 1995, and U.S. patent application Ser. No. 07 / 951,715, filed September 25, 1992 (all of which are incorporated herein by reference).
[0177] As described below, transient expression methods may be used to express the constructs of the present disclosure (see Liu and Lomonossoff, 2002, Journal of Virological Methods, 105:343-348, incorporated herein by reference). Alternatively, vacuum-based transient expression methods may be used, such as those described by Kapila et al., 1997, incorporated herein by reference. These methods may include, but are not limited to, agroinoculation or agroinfiltration, syringe infiltration, although other transient methods may also be used as described above. With agroinoculation, agroinfiltration, or syringe infiltration, a mixture of Agrobacteria containing the desired nucleic acid penetrates the intercellular spaces of tissues, such as leaves, aboveground parts of the plant (including stems, leaves, and flowers), other parts of the plant (stems, roots, flowers), or the entire plant. Agrobacteria infect after penetration of the epidermis and transfer t-DNA copies into the cells, where the t-DNA is episomally transcribed and the mRNA is translated, resulting in the production of the protein of interest in the infected cell, but the passage of the t-DNA into the nucleus is transient.
[0178] To aid in the identification of transformed plant cells, the constructs of the present disclosure can be further engineered to include a plant selectable marker. Useful selectable markers include enzymes that provide resistance to chemicals such as antibiotics, e.g., gentamicin, hygromycin, kanamycin, or herbicides such as phosphinothricin, glyphosate, and chlorosulfuron. Similarly, compounds that can be identified by a color change, such as GUS (β-glucuronidase), or enzymes that provide light emission, such as luciferase or GFP, can be used.
[0179] Transgenic plants, plant cells, or seeds containing the genetic constructs of the present disclosure, which can be used as suitable platform plants for transient protein expression as described herein, are also considered part of the present disclosure. Methods for regenerating whole plants from plant cells are also known in the art (see, for example, Guerineau and Mullineaux (1993, Plant transformation and expression vectors. In: Plant Molecular Biology Labfax (Croy RRD ed) Oxford, BIOS Scientific Publishers, pp 121-148). Generally, transformed plant cells are cultured in an appropriate medium which may contain a selection agent, such as an antibiotic, where a selectable marker is used to facilitate identification of transformed plant cells. Once callus is formed, shoot formation can be encouraged by using appropriate plant hormones according to known methods, and the shoots are transferred to rooting medium for plant regeneration. This plant may then be used to establish repeated generations from seed or using vegetative propagation techniques. Transgenic plants can also be generated without the use of tissue culture. Methods for stable transformation and regeneration of these organisms are well established in the art and known to those skilled in the art. Available techniques are described in detail in Vasil et al. (Cell Culture and Somatic Cell Genetics of Plants, Vol. 1, 11 and III, Laboratory Procedures and Their Applications, Academic Press, 1984), and Weissbach and Weissbach (Methods for Plant Molecular Biology, Academic Press, 1989). The method for obtaining transformed and regenerated plants is not critical to the present disclosure.
[0180] When a plant, plant part, or plant cell is transformed or co-transformed with two or more nucleic acid constructs, the nucleic acid constructs can be introduced into Agrobacterium in a single transfection, so that the nucleic acids are pooled and bacterial cells are transfected. Alternatively, the constructs can be introduced sequentially. In this case, the first construct is introduced into Agrobacterium as described, and the cells are grown under selective conditions (e.g., in the presence of an antibiotic) that allow only a single transformed bacterium to grow. Following this first selection step, the second nucleic acid construct is introduced into Agrobacterium as described, and the cells are grown under double selective conditions that allow only doubly transformed bacteria to grow. The doubly transformed bacterium can then be used to transform a plant, plant part, or plant cell as described herein, or can be subjected to a further transformation step to receive a third nucleic acid construct.
[0181] Alternatively, when a plant, plant part, or plant cell is transformed or co-transformed with two or more nucleic acid constructs, the nucleic acid constructs may be introduced into the plant by co-infiltrating the plant, plant part, or plant cell with a mixture of Agrobacterium cells, each of which may contain one or more of the constructs to be introduced into the plant. The concentrations of the various Agrobacterium populations containing the desired constructs may be varied during the infiltration process to vary the relative expression levels within the plant, plant part, or plant cell of the nucleotide sequences of interest within the constructs.
[0182] A nucleic acid or nucleotide sequence referred to in this disclosure may be "substantially homologous," "substantially similar," or "substantially identical" to a sequence, or the complement of a sequence, if the nucleic acid or nucleotide sequence hybridizes to one or more of the nucleotide sequences defined herein, or the complement of a nucleic acid or nucleotide sequence, under stringent hybridization conditions. A sequence is "substantially homologous," "substantially similar," or "substantially identical" if at least about 60%, or 60-100%, or any amount in between, of the nucleotides, e.g., 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100%, or any amount in between, match over the defined length of the nucleotide sequence, provided that such homologous sequence exhibits one or more of the characteristics of the sequence or encoded product described herein.
[0183] When referring to a particular sequence, the terms "percent similarity," "sequence similarity," "percent identity," or "sequence identity" are used, for example, as described in the University of Wisconsin GCG software program, or by manual alignment and visual inspection (see, for example, Current Protocols in Molecular Biology, Ausubel et al., eds. 1995 supplement). Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be achieved, for example, using the algorithm of Smith & Waterman, (1981, Adv. Appl. Math. 2:482), by the alignment algorithm of Needleman & Wunsch, (1970, J. Mol. Biol. 48:443), by the similarity search method of Pearson & Lipman, (1988, Proc. Natl. Acad. Sci. USA 85:2444), by computer implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group (GCG), 575 Science Dr., Madison, Wis.), or by manual alignment and visual inspection (see, for example, Current Protocols in Molecular Biology (Ausubel et al., eds. 1995 supplement)).
[0184] Examples of algorithms suitable for determining percent sequence identity and sequence similarity include the BLAST and BLAST 2.0 algorithms described in Altschul et al., (1977, Nuc. Acids Res. 25:3389-3402) and Altschul et al., (1990, J. Mol. Biol. 215:403-410), respectively. BLAST and BLAST 2.0 are used with the parameters described herein to determine percent sequence identity of the nucleic acids and proteins of the present disclosure. For example, the BLASTN program (for nucleotide sequences) can use as default a word length (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands. The BLASTP program may use as defaults a word length of 3, and an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, 1989, Proc. Natl. Acad. Sci. USA 89:10915) alignment (B) of 50, expectation (E) of 10, M=5, N=-4, and a comparison of both strands for amino acid sequences. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (see URL: ncbi.nlm.nih.gov / ).
[0185] Many organisms exhibit biases regarding the use of specific codons to encode the insertion of specific amino acids in growing peptide chains. Codon preference or codon bias, which is the difference in codon usage between organisms, is caused by the degeneracy of the genetic code and is well documented among many organisms. Codon bias often correlates with the translation efficiency of messenger RNA (mRNA), which is thought to depend, among other things, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be tailored for optimal gene expression in a given organism based on codon optimization. The process of optimizing the nucleotide sequence encoding a heterologously expressed protein can be a critical step for improving expression yield. Optimization requirements can include steps to improve the host's ability to produce the foreign protein.
[0186] There are various codon optimization techniques known in the art for improving the translational dynamics of protein coding regions that are inefficiently translated.These techniques mainly rely on identifying the codon usage frequency of specific host organisms.When specific gene or sequence is to be expressed in this organism, the coding sequence of such gene and sequence is modified to replace the codon of the sequence of interest with the more frequently used codon of host organisms.
[0187] "Codon optimization" is defined as modifying a nucleic acid sequence by replacing at least one, several, or a significant number of the codons of the native sequence with codons that may be more frequently or most frequently used in the genes of another organism or species for enhanced expression in the host or host cell of interest. Different species exhibit particular biases toward particular codons for particular amino acids.
[0188] The present disclosure includes codon-optimized synthetic polynucleotide sequences, for example, sequences optimized for human codon usage or plant codon usage. The codon-optimized polynucleotide sequences can then be expressed in a host, such as a plant. More specifically, sequences optimized for human codon usage or plant codon usage can be expressed in a plant. Without wishing to be bound by theory, it is believed that sequences optimized for human codons increase the guanine-cytosine content (GC content) of the sequence, improving expression yield when a plant is used as a host.
[0189] As used herein, the terms "construct," "vector," or "expression vector" refer to a recombinant nucleic acid for transferring an exogenous nucleotide sequence (e.g., a nucleotide sequence encoding a modified S protein described herein) into a host cell (e.g., a plant cell) and directing expression of the exogenous nucleic acid sequence within the host cell. An "expression cassette" refers to a nucleic acid containing a nucleotide sequence of interest under the control of, and operably linked to, an appropriate promoter or other regulatory elements for transcription of the nucleic acid of interest within the host cell. As one skilled in the art will appreciate, an expression cassette can include a termination (terminator) sequence, which can be any sequence active within a host cell (e.g., a plant host). For example, in plants, the termination sequence can be derived from the RNA-2 genome segment of a bipartite RNA virus, e.g., a comovirus; the termination sequence can be a NOS terminator; or the terminator sequence can be obtained from the 3'UTR of the alfalfa plastocyanin gene.
[0190] Nucleic acids containing a nucleotide sequence encoding the modified S protein described herein can further comprise a sequence that enhances expression of the S protein in a host, a portion of a host, or a host cell. The sequence that enhances expression can include a 5'UTR enhancer element or a plant-derived expression enhancer operably associated with the nucleic acid encoding the modified viral structural protein. The sequence encoding the modified S protein can also be optimized to increase expression, for example, by optimizing for human codon usage, increased GC content, or a combination thereof.
[0191] "Regulatory region," "regulatory element," or "promoter" typically, but not always, refers to a portion of nucleic acid upstream of the protein-coding region of a gene, which may be composed of either DNA or RNA, or both DNA and RNA. When a regulatory region is active and functionally associated with or operably linked to a nucleotide sequence of interest, it can result in the expression of the nucleotide sequence of interest. Regulatory elements may be capable of mediating organ specificity or controlling developmental or temporal gene activation. "Regulatory region" includes promoter elements, core promoter elements that exhibit basal promoter activity, elements that are inducible in response to external stimuli, and elements that mediate promoter activity, such as negative regulatory elements or transcriptional enhancers. As used herein, "regulatory region" also includes regulatory elements that regulate gene expression, such as post-transcriptionally active elements, such as translational and transcriptional enhancers, translational and transcriptional repressors, upstream activating sequences, and mRNA instability determinants. Some of these latter elements may be located proximal to the coding region.
[0192] In the context of this disclosure, the term "regulatory element" or "regulatory region" typically refers to a sequence of DNA upstream (5') of the coding sequence of a structural gene, which controls expression of the coding region by providing recognition for RNA polymerase and / or other factors necessary for transcription to initiate at a specific site. However, it should be understood that other nucleotide sequences located within introns or within the 3' region of the sequence can also contribute to regulating expression of a coding region of interest. An example of a regulatory element that provides recognition for RNA polymerase or other transcription factors to ensure initiation at a specific site is a promoter element. Most, but not all, eukaryotic promoter elements contain a TATA box, a conserved nucleic acid sequence consisting of adenosine and thymidine nucleotide base pairs that is usually located approximately 25 base pairs upstream of the transcription start site. A promoter element can include a basal promoter element responsible for initiating transcription and other regulatory elements that modify gene expression.
[0193] There are several types of regulatory regions, including developmentally regulated, inducible, or constitutive. Developmentally regulated regulatory regions, or those that control the differential expression of genes under their control, are activated in an organ or tissue at a specific time during the development of that organ or tissue. However, some developmentally regulated regulatory regions may be preferentially active in a particular organ or tissue at a specific developmental stage, and may be active in a developmentally regulated manner or at a basal level in other organs or tissues within the plant as well. Examples of tissue-specific regulatory regions, such as specific regulatory regions, include the napin promoter and the cultiferin promoter (Rask et al., 1998, J. Plant Physiol. 152:595-599; Bilodeau et al., 1994, Plant Cell 14:125-130). An example of a leaf-specific promoter is the plastocyanin promoter (see U.S. Patent No. 7,125,978, incorporated herein by reference).
[0194] An inducible regulatory region is one that can directly or indirectly activate the transcription of one or more DNA sequences or genes in response to an inducer. In the absence of an inducer, the DNA sequence or gene is not transcribed. Typically, the protein factor that specifically binds to the inducible regulatory region and activates transcription may exist in an inactive form, and then be directly or indirectly converted into an active form by the inducer. However, the protein factor does not have to be present. The inducer can be a protein, a metabolite, a growth regulator, a chemical substance such as a herbicide or a phenolic compound, or a physiological stress imposed directly by heat, cold, salt, or a toxic element, or indirectly through the action of a pathogen or disease agent such as a virus. Plant cells containing an inducible regulatory region can be exposed to an inducer by externally applying the inducer to the cells or plant, such as by spraying, watering, heating, or similar methods. Inducible regulatory elements can be derived from either plant or non-plant genes (see, for example, Gatz, C. and Lenk, IRP, 1998, Trends Plant Sci. 3, 352-358).Examples of potential inducible promoters include, but are not limited to, tetracycline-inducible promoters (Gatz, C., 1997, Ann. Rev. Plant Physiol. Plant Mol. Biol. 48, 89-108), steroid-inducible promoters (Aoyama, T. and Chua, NH, 1997, Plant J. 2, 397-404), and ethanol-inducible promoters (Salter, MG, et al., 1998, Plant Journal 16, 127-132; Caddick, MX, et al., 1998, Nature Biotech. 16, 177-180), cytokinin-inducible IB6 and CKI1 genes (Brandstatter, I. and Kieber, JJ, 1998, Plant Cell 10, 1009-1019; Kakimoto, T., 1996, Science 274, 982-985) and the auxin-inducible element, DR5 (Ulmasov, T., et al., 1997, Plant Cell 9, 1963-1971).
[0195] Constitutive regulatory regions direct the expression of a gene throughout different parts of the plant and continuously throughout plant development. Examples of known constitutive regulatory elements include the CaMV 35S transcript. (p35S; Odell et al., 1985, Nature, 313:810-812, incorporated herein by reference), rice actin 1 (Zhang et al., 1991, Plant Cell, 3:1155-1165), actin 2 (An et al., 1996, Plant J., 10:107-121), or tms 2 (U.S. Pat. No. 5,428,147), and triosephosphate isomerase 1 (Xu et al., 1994, Plant Physiol. 106:459-467) genes, maize ubiquitin 1 gene (Cornejo et al., 1993, Plant Mol. Biol. 29:637-646), Arabidopsis ubiquitin 1 and 6 genes (Holtorf et al., 1995, Plant J. Mol. Biol. 29:637-646), the tobacco translation initiation factor 4A gene (Mandel et al., 1995 Plant Mol. Biol. 29:995-1004), the cassava vein mosaic virus promoter, pCAS (Verdaguer et al., 1996); the promoter for the small subunit of ribulose biphosphate carboxylase, pRbcS (Outchkourov et al., 2003), and the promoter associated with pUbi (for monocotyledonous and dicotyledonous plants).
[0196] As used herein, the term "constitutive" does not necessarily indicate that a nucleotide sequence under the control of a constitutive regulatory region is expressed at the same level in all cell types, but rather indicates that the gene is expressed in a wide range of cell types, although variations in abundance are often observed.
[0197] One or more of the genetic constructs of the present disclosure may also optionally contain additional enhancers, either translational or transcriptional enhancers. Enhancers may be located 5' or 3' to the transcribed sequence. Enhancer regions are well known to those skilled in the art and may include an ATG start codon, adjacent sequences, etc. The start codon, if present, may be in phase ("in frame") with the reading frame of the coding sequence to provide for correct translation of the transcribed sequence.
[0198] The terms "5'UTR" or "5' untranslated region," "5' leader sequence," or "5'UTR enhancer element" refer to a region of an mRNA that is not translated. The 5'UTR typically begins at the transcription start site and ends at the translation start site, or immediately before the start codon of the coding region. The 5'UTR can regulate the stability and / or translation of the mRNA transcript.
[0199] As used herein, the term "plant-derived expression enhancer" refers to a nucleotide sequence obtained from a plant, which encodes a 5'UTR. Examples of plant-derived expression enhancers are described in U.S. Provisional Patent Application No. 62 / 643,053 (filed March 14, 2018) and International Application No. PCT / CA2019 / 050319 (filed March 14, 2019), which are incorporated herein by reference, or Diamos AGet et al. (2016, Front Plant Sci. 7:1-15), which are incorporated herein by reference. Plant-derived expression enhancers can be selected from nbEPI42, nbSNS46, nbCSY65, nbHEL40, nbSEP44, nbMT78, nbATL75, nbDJ46, nbCHP79, nbEN42, atHSP69, atGRP62, atPK65, atRP46, nb30S72, nbGT61, nbPV55, nbPPI43, nbPM64 and nbH2A86, as described in U.S. Patent No. 62 / 643,053 and PCT / CA2019 / 050319. Plant-derived expression enhancers can be used in plant expression systems that include a regulatory region operably linked to a plant-derived expression enhancer sequence and a nucleotide sequence of interest, such as a nucleotide sequence encoding a modified S protein.
[0200] RNA stability and / or translation efficiency can be further improved by including a 3' untranslated region (3'UTR). Thus, one or more of the genetic constructs herein can further include a 3'UTR.
[0201] The 3' untranslated region may contain a polyadenylation signal and any other regulatory signals that can effect mRNA processing or gene expression. Polyadenylation signals are usually characterized by the addition of a polyadenylic acid track to the 3' end of the mRNA precursor. Polyadenylation signals are generally recognized by the presence of homology to the canonical form 5' AATAAA-3', although variations are not uncommon. Non-limiting examples of suitable 3' regions include those derived from Agrobacterium tumor-inducing (Ti) plasmid genes, such as nopaline synthase (Nos gene), and plant genes, such as soybean storage protein genes, the small subunit of the ribulose-1,5-bisphosphate carboxylase gene (ssRUBISCO; U.S. Pat. No. 4,962,028, incorporated herein by reference), the promoter used to regulate plastocyanin expression described in U.S. Pat. No. 7,125,978, incorporated herein by reference, the 3' UTR from Arracacha virus B isolate gene (AvB), the 3' UTR from Beet necrotic yellow vein virus (trBNYVV), the 3' UTR from Southern bean mosaic virus (SBMV), the 3' UTR from Turnip ringspot virus (TuRSV), the 3' UTR from Cowpea mosaic virus (Cowpea mosaic virus), and the 3' UTR from Cowpea mosaic virus (Cowpea mosaic virus). Examples of suitable 3' untranslated regions containing polyadenylation signals include the 3' UTR from the Common Bean Virus (CPMV), the 3' UTR from the Broad Bean True Mosaic Virus (BBTMV), or the 3' UTR from the Ourmia melon virus (trOUMV). 3' UTRs can be used in conjunction with 5' UTRs from heterologous sequences to modulate expression levels.
[0202] Thus, "constructs," "vectors," "expression vectors," or "expression cassettes" are provided that include a nucleic acid that is under the control of a 3'UTR and includes a nucleotide sequence of interest (such as a modified viral structural protein) operably (or functionally) linked to the 3'UTR. Additionally, the nucleic acid may include a 3'UTR operably (or functionally) linked to a nucleotide sequence of interest (such as a modified viral structural protein).
[0203] The modified S proteins or VLPs comprising the modified S proteins described herein can be used to elicit an immune response in a subject.
[0204] "Immune response" generally refers to the response of a subject's adaptive immune system. The adaptive immune system generally includes humoral responses and cell-mediated responses. The humoral response is an aspect of immunity mediated by secreted antibodies produced in cells of the B lymphocyte lineage (B cells). Secreted antibodies bind to antigens on the surface of invading microorganisms (such as viruses or bacteria) and mark them for destruction. Humoral immunity is generally used to refer to antibody production and associated processes, as well as antibody effector functions, including Th2 cell activation and cytokine production, memory cell generation, opsonization of phagocytosis, pathogen elimination, and the like. The terms "modulate" or "modulation," etc., refer to an increase or decrease in a particular response or parameter as determined by any of several commonly known or used assays, some of which are exemplified herein.
[0205] Cell-mediated responses are immune responses that do not involve antibodies, but involve the activation of macrophages, natural killer cells (NK), antigen-specific cytotoxic T lymphocytes, and the release of various cytokines in response to antigens. Cell-mediated immunity is commonly used to refer to some Th cell activation, Tc cell activation, and T cell-mediated responses. Cell-mediated immunity can be particularly important in response to viral infections.
[0206] For example, induction of antigen-specific CD8-positive T lymphocytes can be measured using an ELISPOT assay. Stimulation of CD4-positive T lymphocytes can be measured using a proliferation assay. Anti-coronavirus antibody titers can be quantified using an ELISA assay. The isotype of antigen-specific or cross-reactive antibodies can also be measured using anti-isotype antibodies (e.g., anti-IgG, anti-IgA, anti-IgE, or anti-IgM). Methods and techniques for performing such assays are well known in the art.
[0207] The presence or level of cytokines can also be quantified. For example, T helper cell response (Th1 / Th2) can be characterized by measuring IFN-γ secreting cells and IL-4 secreting cells by ELISA (e.g., BD Biosciences OptEIA kit). Peripheral blood mononuclear cells (PBMCs) or splenocytes obtained from subjects can be cultured and the supernatant analyzed. T lymphocytes can also be quantified by fluorescence-activated cell sorting (FACS) using marker-specific fluorescent labels and methods known in the art.
[0208] Microneutralization assays may be performed to characterize a subject's immune response, see, for example, the method of Rowe et al., 1973. Virus neutralization titers may be quantified by several methods, including counting lytic plaques following crystal violet fixation / colorization of cells (plaque assay); microscopic observation of cell lysis in in vitro cultures; and ELISA and spectrophotometric detection of coronaviruses.
[0209] As used herein, the term "epitope" or "epitopes" refers to the structural portion of an antigen to which an antibody specifically binds.
[0210] Methods for producing antibodies or antibody fragments are provided, comprising administering to a subject or host animal a modified S protein, a trimer or trimer-modified S protein, a VLP comprising the modified S protein, a composition, or a vaccine described herein, thereby producing the antibody or antibody fragment. Antibodies or antibody fragments produced by the methods are also provided.
[0211] Thus, the present disclosure also provides the use of a modified coronavirus S protein or a VLP comprising the modified coronavirus S protein described herein to induce immunity to a coronavirus infection in a subject. Also disclosed herein are antibodies or antibody fragments prepared by administering the modified coronavirus S protein or a VLP comprising the modified coronavirus S protein to a subject or host animal.
[0212] Further provided are compositions comprising an effective dose of a modified coronavirus S protein or a VLP comprising the modified coronavirus S protein described herein and a pharmaceutically acceptable carrier, adjuvant, vehicle, or excipient for inducing an immune response in a subject. Also provided are vaccines for inducing an immune response against coronavirus in a subject, the vaccine comprising an effective dose of a modified coronavirus S protein or a VLP comprising the modified coronavirus S protein.
[0213] The composition or vaccine can include VLPs containing a modified S protein from one type of coronavirus family, subgroup, type, subtype, lineage, subgenus, or strain, or the composition or vaccine can include multiple VLP types, each VLP type containing a modified S protein, and the modified S proteins within the same VLP are from one type of coronavirus family, subgroup, type, subtype, lineage, subgenus, or strain; i.e., the composition or vaccine can include a mixture of different coronavirus VLPs, each VLP containing a modified S protein from the same coronavirus family, subgroup, type, subtype, lineage, subgenus, or strain. For example, the composition or vaccine can include a first VLP containing a first modified S protein from a first coronavirus family, subgroup, type, subtype, lineage, or strain, and a second VLP containing a second modified S protein from a second coronavirus family, subgroup, type, subtype, lineage, or strain. Additionally, the composition may also include a third VLP comprising a third modified S protein from a third coronavirus family, subgroup, type, subtype, lineage or strain, and / or the composition or vaccine may include a fourth VLP comprising a fourth modified S protein from a fourth coronavirus family, subgroup, type, subtype, lineage, subgenus or strain.
[0214] The composition or vaccine may further comprise a VLP comprising modified S proteins from multiple types of coronavirus family, subgroup, type, subtype, lineage, subgenus, or strain. For example, a VLP may comprise a first modified S protein from a first coronavirus family, subgroup, type, subtype, lineage, or strain and a second modified S protein from a second coronavirus family, subgroup, type, subtype, lineage, or strain. Additionally, the VLP may comprise a third modified S protein from a third coronavirus family, subgroup, type, subtype, lineage, or strain, and / or a fourth modified S protein from a fourth coronavirus family, subgroup, type, subtype, lineage, subgenus, or strain.
[0215] Thus, the present specification also provides compositions or vaccines that are monovalent (univalent) or multivalent (polyvalent). A monovalent composition or vaccine can immunize a subject against a single type of coronavirus strain, while a multivalent composition or vaccine can immunize a subject against multiple coronavirus strains. For example, a composition or vaccine can be a bivalent composition or vaccine that, upon administration, can immunize a subject against two different types of coronavirus families, subgroups, types, subtypes, lineages, or strains. Furthermore, a composition or vaccine can be a trivalent composition, or a vaccine or composition can be a tetravalent or quadrivalent composition or vaccine. Furthermore, a vaccine can also be multivalent against different types of viruses. For example, a vaccine can immunize a subject against one or more coronavirus strains (a first type of virus) and against a second type of virus, e.g., influenza virus.
[0216] The monovalent or multivalent composition or vaccine may further comprise a pharmaceutically acceptable carrier, adjuvant, vehicle or excipient for inducing an immune response in a subject.
[0217] Adjuvant systems for enhancing a subject's immune response to vaccine antigens are well known and can be used with the vaccines or pharmaceutical compositions described herein. There are many types of adjuvants that can be used. Common adjuvants for human use are aluminum hydroxide, aluminum phosphate, and calcium phosphate. There are also some adjuvants based on oil emulsions (oil-in-water emulsions or water-in-oil emulsions, such as Freund's incomplete adjuvant (FIA), Montanide™, Adjuvant 65, and Lipovant™), bacterial products (or their synthetic derivatives), endotoxins, fatty acids, paraffin or vegetable oils, cholesterol, and aliphatic amines, or natural organic compounds, such as squalene. Non-limiting adjuvants that can be used include, for example, an oil-in-water emulsion of squalene oil (e.g., MF-59 or AS03), an adjuvant composed of the synthetic TLR4 agonist glucopyranosyl lipid A (GLA) incorporated into a stable emulsion (SE) (GLA-SE), or the Toll-like receptor (TLR9) agonist adjuvant CpG 1018.
[0218] Therefore, the vaccine or pharmaceutical composition may contain one or more adjuvants. For example, the vaccine or pharmaceutical composition may contain aluminum hydroxide, aluminum phosphate, calcium phosphate, an oil-in-water emulsion or a water-in-oil emulsion, an emulsion containing squalene (e.g., MF-59 or AS03), an emulsion containing GLA-SE, or a CpG 1018 adjuvant. In one embodiment, the adjuvant is AS03.
[0219] The pharmaceutical compositions, vaccines or formulations herein may be manufactured in a manner that is itself known, for example by means of conventional mixing, dissolving, granulating, dragee-making, levigating, emulsifying, encapsulating, entrapping or tabletting processes.
[0220] A pharmaceutical composition, vaccine or formulation may be produced by mixing or premixing any of the components prior to administration, for example by manual or mechanically assisted mixing of two or more vaccine suspensions, pharmaceutically acceptable carriers, adjuvants, vehicles or excipients as a step performed before the final formulation, vaccine or pharmaceutical composition is administered.
[0221] The pharmaceutical composition, vaccine or formulation may be administered to a subject orally, intradermally, intranasally, intramuscularly, intraperitoneally, intravenously or subcutaneously.
[0222] Injectables can be prepared in conventional forms, as liquid solutions or suspensions, as solid forms suitable for solution or suspension in liquid prior to injection, or as emulsions. Suitable excipients include, for example, water, saline, dextrose, mannitol, lactose, lecithin, albumin, sodium glutamate, cysteine hydrochloride, and the like. In addition, if desired, injectable pharmaceutical compositions may contain small amounts of non-toxic auxiliary substances, such as wetting agents, pH buffering agents, and the like. Physiologically compatible buffers include, but are not limited to, Hank's solution, Ringer's solution, or physiological saline buffer. If desired, absorption-enhancing preparations (e.g., liposomes) may be utilized.
[0223] The composition or vaccine can be administered to a subject once (single dose). Furthermore, the vaccine or composition can be administered to a subject multiple times (multiple doses). Thus, the composition, formulation, or vaccine can be administered to a subject in a single dose to induce an immune response, or the composition, formulation, or vaccine can be administered multiple times (multiple doses). For example, a dose of the composition or vaccine can be administered 2, 3, 4, or 5 times. Thus, the composition or vaccine can be administered to a subject with an initial dose, and one or more doses can be administered to the subject thereafter. The administration of the doses can be separated in time from each other. For example, after the administration of the initial dose, one or more subsequent doses can be administered 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, or any time therebetween from the administration of the initial dose. Furthermore, the composition or vaccine can be administered annually. For example, the composition or vaccine may be administered as a seasonal vaccine.
[0224] The disclosure further provides the following sequences: [Table 1-1] [Table 1-2]
[0225] The present invention is further illustrated by the following examples.
[0226] example Example 1: Construct Modified coronavirus S proteins in the 2X35S-nbHEL40 / AvB-NOS terminator expression system (construct numbers 9125, 9588, 9801, 9802, 9808, 9513, 10011, 10090, 10092, 10346, 10351, 10502, 10503, 10504, 10505, 10506, 10507, 10508, and 10509). The following PCR-based method was used to clone the Spike (S) protein coding sequence from B-strain SARS-CoV-2 hCoV-19 / USA / CA2 / 2020, in which the native signal peptide was replaced with that of alfalfa protein disulfide isomerase (PDISP / CoV S), into the 2X35S / nbHEL40-AvB / NOS expression system. The PDISP / CoV S sequence (SEQ ID NO: 3) was used as a template to amplify a fragment containing the PDISP / CoV S coding sequence using primers IF(PDI)S-3-NY-CoV.c (SEQ ID NO: 23) and IF(Avb)-H5I.r (SEQ ID NO: 24). The PCR product was cloned into the 2X35S / nbHEL40-AvB / NOS expression system using the In-Fusion cloning system (Clontech, Mountain View, CA). Construct #8716 (Figure 1A) was digested with AatII and StuI restriction enzymes, and the linearized plasmid was used in an In-Fusion assembly reaction. Construct #8716 is an acceptor plasmid intended for "In-Fusion" cloning of a gene of interest in a 2X35S / nbHEL40-AvB / NOS-based expression cassette. This acceptor plasmid also contains a gene construct for coexpression of the TBSV P19 suppressor of silencing under the alfalfa plastocyanin gene promoter and terminator. The backbone is a pCAMBIA binary plasmid, and the sequence from the left t-DNA border to the right t-DNA border is shown in SEQ ID NO: 21. The resulting construct was designated #9125 (SEQ ID NO: 22). The amino acid sequence of the complete B-strain SARS-CoV-2 spike fused to PDISP is shown in SEQ ID NO: 4. An image of plasmid #9125 is shown in Figure 1B.
[0227] [Table 2]
[0228] Constructs 9588, 9801, 9802, 9808, 9513, 10011, 10090, 10092, 10346, 10351, 10502, 10503, 10504, 10505, 10506, 10507, 10508 and 10509 were cloned using the same method as plasmid 9125, and a summary of the primers, templates, recipient vectors and products is shown in Table 3 below. [Table 3]
[0229] Example 2: Method Agrobacterium tumefaciens transfection Agrobacterium tumefaciens strain AGL1 was transfected by electroporation with a SARS-CoV-2 modified S protein expression vector using the method described by D'Aoust et al., 2008 (Plant Biotech. J. 6:930-40).
[0230] Plant biomass preparation, inoculum, and agroinfiltration N. benthamiana plants were grown from seeds in flats filled with a peat moss substrate. Plants were grown in a greenhouse under a controlled photoperiod and temperature regime. Three weeks after sowing, individual plantlets were removed, transplanted into pots, and grown in the greenhouse for an additional three weeks under the same environmental conditions.
[0231] OD of 0.6 to 1.6 600Agrobacteria transfected with each expression vector were grown until a total number of cells reached 1000. The Agrobacterium suspension was centrifuged before use, resuspended in infiltration medium (10 mM MgCl2 and 10 mM MES pH 5.6), and stored overnight at 4°C. On the day of infiltration, the culture batch was diluted to 2.5 culture volumes and warmed before use. Whole N. benthamiana plants were placed upside down in the bacterial suspension in an airtight stainless steel tank under vacuum for 2 minutes. The plants were returned to the greenhouse for an incubation period of 6–9 days before harvest.
[0232] Leaf harvesting and extraction of total protein and VLPs After incubation, the aerial parts of the plants were harvested, frozen at -80°C, and crushed into pieces. Total soluble protein was extracted by mechanically homogenizing (Polytron) each sample of freeze-ground plant material in two volumes of cold 50 mM Tris buffer, pH 8.0 + 500 mM NaCl, 0.4 μg / ml metadisulfite, and 1 mM phenylmethanesulfonyl fluoride. After homogenization, the slurry was centrifuged at 10,000 g for 10 minutes at 4°C, and these clarified crude extracts (supernatants) were saved for analysis.
[0233] The percentage of intact S protein in planta was assessed for clarified crude extracts and analyzed using capillary-based electrophoresis (Protein Simple, BioTechne) technology and a WES analysis system. Briefly, soluble proteins from crude extracts were separated by molecular weight in a capillary and immobilized on a matrix. Anti-S2 antibody (Novus biological, catalog number NB100-56578) was used for detection according to the manufacturer's instructions. In this method, two bands (intact S protein and a smaller fragment) are recognized by the anti-S2 antibody. The percentage of intact S protein was measured to assess the % intact protein. The percentage of intact S protein in planta for coronavirus S protein and modified coronavirus S protein is shown in Figures 3A, 3B, 6B, and 7.
[0234] Electron microscopy To determine whether the expressed S protein assembled into VLPs, immunocapture particle transmission electron microscopy (TEM) of purified VLPs was performed. Glow-discharged carbon / copper grids (10 s, 0.3 mbar) were placed on 20 μL of purified VLPs (100 μg / mL) for 5 min and then washed four times with sterile distilled water. The grids were floated on 20 μL of 2% uranyl acetate for 1 min, then the excess solution was removed by touching them to wet filter paper. The grids were then allowed to dry on the filter paper for 24 h before being observed under a TEM (Tecnai Microscope). TEM images of VLPs formed by coronavirus S protein and modified coronavirus S protein, as further described in this application, are shown in Figures 2A-2E and 6C.
[0235] Example 3: Evaluation of the complete S protein The percentage (%) of intact S protein in the implant was assessed by capillary Western blot using crude extracts as described above, and the modified coronavirus S protein was detected using antibodies and quantified using a standard curve. The implant percentage of full-length coronavirus S protein versus truncated coronavirus S protein ("purity" or "% intact S protein") is shown in Figures 3A, 3B, 6B, and 7 for the coronavirus S protein and modified coronavirus S protein, as further described herein.
[0236] Percent Drug Substance (DS) purity was assessed by densitometric analysis of Coomassie-stained proteins on SDS gels after small-scale clarification and purification to remove impurities, and immunologically related products were included in the quantification and purity measurements. Percent Drug Substance (DS) purity is shown in Figures 4A and 4B for coronavirus S protein and modified coronavirus S protein, as further described in this application.
[0237] Example 4: Stability assessment of S protein Modified and unmodified S proteins were extracted using enzymatic extraction at pH 5.5 (Enz. pH 5.5) or pH 6.1 (Enz. pH 6.1) or by mechanical extraction (Mech) (Figure 5). For enzymatic digestion, diced fresh leaves were incubated in a plastic container containing a pectinase mixture in 200 mM mannitol + 125 mM citrate + 500 mM NaCl + 25 mM EDTA + 0.04% (w / v) disulfite (pH 6.2), and the pH was adjusted to 5.5 or maintained at 6.1 for 5.25 h at 20 °C. For mechanical extraction, diced fresh biomass was extracted for 2 min using two volumes of the same buffer in a blender, and the pH was adjusted to 5.5. The extract was clarified using a 400 μm mesh and centrifuged at 4500 g for 15 min at 4 °C. The clarified extract was then incubated at 20°C for 20 hours, and the intact S protein percentage was assessed before and after overnight incubation using the method described above.
[0238] Example 5: Arrays SEQ ID NO: 1 Natural SARS-CoV-2 S protein wtTM / CT AA(P0DTC2)
[0239] SEQ ID NO: 2 Native SARS-CoV-2 S protein wtTM / CT AA (P0DTC2) without signal peptide (SP)
[0240] SEQ ID NO: 3 B CoV S(GSAS-2P) DNA
[0241] Allocation number 4 B CoV S(GSAS-2P)AA
[0242] SEQ ID NO:5 C.37 CoV S(GSAS-2P) DNA
[0243] Allocation number 6 C.37 CoV S(GSAS-2P)AA
[0244] SEQ ID NO:7 B CoV S(GSAS-2P+del246-252)DNA
[0245] SEQ ID NO:8 B CoV S(GSAS-2P+del246-252)AA
[0246] SEQ ID NO:9 B CoV S(GSAS-2P+D253N)DNA
[0247] SEQ ID NO: 10 B CoV S(GSAS-2P+D253N)AA
[0248] SEQ ID NO: 11 B CoV S(GSAS-2P+del246-252+D253N)DNA
[0249] SEQ ID NO: 12 B CoV S(GSAS-2P+del246-252+D253N)AA
[0250] SEQ ID NO: 13 B.1.617.2 CoV S(GSAS-2P) DNA
[0251] Allocation number 14 B.1.617.2 CoV S(GSAS-2P)AA
[0252] SEQ ID NO: 15 B.1.617.2 CoV S(GSAS-2P+D253N)DNA
[0253] SEQ ID NO: 16 B.1.617.2 CoV S(GSAS-2P+D253N)AA
[0254] SEQ ID NO: 17 B.1.1.529 CoV S(GSAS-2P) DNA
[0255] Allocation number 18 B.1.1.529 CoV S(GSAS-2P)AA
[0256] SEQ ID NO: 19 B.1.1.529 CoV S(GSAS-2P+D253N)DNA
[0257] SEQ ID NO: 20 B.1.1.529 CoV S(GSAS-2P+D253N)AA
[0258] SEQ ID NO: 21 Left T-DNA to right T-DNA cloning vector 8716
[0259] SEQ ID NO: 22 2X35S promoter to NOS terminator construct 9125
[0260] SEQ ID NO: 23 IF(PDI)S-3-NY-CoV.c TCTCAGATCTTCGCGTCCCAATGCGTGAATCTTACGACGCGAACACAGT
[0261] SEQ ID NO: 24 IF(Avb)-H5I.r acgacacgactaaggcctttaaatgcaaattctgcattgtaacgatcc
[0262] SEQ ID NO: 25 Mature B CoV S(GSAS-2P+del246-252)AA without SP
[0263] SEQ ID NO: 26 Mature B CoV S(GSAS-2P+D253N)AA without SP
[0264] SEQ ID NO: 27 Mature B CoV S(GSAS-2P+del246-252+D253N)AA without SP
[0265] SEQ ID NO: 28 Mature B CoV S(GSAS-2P+S13+G252N)AA without SP
[0266] SEQ ID NO: 29 Mature B CoV S(GSAS-2P+S13+P251N+D253T)AA without SP
[0267] SEQ ID NO: 30 Mature B CoV S(GSAS-2P+S13+Δ247-250)AA without SP
[0268] SEQ ID NO: 31 Mature B CoV S(GSAS-2P+S13+Δ248-251)AA without SP
[0269] SEQ ID NO: 32 Mature B CoV S(GSAS-2P+S13+Δ249-252)AA without SP
[0270] SEQ ID NO: 33 Mature B CoV S(GSAS-2P+S13+Δ246-250)AA without SP
[0271] SEQ ID NO: 34 Mature B CoV S(GSAS-2P+S13+Δ247-251)AA without SP
[0272] SEQ ID NO: 35 Mature B CoV S(GSAS-2P+S13+Δ248-252)AA without SP
[0273] SEQ ID NO: 36 Mature B CoV S(GSAS-2P+S13+Δ246-251)AA without SP
[0274] SEQ ID NO: 37 Mature B CoV S(GSAS-2P+S13+Δ247-252)AA without SP
[0275] SEQ ID NO: 38 B CoV S(GSAS-2P+S13+G252N) DNA
[0276] SEQ ID NO: 39 B CoV S(GSAS-2P+S13+G252N)AA
[0277] Allocation number 40 B CoV S(GSAS-2P+S13+P251N+D253T)DNA
[0278] Allocation number 41 B CoV S(GSAS-2P+S13+P251N+D253T)AA
[0279] Allocation number 42 B CoV S(GSAS-2P+S13+Δ247-250)DNA
[0280] Allocation number 43 B CoV S(GSAS-2P+S13+Δ247-250)AA
[0281] Allocation number 44 B CoV S(GSAS-2P+S13+Δ248-251)DNA
[0282] Allocation number 45 B CoV S(GSAS-2P+S13+Δ248-251)AA
[0283] SEQ ID NO: 46 B CoV S(GSAS-2P+S13+Δ249-252)DNA
[0284] SEQ ID NO: 47 B CoV S(GSAS-2P+S13+Δ249-252)AA
[0285] SEQ ID NO: 48 B CoV S(GSAS-2P+S13+Δ246-250)DNA
[0286] SEQ ID NO: 49 B CoV S(GSAS-2P+S13+Δ246-250)AA
[0287] Allocation number 50 B CoV S(GSAS-2P+S13+Δ247-251)DNA
[0288] Allocation number 51 B CoV S(GSAS-2P+S13+Δ247-251)AA
[0289] Allocation number 52 B CoV S(GSAS-2P+S13+Δ248-252)DNA
[0290] Allocation number 53 B CoV S(GSAS-2P+S13+Δ248-252)AA
[0291] SEQ ID NO:54 B CoV S(GSAS-2P+S13+Δ246-251)DNA
[0292] SEQ ID NO: 55 B CoV S(GSAS-2P+S13+Δ246-251)AA
[0293] Allocation number 56 B CoV S(GSAS-2P+S13+Δ247-252)DNA
[0294] Allocation number 57 B CoV S(GSAS-2P+S13+Δ247-252)AA
Claims
1. A modified coronavirus S protein comprising one or more amino acid sequence modifications compared to the corresponding parent amino acid sequence, wherein the one or more modifications stabilize the modified coronavirus S protein, i) Substitutions of one or more amino acids to introduce an N-glycosylation site at a position corresponding to position 251, 252, or 253 of reference sequence SEQ ID NO: 1, wherein the N-glycosylation site is asparagine (N) within the consensus sequence N-X-(S or T), or ii) A deletion of at least four consecutive amino acid residues, the deletion comprising at least the residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO:
1. Modified coronavirus S protein, including
2. i) The modified coronavirus S protein according to claim 1, wherein the modified coronavirus S protein comprising the substitution of one or more amino acids for introducing the N-glycosylation site further comprises the deletion of one or more amino acids.
3. The modified coronavirus S protein according to claim 1, wherein the deletion includes the following: i) At least amino acid residues corresponding to positions 247, 248, 249, and 250 of reference sequence SEQ ID NO: 1; ii) At least amino acid residues corresponding to positions 248, 249, 250, and 251 of reference sequence SEQ ID NO: 1; iii) At least amino acid residues corresponding to positions 249, 250, 251, and 252 of reference sequence SEQ ID NO: 1; iv) At least amino acid residues corresponding to positions 246, 247, 248, 249, and 250 of reference sequence SEQ ID NO: 1; v) At least amino acid residues corresponding to positions 247, 248, 249, 250, and 251 of reference sequence SEQ ID NO: 1; vi) At least amino acid residues corresponding to positions 248, 249, 250, 251, and 252 of reference sequence SEQ ID NO: 1; vii) At least amino acid residues corresponding to positions 246, 247, 248, 249, 250, and 251 of reference sequence SEQ ID NO: 1; viiii) At least amino acid residues corresponding to positions 247, 248, 249, 250, 251 and 252 of reference sequence SEQ ID NO: 1; or ix) At least amino acid residues corresponding to positions 246, 247, 248, 249, 250, 251, and 252 of reference sequence SEQ ID NO:
1.
4. i) The N-glycosylation site is introduced into the amino acid corresponding to position 251 of reference sequence SEQ ID NO: 1, and the deletion comprises at least amino acid residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1; ii) The N-glycosylation site is introduced into the amino acid corresponding to position 252 of reference sequence SEQ ID NO: 1, and the deletion comprises at least amino acid residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1; iii) The N-glycosylation site is introduced into the amino acid corresponding to position 253 of reference sequence number 1, and the deletion comprises at least amino acid residues corresponding to positions 249 and 250 of reference sequence number 1; iv) The N-glycosylation site is introduced into the amino acid corresponding to position 253 of reference sequence number 1, the deletion comprises at least four consecutive amino acid residues, and the deletion comprises at least residues corresponding to positions 249 and 250 of reference sequence number 1; v) The N-glycosylation site is introduced into the amino acid corresponding to position 253 of reference sequence SEQ ID NO: 1, and the deletion includes at least amino acid residues corresponding to positions 246, 247, 248, 249, 250, 251 and 252 of reference sequence SEQ ID NO: 1, The modified coronavirus S protein according to claim 2.
5. The modified coronavirus S protein according to claim 1, wherein the amino acid sequence modification includes a substitution to asparagine (N) at a position corresponding to position 252 or 253 of reference sequence number 1, or the amino acid sequence modification includes a substitution to asparagine (N) at a position corresponding to position 251 of reference sequence number 1 and a substitution to threonine (T) at a position corresponding to position 253.
6. The modified coronavirus S protein according to claim 1, comprising 80% to 100% identity with the sequence of sequence numbers 8, 10, 12, 16, 20, 39, 41, 43, 45, 47, 49, 51, 53, 55, or 57.
7. The modified coronavirus S protein according to claim 1, wherein the modified coronavirus S protein is a chimeric S protein, and the chimeric S protein includes a cytoplasmic tail derived from influenza hemagglutinin.
8. The modified coronavirus S protein according to claim 1, wherein the parent amino acid sequence is derived from betacoronavirus.
9. The modified coronavirus S protein according to claim 8, wherein the beta-coronavirus is derived from lineage A, B, C, or D of beta-coronaviruses.
10. The modified coronavirus S protein according to claim 1, comprising a plant-specific N-glycan.
11. A nucleic acid comprising a nucleotide sequence encoding the modified coronavirus S protein described in claim 1.
12. A virus-like particle (VLP) containing the modified coronavirus S protein described in claim 1.
13. A vaccine for inducing an immune response, comprising an effective dose of the VLP described in claim 12.
14. The vaccine according to claim 13, which is a polyvalent vaccine containing a mixture of VLPs.
15. A method for inducing immunity against coronavirus infection in a subject, comprising administering the vaccine described in claim 13 to the subject.
16. A host or host cell comprising the VLP described in Claim 12.
17. A method for producing virus-like particles (VLPs) in a host or within host cells, a) Introducing the nucleic acid described in claim 11 into the host or host cell, or providing the host or host cell containing the nucleic acid described in claim 11, and b) A method comprising incubating the host or host cells under conditions that enable the expression of the nucleic acid, thereby producing the VLP, and optionally extracting and purifying the VLP from the host or host cells.
18. VLP produced by the method of claim 17.
19. A method for increasing the production of full-length coronavirus S protein in a host or host cell by modifying the parent coronavirus S protein, a) Introducing the nucleic acid described in claim 11 into the host or host cell, or providing the host or host cell containing the nucleic acid described in claim 11, and b) A method comprising incubating the host or host cells under conditions that enable the expression of the nucleic acid, thereby producing a modified coronavirus S protein, wherein a greater amount or a higher proportion of the modified coronavirus S protein is a fully modified coronavirus S protein compared to the parent coronavirus S protein produced in the host or host cells under similar conditions, and optionally extracting and purifying the modified coronavirus S protein from the host or host cells.
20. A method for producing a modified coronavirus S protein having one or more amino acid sequence modifications by modifying the coronavirus S protein, wherein the one or more amino acid modifications stabilize the modified coronavirus S protein, and the method i) Introducing one or more amino acid substitutions to the coronavirus S protein to introduce an N-glycosylation site at a position corresponding to position 251, 252, or 253 of reference sequence SEQ ID NO: 1, wherein the N-glycosylation site is asparagine (N) within the consensus sequence N-X-(S or T), or ii) A method comprising modifying the coronavirus S protein by introducing a deletion of at least four consecutive amino acid residues, the deletion of which includes at least four residues corresponding to positions 249 and 250 of reference sequence SEQ ID NO: 1.