Methods and compositions for enhancing recombinant protein production in plants using endoplasmic reticulum-resident molecular chaperones

Co-expression of ER-resident molecular chaperones in plants addresses the low yield and misfolding issues of recombinant proteins by enhancing protein folding and accumulation, providing a sustainable solution for industrial-scale production of casein and other mammalian proteins.

WO2026076437A1PCT designated stage Publication Date: 2026-04-09MOZZA FOODS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-04
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

The challenge of efficiently producing biologically active recombinant proteins, particularly mammalian proteins like casein, in plant hosts is hindered by the limitations of the endoplasmic reticulum (ER) protein folding machinery, leading to low yields and misfolding, aggregation, and degradation, which are not adequately supported in plant cells.

Method used

Co-expression of specific endoplasmic reticulum (ER)-resident molecular chaperones, such as Glyma.01G003700 and bovine DNAJB12, with heterologous proteins in transgenic plants to enhance protein folding and accumulation, resulting in up to 4-fold increases in yield.

Benefits of technology

The co-expression of selected ER-resident molecular chaperones significantly enhances the production of complex proteins like casein in plants, offering a sustainable and scalable alternative for industrial-scale production of non-animal-derived proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025049544_09042026_PF_FP_ABST
    Figure US2025049544_09042026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and compositions for enhancing heterologous protein expression in plants are provided. The invention is directed to the co-expression of a heterologous protein of interest with a specific endoplasmic reticulum (ER)-resident molecular chaperone protein. Co-expressing chaperones such as a native soybean protein disulfide isomerase-like protein (Glyma.01G003700, Glyma.04g247900, Glyma.03G218300, or Glyma.18G204000), or a heterologous mammalian chaperone (bovine DNAJB12) with casein proteins in soybean cells resulted in a significant increase in casein accumulation, with up to 4-fold enhancement observed. The methods are broadly applicable for improving the yield of complex recombinant proteins, including milk proteins, therapeutic proteins, and industrial enzymes, in plant-based production systems, thereby overcoming common limitations related to protein folding and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. 719.001.002.PCTPCT PATENT APPLICATIONMETHODS AND COMPOSITIONS FOR ENHANCING RECOMBINANT PROTEIN PRODUCTION IN PLANTS USING ENDOPLASMIC RETICULUM-RESIDENT MOLECULAR CHAPERONESInventor(s): Cory J. TOBINAssignee: Mozza Foods, Inc.1927 Zonal Avenue, Los Angeles, CA 90033Entity: Small business concernFiled Electronically on: October 4, 2025Docket No. 719.001.002.PCTMETHODS AND COMPOSITIONS FOR ENHANCING RECOMBINANT PROTEIN PRODUCTION IN PLANTS USING ENDOPLASMIC RETICULUM-RESIDENT MOLECULAR CHAPERONESCROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 703,899, filed on October 4, 2024; and 63 / 728,675, filed on December 5, 2024, wherein each application is incorporated herein by reference in its entirety.INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 719001002PCT SEQLIST, created on October 4, 2025, which is 24 kilobytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety.INCORPORATION BY REFERENCE

[0003] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BACKGROUND

[0004] The present disclosure relates to the field of genetic engineering and biotechnology.More specifically, it pertains to methods for optimizing and enhancing the expression of heterologous proteins, such as milk proteins, in transgenic plants through the co-expression of specific molecular chaperone proteins.

[0005] The efficient production of complex recombinant proteins in heterologous host systems, such as plants, remains a significant challenge in biotechnology. This is particularly true for mammalian proteins, which often require specific cellular machinery for proper folding, assembly, and post-translational modification to become biologically active. When expressed in plant hosts, differences in these cellular pathways can lead to protein misfolding, aggregation, and subsequentDocket No. 719.001.002.PCT degradation by the plant cell's endogenous quality control mechanisms, resulting in low yields and reduced functionality.

[0006] A protein class of significant commercial and nutritional interest is casein. Casein proteins are the primary protein component of mammalian milk, accounting for over 80% of the protein in bovine milk and forming the structural basis of casein micelles, which are essential for cheese production. The conventional production of casein relies on industrial-scale dairy farming. However, this model presents significant drawbacks. Environmentally, dairy farming is a major contributor to greenhouse gas emissions, consumes vast quantities of water, and contributes to land degradation. From an animal welfare perspective, common practices such as dehorning, routine forced insemination, and the separation of calves from mothers are increasingly viewed as ethically problematic.

[0007] Accordingly, there is a clear and long-felt need for a sustainable, scalable, and ethical alternative for producing biologically active casein proteins. An in vivo plant-based expression system offers a promising solution, potentially providing a cost-effective and environmentally friendly platform for industrial-scale production of caseins for use in food products, such as non- animal-derived dairy alternatives.

[0008] However, a major impediment to realizing this goal has been the difficulty of expressing caseins and other milk proteins at commercially viable levels in plants. Previous attempts have often resulted in low yields, highlighting a fundamental technical barrier. This low efficiency is largely attributed to the complex folding and assembly requirements of casein proteins, which are naturally optimized in the specialized environment of mammary epithelial cells but are not adequately supported in plant cells.

[0009] The endoplasmic reticulum (ER) is the primary cellular organelle responsible for the folding, modification, and quality control of secretory and transmembrane proteins. The ER contains a host of molecular chaperone proteins, such as protein disulfide isomerases (PDIs) and members of the DnaJ / Hsp40 family, that assist in the proper folding of newly synthesized polypeptides and prevent their aggregation. Recent genomic studies have revealed that the expression of specific ER- resident chaperones is significantly upregulated in mammary tissue during lactation, suggesting their critical role in facilitating the high-volume production of milk proteins like casein.

[0010] This observation suggests that the limited capacity of the ER protein folding machinery in plant cells may be a key bottleneck in the production of recombinant milk proteins. ItDocket No. 719.001.002.PCT has been hypothesized that supplementing the plant cell's endogenous chaperone system by coexpressing specific, highly effective chaperones could alleviate this bottleneck. However, the identification of which specific chaperones are most effective for enhancing the expression of a particular heterologous protein, such as casein, has remained a significant challenge. Furthermore, it was not known whether native plant chaperones or chaperones from heterologous systems (such as the native mammalian system) would be more effective, or how to best optimize their co-expression in a plant host.

[0011] Therefore, technical solutions to the problem of low recombinant protein yield in plants have been long sought but have largely eluded those skilled in the art. There remains a critical need for methods and compositions that can overcome these limitations and substantially enhance the accumulation of functional heterologous proteins in plants.SUMMARY

[0012] The present invention is based on the discovery that the co-expression of specific endoplasmic reticulum (ER)-resident molecular chaperones significantly enhances the accumulation and production of heterologous proteins, particularly complex mammalian proteins, in transgenic plant cells. This technical solution overcomes a long-standing challenge in biotechnology and molecular farming, namely the difficulty of achieving commercially viable expression levels of proteins that require extensive folding and assembly assistance. Some aspects of the disclosure provide: (i) nucleotide sequences inserted into plants, thereby leading to higher levels of expression of milk proteins, such as caseins in plants; and (ii) food products that comprise proteins produced by the nucleotide sequences.

[0013] The invention was demonstrated through the co-expression of selected molecular chaperones with casein proteins in soybean cells. It was discovered that overexpressing specific chaperones, including the native soybean protein Glyma.01G003700 (SEQ ID NO: 2) and the heterologous mammalian chaperone bovine DNAJB12 (SEQ ID NO: 4), led to a significant and reproducible increase in the accumulation of both alpha and kappa casein proteins, with observed increases of up to 4-fold compared to control expressions lacking the co-expressed chaperone.

[0014] Accordingly, in one aspect, the invention provides methods for increasing the expression of a heterologous protein in a plant or plant cell. In certain embodiments, the method comprises introducing into the plant or plant cell a first nucleic acid molecule encoding the heterologous protein and a second nucleic acid molecule encoding an ER -resident molecularDocket No. 719.001.002.PCT chaperone protein, and co-expressing said proteins, wherein the chaperone protein enhances the accumulation of the heterologous protein.

[0015] In certain embodiments, the molecular chaperone protein is selected from the group consisting of Glyma.01G003700 (SEQ ID NO: 2), bovine DNAJB12 (SEQ ID NO: 4), Glyma.04g247900 (SEQ ID NO: 6), and homologs or functional variants thereof. The invention also provides for the synergistic enhancement of protein production through the co-expression of multiple different molecular chaperones simultaneously.

[0016] In another aspect, the invention provides compositions for practicing these methods.These compositions include nucleic acid molecules, recombinant DNA constructs, vectors, and transgenic plant cells and plants.

[0017] In certain embodiments, a nucleic acid molecule of the invention comprises a nucleotide sequence that is at least 85%, 90%, 95%, 98%, or 100% identical to a sequence selected from the group consisting of SEQ ID NO: 1 (encoding Glyma.01G003700), SEQ ID NO: 3 (encoding bovine DNAJB12), and SEQ ID NO: 5 (encoding Glyma.04g247900). These nucleotide sequences may be codon-optimized for expression in a particular plant host, such as soybean.

[0018] In another embodiment, a recombinant DNA construct is provided, comprising a first gene encoding a heterologous protein of interest operably linked to a promoter, and a second gene encoding an ER-resident molecular chaperone protein operably linked to a promoter. The heterologous protein can be any protein that benefits from enhanced folding or stability in the ER, with a particular focus on milk proteins.

[0019] In preferred embodiments, the heterologous protein is a casein protein, such as uSl - casein, aS2-casein, P-casein, or K-casein, or combinations thereof. However, the methods and compositions are broadly applicable to other valuable proteins, including whey proteins, egg white proteins, therapeutic proteins (e.g., monoclonal antibodies), and industrial enzymes.

[0020] In a further aspect, the invention provides a transgenic plant or plant cell comprising the nucleic acid molecules and constructs described herein. While soybean (Glycine max) is a preferred plant host, the invention is applicable to a wide range of monocot and dicot plants, including but not limited to rice, corn, wheat, tobacco, alfalfa, and potato. In a preferred embodiment, the transgenic plant produces casein proteins that assemble into a micellar form.

[0021] In a variant, a method for increasing the expression of casein protein in a plant, comprises: inserting a plasmid containing SEQ ID NO: 7 or a plasmid containing SEQ ID NO: 8;Docket No. 719.001.002.PCT and overexpressing SEQ ID NO: 7 to produce SEQ ID NO: 9 or SEQ ID NO: 8 to produce SEQ ID NO: 10, thereby co-expressing SEQ ID NO: 9 or SEQ ID NO: 10 with one or more casein proteins, respectively.

[0022] In a variant, the plant is a soybean plant (Glycine max).

[0023] In a variant, plant comprises: a plasmid containing SEQ ID NO: 7 or a plasmid containing SEQ ID NO: 8, thereby introducing a transgene encoding a casein protein and a transgene encoding for SEQ ID NO: 9, or SEQ ID NO: 10, or homologs of the transgenes encoding for SEQ ID NO: 9 or SEQ ID NO: 10.

[0024] In a variant, the plant is a transgenic soybean plant.

[0025] In a variant, a recombinant DNA construct comprises a casein-encoding gene and one or more genes selected from SEQ ID NO: 7 or SEQ ID NO: 8, or homologues thereof.

[0026] The invention also provides food products and food compositions, such as dairy product substitutes, that comprise one or more heterologous proteins and / or one or more molecular chaperone proteins produced by the methods and transgenic plants disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Fig. l is a bar graph depicting relative expression levels of alpha and kappa casein proteins, when co-expressed with different molecular chaperones in soybean cells. Blue bars represent V5-tagged alpha casein normalized to RNA levels, while purple bars represent FLAG- tagged kappa casein normalized to RNA levels. Sample 3617 is the control expressing mScarlet fluorescent protein. Sample 3781 contains Glyma.04g247900 (PDI), Sample 3784 contains Glyma.01G003700 (PDI-like protein), and Sample 3786 contains bovine DNAJB12. Co-expression with bovine DNAJB12 (3786) shows the highest enhancement of alpha casein expression (~3.5-fold increase) and increased kappa casein expression (~2-fold), in comparison to the control.

[0028] Fig. 2 is a bar graph depicting normalized expression of V5-tagged alpha casein (blue) and FLAG-tagged kappa casein (purple) in soybean cells. Sample 3617 is the mScarlet control. Sample 3784, which is or contains, Glyma.01G003700, shows highest enhancement (~4- fold) for alpha casein and kappa casein, in comparison to Sample 3786 and Sample 3781. Sample 3786, which is or contains bovine DNAJB12, shows ~3-fold increase for alpha and 2-fold for kappa casein. Sample 3781, which is or contains Glyma.04g247900 / PDI, shows increases for alpha casein and kappa casein.Docket No. 719.001.002.PCT

[0029] Fig. 3 is a graph depicting alpha casein expression of two replicates (light blue: Repl, dark blue: Rep2), across four samples. Control Sample 3617 is a control sample corresponding to a baseline expression of alpha casein. Sample 3784, which is or contains Glyma.01G003700, shows the strongest enhancement at ~2-fold (Repl) and ~4-fold (Rep2), in comparison to Sample 3786 and Sample 3781. Sample 3786, which is or contains bovine DNAJB12, shows consistent ~3-fold enhancement in both replicates. Sample 3781, which is or contains PDI, shows moderate enhancement of ~2-fold and -1.5-fold in Repl and Rep2, respectively.

[0030] Fig. 4 is a graph depicting the kappa casein expression of replicates (light red: Repl, dark red: Rep2). Sample 3617 is the control. Sample 3784, which is or contains Glyma.01G003700, shows the highest enhancement at -1.7-fold (Repl) and -2.3-fold (Rep2), in comparison to Sample 3786 and Sample 3781. Sample 3786, which is or contains bovine DNAJB12, shows consistent -2- fold enhancement in both replicates. Sample 3781, which is or contains PDI), shows moderate -1.7- fold and -1.5-fold increases in Repl and Rep2.

[0031] Fig. 5 is a Western blot of alpha-Sl casein protein from transformed soybeans for different plasmids.

[0032] Fig. 6 is a Western blot of kappa casein protein from transformed soybeans for different plasmids.

[0033] Fig. 7 is a graph of RNA corrected accumulation of alpha and kappa casein protein.

[0034] Fig. 8 is a diagram of the base plasmid used for pMOZ3478, pMOZ3479, and pMOZ3617, where the GOI is SEQ ID NO: 7 for pMOZ3478 (Glyma.03G218300), SEQ ID NO: 8 for pMOZ3479 (Glyma.18G204000), and SEQ ID NO: 11 for pMOZ3617. DNA sequences for expressing alpha, kappa, and beta caseins are also included in each plasmid as indicated.DETAILED DESCRIPTION

[0035] The present disclosure provides a significant technical advance in the field of agricultural biotechnology and recombinant protein production. The invention is directed to novel methods and compositions for enhancing the expression, folding, and accumulation of heterologous proteins in plants, particularly in plant cells. The core of this invention lies in the discovery that the co-expression of specific, carefully selected endoplasmic reticulum (ER)-resident molecular chaperones with a heterologous protein of interest results in a substantial and reproducible increase in the yield of the target protein. This approach addresses a long-standing bottleneck in molecularDocket No. 719.001.002.PCT farming, where the folding capacity of the host plant's ER often limits the production of complex proteins.

[0036] While the invention is broadly applicable, it is exemplified herein by the enhanced production of mammalian milk proteins, specifically caseins, in soybean (Glycine max). As demonstrated in the examples, co-expression of casein proteins with either the native soybean chaperone Glyma.01G003700 or the heterologous bovine chaperone DNAJB12 resulted in protein accumulation increases of up to 4-fold compared to controls. This method holds significant potential for the cost-effective, industrial-scale production of plant-based caseins for use in food products, such as vegan dairy alternatives, thereby meeting a growing demand for sustainable, non-animal- derived proteins.Molecular Chaperones of the Invention

[0037] A key aspect of the invention is the identification and use of specific ER -resident molecular chaperones that are particularly effective at enhancing heterologous protein expression. These chaperones were identified through both gene expression analysis in the host plant and comparative genomic analysis of genes upregulated during lactation in mammals.

[0038] In certain embodiments, the molecular chaperone is encoded by a nucleotide sequence that is at least 85%, 90%, 95%, 98%, or 100% identical to a sequence selected from the group consisting of SEQ ID NO: 1 (encoding Glyma.01G003700), SEQ ID NO: 3 (encoding bovine DNAJB12), and SEQ ID NO: 5 (encoding Glyma.04g247900). The invention also encompasses the protein products of these nucleotide sequences, namely SEQ ID NO: 2 (Glyma.01 G003700), SEQ ID NO: 4 (bovine DNAJB12), and SEQ ID NO: 6 (Glyma.04g247900), as well as functional homologs and variants thereof. In other embodiments, the invention contemplates the use of molecular chaperone proteins that have a lower degree of sequence identity but retain the functional activity of enhancing heterologous protein expression. Accordingly, the molecular chaperone may be encoded by a nucleotide sequence that is greater than 72% but less than 80% identical to a sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 3, and SEQ ID NO: 5. Such sequences represent functional homologs from other species that may be used in the methods and compositions of the invention.

[0039] One such effective chaperone is the soybean protein Glyma.01G003700 (SEQ ID NO: 2), a putative protein disulfide isomerase (PDI)-like protein. This protein is believed to facilitate the correct formation and rearrangement of disulfide bonds, a critical step in the folding of manyDocket No. 719.001.002.PCT complex proteins. Overexpression of this endogenous protein enhances the natural protein folding capacity of the soybean cell's ER.

[0040] Another highly effective chaperone identified is bovine DNAJB12 (SEQ ID NO: 4). This protein is a member of the DnaJ / Hsp40 family of molecular chaperones. Its selection was guided by the observation that its expression is upregulated in mammary tissue during lactation, suggesting a specialized role in high-volume milk protein production. Specifically, DNAJB12 is understood to function as a co-chaperone that delivers unfolded or misfolded client proteins to Hsp70 chaperones and stimulates their ATPase activity, thereby facilitating the protein folding process. Its demonstrated effectiveness in a plant system highlights the conserved and cross-species functionality of this crucial chaperone mechanism. Furthermore, it was observed that certain chaperones exhibit preferential enhancement for specific proteins. For example, co-expression with bovine DNAJB12 (SEQ ID NO: 4) resulted in an accumulation of kappa casein that was approximately four times higher than the accumulation of alpha-Sl casein, demonstrating the potential for using specific chaperones to selectively enhance the production of individual components within a multi-protein expression system.

[0041] The invention also contemplates the use of other chaperones, such as the known soybean PDI Glyma.04g247900 (SEQ ID NO: 6), and splice isoforms of the aforementioned chaperones. Furthermore, the invention includes the use of homologous chaperone proteins from other species, including other dicots, monocots, bacteria, or yeast. These homologs may be identified based on sequence similarity to the specified chaperones and can be introduced into the plant host for co-expression with the heterologous protein.

[0042] In some embodiments, the invention provides for the co-expression of two or more different molecular chaperones to achieve an additive or synergistic enhancement of protein production. For instance, the simultaneous expression of a PDI-like protein (such as Glyma.01G003700) and a DnaJ family protein (such as bovine DNAJB12) can address different bottlenecks in the protein folding pathway, leading to a greater overall increase in protein yield than the expression of either chaperone alone. In addition to the PDI-like and DnaJ / Hsp40 family members exemplified herein, the invention contemplates the use of other functional classes of ER- resident molecular chaperones. In some embodiments, the molecular chaperone can be a member of the calnexin / calreticulin family, which are lectin-like chaperones involved in the folding of glycoproteins. In other embodiments, the chaperone can be a member of the Hsp70 or Hsp90Docket No. 719.001.002.PCT families, such as the ER-resident Hsp70 protein known as BiP (Binding immunoglobulin Protein) or Grp78. The co-expression of these or other chaperones, either alone or in combination with those of the present invention, can provide alternative or synergistic benefits for specific heterologous proteins. The source of the molecular chaperone is not limited to soybean or bovine. In alternative embodiments, the nucleotide sequence encoding the chaperone can be derived from other plant species (monocots or dicots), mammals, yeast, bacteria, or may be a synthetic or chimeric sequence designed for optimal activity in a plant ER environment. The chaperone protein may also be modified, for example, by adding or modifying targeting signals to ensure ER retention, or by including affinity tags for analysis or co-purification.

[0043] The disclosed compositions herein are directed to an advance in increasing casein protein expression in plants, specifically soybean, by overexpressing endogenous or homologous proteins. This method holds significant potential for producing plant -based casein for use in food products, thus meeting the growing demand for sustainable, non-animal-derived proteins.

[0044] In the compositions herein, nucleotide sequences can be obtained, isolated, and codon optimized, which are at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NO: 7, which can be transcribed from pMOZ3478 (see Fig. 8 where the GOI is SEQ ID NO: 7); or SEQ ID NO: 8, which is transcribed from pMOZ3479 (see Fig. 8 where the GOI is SEQ ID NO: 2). DNA sequences for expressing alpha, kappa, and beta caseins are included in the plasmid, such as pMOZ3478 and pMOZ3479, which encode for the expression of casein proteins (e.g., alpha-Sl casein, alpha-S2 casein, beta casein, and kappa casein). Additionally, a plasmid can contain a sequence for encoding a bright monomeric red fluorescent protein, such as SEQ ID NO. 11.

[0045] In the compositions herein, nucleotide sequences can be obtained, isolated, and codon optimized, which are at greater than 72% but less than 80% identical to any one of SEQ ID NO: 7, which can be transcribed from pMOZ3478 (see Fig. 8 where the GOI is SEQ ID NO: 7); or SEQ ID NO: 8, which is transcribed from pMOZ3479 (see Fig. 8 where the GOI is SEQ ID NO: 8). DNA sequences for expressing alpha, kappa, and beta caseins are included in the plasmid, such as pMOZ3478 and pMOZ3479, which encode for the expression of casein proteins (e.g., alpha-Sl casein, alpha-S2 casein, beta casein, and kappa casein). Additionally, a plasmid can contain a sequence for encoding a bright monomeric red fluorescent protein, such as SEQ ID NO. 11.Docket No. 719.001.002.PCT

[0046] In the compositions herein, pMOZ3478 and pMOZ3479 are a plasmid, vector, or other type of delivery vehicle that can be inserted or added into a plant cell, thereby integrating SEQ ID NO: 7 or SEQ ID NO: 8, respectively, into the genome of the plant cell separately or together.

[0047] In the compositions herein, the introduction of SEQ ID NO: 7 into the plant cell via the insertion of pMOZ3478, which co-expresses casein proteins, and which is followed by integration of SEQ ID NO: 7 into the genome, transforms the plant cell and leads to enhanced casein expression levels. Preferably, SEQ ID NO: 7 is overexpressed leading to higher than normal amounts of SEQ ID NO: 9. Stated another way, SEQ ID NO. 7 is a coding sequence for expressing SEQ ID NO: 9. The insertion of pMOZ3478 also leads to the expression of casein, thereby the plant cell can exhibit co-expression of casein and SEQ ID NO: 9. The overexpression of SEQ ID NO: 7 is the production of abnormally large amounts of SEQ ID NO: 9, which is coded for by SEQ ID NO: 1 in pMOZ3478, wherein pMOZ3478 also has, for example, cassettes for expressing caseins, wherein the casein proteins comprise alpha-Sl casein, alpha- S2 casein, beta casein, and kappa casein. This leads to increased accumulation of casein in tissues where SEQ ID NO: 9 and casein proteins are coexpressed.

[0048] In the compositions herein, the introduction of SEQ ID NO: 8 into the plant cell via the insertion of pMOZ3479, which co-expresses casein proteins, and which is followed by integration of SEQ ID NO: 8 into the genome, transforms the plant cell and leads to enhanced casein expression levels. Preferably, SEQ ID NO: 8 is overexpressed leading to higher than normal amounts of SEQ ID NO: 10. Stated another way, SEQ ID NO. 8 is a coding sequence for expressing SEQ ID NO: 10. The insertion of pMOZ3479 also leads to the expression of casein proteins, thereby the plant cell can exhibit co-expression of casein and SEQ ID NO: 10. The overexpression of SEQ ID NO: 8 is the production of more SEQ ID NO: 10 than normal (inclusive of when SEQ ID NO: 10 is not natively present), which is coded for by SEQ ID NO: 8 in pMOZ3479, wherein pMOZ3479 also has, for example, cassettes for expressing casein proteins, wherein the casein proteins comprise alpha-Sl casein, alpha- S2 casein, beta casein, and kappa casein. This leads to increased accumulation of casein in tissues where SEQ ID NO: 10 and casein proteins are co-expressed.

[0049] In the compositions herein, a nucleotide sequence, SEQ ID NO: 7 or SEQ ID NO: 8, when overexpressed in tissues where casein is expressed (pMOZ3478 or pMOZ3479), can increase the overall yield of casein protein, in comparison to a plasmid lacking SEQ ID NO: 7 and SEQ ID NO: 8 (pMOZ3617). Splice isoforms of SEQ ID NO: 7 and SEQ ID NO: 8 can also have this effect.Docket No. 719.001.002.PCT

[0050] In the compositions herein, SEQ ID NO: 7 and SEQ ID NO: 8, as applied to develop transgenic soybean plants that produce high levels of casein protein, while leading to co-expression of SEQ ID NO: 9 and SEQ ID NO: 10, respectively, can be harvested and used in the production of food products, such as vegan dairy alternatives. This method could also be applied to other plant species for the expression of various heterologous proteins.

[0051] In the compositions herein, casein expression is targeted to specific tissues in the soybean plant. The SEQ ID NO: 7 or SEQ ID NO: 8 may already be endogenously expressed in these tissues. Overexpressing these genes in these locations enhances the gene products natural function, leading to higher casein levels. Proteins produced by expression or overexpression of SEQ ID NO: 7 or SEQ ID NO: 8 genes can also be targeted to tissues where these genes are not natively present using vacuolar sorting determinants or other organelle targeting peptides known in the art.

[0052] In the compositions herein, in addition to SEQ ID NO: 9 and SEQ ID NO: 10, homologous proteins from other species, such as other dicots or monocots can be used in the same system; and in addition to SEQ ID NO: 7 and SEQ ID NO: 8, homologous genes from other species, such as other dicots, monocots, bacteria, or yeast can be used in the same system. These genes may be identified based on sequence similarity to SEQ ID NO: 7 and SEQ ID NO: 8 and can be introduced into soybean via a plasmid, whereby the resulting protein is co-expressed with casein, for enhanced casein production.

[0053] In the compositions herein, plant systems treated or incorporated with SEQ ID NO: 7, from pMOZ3478; or SEQ ID NO: 8, from pMOZ3479, exhibit higher accumulation of casein (alpha- S1 and kappa casein) than genes, from pMOZ3617 which does not contain either gene SEQ ID NO: 7 or SEQ ID NO: 8. Co-expression of casein and SEQ ID NO: 9 or SEQ ID NO: 10 can lead to increased levels of casein in a plant cell. In comparison to pMOZ3617, pMOZ3478 and pMOZ3479 have increased levels of alpha-Sl and kappa casein, whereby: (i) the level of alpha-Sl casein is ~4 times higher than the control pMOZ3617, in the instance of pMOMZ3478; (ii) the level of kappa casein is ~4 times higher than the control pMOZ3617, in the instance of pMOMZ3478; (iii) the level of alpha-Sl casein is ~2.5 times higher than the control pMOZ3617, in the instance of pMOMZ3479; and (iv) the level of kappa casein is ~9.5 times higher than the control pMOZ3617, in the instance of pMOMZ3479. The level of alpha-Sl casein is similar to the level of kappa casein, in the instance of pMOZ3478, which contains SEQ ID NO: 9. The level of kappa casein, however, isDocket No. 719.001.002.PCT~4 times higher than the level of alpha-Sl casein, in the instance of pMOZ3479, which contains SEQ ID NO: 10. (See Fig. 7.)

[0054] In the composition herein, SEQ ID NO: 7 or SEQ ID NO: 8, which can be expressed to yield SEQ ID NO: 9 and SEQ ID NO: 10, respectively, can be codon-optimized for a plant, wherein the plant can be any one of the following, for example, angiosperms and gymnosperms such as Arabidopsis, potato, tomato, tobacco, alfalfa, lemice, carrot, strawberry, sugar beet, cassava, sweet potato, soybean, lima bean, pea, chick pea, maize (com), turf grass, wheat, rice, barley, sorghum, oat, oak, eucalyptus, walnut, palm and duckweed as well as fern and moss. In some aspects, the nucleotide sequences provided herein are codon-optimized for a plant, wherein the plant is a monocot, a dicot, or a vascular plant reproduced from spores such as fern or a nonvascular plant such as moss, liverwort, hornwort, and algae. In some aspects, the nucleotide sequences provided herein are codon-optimized for a dicot plant, include for example Arabidopsis, tobacco, tomato, potato, sweet potato, cassava, alfalfa, lima bean, pea, chick pea, soybean, carrot, strawberry, lettuce, oak, maple, walnut, rose, mint, squash, daisy, quinoa, buckwheat, mung bean, cow pea, lentil, lupin, peanut, fava bean, French beans, mustard, or cactus. In some aspects, the nucleotide sequences provided herein are codon-optimized for a monocot plant, including for example, turf grass, maize (corn), rice, oat, wheat, barley, sorghum, orchid, iris, lily, onion, palm, and duckweed. In some cases, the nucleotide sequences provided herein are codon-optimized for a soybean (i.e., glycine max).

[0055] In the compositions herein, a food composition (such as a dairy product) can comprise: (i) a casein protein with (ii) SEQ ID NO: 9 or SEQ ID NO: 10, as obtained from transfection of a plant by SEQ ID NO: 7 or SEQ ID NO: 8; and a nucleic acid molecule comprising SEQ ID NO: 7 and SEQ ID NO: 8. In a preferred embodiment, the food composition comprises: (i) casein proteins and (ii) SEQ ID NO: 9 or SEQ ID NO: 10, wherein SEQ ID NO: 9 or SEQ ID NO: 10 is present at levels greater than 0.001%, 0.01%, 0.1% or 1% of the total weight of the food composition.Heterologous Proteins of Interest

[0056] The methods and compositions of the invention are broadly applicable to enhancing the expression of a wide variety of heterologous proteins in plants. The primary examples provided herein relate to milk proteins, specifically caseins. The caseins may be selected from aS 1 -casein,Docket No. 719.001.002.PCT aS2-casein, P-casein, K-casein, or any combination thereof, to facilitate the formation of functional casein micelles in planta.

[0057] Beyond caseins, the invention is suitable for enhancing the production of other valuable proteins, including but not limited to: (a) Other Milk Proteins, such as whey proteins including P-lactoglobulin, a-lactalbumin, bovine serum albumin, and lactoferrin; (b) Therapeutic Proteins, such as monoclonal antibodies, growth factors, blood factors, cytokines, hormones, enzymes, and vaccine antigens; (c) Industrial Enzymes, such as cellulases, amylases, proteases, lipases, oxidoreductases, and phytases; (d) Structural Proteins, such as collagens, elastins, and silks; (e) Egg White Proteins, including ovalbumin, ovotransferrin, ovomucoid, lysozyme, and ovomucin; and (f) Seed Storage Proteins, including zeins, glutelins, albumins, and globulins.

[0058] The system is particularly advantageous for proteins that are difficult to express in heterologous systems due to their complex structure, requirement for multiple disulfide bonds, or tendency to misfold and aggregate. The enhancement provided by the invention is particularly valuable for heterologous proteins that are considered “difficult-to-express” due to their structural complexity. Such proteins may include, but are not limited to, multi-subunit proteins that require proper assembly, proteins that form oligomers (e.g., dimers, tetramers), proteins with multiple or complex disulfide bonds, or proteins that require specific post -translational modifications such as glycosylation that are initiated in the ER. The invention also encompasses the production of fusion proteins, where the chaperone-assisted folding improves the functionality of one or both fusion partners.Plant Host Systems and Transformation

[0059] The invention can be practiced in a wide range of plant species. In a preferred embodiment, the plant is a soybean plant (Glycine max)', which serves as an excellent host for protein production due to its high seed protein content. However, the methods are applicable to virtually any transformable plant species. Suitable plant hosts include, but are not limited to, dicot plants such as Arabidopsis, tobacco, tomato, potato, sweet potato, cassava, alfalfa, lima bean, pea, chick pea, carrot, strawberry, lettuce, oak, maple, walnut, rose, mint, squash, daisy, quinoa, buckwheat, mung bean, cow pea, lentil, lupin, peanut, fava bean, French beans, mustard, or cactus; and monocot plants such as turf grass, maize (com), rice, oat, wheat, barley, sorghum, orchid, iris, lily, onion, palm, and duckweed. The invention is also applicable to non-vascular plants such as moss, liverwort, hornwort, and algae, as well as spore-reproducing vascular plants like ferns.Docket No. 719.001.002.PCTFurthermore, the term "plant" as used herein also encompasses isolated plant cells or tissues grown in culture, including plant cell suspension cultures or hairy root cultures. These contained systems offer alternative production platforms that can benefit from the chaperone-mediated expression enhancement described herein.

[0060] The nucleic acid molecules encoding the chaperone(s) and the heterologous protein(s) can be introduced into plant cells using any suitable transformation method known in the art. These methods include, but are not limited to, Agrobacterium-mediated transformation, biolistic particle delivery (gene gun), electroporation of protoplasts, polyethylene glycol (PEG) -mediated transformation, and viral vector-mediated delivery. The result is a transgenic plant or plant cell that is stably transformed with the introduced nucleic acids integrated into its genome.Genetic Constructs and Expression Strategies

[0061] To practice the invention, recombinant DNA constructs are designed for the coexpression of the molecular chaperone and the heterologous protein. These constructs typically comprise one or more expression cassettes, each containing a gene of interest (either the chaperone or the heterologous protein) operably linked to regulatory elements that drive its expression in the plant cell.

[0062] The nucleotide sequences encoding the chaperone and / or the heterologous protein may be codon-optimized for efficient expression in the chosen plant host. Codon optimization involves modifying the nucleotide sequence to match the codon usage preference of the host organism without altering the encoded amino acid sequence, thereby enhancing translational efficiency and protein yield.

[0063] Expression of the genes is controlled by promoters that are active in plant cells. The choice of promoter can be used to control the timing, location, and level of expression. Promoters may include: (1) Constitutive promoters (e.g., CaMV 35S, Ubiquitin) that drive expression in most or all tissues throughout the plant's life; (2) Tissue-specific promoters (e.g., seed-specific promoters like glycinin or phaseolin) that restrict expression to specific tissues, such as the seed, which can be advantageous for protein accumulation and downstream processing; (3) Inducible promoters that allow for controlled expression in response to an external chemical or physical stimulus. Examples of such promoters include, but are not limited to, heat-shock promoters, ethanol-inducible promoters, or hormone-inducible promoters (e.g., glucocorticoid-inducible systems). This allows for temporal control over chaperone expression, which can be timed to coincide with the peak synthesis of theDocket No. 719.001.002.PCT heterologous protein, thereby maximizing the enhancement effect while minimizing metabolic load on the plant cell.

[0064] The constructs may also include subcellular targeting sequences to direct the synthesized proteins to the appropriate cellular compartment. Since the chaperones of this invention are ER-resident, they may naturally contain, or be engineered to contain, an ER retention signal such as the KDEL sequence. Similarly, the heterologous protein is typically engineered with a signal peptide to direct its synthesis into the ER lumen, where it can interact with the co-expressed chaperones. In other embodiments, proteins may be targeted to other organelles, such as protein storage vacuoles, using appropriate targeting peptides.

[0065] The introduction of the nucleic acid sequences results in the overexpression of the chaperone protein, leading to higher than normal amounts of the chaperone in the ER. This coexpression and co-localization of the chaperone and the heterologous protein enhances protein folding, improves stability, reduces degradation, and ultimately leads to a higher accumulation of the functional heterologous protein in the plant tissue. The arrangement of the expression cassettes for the heterologous protein and the molecular chaperone(s) can be varied. In one embodiment, the expression cassettes are located on the same T-DNA within a single vector and are transformed into the plant cell simultaneously. In another embodiment, the expression cassettes may be located on separate vectors, which are then co-transformed into the same plant cell. In yet another embodiment, transgenic plants expressing the heterologous protein can be crossed with transgenic plants expressing the molecular chaperone to generate progeny that co-express both. Furthermore, multiple genes could be linked in a polycistronic arrangement, using elements such as Internal Ribosome Entry Sites (IRES) or self-cleaving 2A peptides, to drive their expression from a single promoter.Food Products

[0066] In one embodiment, the invention provides a food composition, such as a dairy product substitute, comprising the heterologous proteins produced by the transgenic plants described herein. For example, casein proteins can be extracted from the plant biomass and formulated into food products. In some variants, the food composition may comprise both the expressed casein proteins and the co-expressed chaperone protein (e.g., SEQ ID NO: 2 or SEQ ID NO: 4), which may be co-purified or retained in the final product. In a preferred embodiment, the food composition comprises both the heterologous protein (e.g., casein) and the co-expressed molecular chaperoneDocket No. 719.001.002.PCT protein (e.g., SEQ ID NO: 2 or SEQ ID NO: 4), wherein the molecular chaperone protein is present at a level greater than 0.001%, 0.01%, 0.1%, or 1% of the total weight of the food composition.Characterization of Enhanced Protein Expression

[0067] As used herein, an “increase” or “enhancement” in protein expression or accumulation refers not only to a greater quantity of the protein but also to qualitative improvements. In various embodiments, the enhancement provided by co-expression of the molecular chaperone is characterized by one or more of the following attributes: (a) increased total protein accumulation of at least 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, or greater; (b) increased protein solubility and a corresponding reduction in protein aggregation or inclusion body formation; (c) enhanced biological activity or specific activity of the protein; (d) improved fidelity of post -translational modifications, such as disulfide bond formation or glycosylation; (e) reduced proteolytic degradation of the protein; and (f) improved protein secretion efficiency if the protein is targeted for export from the cell. These improvements collectively contribute to a higher yield of functional, high-quality heterologous protein.Definitions

[0068] The following embodiments are described in sufficient detail to enable those skilled in the art to make and use the invention. It is to be understood that other embodiments would be evident based on the present disclosure, and that system, process, or mechanical changes can be made without departing from the scope of an embodiment of the present disclosure.

[0069] In the following description, numerous specific details are given to provide a thorough understanding of the invention. However, it will be apparent that the invention can be practiced without these specific details. In order to avoid obscuring an embodiment of the present disclosure, some well-known techniques, system configurations, and process steps are not disclosed in detail. Throughout this disclosure, various publications, patents and published patent specifications are referenced by an identifying citation. The disclosures of these publications, patents and published patent specifications are hereby incorporated by reference into the present disclosure to more fully describe the state of the art to which this invention pertains.DefinitionsDocket No. 719.001.002.PCT

[0070] A specified nucleic acid is “derived from” a given nucleic acid when it is constructed using the given nucleic acid’s sequence, or when the specified nucleic acid is constructed using the given nucleic acid. For example, a cDNA or EST is derived from an expressed mRNA.

[0071] As used herein, the term “plant” includes whole plant, plant organ, plant tissues, and plant cell and progeny of same, but is not limited to angiosperms and gymnosperms such as Arabidopsis, potato, tomato, tobacco, alfalfa, lemice, carrot, strawberry, sugarbeet, cassava, sweet potato, soybean, lima bean, pea, chick pea, maize (com), turf grass, wheat, rice, barley, sorghum, oat, oak, eucalyptus, walnut, palm and duckweed as well as fern and moss. Thus, a plant may be a monocot, a dicot, a vascular plant reproduced from spores such as fern or a nonvascular plant such as moss, liverwort, hornwort and algae. The term “plant,” as used herein, also encompasses plant cells, seeds, plant progeny, propagule whether generated sexually or asexually, and descendants of any of these, such as cuttings or seed. Plant cells include suspension cultures, callus, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen, seeds and microspores. Plants may be at various stages of maturity and may be grown in liquid or solid culture, or in soil or suitable media in pots, greenhouses or fields.

[0072] As used herein, the term “dicof ’ refers to a flowering plant whose embryos have two seed leaves or cotyledons. Examples of dicots include Arabidopsis, tobacco, tomato, potato, sweet potato, cassava, alfalfa, lima bean, pea, chick pea, soybean, carrot, strawberry, lettuce, oak, maple, walnut, rose, mint, squash, daisy, quinoa, buckwheat, mung bean, cow pea, lentil, lupin, peanut, fava bean, French beans, mustard, or cactus.

[0073] As used herein, the term “monocot” refers to a flowering plant whose embryos have one cotyledon or seed leaf. Examples of monocots include turf grass, maize (corn), rice, oat, wheat, barley, sorghum, orchid, iris, lily, onion, palm, and duckweed.

[0074] The term “seed” is meant to encompass the whole seed and / or all seed components, including, for example, the coleoptile and leaves, radicle and coleorhiza, scutellum, starchy endosperm, aleurone layer, pericarp and / or testa, either during seed maturation and seed germination.

[0075] As used herein, the term “transgenic plant” means a plant that has been transformed with one or more exogenous nucleic acids. “Transformation” refers to a process by which a nucleic acid is stably integrated into the genome of a plant cell. “Stably transformed” refers to the permanent, or non-transient, retention, expression, or a combination thereof of a polynucleotide inDocket No. 719.001.002.PCT and by a cell genome. A stably integrated polynucleotide is one that is a fixture within a transformed cell genome and can be replicated and propagated through successive progeny of the cell or resultant transformed plant. Transformation can occur under natural or artificial conditions using various methods. Transformation can rely on any method for the insertion of nucleic acid sequences into a prokaryotic or eukaryotic host cell, including Agrobacterium-mediated transformation as illustrated in U.S. Pat. Nos. 5,159,135; 5,824,877; 5,591,616 and 6,384,301, all of which are incorporated herein by reference in its entirety. Methods for plant transformation also include microprojectile bombardment as illustrated in U.S. Pat. Nos. 5,015,580; 5,550,318; 5,538,880; 6,153,812; 6,160,208; 6,288,312 and 6,399,861, all of which are incorporated herein by reference in its entirety. Recipient cells for the plant transformation include meristem cells, callus, immature embryos, hypocotyls explants, cotyledon explants, leaf explants, and gametic cells such as microspores, pollen, sperm and egg cells, and any cell from which a fertile plant can be regenerated, as described in U.S. Pat. Nos. 6,194,636; 6,232,526; 6,541,682 and 6,603,061 and U.S. Patent Application publication US 2004 / 0216189 Al, all of which are incorporated herein by reference in its entirety.

[0076] The term “plant tissue” refers to any part of a plant, such as a plant organ. Examples of plant organs include, but are not limited to the leaf, stem, root, tuber, seed, branch, pubescence, nodule, leaf axil, flower, pollen, stamen, pistil, petal, peduncle, stalk, stigma, style, bract, fruit, trunk, carpel, sepal, anther, ovule, pedicel, needle, cone, rhizome, stolon, shoot, pericarp, endosperm, placenta, berry, stamen, and leaf sheath.

[0077] As used herein, the term “stably expressed” refers to expression and accumulation of a protein in a plant cell over time. As an example, a recombinant protein may accumulate because it is not degraded by endogenous plant proteases. As a further example, a recombinant protein is considered to be stably expressed in a plant if it is present in the plant in an amount of 1% or higher per total protein weight of soluble protein extractable from the plant.

[0078] As used herein, the term “recombinant” refers to nucleic acids or proteins formed by laboratory methods of genetic recombination (e.g., molecular cloning) to bring together genetic material from multiple sources, creating sequences that would otherwise not be found in the genome. Recombinant proteins may be expressed in vivo in various types of host cells, including plant cells, bacterial cells, fungal cells, avian cells, and mammalian cells. Recombinant proteins may also be generated in vitro. As used herein, the term “tagged protein” refers to a recombinant protein thatDocket No. 719.001.002.PCT includes additional peptides that are not part of the native protein and that remain after post- translational processing.

[0079] These and other valuable aspects of the embodiments of the present disclosure consequently further the state of the technology to at least the next level. While the disclosure has been described in conjunction with a specific best mode, it is to be understood that many alternatives, modifications, and variations will be apparent to those skilled in the art in light of the descriptions herein. Accordingly, it is intended to embrace all such alternatives, modifications, and variations that fall within the scope of the included claims. All matters set forth herein or shown in the accompanying drawings are to be interpreted in an illustrative and non-limiting sense.

[0080] As used herein, the phrases “at least one”, “one or more”, and “and / or” are open- ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions “at least one of A, B and C”, “at least one of A, B, or C”, “one or more of A, B, and C”, “one or more of A, B, or C” and “A, B, and / or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B and C together.

[0081] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0082] Use of absolute or sequential terms, for example, “will,” “will not,” “shall,” “shall not,” “must,” “must not,” “first,” “initially,” “next,” “subsequently,” “before,” “after,” “lastly,” and “finally,” are not meant to limit scope of the present embodiments disclosed herein but as exemplary.

[0083] As used herein, “or” may refer to “and”, “or,” or “and / or” and may be used both exclusively and inclusively. For example, the term “A or B” may refer to “A or B”, “A but not B”, “B but not A”, and “A and B”. In some cases, context may dictate a particular meaning.

[0084] Any systems, methods, software, and platforms described herein are modular and not limited to sequential steps. Accordingly, terms such as “first” and “second” do not necessarily imply priority, order of importance, or order of acts.

[0085] As used herein, the term “about” when referring to a number or a numerical range means that the number or numerical range referred to is an approximation within experimentalDocket No. 719.001.002.PCT variability (or within statistical experimental error), and the number or numerical range may vary from, for example, from 1% to 10% of the stated number or numerical range. Unless otherwise indicated by context, the term “about” refers to ±10% of a stated number or value.

[0086] As used herein, the term “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “approximately” can mean within 1 or more than 1 standard deviation, per the practice in the given value. Where particular values are described in the application and claims, unless otherwise stated the term “approximately” should be assumed to mean an acceptable error range for the particular value.

[0087] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0088] As used herein, “percent (%) sequence identity” with respect to the nucleic acid or amino acid sequences identified herein is defined as the percentage of nucleic acid or amino acid residues in a candidate sequence that are identical with the amino acid residues in the polypeptide being compared, after aligning the sequences considering any conservative substitutions as part of the sequence identity.

[0089] All ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, and so forth. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, and the like. All language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3Docket No. 719.001.002.PCT articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.

[0090] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multistranded form. In some cases, a polynucleotide is exogenous (e.g. a heterologous polynucleotide). In some cases, a polynucleotide is endogenous to a cell. In some cases, a polynucleotide can exist in a cell-free environment. In some cases, a polynucleotide is a gene or fragment thereof. In some cases, a polynucleotide is DNA. In some cases, a polynucleotide is RNA. A polynucleotide can have any three-dimensional structure, and can perform any function, known or unknown. In some cases, a polynucleotide comprises one or more analogs (e.g. altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include: 5 -bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g. rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), non-coding RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. In some cases, the sequence of nucleotides is interrupted by non-nucleotide components.

[0091] As used herein, the term “ / / / vitro " is used to describe an event that takes place contained in a container for holding a laboratory reagent such that it is separated from the living biological source organism from which the material is obtained. In vitro assays can encompass cellbased assays in which live or dead cells or other biological materials are employed. In vitro assays can also encompass a cell-free assay in which no intact cells are employed.Docket No. 719.001.002.PCT

[0092] As used herein, the term “zzz vivo" is used to describe an event that takes place in a plant, for example, inside a plant cell.

[0093] A “plasmid” as used herein, generally refers to a non-viral expression vector, e.g., a nucleic acid molecule that encodes for genes and / or regulatory elements necessary for the expression of genes. The term “vector,” as used herein, generally refers to a nucleic acid molecule capable of transferring or transporting a payload nucleic acid molecule. The payload nucleic acid molecule can be generally linked to, e.g., inserted into, the vector nucleic acid molecule. A vector can include sequences that direct autonomous replication in a cell, or can include sequences sufficient to allow integration into host cell genes (e.g., host cell DNA). Examples of a vector can include, but are not limited to, plasmids (e.g., DNA plasmids or RNA plasmids), transposons, cosmids, bacterial artificial chromosomes, and viral vectors. A vector of any of the aspects of the present disclosure can comprise exogenous, endogenous, or heterologous control sequences such as promoters and / or enhancers. For example, a “vector” can be a plasmid comprising operably linked polynucleotide sequences that facilitate expression of a coding sequence in a particular host organism (e.g., a bacterial expression vector or a plant expression vector). Polynucleotide sequences that facilitate expression in prokaryotes can include, e.g., a promoter, an enhancer, an operator, and a ribosome binding site, often along with other sequences. Eukaryotic cells can use promoters, enhancers, termination and polyadenylation signals and other sequences that are generally different from those used by prokaryotes.

[0094] Whenever the term “at least,” “greater than,” “greater than or equal to”, or a similar phrase precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than,” “greater than or equal to” or similar phrase applies to each of the numerical values in that series of numerical values. For example, “at least 1, 2, or 3” is equivalent to “at least 1, at least 2, and / or at least 3.”

[0095] Whenever the term “no more than,” “less than,” “less than or equal to,” “no greater than,” “at most” or a similar phrase, precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” “less than or equal to,” “no greater than,” “at most,” or similar phrase applies to each of the numerical values in that series of numerical values. For example, “less than 3, 2, or 1” is equivalent to “less than 3, less than 2, and / or less than 1.”

[0096] As used herein, the term “essentially free” means a component, if present, is present in an amount that does not contribute, or contributes only in a de minimus fashion, to the propertiesDocket No. 719.001.002.PCT or function of the composition. In some cases, where a composition is essentially free of a particular component, the component is present in trace amounts, for example, less than 5% by weight, less than 4% by weight, less than 3% by weight, less than 2% by weight, less than 1% by weight, less than 0.5% by weight, less than 0.1% by weight, less than 0.05%, or less than 0.01% by weight. As used herein, “in trace amounts” means that the component is detectable using a testing method known to a person of ordinary skill in the field.EXAMPLES

[0097] It will be more readily understood by reference to the following examples, which are included merely for purposes of illustration of certain aspects and embodiments of the present invention, and are not intended to limit the invention. The following example describes the general materials and methods used to construct expression vectors, transform soybean cells, and analyze recombinant protein expression.Example 1. General Materials and Methods for Recombinant Protein Expression in Soybean

[0098] Multigene expression vectors were designed to co-express casein proteins with either a molecular chaperone of interest or a control protein. The vectors were assembled using the Golden Gate modular cloning system MoClo, as described by Engler et al. (ACS Synthetic Biology 3, no. 11 (2014): 839-43), which is incorporated herein by reference. The base vectors contained expression cassettes for the following bovine proteins: a-Sl -casein (UniProt Accession P02662), -casein (UniProt P02666), K-casein (UniProt P02668), and bovine FAM20C kinase (uniprot accession number # F1MXQ3). Each vector was further engineered to include an expression cassette for a gene of interest (GOI), which in experimental vectors was a nucleotide sequence encoding a molecular chaperone (e.g., Glyma.01G003700, Glyma.04g247900, Glyma.03G218300, Glyma.l8G204000, or bovine DNAJB12) and in control vectors was a nucleotide sequence encoding a biomarker protein (e.g., mScarlet fluorescent protein). All proteins were expressed under the control of constitutively active, seed-specific, or synthetic plant promoters, and a subset of these proteins possessed translationally-fused epitope tags and / or subcellular targeting peptide sequences.

[0099] The gene sequences encoding the bovine casein proteins were derived from UniProt and modified to include specific C-terminal sequences for purification and subcellular retention. These sequences included a two-serine spacer for proper folding, dual affinity purification tags (FLAG-tag and 6-Histidine tag), and an HDEL target peptide for retention of soluble proteins in theDocket No. 719.001.002.PCT endoplasmic reticulum (ER). To ensure translocation into the secretory pathway, an N-terminal signal peptide from soybean glycinin 1 (GY1, UniProt P04776) was added to all three casein sequences. The recombinant casein sequences were expressed under the constitutive Agrobacterium turn efaci ens mannopine synthase (AtuMas) promoter and 5’ untranslated region (UTR) (https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC287101 / ).Bacterial Transformation and Plasmid Verification

[0100] The assembled plasmids were first transformed into E. coli for amplification. Lucigen Ecloni 10G thermo-competent cells were thawed on ice and spiked with 1 pg-100 ng of plasmid DNA, followed by a 10-minute incubation on ice. The mixture was then heat-shocked at 42°C for 90 seconds and immediately returned to ice for 5 minutes. Transformed cells were recovered in liquid broth (LB) and cultured in a 37°C shaker at 225 rpm for 45 minutes. Following recovery, the bacteria were plated on LB agar plates containing the appropriate antibiotic for selection and incubated overnight at 37°C. A MoClo-compatible blue-white selection system was used to identify colonies containing recombinant plasmids. White colonies, indicating successful DNA insertion, were selected for further verification.

[0101] Picked colonies were cultured overnight in 5 mL of LB with the appropriate selection antibiotic. Plasmids were purified from the bacteria using a NucleoSpin miniprep kit (Takara Bio Inc.). The integrity and correctness of each plasmid were confirmed by restriction enzyme digestion to verify the presence of the DNA insert, followed by Sanger and nanopore sequencing to confirm sequence accuracy. Colonies containing the correct plasmids were preserved as 20% glycerol stocks for long-term storage.Agrobacterium- e iateA Transformation of Soybean Embryos

[0102] For plant transformation, the verified plasmids were introduced into Agrobacterium tumefaciens strain EHA105 electrocompetent cells. Purified plasmid DNA (30 ng) was mixed with ice-thawed electrocompetent cells and transferred to a pre-chilled 0.2 cm electroporation cuvette. An electric pulse of 2.5 kV was applied, and the transformed cells were recovered in 500 pL of LB medium and cultured for 2-4 hours in a 28°C shaking incubator. Cultured cells were then plated on LB agar plates with the appropriate antibiotic selection and incubated at 28°C for two days until colonies formed. A minimum of three colonies were picked and verified via plasmid purification, digestion, and sequencing.Docket No. 719.001.002.PCT

[0103] Zygotic soybean embryos were transiently transformed using the verified Agrobacterium strains. A 5 mL starter culture of Agrobacterium containing the plasmid of interest was incubated for two days at 28°C, and this starter was used to inoculate a larger 200 mL overnight culture. The large-scale culture was centrifuged, and the bacterial pellet was washed and resuspended in LCCM medium to a final concentration corresponding to an OD600 of 1.4. The final agrobacterium solution was supplemented with acetosyringone to a final concentration of 100 pM and incubated at room temperature with shaking for 1 -2 hours prior to inoculation.

[0104] Soybean pods containing 8-10 mm embryos were surface-sterilized using a sequential treatment of 70% ethanol, 10% bleach, and sterile deionized water. Seeds were aseptically removed from the pods, dissected from their seed coats, and split in half to produce cotyledonary explants. The prepared explants (30 cotyledon halves) were placed in the Agrobacterium culture, which had been supplemented with Silwet-77 (0.03% v / v). The mixture was sonicated for 20 seconds at 20% amplitude, followed by vacuum infdtration for 5 minutes. After infiltration, the explants were incubated for two hours on a rotator at room temperature. The explants were then removed from the bacterial culture and placed adaxial side down on SCCM plates overlaid with sterile filter paper, wrapped with micropore tape, and incubated for 3 days in a dark 24°C chamber. Following this cocultivation period, the explants were washed three times with sterile water containing antibiotics (Rifampicin, Carbenicillin, and Cefotaxime) to remove excess Agrobacterium and were replated on SCCM plates to incubate for an additional 7-9 days.Protein Extraction, Purification, and Characterization

[0105] After the incubation period, the transformed soybean cotyledons were flash-frozen in liquid nitrogen and ground into a fine powder. Approximately 100 grams of the powdered tissue was mixed with a Tris-based protein extraction buffer (50 mM Tris, 300 mM KC1, 0.5% Tween-20, 3.65% glycerol, pH 8.6) containing a plant protease inhibitor cocktail (Sigma). The mixture was rotated for 1 hour at 4°C and then clarified by centrifugation at 1300 rpm for 30 minutes at 4°C.

[0106] The crude protein extract was further clarified using a 0.45 pm filter. The His-tagged casein proteins were purified from the clarified extract using nickel resin affinity chromatography. The extract was mixed with nickel resin and rotated at 4°C for 4 hours to allow binding. The resin was then collected by centrifugation, and the supernatant was removed. The resin was washed four times with a wash buffer (50 mM Tris base, 300 mM KC1, 20 mM imidazole, pH 7.4) to removeDocket No. 719.001.002.PCT non-specifically bound proteins. The bound proteins were subsequently eluted from the resin using an elution buffer containing a higher concentration of imidazole.

[0107] To confirm the successful expression and purification of the casein proteins, eluted samples were analyzed by SDS-PAGE on a Bio-Rad TGX AnykD Mini-PROTEAN gel. The proteins were then transferred to a membrane for Western blot analysis using antibodies against the FLAG epitope tag. The presence of bands at the expected molecular weight of approximately 27 kDa in the elution fractions confirmed the presence of the tagged casein proteins. To confirm the presence and identity of all individual casein proteins (a-Sl, P, and K-casein), mass spectrometry was used for further characterization.Example 2: Enhancement of Casein Expression by Co-expression of Single ER-Resident Molecular Chaperones

[0108] To validate the hypothesis that co-expression of specific ER-resident molecular chaperones can enhance heterologous protein production, an experiment was conducted in soybean cells. A series of expression constructs were designed to co-express alpha and kappa casein proteins alongside either a candidate molecular chaperone or a neutral control protein. The general methods for vector construction, plant transformation, and protein analysis are described in Example 1.

[0109] The specific expression constructs tested were as follows: plasmid pMOZ3784 expressing the native soybean chaperone Glyma.01G003700 (a putative protein disulfide isomerase- like protein, SEQ ID NO: 2); plasmid pMOZ3786 expressing the heterologous mammalian chaperone bovine DNAFB12 (SEQ ID NO: 4); and plasmid pMOZ3781 expressing a known soybean protein disulfide isomerase, Glyma.04g247900 (SEQ ID NO: 6). A control plasmid, pMOZ3617, expressing the fluorescent protein mScarlet, was used to establish a baseline for casein expression levels in the absence of a co-expressed chaperone. Each of these constructs was co-expressed with constructs for V5-tagged alpha casein and FLAG-tagged kappa casein.

[0110] Following transient transformation of soybean cotyledons, total protein and RNA were extracted from the tissue samples. The accumulation of alpha and kappa casein was quantified by Western blot analysis using antibodies against the V5 and FLAG tags, respectively. To ensure that the observed differences were due to post-transcriptional effects such as improved protein folding and stability, the protein expression levels were normalized against the corresponding mRNA levels, which were determined by quantitative PCR (qPCR). The experiment was conducted in two independent biological replicates to confirm the reproducibility of the results.Docket No. 719.001.002.PCT

[0111] The experimental results, shown in Figures 1-4, demonstrate a significant and reproducible enhancement of casein protein accumulation when co-expressed with specific molecular chaperones. Co-expression with the soybean chaperone Glyma.01G003700 (sample 3784) resulted in a substantial increase in the accumulation of both casein types, with enhancement reaching up to approximately 4-fold for both alpha and kappa casein compared to the mScarlet control (sample 3617), as particularly evident in the second replicate (Fig. 2 and Fig. 3).

[0112] Similarly, co-expression with the heterologous bovine chaperone bovine DNAJB12 (sample 3786) also yielded a strong and consistent enhancement. This construct produced an approximately 3.5-fold increase in alpha casein accumulation and a 2-fold increase in kappa casein accumulation across both replicates (Fig. 1, Fig. 3, and Fig. 4). This differential enhancement suggests that certain chaperones may be particularly effective for specific target proteins. Coexpression with the known PDI, Glyma.04g247900 (sample 3781), also showed a moderate but consistent increase for both proteins.

[0113] These results clearly indicate that the overexpression of select ER-resident molecular chaperones leads to a measurable and significant increase in the accumulation of co-expressed casein proteins in soybean tissues. This increase is attributed to post -transcriptional mechanisms, likely improved protein folding, assembly, and stability, which reduce protein degradation by the cell's quality control systems. The success of both a native soybean chaperone and a heterologous chaperone from a lactating mammal validates the broad applicability of this strategy for enhancing the production of complex recombinant proteins in plants.

[0114] Experimental results indicated that overexpression of SEQ ID NO: 9 and SEQ ID NO: 10 leads to a measurable increase in casein protein levels in soybean tissues as can be seen in Fig. 7. This increase may result from improved protein folding and stabilization due to the activity of the overexpressed protein. Protein expression levels in Fig. 7 were further normalized against transformation efficiency and RNA expression levels, thereby ensuring accurate determination of the levels of casein accumulated in the plant tissue. The encoding of SEQ ID NO: 7 was accompanied with overexpression of SEQ ID NO: 9, which was also co-expressed with casein in soybean; and the encoding of SEQ ID NO: 8 was accompanied with overexpression of SEQ ID NO: 10, which was also co-expressed with casein in soy bean. The proteins isolated from the soybean transformed with SEQ ID NO: 7 or SEQ ID NO: 8 were caseins, as confirmed by the Western Blot analysis depicted in Fig. 5 and 6.Docket No. 719.001.002.PCTExample 3: Synergistic Enhancement of Protein Expression by Co-expression of Multiple Molecular Chaperones

[0115] To investigate potential synergistic effects for protein expression enhancement, combinations of the most effective molecular chaperones identified in Example 2 will be coexpressed with casein proteins in soybean cells. This approach is based on the hypothesis that combining chaperones with different, complementary functions will provide an additive or synergistic benefit to protein folding and accumulation, surpassing the enhancement achieved by any single chaperone.

[0116] The experimental design will parallel the methods described in Example 1. New multigene expression constructs will be generated for the simultaneous expression of two distinct molecular chaperones alongside alpha and kappa casein. Based on the strong individual performance observed in Example 2, the primary combinations to be tested will include: (1) Glyma.01G003700 (SEQ ID NO: 2) co-expressed with bovine DNAJB12 (SEQ ID NO: 4); (2) Glyma.04g247900 (PDI, SEQ ID NO: 6) co-expressed with bovine DNAJB12 (SEQ ID NO: 4); (3) Glyma.01G003700 (SEQ ID NO: 2) co-expressed with Glyma.04g247900 (SEQ ID NO: 6).

[0117] The hypothesis is that combining chaperones with distinct mechanisms — for instance, the predicted protein disulfide isomerase activity of Glyma.01G003700 and the Hsp70-recruiting activity of the DnaJ family member DNAJB12 — will address multiple potential bottlenecks in the protein folding pathway. This concerted action is expected to result in a more robust folding environment in the endoplasmic reticulum.

[0118] Analysis of protein accumulation will be conducted as described previously. V5- tagged alpha casein and FLAG-tagged kappa casein expression levels will be quantified by Western blot and normalized to their respective mRNA levels determined by qPCR. The results for each chaperone combination will be compared not only to the control construct expressing mScarlet (pMOZ3617) but also to the results from the individual chaperone expressions detailed in Example 2. This comparison will determine whether specific combinations provide a synergistic enhancement, where the combined effect is greater than the sum of the individual effects. Such an investigation will provide valuable insights into optimizing chaperone networks for heterologous protein production and may lead to the development of highly efficient, commercially viable protein expression systems in plants.

Claims

Docket No. 719.001.002.PCTCLAIMS1. A method for increasing the accumulation of a heterologous protein in a plant cell, the method comprising: introducing into the plant cell one or more nucleic acid molecules comprising:(a) a first nucleotide sequence encoding the heterologous protein; and(b) a second nucleotide sequence encoding an endoplasmic reticulum (ER)-resident molecular chaperone protein having at least 85% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 4, and SEQ ID NO: 6; and co-expressing the heterologous protein and the molecular chaperone protein in the plant cell, thereby increasing the accumulation of the heterologous protein compared to a plant cell not coexpressing the molecular chaperone protein.

2. The method of claim 1, wherein the molecular chaperone protein has at least 95% sequence identity to the amino acid sequence set forth in SEQ ID NO: 2.

3. The method of claim 1, wherein the molecular chaperone protein has at least 95% sequence identity to the amino acid sequence set forth in SEQ ID NO: 4.

4. The method of claim 1, wherein the method comprises introducing a third nucleotide sequence encoding a second, different ER-resident molecular chaperone protein, and co-expressing the heterologous protein with both molecular chaperone proteins.

5. The method of claim 1, wherein the heterologous protein is a milk protein.

6. The method of claim 5, wherein the milk protein is a casein protein selected from the group consisting of aSl-casein, aS2-casein, 0-casein, K-casein, and combinations thereof.

7. The method of claim 1, wherein the heterologous protein is selected from the group consisting of a whey protein, an egg white protein, a therapeutic protein, and an industrial enzyme.

8. The method of claim 1, wherein the plant cell is a dicot plant cell.

9. The method of claim 8, wherein the plant cell is a soybean (Glycine max) cell.Docket No. 719.001.002.PCT10. The method of claim 1, wherein the first and second nucleotide sequences are codon-optimized for expression in the plant cell.

11. The method of claim 1, wherein the first and second nucleotide sequences are operably linked to promoters that are active in the plant cell.

12. The method of claim 11, wherein at least one of the promoters is a seed-specific promoter.

13. The method of claim 1, wherein the first nucleotide sequence further comprises a sequence encoding an ER signal peptide.

14. The method of claim 1, wherein the accumulation of the heterologous protein is increased by at least 1.5-fold.

15. A recombinant DNA construct for enhancing heterologous protein expression in a plant cell, the construct comprising:(a) a first expression cassette comprising a first nucleotide sequence encoding a heterologous protein, operably linked to a first promoter; and(b) a second expression cassette comprising a second nucleotide sequence encoding an endoplasmic reticulum (ER)-resident molecular chaperone protein having at least 85% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 4, and SEQ ID NO: 6, operably linked to a second promoter.

16. The recombinant DNA construct of claim 15, wherein the second nucleotide sequence encodes a molecular chaperone protein having at least 95% sequence identity to SEQ ID NO: 2 or SEQ ID NO: 4.

17. The recombinant DNA construct of claim 15, wherein the heterologous protein is a casein protein.

18. A transgenic plant cell comprising one or more exogenous nucleic acid molecules encoding:(a) a heterologous protein; and(b) an endoplasmic reticulum (ER)-resident molecular chaperone protein having at least 85% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 4, and SEQ ID NO: 6.Docket No. 719.001.002.PCT19. The transgenic plant cell of claim 18, wherein the plant cell is from a soybean plant.

20. The transgenic plant cell of claim 18, wherein the heterologous protein is a casein protein and is expressed in a micellar form.

21. A transgenic plant, or part thereof, comprising a plurality of the transgenic plant cells of claim 18.

22. The transgenic plant of claim 21, wherein the plant is a soybean plant.

23. A seed produced by the transgenic plant of claim 21, wherein the seed comprises the one or more exogenous nucleic acid molecules.

24. A food composition comprising a protein extract obtained from the transgenic plant of claim 21.

25. The food composition of claim 24, wherein the food composition is a dairy product substitute and the protein extract comprises one or more casein proteins.

26. The food composition of claim 25, wherein the protein extract further comprises the ER-resident molecular chaperone protein.

27. A method for increasing the expression of casein protein in a plant, comprising: inserting a plasmid containing SEQ ID NO: 7 or a plasmid containing SEQ ID NO: 8; overexpressing SEQ ID NO: 7 to produce SEQ ID NO: 9 or SEQ ID NO: 8 to produce SEQ ID NO: 9, thereby co-expressing SEQ ID NO: 9 or SEQ ID NO: 10 with one or more casein proteins.

28. The method of claim 27, wherein the plant is a soybean plant (Glycine max).

29. A plant comprising: a plasmid containing SEQ ID NO: 7 or a plasmid containing SEQ ID NO: 8, thereby introducing a transgene encoding a casein protein and a transgene encoding for SEQ ID NO: 9, or SEQ ID NO: 10, or homologs of the transgenes encoding for SEQ ID NO: 9 or SEQ ID NO: 10.

30. The plant of claim 29, wherein the plant is a transgenic soybean plant.Docket No. 719.001.002.PCT31. A recombinant DNA construct, comprising a casein-encoding gene and one or more genes selected from SEQ ID NO: 7 or SEQ ID NO: 8, or homologues thereof.