Method for enhanced expression of casein proteins in transgenic SOY using inserted regulatory sequences

By inserting regulatory sequences and optimizing GC content within the coding sequence of casein proteins, the expression levels in soybean plants are enhanced, addressing the challenge of high-yield, cost-effective, and sustainable production of casein proteins.

WO2026096477A1PCT designated stage Publication Date: 2026-05-07MOZZA FOODS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MOZZA FOODS INC
Filing Date
2025-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Achieving high expression levels of foreign proteins, particularly casein proteins, in soybean systems remains challenging, limiting their industrial and nutritional applications.

Method used

Strategic insertion of regulatory sequences at defined positions within the coding sequence of casein proteins, combined with GC content optimization, enhances protein accumulation by improving mRNA stability and translation efficiency.

Benefits of technology

Significantly increases casein protein expression levels in transgenic soybean plants, offering cost-effective and sustainable production of milk proteins with higher yields and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052863_07052026_PF_FP_ABST
    Figure US2025052863_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides genetic constructs and methods for enhanced production of milk proteins in transgenic plants. The invention comprises novel nucleic acid constructs wherein regulatory sequences derived from plant genes are inserted into bovine casein coding sequences, resulting in increased protein accumulation when expressed in plant cells. In particular embodiments, the insertion of these regulatory sequences into casein coding sequences leads to enhanced protein accumulation in transgenic soybean plants compared to constructs lacking the regulatory sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. 720.001.003.PCTPatent Cooperation TreatyPatent ApplicationMETHOD FOR ENHANCED EXPRESSION OF CASEIN PROTEINS IN TRANSGENIC SOY USING INSERTED REGULATORY SEQUENCESInventor(s): Cory Tobin, PhDAssignee: Mozza Foods, Inc.1927 Zonal AvenueLos Angeles, CA 90033 a Delaware CorporationEntity: Small business concernFILED ELECTRONICALLY ON October 28, 2025Docket No. 720.001.003.PCTCROSS-REFERENCE INFORMATION

[0001] This application claims priority to US Provisional Patent Application Serial No. 63 / 713,028 filed on October 28, 2024; and US Provisional Patent Application Serial No. 63 / 748,848 filed on lanuary 23, 2025, each of are incorporated herein in their entirety.Sequence Listing

[0002] This application contains a sequence listing in electronic format. The Sequence Listing is provided as a file entitled 720001003PCT. xml, created on October 28, 2025 which is 35 kilobytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety.TECHNICAL FIELD

[0003] The present invention relates to the field of genetic engineering, specifically to methods for enhanced protein expression in transgenic plants. More particularly, the invention relates to genetic constructs comprising inserted regulatory sequences that improve the expression of milk proteins (caseins) in soybean plants.BACKGROUND

[0004] Plant-based protein production systems have become increasingly important for sustainable protein production. Soybeans (Glycine max)' are well-established platforms for recombinant protein expression. However, achieving high expression levels of foreign proteins remains challenging. Casein proteins, particularly alpha si (asi), alpha s2 (as ), beta (P), and kappa (K) casein, are valuable milk proteins with numerous industrial and nutritional applications. The ability to produce these proteins in soybean systems could provide significant advantages in terms of scalability, cost-effectiveness, and sustainability compared to traditional dairy -based production methods.

[0005] A significant challenge in expressing heterologous proteins in plant systems is achieving commercially viable production levels. Various genetic regulatory elements and DNA sequences have been explored to enhance protein expression in transgenic plants. Of particular interest areDocket No. 720.001.003.PCT specific DNA sequences that, when inserted into genetic constructs, can significantly increase the expression and accumulation of target proteins. These regulatory sequences can enhance gene expression through various mechanisms, including improved transcript processing, increased mRNA stability, and enhanced protein accumulation. Such approaches offer promising solutions for improving recombinant protein yields in transgenic plants, potentially enabling economically viable production of valuable proteins such as caseins in soybean-based systems.SUMMARY OF INVENTION

[0006] The present invention provides novel genetic constructs and methods for enhancing casein protein expression in transgenic soybean plants through strategic modification of protein coding sequences. The invention comprises specific genetic sequences designed to optimize the expression of casein proteins through the insertion of regulatory DNA sequences at defined positions within the coding sequence. These constructs incorporate carefully positioned nucleic acid sequences that enhance protein accumulation, resulting in significantly improved expression levels of the target proteins in transgenic plants.Brief Description of Drawings

[0007] FIG. 1 is a depiction of a graph of expression levels of plasmids which increase expression of asi-casein.

[0008] FIG. 2 is a depiction of a Western blot corresponding to the proteins expressed by the plasmids in FIG. 1.

[0009] FIG. 3 is a depiction of a graph of expression levels of plasmids which increase expression of K-casein.

[0010] FIG. 4 is a depiction of a Western blot corresponding to the proteins expressed by the plasmids in FIG. 3.

[0011] FIG. 5 is a depiction of a graph of expression levels of plasmids which increase expression of [3-casein.

[0012] FIG. 6 is a depiction of a Western blot corresponding to the proteins expressed by the plasmids in FIG. 5.

[0013] FIG. 7 is a plasmid map depicting pMOZ4334, a construct optimized for high expression through GC content optimization and inserted regulatory sequences.Docket No. 720.001.003.PCT

[0014] FIG. 8 is a plasmid map depicting pMOZ2023, a control construct used for comparison in Example 5.

[0015] FIG. 9 is a depiction of a Western blot comparing asi-casein protein accumulation in transgenic soybean lines transformed with pMOZ4334 against a control construct (pMOZ2023). Lanes 1 and 2 contain extracts from independent pMOZ4334 events, and Lane 3 contains an extract from the pMOZ2023 control. Lanes 4 through 8 show a V5-tagged protein standard curve (10 ng, 20 ng, 30 ng, 45 ng, and 50 ng, respectively) used for quantification.DETAILED DESCRIPTION

[0016] The present invention encompasses genetic constructs specifically designed for enhanced expression of casein proteins in soybean plants. These constructs incorporate several key components working in concert to achieve optimal protein expression. At the 5' end of the construct, a promoter sequence suitable for robust expression in soybean plants initiates transcription. The promoter may be constitutive, tissue-specific, synthetic, or inducible, depending on the desired expression pattern.

[0017] The invention provides novel genetic constructs comprising regulatory sequences (SEQ ID NOs: 2 and 4-7) that can be inserted into casein coding sequences to enhance protein accumulation in transgenic soybean plants. These regulatory sequences may be positioned at various locations within the coding sequence of the target casein protein. The specific positioning of these regulatory sequences has been determined to significantly impact their enhancement capability. SEQ ID NO: 1 or SEQ ID NO: 3 represent a specific embodiment wherein one regulatory sequence, SEQ ID NO: 2 or SEQ ID NO: 9, respectively, is inserted at a defined position within the asi-casein coding sequence, resulting in enhanced protein accumulation. Regulatory sequences (SEQ ID NOs: 2 and 4-8) may enhance expression when properly positioned within different casein coding sequences, including asi-casein, P-casein, and K-casein genes. The regulatory sequences, which are inserted into asi-casein, P-casein, and K-casein genes, are introns derived from various sources, including both soybean genes (e.g., Glyma.17G186600, Glyma.08G28700, Glyma.12G074800, Glyma.15G163600, Glyma.12G038100, Glyma.17G186600) and Arabidopsis genes (e.g., AT4G05320, AT2G47600, AT5G17990), and range in size from approximately 297 to 2500 base pairs.Docket No. 720.001.003.PCT

[0018] The coding sequence portion of the construct contains sequences encoding one or more casein proteins. These may include any or all of asi-casein, as2-casein, P-casein, and K-casein. The coding sequences have been optimized for expression in soybean plants, taking into account codon usage preferences and other factors that may affect expression efficiency. Increased GC content in coding sequences is generally known to correlate with improved stability and translational initiation efficiency in plant expression systems. The construct is terminated by appropriate terminator sequences that ensure proper processing of the transcript.

[0019] In specific and preferred embodiments, the nucleic acid sequences encoding the bovine casein proteins (asi-casein, as2-casein, P-casein, K-casein) or other target proteins such as FAM20C kinase, are further optimized by modification of the codon usage to increase the overall Guanine-Cytosine (GC) content of the coding sequence. This optimization is performed while maintaining the specific amino acid sequence of the target protein. Synonymous codons are selected to preferentially incorporate G and C nucleotides, maximizing the GC percentage within the coding region. This approach is independent of, but complementary to, the insertion of the regulatory sequences inserted into bovine caseins, wherein the regulatory sequences (e.g., SEQ ID NOs: 2 and 4-9) are defined herein.

[0020] The resulting increased GC content is utilized to enhance expression by influencing mechanisms such as mRNA secondary structure, transcript stability, and ribosome binding affinity, leading to improved translation efficiency within the soybean host cell. In certain embodiments, the GC-optimized coding sequence is engineered to achieve an overall GC content of at least 55%, or preferably at least 60% or 65%, throughout the coding sequence. This optimization acts synergistically with the 5' proximal insertion of the regulatory sequences (e.g., as introns) to achieve significantly enhanced accumulation of the target protein.

[0021] The enhancement mechanism of the invention operates through several coordinated processes. The inserted regulatory sequence functions to increase gene expression through mechanisms that may include improved mRNA stability, enhanced nuclear export, or increased transcription rates. The specific sequence elements have been selected based on their demonstrated ability to facilitate enhanced expression in soybean plants.

[0022] The invention further encompasses methods for introducing these genetic constructs into soybean plants and expressing the casein proteins. Transformation of soybean plants may be accomplished through various methods known in the art, including but not limited toDocket No. 720.001.003.PCTAgrobacterium-mediated transformation or biolistic transformation. Following transformation, plants are selected based on appropriate markers and screened for casein protein expression.

[0023] The transformed plants are cultivated under conditions suitable for growth and protein expression. The expressed casein proteins may be extracted from the plant tissue using various methods appropriate for protein isolation. These methods may include aqueous extraction, followed by purification steps such as chromatography or filtration. The specific extraction and purification methods may be optimized based on the particular casein protein being produced and the desired level of purity.

[0024] The present invention provides several significant advantages over existing protein production methods. Primarily, the strategic insertion of the regulatory sequence significantly increases the expression levels of casein proteins in transgenic soybean plants, leading to higher protein yields per plant. This enhanced expression efficiency translates directly to improved production economics in plant-based protein systems. The invention offers a particularly cost- effective approach to producing milk proteins, as it leverages the well-established and economical soybean agricultural system rather than requiring specialized bioreactor facilities or animal-based production methods. Furthermore, this plant-based production system represents a more sustainable alternative to traditional animal-derived protein production, requiring fewer resources and generating a smaller environmental footprint. The scalability of soybean agriculture also means that this system can be readily expanded to meet increasing market demands for milk proteins.

[0025] In some embodiments, the regulatory sequence may be operably linked to genes encoding other milk proteins to enhance their expression in transformed soybean cells. In certain embodiments, the milk proteins are selected from the group consisting of caseins, whey proteins, and milk fat globule membrane proteins. In particular embodiments, the caseins are selected from the group consisting of asi-casein, as2-casein, [3-casein, K-casein, and gamma-casein. In further embodiments, the whey proteins are selected from the group consisting of beta-lactoglobulin, alpha-lactalbumin, serum albumin, immunoglobulins, lactoferrin, lactoperoxidase, glycomacropeptide, and proteose peptones. In additional embodiments, the milk fat globule membrane proteins are selected from the group consisting of butyrophilin, xanthine oxidase, CD36, mucin 1, mucin 15, lactadherin (PAS-6 / 7), and adipophilin. In certain embodiments, the milk proteins may be derived from various species, including but not limited to bovine, caprine,Docket No. 720.001.003.PCT ovine, buffalo, camelid, equine, or human sources. In particular embodiments, the regulatory sequence enhances accumulation of these milk proteins in transformed soybean cells by at least 2- fold, at least 3-fold, at least 5-fold, or at least 10-fold compared to constructs lacking the regulatory sequence.

[0026] The enhancement effect of the inserted regulatory sequence is not dependent on the specific protein being expressed. Rather, the enhancement is a function of the regulatory sequence itself and its positioning within the construct. This allows for flexible application across different proteins of interest, including but not limited to all casein types, whey proteins, milk fat globule membrane proteins, and other valuable recombinant proteins (e.g., kinases like FAM20C).

[0027] The regulatory sequence of the present invention may be positioned at various locations within the coding sequence to achieve different enhancement effects. The insertion position may be within any region of the coding sequence, provided it maintains the proper reading frame and does not disrupt critical functional domains of the protein. Various positions within the coding sequence may be utilized, including positions near the 5' end of the coding sequence, in the middle region, or toward the 3' end.

[0028] In some embodiments, the construct utilizes multiple copies of the regulatory sequence or functionally active fragments thereof inserted at different positions within the coding sequence. In certain embodiments, the combination of multiple insertions may result in additive or synergistic enhancement effects.

[0029] The present invention further encompasses synthetic regulatory sequences constructed from functional sub-fragments of SEQ ID NOs: 4-7. These synthetic sequences may contain multiple tandem copies or repeats (monomeric, dimeric, trimeric, or multimeric) of the core enhancing motif derived from the regulatory sequences. Construction and insertion of such synthetic variants may be utilized to maximize the Intron-Mediated Enhancement effect achieved in the transgenic soybean plants.

[0030] The present invention further encompasses nucleic acid sequences that are functionally equivalent to SEQ ID NOs: 2, 4, 5, 6, 7, 8, or 9, and which, when inserted into the coding sequence of a target protein (such as a bovine casein), enhance protein accumulation in a transgenic plant cell by at least 1.5-fold compared to an otherwise identical construct lacking the inserted sequence. Functionally equivalent sequences include variants having at least 80%, at least 90%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 2, 4, 5, 6, 7, 8, or 9, wherein saidDocket No. 720.001.003.PCT variants retain the ability to function as an expression enhancer when positioned 5' proximally within the coding sequence.

[0031] In specific embodiments, the regulatory sequence is a functionally active fragment of SEQ ID NOs: 4-9, wherein said fragment is at least 50 base pairs (bp), at least 100 bp, at least 200 bp, or at least 300 bp in length, and retains the capacity to mediate enhancement of gene expression when inserted into the target coding sequence. Such functional fragments can be identified by standard deletion and insertion assays known in the art.

[0032] The position of the regulatory sequence insertion relative to specific protein domains may be optimized for maximum enhancement. The optimal insertion position may vary depending on the specific protein being expressed and the desired expression characteristics. The invention encompasses all such positional variations that achieve the enhancement effect.

[0033] The insertion point of SEQ ID NOs 2 and 4-9 within the target casein gene is the position of insertion within the DNA sequence of the casein gene that may directly affect multiple critical protein characteristics. In some embodiments, the insertion position fundamentally influences protein folding, structural integrity, expression levels, and protein accumulation in transformed plants.

[0034] Experimental determination of optimal insertion points may require structural analysis of the casein proteins to identify suitable regions, assessment of conserved domains that should remain unperturbed, and testing of multiple insertion constructs to identify positions that maintain both protein functionality and enhanced accumulation. SEQ ID NO: 1 and 3, which represent a successful embodiment achieving enhanced protein accumulation, via optimal insertion points in a nucleotide sequence, wherein the optimal insertion points are: (1) regions of the nucleotide sequence appended to an intron for modulating the expression of genes (a regulatory sequence); and (2) proximal to the 5’ side of the nucleotide sequence (5’ proximal insertion).

[0035] For clarity and consistency throughout this disclosure, the common names of the target casein proteins (e.g., alpha si casein, alpha s2, beta casein, kappa casein) are used interchangeably with their corresponding standard scientific nomenclature utilizing Greek symbols (asi-casein, aS2- casein, 0-casein, K-casein, etc.). Both the text-based names and the Greek symbolic representations refer to the identical casein protein species unless otherwise specified.

[0036] As used herein, a “regulatory sequence” comprises a nucleic acid segment that modulates the expression of one or more genes through direct or indirect mechanisms, wherein said sequenceDocket No. 720.001.003.PCT does not encode a protein product. Such regulatory sequences may be positioned upstream, downstream, or within the gene unit being regulated, and may exert their regulatory function through interaction with cellular factors and / or through inherent structural characteristics of the nucleic acid segment. The regulatory activity may manifest in the modulation of any step of gene expression, including but not limited to transcription initiation, transcription rate, RNA processing, RNA stability, translation efficiency, or post-translational modifications. The sequence may function through recruitment of regulatory factors, alteration of local chromatin architecture, or other molecular mechanisms that influence gene expression.

[0037] The enhancement effect described herein is frequently referred to as Intron-Mediated Enhancement (IME), which defines the process by which an intron, typically located near the 5' end of a gene's transcribed region, significantly increases gene expression relative to an identical construct lacking the intron. IME operates through mechanisms that are generally coupled with intron splicing but are independent of the intron’s nucleotide sequence simply encoding a protein product.

[0038] As used herein, “5' proximal insertion” refers to the positioning of a regulatory sequence within the first 200 base pairs downstream of the start codon of a coding sequence. The start codon is designated as positions 1-3 of the coding sequence, with subsequent positions numbered sequentially.

[0039] In specific embodiments, the regulatory sequence functions as a plant intron, and the 5' proximal insertion is designed to maximize Intron -Mediated Enhancement (IME). The optimal 5' proximal position for achieving IME is typically within the first 150 codons of the coding sequence. Therefore, in preferred embodiments, the regulatory sequence (intron) is inserted such that its 5’ splice site is located between codon 1 and codon 150 of the bovine casein coding sequence.

[0040] The regulatory sequence, when inserted 5' proximally, contains sequences enabling proper splicing in the host plant cell, characterized by the presence of canonical 5' splice donor sites (GT) and 3' splice acceptor sites (AG), and an internal branch point sequence, allowing it to be effectively excised from the nascent mRNA transcript during processing. This mechanism ensures that the final mature mRNA encodes the full, unmodified amino acid sequence of the target casein protein.Docket No. 720.001.003.PCT

[0041] Transgenic soybean plants expressing casein proteins can be generated using transformation methods known in the art (see e.g., U.S. Patent No. 11,326,176). Such established methods can be employed to introduce casein modified by SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9, into soybean plants. The nucleic acid constructs used for transformation may further comprise additional regulatory elements known in the art, including but not limited to promoters, terminators, 5' untranslated regions (UTRs), 3' UTRs, signal peptides, transit peptides, selection markers, and combinations thereof, wherein, in some embodiments, said regulatory elements are operably linked to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9 to facilitate expression in transgenic soybean plants.

[0042] In certain embodiments, the regulatory sequences (e.g., SEQ ID NOs: 2 and 4-9) may be inserted within the coding sequence at various positions relative to the start codon, including but not limited to within the first 50, 100, 150, 200, 250, 300, 400, or 500 base pairs of the casein coding sequence. The positioning of intron sequences near the 5' end of coding sequences has been demonstrated to enhance gene expression in plant systems. In some embodiments, the regulatory sequences may be inserted within the first third, first half, or first two-thirds of the coding sequence. In particular embodiments, the regulatory sequences function as introns when inserted at these positions within casein coding sequences. In preferred embodiments, the regulatory sequences are inserted within the first 200 base pairs of the casein coding sequence, as this 5' proximal positioning allows the regulatory sequences to enhance expression through intron- mediated enhancement mechanisms while maintaining proper protein structure and function. The optimal positioning within the first 200 base pairs may be determined through standard optimization procedures known in the art.

[0043] The enhancement effect observed with 5' proximal insertions may be attributed to multiple mechanisms, including but not limited to: improved recruitment of splicing machinery, enhanced mRNA export from the nucleus, increased mRNA stability, and more efficient recruitment of translation initiation factors. The proximity to the start codon may facilitate these processes by allowing the regulatory sequences to interact more effectively with the core promoter elements and the translation initiation complex, thereby maximizing the Intron-Mediated Enhancement (IME) effect.Docket No. 720.001.003.PCT

[0044] In highly preferred embodiments, the present invention utilizes a synergistic optimization strategy combining GC content modification and 5' proximal regulatory sequence insertion to achieve maximal protein accumulation. A construct utilizing this strategy comprises: (i) a coding sequence for a target protein, optimized to achieve an overall Guanine-Cytosine (GC) content of at least 55%, and (ii) a regulatory sequence (e.g., an intron selected from SEQ ID NOs: 4-9) inserted 5' proximally within the first 200 base pairs of said GC-optimized coding sequence.

[0045] The combined effect of increased GC content (enhancing translation efficiency and stability) and 5' proximal intron insertion (enhancing transcription and RNA processing via IME) provides a statistically significant synergistic increase in protein accumulation compared to constructs utilizing either optimization method alone. As demonstrated in Example 5, this dualoptimized strategy (pMOZ4334) can lead to superior accumulation of target proteins (e.g., casein) compared to baseline controls.

[0046] The expression of asi-casein can be attributed to the identity of the intron and location of the intron in the 5’-end of the asi-casein (i.e., point of insertion of the intron). SEQ ID NO: 2 and SEQ ID NO: 4 can be inserted into a region proximal to the 5’-end of a nucleotide sequence for expressing asi-casein. The resulting nucleotide sequence from insertion of SEQ ID NO: 2 into the region proximal to the 5’-end of the nucleotide sequence for expressing asi-casein is SEQ ID NO: 1. The resulting nucleotide sequence from insertion of SEQ ID NO: 9 into the region proximal to the 5’-end of the nucleotide sequence for expressing asi-casein is SEQ ID NO: 3.

[0047] SEQ ID NO: 1 and SEQ ID NO: 3 are nucleotides that can be encoded for the expression of recombinant asi-casein casein in the plant cell. Stated another way, SEQ ID NO: 2 and SEQ ID NO: 9 do not interfere with the expression of recombinant asi-casein in the plant cell. SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO:7, SEQ ID NO: 8, and SEQ ID NO: 9 are other introns that can be inserted into a nucleotide sequence for expressing asi-casein in the plant cell, that can: increase expression of recombinant asi-casein in the plant cell, in comparison to plasmids absent of said introns; and not interfere with the expression of functional recombinant aS 1 -casein casein in the plant cell.

[0048] A plasmid can contain: a nucleotide sequence for expressing asi-casein and an intron inserted into the nucleotide sequence for expressing asi-casein, wherein: the plasmids are selected among plasmids 3547, 3546, 3545, 3544, 3543, 3539, 3537, 3535, 3540, and 3536. The intron is selected from: SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7,Docket No. 720.001.003.PCTSEQ ID NO: 8, or SEQ ID NO: 9; and their associations with the plasmids are described below. Plasmid 3548 contains a nucleotide sequence for expressing asi-casein, but does not have any introns inserted into the nucleotide sequence for expressing asi-casein.

[0049] The expression of K-casein can be attributed to the identity of the intron and location of the intron in the 5’-end of the K-casein (i.e., point of insertion of the intron). SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO:7, SEQ ID NO: 8, and SEQ ID NO: 9 are introns that can be inserted into a region proximal to the 5 ’-end of nucleotide sequence for expressing K-casein in the plant cell, that can: increase expression of recombinant K-casein in the plant cell, in comparison to plasmids absent of said introns; and not interfere with the expression of functional recombinant K-casein in the plant cell.

[0050] A plasmid can contain: a nucleotide sequence for expressing K-casein and an intron inserted into the nucleotide sequence for expressing K-casein, wherein: the plasmids are selected among plasmids 3866, 3865, 3864, 3863, 3862, 3861, 3860, 3859, 3858, and 3540. The intron is selected from: SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9; and the respective association within a plasmid are described herein. Plasmid 3548 contains a nucleotide sequence for expressing K-casein, but does not have any introns inserted into the nucleotide sequence for expressing K-casein.

[0051] The expression of P-casein can be attributed to the identity of the intron and location of the intron in the 5’-end of the P-casein (i.e., point of insertion of the intron). SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO:7, SEQ ID NO: 8, and SEQ ID NO: 9 are introns that can be inserted into a region proximal to the 5 ’-end of nucleotide sequence for expressing P-casein in the plant cell, that can: increase expression of recombinant P-casein in the plant cell, in comparison to plasmids absent of said introns; and not interfere with the expression of functional recombinant P-casein in the plant cell.

[0052] A plasmid can contain: a nucleotide sequence for expressing P-casein and an intron inserted into the nucleotide sequence for expressing P-casein, wherein: the plasmids are selected among plasmids 3857, 3856, 3855, 3854, 3853, 3852, 3852, 3851, 3850, and 3849. The intron is selected from: SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9; and the respective association within a plasmid are described herein. Plasmid 3868 contains a nucleotide sequence for expressing P-casein, but does not have any introns inserted into the nucleotide sequence for expressing P-casein.Docket No. 720.001.003.PCT

[0053] The target protein utilized in this synergistic strategy may be any casein protein, whey protein, or other recombinant proteins (e.g., FAM20C kinase) for which high-level expression in plants is desired. In particular embodiments, the dual-optimized construct results in protein accumulation of the target protein at levels exceeding 0.1% of the total soluble protein (TSP) in the transformed soybean cells.

[0054] In some embodiments, insertion of the nucleic acid sequence at specific positions within the asi-casein sequence provides at least a three-fold increase in protein accumulation compared to an unmodified asi-casein sequence. In certain embodiments, specific positioning of the inserted sequence increases accumulation of asi-casein by at least 200% compared to control constructs. In particular embodiments, the presence of the inserted sequence at defined positions results in protein accumulation levels that are at least three times higher than those achieved with unmodified asi-casein sequences.

[0055] In some embodiments, the regulatory sequence may be used to enhance milk protein expression in plants other than soybean. In certain embodiments, the plants are selected from the group consisting of other legumes, cereals, oilseeds, and leafy plants. In particular embodiments, the legumes are selected from the group consisting of pea, chickpea, lentil, common bean, peanut, alfalfa, and clover. In further embodiments, the cereals are selected from the group consisting of rice, maize, wheat, barley, oats, sorghum, millet, and rye. In additional embodiments, the oilseeds are selected from the group consisting of canola, sunflower, safflower, flax, hemp, and camelina. In other embodiments, the plants may include tobacco, potato, tomato, lettuce, spinach, and other leafy greens. In certain embodiments, the regulatory sequence may be used in both monocotyledonous and dicotyledonous plants. In particular embodiments, the plants may be selected for their protein content, biomass yield, ease of transformation, established agricultural practices, or industrial processing capabilities.EXAMPLES

[0056] It will be more readily understood by reference to the following examples, which are included merely for purposes of illustration of certain aspects and embodiments of the present invention, and are not intended to limit the invention. Aspects of the present invention are now being generally described, namely: the determination of the nucleotides in sequences corresponding to plasmids described herein; the isolation of RNA, DNA and proteins produced byDocket No. 720.001.003.PCT these sequences and plasmids; the insertion of sequences of DNA into relevant plasmids, and the subsequent insertion into a soybean plant; and the determination of casein levels in the soybean after insertion of the plasmids.EXAMPLE 1: Plasmids for Transfection of Plants Leading to In-Vivo Casein Micelles

[0057] Vectors: Multigene vectors, such as the plasmids described above, were isolated, wherein the vectors expressed the coding regions for the following proteins: (1) bovine asi-casein (Uniprot accession # P02662), (2) green fluorescent protein (GFP, Uniprot accession # P42212), 3) bovine P casein (Uniprot accession # P02666), (4) bovine K casein (Uniprot accession # P02668), and (5) bovine FAM20C kinase (uniprot accession number # F1MXQ3). The vectors herein investigated for transfecting soybean include: constructs 2023, 3535, 3536, 3537, 3539, 3540, 3543, 3544, 3545, 3546, 3547, 3548, 3859, 3860, 3861, 3862, 3863, 3864, 3865, 3866, 3849, 3850, 3851, 3852, 3853, 3854, 3855, 3855, 3856, 3857, 3868, and 4334. The vectors were assembled using the modular cloning system MoClo (Engler, C., Youles, M., Gruetzner, R , Ehnert, T.-M., Werner, S., Jones, J. D. G., Patron, N. J., & Marillonnet, S. (2014). A Golden Gate Modular Cloning Toolbox for Plants. ACS Synthetic Biology’, 3(11), 839-843.). All proteins were expressed under constitutively active, seed specific, or synthetic plant promoters, with a subset of these proteins possessing translationally-fused epitope tags and / or target-peptide sequences on their C-termini.

[0058] Caseins: Gene sequences that encode for bovine asi-casein (P02662), bovine P-casein (P02666), and bovine K-casein (P02668) were derived from Uniprot and modified to include: SEQ ID NO: 10 (i.e., tags) in the respective C-terminus thereof. The first sequence encodes for a two- serine spacer that functions to provide enough space for each casein protein to properly fold without steric hindrance or interference from other tags. The second and third sequences encode for two affinity purification tags (flag-tag and 6-Histidine tag) that are commonly used for protein identification and purification. The fourth sequence encodes for an HDEL target peptide that functions to retain soluble casein proteins in the endoplasmic reticulum (ER). Additionally, an N- terminus signal peptide, gmGlycininl (GY1, P04776), was added to all three caseins in order to target them towards the ER and vacuoles. Recombinant casein sequences were expressed under the constitutive AtuMas promoter and 5’ untranslated region (UTR). (See Langridge, W. H. R., Fitzgerald, K. J., Koncz, C., Schell, J., & Szalay, A. A. (1989). Dual promoter of AgrobacteriumDocket No. 720.001.003.PCT tumefaciens mannopine synthase genes is regulated by plant growth hormones. Proceedings of the National Academy of Sciences, 86(9), 3219-3223).

[0059] E. coJi transformation: The plasmids, above, were each transformed into a Lucigen Ecloni 10G bacterium using the following chemical transformation protocol. First, Lucigen Ecloni 10G E. coli thermo-competent cells were thawed on ice for approximately 5 minutes. Then, competent cells were spiked with 1 pg-100 ng of plasmid DNA and left to incubate on ice for 10 minutes. Once complete, the plasmid-bacterial mixture was heat shocked at 42°C for 90 seconds and immediately replaced on ice for 5 minutes. Afterwards, transformed cells were mixed with lOOpl of liquid broth (LB) and cultured in a 37°C shaker set to 225 rpm for 45 minutes. Following incubation, cultured bacteria was plated on LB agar plates with the appropriate antibiotic for selection (1 :1000 concentration) and left to further incubate overnight in a 37°C growth chamber.

[0060] High throughput blue-white selection: Following overnight incubation, plates were checked for bacterial colony growth. To increase recombinant bacteria selection, a MoClo compatible blue-white selection system was used. In brief, blue-white selection plasmid vectors carry the “lacZ” operon sequence within their multiple cloning site. In the absence of recombinant DNA, the lacZ operon will enable a biochemical reaction that turns the colony blue. Whereas, when recombination occurs, lacZ operon activity will be disrupted leaving the colony to be white. For each cloning reaction, at least two white colonies were chosen for further verification.

[0061] Plasmid selection and verification: Picked colonies were placed in 5 mLs of LB plus their respective selection antibiotic (1 : 1000 concentration) and cultured overnight in a 37°C shaker (225 rpm). Once cloudy, plasmids were purified out of bacteria using the NucleoSpin miniprep kit (Takara bio inc.). Then, each plasmid was digested using the appropriate restriction enzymes in order to confirm the presence of a DNA insert and sent for Sanger and nanopore sequencing to confirm sequence correctness. Colonies containing the correct plasmids were made into frozen 20% glycerol stocks.

[0062] Electroporation into agrobacterium: The completed plasmids were each transformed into a respective EHA105 electrocompetent agrobacterium cell. To achieve this, 30ng of the purified plasmid was mixed into ice-thawed electrocompetent cells and swiftly transferred into a prechilled 0.2 cm Gene pulser cuvette. The cuvette was then loaded into an electroporation chamber and given an electric pulse of 2.5kV. The resulting transformed cells were mixed with 500 pl of LB and left to culture in a 28°C shaking incubator (120rpm) for 2-4 hours. Afterwards, culturedDocket No. 720.001.003.PCT cells were plated onto LB agar plates containing the appropriate antibiotic selection media and placed in a 28°C incubator for two days. Once colonies formed, a minimum of three were picked for further verification via purification, digestion, and sequencing.

[0063] Transient agrobacterium transformation into zygotic soy embryos: Zygotic soybeans were transiently transformed using agrobacterium. Specifically, a 5mL starter culture of agrobacterium, which contains the plasmid, was started from either a glycerol stock or bacterial colony and incubated for two days in a 28°C shaker (120rpm). Once cloudy, the 5mL starter culture was used to inoculate a larger 200mL overnight culture.

[0064] Bacteria preparation: Large plasmid agrobacterium cultures were spun down in a large centrifuge at 3400g for 10 minutes. The supernatant was then removed and the remaining agrobacterium pellet was resuspended in 30mL of LCCM. Post resuspension, the agrobacterium was centrifuged at 3400g for 6 minutes, and then subjected to one more round of supernatant removal, pellet resuspension and centrifugation. After the final spin down, the remaining supernatant was removed and the bacterial pellet was resuspended in lOmL of LCCM. An OD600 measurement was then taken and the agrobacterium solution was diluted to a final concentration of 1.4 OD600. The resulting agrobacterium solution was spiked with fresh acetosyringone (lOOpM final concentration) and subsequently incubated at room temperature while shaking (120rpm) for 1-2 hours.

[0065] Seed sterilization and preparation: During the agrobacterium incubation period, pods were picked from soy plants containing 8-10 mm embryos (~8 weeks old) and sterilized by the following method: 70% ethanol bath for 30 seconds, 10% bleach bath for 10 minutes, three sequential sterilized deionized water bath for 5 minutes each. After sterilization, seeds were aseptically removed from their pods, dissected from their seed coats, and split in half.

[0066] Explant inoculation: Once agrobacterium cultures finished incubating, silwet-77 (0.03% v / v) was added and mixed until dissolved. Then, 30 cotyledon halves (15 explants) were placed in the agrobacterium culture (which contained the vectors), and sonicated for 20 seconds at a 20% amplitude with 5s / l / s on / off pulse cycles. The contact of the cotyledon halves and the agrobacterium culture (which contained the vectors) allowed for the suppression of autophagy in the cotyledon. After sonication, explants were vacuum infiltrated for 5 minutes and left to incubate for two hours on a room temperature rotator.Docket No. 720.001.003.PCT

[0067] Plating explants: Explants were removed from bacterial culture and placed flat down (adaxial side down) on SCCM plates with a layer of sterile filter paper. Plates were then wrapped with micropore tape and incubated for 3 days in a dark 24°C chamber.

[0068] Washing explants: After plants incubated in the dark for 3 days, they were washed three times for five minutes each with sterile water containing Rif, Carb+Cef. Then, they were replated on SCCM plates containing filter paper and left to incubate in a dark 24°C chamber for another 7- 9 days.

[0069] Protein Crude Extraction: Roughly 40 plates containing the plasmids - transformed cotyledons were flash frozen in liquid nitrogen and crushed into a fine powder. Then, 100 grams of powder was measured out and mixed with a tris protein extraction buffer (50 mM Tris, 300 mM KC1, 0.5% Tween-20, 3.65% glycerol, Sigma plant protease inhibitor, pH 8.6). The resulting mixture rotated for 1 hour at 4°C and then spun down at 1300rpm for 30 minutes at 4°C.

[0070] Casein purification: Crude protein extract was first clarified using a ,45pm filter. Then, the sample was mixed with nickel resin, rotated at 4°C for 4 hours, and subsequently centrifuged for 2 minutes at 1000g. After centrifugation, the supernatant was removed and the remaining sample was washed with a wash buffer (50mM Tris base, 300mM KCL, 20mM imidazole, pH 7.4). To remove impurities, the wash step was repeated four times. After the last wash, the proteins were rotated in an elution buffer at 4°C for 15 minutes. Then, they were centrifuged at 700g for two minutes. The supernatant was saved for later analysis and the elution step was repeated another four times.

[0071] In preferred embodiments, the expressed bovine casein proteins include specific C- terminal modifications to facilitate cellular localization, identification, and purification. These modifications include a C-terminus fusion comprising a two-serine spacer, at least one affinity purification tag selected from a FLAG-tag or a 6-Histidine tag, and an HDEL target peptide sequence for endoplasmic reticulum (ER) retention. The presence of these tags and the expected molecular weight (approximately 27 kDa) are used throughout the examples for identification and quantification via Western blot analysis.

[0072] Casein characterization: To confirm successful purification of the caseins, the first three elution samples were run on a BioRad TGX AnykD Mini Protean SDS-PAGE gel and blotted (Western blot) with antibodies against FLAG to check for the presence of the caseins, whereby: “WT” is wildtype control. “Flag” is a positive control for flag tag. “Wl” is wash 1. “El” (elutionDocket No. 720.001.003.PCT1), “E2” (elution 2) and “E3” (elution 3) are serial elution. This blot was compared to the initial flow through and wash supernatants, which were expected to have minimal FLAG or V5 detection. Positive bands were enriched in the elution samples and present at the expected size of 27kDa. As all three caseins have a Flag or a V5 tag, mass spectrometry was used to confirm the presence of all three proteins.EXAMPLE 2: Enhanced Expression of asiCasein in Transgenic Soybean

[0073] Experimental Design and Quantitative Analysis. The effect of the inserted sequence on otsi-casein accumulation was evaluated across multiple genetic constructs. As shown in FIG. 1, eleven different transgenic soybean samples were generated and analyzed. The control sample (identified as 3548) contained an unmodified bovine asi-casein sequence. Ten additional samples were created, each containing the same asi-casein sequence but modified to include the inserted sequence of SEQ ID NO: 2 or SEQ ID NO: 4 at different positions within the asi-casein sequence.

[0074] Plasmids. The following plasmids, 3547, 3546, 3545, 3544, 3543, 3539, 3537, 3535, 3540, and 3536, each contained: (i) a nucleotide sequence for expressing asi-casein; and (ii) an intron inserted into the nucleotide sequence for expressing asi-casein. Plasmid 3548 served as a control and contained: (i) a nucleotide sequence for expressing asi-casein in, for example, soy plant cells; and (ii) no intron inserted into the nucleotide sequence for expressing asi-casein. The impact of the inserted introns on asi-casein expression levels was investigated. Protein fold change of asi-casein was calculated by: (i) first normalizing the protein expression against the total aSl-casein RNA found in each sample, (ii) measuring V5-tagged asi-casein levels, and (iii) comparing the V5-tagged asi-casein level to the normalized asi-casein RNA.

[0075] Introns. Each plasmid contained an intron derived from a specific gene: plasmid 3547 contained an intron derived from Glyma.16G077300 (SEQ ID NO. 4); plasmid 3546 contained an intron derived from Glyma.08G286700 (SEQ ID NO: 5); plasmid 3545 contained an intron derived from Glyma.12G074800 (SEQ ID NO: 8); plasmid 3544 contained an intron derived from Glyma.l5G163600 (SEQ ID NO: 6); plasmid 3543 contained an intron derived from Glyma.12G038100 (SEQ ID NO: 7); plasmid 3539 contained an intron derived from AT2G47600; plasmid 3537 contained an intron derived from AT5G17990; plasmid 3535 contained an intron derived from AT4G05320 (SEQ ID NO: 9); plasmid 3540 contained anDocket No. 720.001.003.PCT intron derived from Os07G066520; and plasmid 3536 contained an intron derived from Glyma.17G186600 (SEQ ID NO: 2).

[0076] Plasmid 3548 did not contain any of intron selected from SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9.

[0077] Results and Expression Analysis. Expression levels of asi-casein were measured using quantitative Western blot analysis. To generate the quantitative data presented in FIG. 1, densitometric analysis was performed on the Western blot shown in FIG. 2. The asi-casein band intensity for each construct was normalized to the amount of RNA produced in each sample. These normalized values were used to calculate relative expression levels, with the control construct (3548) set as the reference point. The results, summarized in FIG. 1, demonstrate that the position of the inserted sequence significantly affects asi-casein accumulation. Notably, the plasmid 3536, which comprises SEQ ID NO: 1 (i.e., a nucleotide where SEQ ID NO: 2 is inserted into nucleotide for expressing asi-casein), showed asi-casein accumulation levels that were more than three-fold higher than those observed in the control line (3548); and plasmid 3535, which comprises SEQ ID NO: 3 (i.e., a nucleotide where SEQ ID NO: 9 is inserted into nucleotide for expressing asi- casein). This experiment was replicated multiple times with consistent results across each replicate.

[0078] The introns corresponding to SEQ ID NO: 8 (derived from Glyma.12G074800 in plasmid 3545), SEQ ID NO: 9 (derived from AT4G05320 in plasmid 3535), SEQ ID NO: 2 (derived from Glyma.17G186600 in plasmid 3536), and SEQ ID NO: 4 (derived from Glyma.16G077300 in plasmid 3547), when inserted into a nucleotide sequence for expressing asi-casein, exhibited increased expression of asi-casein. Specifically, they resulted in a 2.42-fold, 1.92-fold, 3.07-fold, and 2.06-fold increase, respectively, in comparison to plasmid 3548.

[0079] However, other introns, when inserted into a nucleotide sequence for expressing asi- casein, did not lead to any expression of asi-casein. These included introns derived from Os07G066520 (in plasmid 3540), AT5G17990 (in plasmid 3537), AT2G47600 (in plasmid 3539), Glyma.08G286700 (SEQ ID NO: 5 in plasmid 3546), and Glyma.l5G163600 (SEQ ID NO: 6 in plasmid 3544).

[0080] Western Blot Confirmation 1. The enhanced accumulation of asi-casein was further confirmed by Western blot analysis, as shown in FIG. 2. Western blot analysis was performed on protein extracts from transgenic soybean plants using antibodies specific to bovine asi-casein. FIG.Docket No. 720.001.003.PCT2 shows a representative immunoblot comparing protein extracts from multiple independent lines. Lane 3 contains extracts from control plants (3548) expressing unmodified asi-casein, while lanes 2-11 contain extracts from plants transformed with constructs containing the inserted sequence at various positions. The immunoblot demonstrates a clear band corresponding to asi-casein at the expected molecular weight. Notably, lane 12, corresponding to construct 3536 (SEQ ID NO: 1), shows a significantly more intense band compared to the control lane, providing visual confirmation of the enhanced accumulation levels described in FIG. 1. Lanes 1-11 corresponded to: ladder, V5+, plasmid 3548, plasmid 3547, plasmid 3546, plasmid 3545, plasmid 3544, plasmid 3543, plasmid 3540, plasmid 3539, plasmid 3537, plasmid 3536, and plasmid 3535, respectively.

[0081] Western Blot Confirmation 2. The enhanced accumulation of asi-casein was yet further confirmed by another Western blot analysis, as shown in FIG. 2. Western blot analysis was performed on protein extracts from transgenic soybean plants using antibodies specific to bovine asi-casein. FIG. 2 shows a representative immunoblot comparing protein extracts from multiple independent lines. Lane 3 contains extracts from control plants (3548) expressing unmodified asi-casein, while lanes 2-11 contain extracts from plants transformed with constructs containing the inserted sequence at various positions. The immunoblot demonstrates a clear band corresponding to asi-casein at the expected molecular weight. Notably, lane 12, corresponding to construct 3536 (SEQ ID NO: 1), exhibited a significantly more intense band compared to the control lane (construct 3548), providing visual confirmation of the enhanced accumulation levels described in FIG. 2. Additionally, lanes 4, 6, and 13, corresponding to constructs 3547, 3545, and 3535, respectively, exhibited significantly more intense bands compared to the control lane (construct 3548), providing visual confirmation of the enhanced accumulation levels described in Figure 1. Lanes 1-13 corresponded to: ladder, V5+, plasmid 3548, plasmid 3547, plasmid 3546, plasmid 3545, plasmid 3544, plasmid 3543, plasmid 3540, plasmid 3539, plasmid 3537, plasmid 3536, and plasmid 3535, respectively.EXAMPLE 3: Enhanced Expression of K-Casein in Transgenic Soybean

[0082] Experimental Design and Quantitative Analysis. The effect of inserted intron sequences on K-casein accumulation was evaluated across multiple genetic constructs. Nine transgenic soybean samples were generated and analyzed, as shown in FIG. 3. The control sample (identified as 3548) contained an unmodified bovine K-casein sequence. Additional samples (constructs 3859,Docket No. 720.001.003.PCT3860, 3861, 3862, 3863, 3864, 3865, and 3866) were created containing the same K-casein nucleotide sequence, but modified to include various intron sequences at different positions within the K-casein nucleotide sequence. Construct 3859 contained the Glyma.l7Gl 86600 intron (SEQ ID NO: 2) inserted at an optimized position for K-casein, which increased the expression of K- casein. Glyma.l7G186600 intron (SEQ ID NO: 2) was also inserted at a non-optimized position for K-casein (distal from the 5 ’-end of a nucleotide encoding for K-casein). At an optimized position SEQ ID NO: 2, there was an appreciable fold increase in K-casein, that was not observed at an nonoptimized position of insertion for K-casein.

[0083] Plasmids: The following plasmids, 3866, 3865, 3864, 3863, 3862, 3861, 3860, 3859, 3858, and 3540, each contained: (i) a nucleotide sequence for expressing K-casein; and (ii) an intron inserted into the nucleotide sequence for expressing K-casein. Plasmid 3548 served as a control and contained: (i) a nucleotide sequence for expressing K-casein in, for example, soy plant cells; and (ii) no intron inserted into the nucleotide sequence for expressing K-casein. The impact of the inserted introns on K-casein expression levels was investigated. Protein fold change of K-casein was calculated by: (i) first normalizing the protein expression against the total K-casein RNA found in each sample, (ii) measuring FLAG-tagged K-casein levels, and (iii) comparing the FEAG-tagged K-casein level to the normalized K-casein RNA.

[0084] Introns: Each plasmid contained an intron derived from a specific gene: Plasmid 3864 contained an intron derived from Glyma.l 6G077300 (SEQ ID NO: 4); Plasmid 3866 contained an intron derived from Glyma.12G074800 (SEQ ID NO: 8); Plasmid 3865 contained an intron derived from Glyma. 15G163600 (SEQ ID NO: 6); Plasmid 3863 contained an intron derived from Glyma.12G038100 (SEQ ID NO: 7); Plasmid 3861 contained an intron derived from AT2G47600; Plasmid 3860 contained an intron derived from AT5G17990; Plasmid 3858 contained an intron derived from AT4G05320 (SEQ ID NO: 9); Plasmid 3862 contained an intron derived from Os07G0665200; Plasmid 3859 contained an intron derived from Glyma.17G186600 (SEQ ID NO: 2); and Plasmid 3548 did not contain any intron selected from SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9.

[0085] Results and Expression Analysis. Expression levels of K-casein were measured using quantitative Western blot analysis of the V5-tagged protein, with results normalized to K- casein RNA levels for each sample. As shown in FIG. 3, these normalized values demonstrate that theDocket No. 720.001.003.PCT position and type of inserted sequence significantly affect K-casein accumulation. Notably, construct 3859 showed approximately 1.15-fold higher normalized protein accumulation compared to the control line (3548), while other constructs (3860-3866) showed reduced expression. Only constructs 3860 and 3864 showed measurable expression at approximately 0.05- fold, while constructs 3866, 3863, 3862, 3861, and 3858 showed negligible protein accumulation relative to their RNA levels.

[0086] The introns corresponding to SEQ ID NO: 2 (derived from Glyma.l7Gl 86600 in plasmid 3859), when inserted into an optimized position within the nucleotide sequence for expressing K- exhibited increased expression of K-casein. Specifically, the intron corresponding to SEQ ID NO: 2 resulted in a 1.13 -fold increase, in comparison to plasmid 3548. There was expression of kappacasein when introns corresponding to SEQ ID NO: 4 (derived from Glyma.16G077300 in plasmid 3864) and AT5G17990 (in plasmid 3860) was inserted into a nucleotide for expressing K-casein.

[0087] However, other introns, when inserted into a nucleotide sequence for expressing K-casein, did not lead to any expression of K-casein. These included introns derived from Os07G0665200 (in plasmid 3862), AT2G47600 (in plasmid 3861), Glyma.15G163600 (in plasmid 3865), or introns in plasmids 3866, 3863, and 3858.

[0088] Western Blot Confirmation. The enhanced accumulation of K-casein in construct 3859 was confirmed by Western blot analysis using antibodies specific to the V5 tag. The quantitative data presented in FIG. 3 was derived from densitometric analysis of these Western blots, with band intensities normalized to K-casein RNA levels as measured by RT-PCR. This normalization accounts for any variations in transcription levels between constructs, allowing direct comparison of protein accumulation efficiency. The results demonstrate that the Glyma.17G186600 intron sequence, when properly positioned in the optimal position, can enhance protein accumulation compared to the unmodified control sequence. Lanes 1-11 in FIG. 4 corresponded to: ladder, FLAG, 3548, 3859, 3860, 3861, 3862, 3863, 3864, 3865, and 3866, respectively.EXAMPLE 4: Enhanced Expression of P-Casein in Transgenic Soybean

[0089] Experimental Design and Quantitative Analysis. The effect of various intron sequences on P-casein RNA expression was evaluated across multiple genetic constructs. Six transgenic soybean samples were generated and analyzed. The control sample (identified as 3868) containedDocket No. 720.001.003.PCT an unmodified bovine P-casein sequence. Five additional samples were created, each containing the same P-casein nucleotide sequence modified to include different intron sequences selected from AT4G05320 intron (construct 3849, SEQ ID NO: 9), Glyma.l7G186600 intron (construct 3850, SEQ ID NO: 2), AT5G17990 intron (construct 3851), GLYMA.12G038100 intron (construct 3854, SEQ ID NO: 7), GLYMA.16G077300 intron (construct 3855, SEQ ID NO: 4), and GLYMA.12G074800 intron (construct 3857, SEQ ID NO: 8), wherein said introns are each inserted in the same position in the P-casein nucleotide sequence.

[0090] Plasmids: The following plasmids, 3857, 3856, 3855, 3854, 3853, 3852, 3852, 3851, 3850, and 3849, each contained: (i) a nucleotide sequence for expressing P-casein; and (ii) an intron inserted into the nucleotide sequence for expressing P-casein. Plasmid 3868 served as a control and contained: (i) a nucleotide sequence for expressing P-casein in, for example, soy plant cells; and (ii) no intron inserted into the nucleotide sequence for expressing P-casein. The impact of the inserted introns on P-casein expression levels was investigated. Protein fold change of P-casein was calculated by: (i) first normalizing the protein expression against the total P- casein RNA found in each sample, (ii) measuring V5-tagged P-casein levels, and (iii) comparing the V5-tagged P-casein level to the normalized P-casein RNA.

[0091] Introns: Each plasmid contained an intron derived from a specific gene: Plasmid 3855 contained an intron derived from Glyma.16G077300 (SEQ ID NO: 4); Plasmid 3857 contained an intron derived from Glyma.12G074800 (SEQ ID NO: 8); Plasmid 3856 contained an intron derived from Glyma.15G163600 (SEQ ID NO: 6); Plasmid 3854 contained an intron derived from Glyma.l2G038100 (SEQ ID NO: 7); Plasmid 3852 contained an intron derived from AT2G47600; Plasmid 3851 contained an intron derived from AT5G17990; Plasmid 3849 contained an intron derived from AT4G05320 (SEQ ID NO: 9); Plasmid 3853 contained an intron derived from Os07G0665200; Plasmid 3850 contained an intron derived fromGlyma.17G186600 (SEQ ID NO: 2); and Plasmid 3868 did not contain any intron selected from SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9.

[0092] Results and Expression Analysis. Expression levels were analyzed at both RNA and protein levels. RNA expression was quantified using qPCR analysis, with results normalized to the control construct (3868). All intron-containing constructs showed enhanced RNA expression compared to the control, with construct 3850 (Glyma.17G186600 intron) showing the highestDocket No. 720.001.003.PCT expression levels, followed by construct 3849 (AT4G05320 intron) and construct 3857 (GLYMA.12G074800 intron). Constructs 3851 and 3854 showed more modest increases of 1.32- fold and 1.55-fold, respectively.

[0093] In another trial is depicted as FIG. 5. The introns corresponding to SEQ ID NO: 8 (derived from Glyma.12G074800 in plasmid 3545), SEQ ID NO: 9 (derived from AT4G05320 in plasmid 3849), SEQ ID NO: 2 (derived from Glyma.l7G186600 in plasmid 3536), and SEQ ID NO: 4 (derived from Glyma.16G077300 in plasmid 3547), when inserted into a nucleotide sequence for expressing P-casein, exhibited increased expression of P-casein. Specifically, introns inserted into nucleotides expressing casein in plasmids 3849, 3850, 3851, 3854, 3855, and 3857 resulted in a 2.98-fold, 2.82-fold, 0.92-fold, 0.77-fold, 1.26 fold, and 0.85 changes, respectively, in comparison to plasmid 3868.

[0094] However, other introns, when inserted into a nucleotide sequence for expressing P-casein, did not lead to any expression of P-casein. These included introns derived from Os07G066520 (in plasmid 3853), AT2G47600 (in plasmid 3852), and Glyma.15G163600 in plasmid 3856

[0095] Protein expression was measured using quantitative Western blot analysis and normalized to P-casein RNA levels for each sample. Three constructs showed particularly notable results: construct 3850 (Glyma.17G186600 intron), construct 3849 (AT4G05320 intron), and construct 3855 (GLYMA.16G077300 intron) all demonstrated increased protein accumulation compared to the control construct 3868. The enhanced protein accumulation, combined with increased RNA expression, suggests that these introns improve both transcription and translation efficiency of the P-casein gene.

[0096] Western Blot Confirmation. The enhanced accumulation of P-casein was verified through Western blot analysis performed on protein extracts from transgenic soybean plants using antibodies specific to bovine P-casein. The immunoblot analysis confirmed the quantitative measurements, which exhibited more intense bands for constructs containing the Glyma.17G186600 intron (construct 3850), AT4G05320 intron (construct 3849), and GLYMA. 16G077300 intron (construct 3855) compared to the control (construct 3868). These results demonstrate that specific intron sequences can significantly enhance both RNA expression and protein accumulation of P-casein in transgenic soybean plants. The lanes in FIG. 6 corresponded from left to right: the ladder, V5-tag, plasmid 3849, plasmid 3850, plasmid 3851,Docket No. 720.001.003.PCT plasmid 3852, plasmid 3853, plasmid 3854, plasmid 3855, plasmid 3856, plasmid 3857, and plasmid 3868.EXAMPLE 5: Synergistic Enhancement of Casein Expression Using GC Optimization and Intron Regulatory Sequences

[0097] This example demonstrates the synergistic effect of combining GC content optimization with the insertion of 5' proximal intron regulatory sequences on the expression and accumulation of bovine casein proteins in transgenic soybean plants. The construct utilized, pMOZ4334, was engineered to express multiple milk proteins (asi-casein, P-casein, and K-casein) and the essential kinase FAM20C, optimized for maximal expression. See FIG. 7 for a plasmid diagram for the construct, pMOZ4334.

[0098] Construct Design and Dual Optimization. In the pMOZ4334 construct, the coding sequences for asi, , and K-caseins were first subjected to high GC content optimization, as described above, to enhance translational efficiency. Second, these GC-optimized sequences were modified by the insertion of specific regulatory sequences (introns), positioned 5' proximally within the coding sequence to maximize Intron-Mediated Enhancement (IME). The cisi-casein coding sequence incorporated the regulatory sequence (intron) of SEQ ID NO: 8; the P-casein coding sequence incorporated the regulatory sequence (intron) of SEQ ID NO: 9 (derived from AT4G05320); and the K-casein sequence incorporated the regulatory sequence (intron) of SEQ ID NO: 2 (derived from Glyma.17G186600). These dual modifications — codon usage optimization for translation, and regulatory sequence insertion for transcription and RNA stability — are designed to act in concert to achieve protein accumulation levels significantly higher than those achieved by expression constructs lacking these features.

[0099] For example, the asi-casein coding sequence incorporated a regulatory sequence derived from GLYMA.12G074800 (SEQ ID NO: 8); the P-casein coding sequence incorporated the regulatory sequence derived from AT4G05320 (SEQ ID NO: 9); and the K-casein sequence incorporated the regulatory sequence derived from Glyma.17G186600 (SEQ ID NO: 2). These dual modifications — codon usage optimization for translation, and regulatory sequence insertion for transcription and RNA stability — are designed to act in concert to achieve protein accumulation levels significantly higher than those achieved by expression constructs lacking these features.Docket No. 720.001.003.PCT

[0100] Transformation and Analysis. The plasmid pMOZ4334 (SEQ ID NO: 11) was transformed into soybean cells using the methods described in Example 1. Expression levels were quantified using Western blot analysis and compared against Control Construct 2 (pMOZ2023), a highly relevant baseline construct containing the same casein proteins but lacking both the GC content optimization and the regulatory sequence insertions (see FIG. 8).

[0101] Quantitative Results and Synergistic Enhancement. Protein accumulation analysis of the V5-tagged asi-casein demonstrated superior performance for the dual -optimized pMOZ4334 construct compared to the non-optimized baseline control pMOZ2023. Control Construct 2 (pMOZ2023), which lacked both the GC optimization and the intron regulatory sequence insertion, showed an accumulation of 0.1456% of total soluble protein (TSP). In contrast, the two independent transgenic lines for pMOZ4334 showed accumulation levels significantly higher, averaging 0.1623% of TSP. This result demonstrates an average increase in asi-casein accumulation of approximately 11.5% in the pMOZ4334 construct compared to the control pMOZ2023. This enhancement, achieved through the synergistic application of 5' proximal intron insertion and GC codon optimization, was determined to be statistically significant, confirming that the dual optimization strategy achieves superior protein accumulation.

[0102] The enhanced accumulation of asi-casein from the synergistic construct pMOZ4334 was verified using quantitative Western blot analysis, as shown in the accompanying Figure 9. This immunoblot specifically targets the V5 epitope tag, which is translationally fused to the C- terminus of the asi-casein protein.

[0103] The Western blot (see FIG. 9) compares protein extracts harvested from independent transgenic soybean lines: pMOZ4334 events (lanes 1 and 2), and the pMOZ2023 control (lane 3). Lanes 4 through 8 represent a quantitative standard curve using purified V5-tagged protein (10 ng, 20 ng, 30 ng, 45 ng, and 50 ng, respectively) to allow for protein quantification.

[0104] The intense bands corresponding to asi-casein visible in the pMOZ4334 lanes (1 and 2) are visibly more intense than the band corresponding to the pMOZ2023 control (lane 3). Densitometric comparison and subsequent quantification confirmed that the 0.1623% average accumulation achieved by pMOZ4334 is significantly superior to the accumulation observed in the relevant control construct pMOZ2023. The minimal accumulation observed in Control Construct 1 (pMOZ4333, data not shown) further demonstrates that the specific geneticDocket No. 720.001.003.PCT modifications claimed by the present invention are necessary to achieve high protein accumulation levels in soybean.

Claims

Docket No. 720.001.003.PCTCLAIMSWHAT IS CLAIMED IS:

1. A genetic construct comprising a nucleic acid sequence encoding a bovine casein protein, wherein said sequence comprises an inserted regulatory sequence selected from SEQ ID NO: 2 and SEQ ID NOs: 4-9.

2. The genetic construct of claim 1, wherein the bovine casein protein is selected from the group consisting of asi-casein„ P-casein, and K-casein.

3. The genetic construct of claim 1, further comprising: (a) a plant-compatible promoter; and (b) a terminator sequence.

4. The genetic construct of claim 3, wherein the plant-compatible promoter is a soybeancompatible promoter.

5. The genetic construct of claim 1, further comprising SEQ ID NO: 1, wherein SEQ ID NO: 1 is a combination of SEQ ID NO: 2 inserted into a nucleotide for expressing asi- casein.

6. The genetic construct of claim 1, further comprising SEQ ID NO: 3, wherein SEQ ID NO: 3 is a combination of SEQ ID NO: 9 inserted into a nucleotide for expressing asi- casein.

7. A method of producing a bovine casein protein in a plant comprising: (a) introducing into a plant cell a genetic construct comprising a nucleic acid sequence encoding a bovine casein protein, wherein said sequence comprises an inserted regulatory sequence comprising SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO:Docket No. 720.001.003.PCT7, SEQ ID NO: 8, or SEQ ID NO: 9; and (b) expressing said bovine casein protein in said plant cell.

8. The method of claim 6, wherein the plant cell is a soybean cell.

9. A method of enhancing bovine casein protein accumulation in plants comprising: (a) inserting SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9 into a nucleic acid sequence encoding a bovine casein protein; and (b) expressing the resulting construct in a plant cell.

10. The method of claim 8, wherein the bovine casein protein shows increased accumulation compared to expression of the same bovine casein protein without the inserted regulatory sequence selected from SEQ ID NO: 2 and SEQ ID NOs: 4-9.

11. A nucleic acid construct comprising: a dsi-casein coding sequence incorporating a regulatory sequence derived from SEQ ID NO: 8; a P-casein coding sequence incorporated the regulatory sequence derived from SEQ ID NO: 9; and a K-casein sequence incorporated the regulatory sequence derived from SEQ ID NO: 2.

12. A genetic construct comprising: one or more sequences encoding for a recombinant casein protein; and a regulatory sequence selected from the group consisting of SEQ ID NO: 2, 4, 5, 6, 7, 8, and 9, wherein the regulatory sequence is inserted in a region at a 5’ region of the one or more sequences encoding for the recombinant casein protein.

13. A genetic construct comprising a nucleic acid sequence for encoding for asi-casein, wherein said sequence further comprises SEQ ID NO: 2.

14. A genetic construct comprising a nucleic acid sequence for encoding for P-casein, wherein said sequence further comprises SEQ ID NO: 2.Docket No. 720.001.003.PCT15. A genetic construct comprising a nucleic acid sequence for encoding for K-casein, wherein said sequence further comprises SEQ ID NO: 2.

16. A genetic construct comprising:(a) a GC-optimized asi-casein coding sequence incorporating a regulatory sequence derived from SEQ ID NO: 8;(b) a GC-optimized P-casein coding sequence incorporating a regulatory sequence derived from SEQ ID NO: 9; and(c) a GC-optimized K-casein sequence coding sequence incorporating a regulatory sequence derived from SEQ ID NO: 2.