Compositions and methods for enhancing protein production in Bacillus cells

By introducing a variant ilvE gene with an SNP in the 5'-UTR, recombinant Bacillus strains enhance protein production, addressing yield challenges in industrial applications.

JP2025535419APending Publication Date: 2025-10-24DANISCO US INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522831
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-24
Filing Date
2023-10-13
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing Bacillus strains used for industrial protein production face challenges in achieving predictable and enhanced yields of heterologous proteins, with optimization methods being inadequate for large-scale industrial applications.

Method used

Introduction of a variant ilvE gene with a single nucleotide polymorphism (SNP) in the 5'-untranslated region (5'-UTR) of the ilvE gene, enhancing ilvE messenger RNA (mRNA) levels and carbon yield efficiency in recombinant Bacillus cells, leading to increased protein production.

Benefits of technology

The modified Bacillus strains exhibit improved protein production capabilities, achieving higher yields of proteins of interest compared to wild-type strains, making them suitable for industrial-scale protein production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535419000001_ABST
    Figure 2025535419000001_ABST
Patent Text Reader

Abstract

Certain embodiments of the present disclosure relate to, inter alia, mutated and / or modified Bacillus cells (strains) comprising an enhanced protein production phenotype, mutated and / or modified cells comprising enhanced / increased ilvE messenger RNA (mRNA) levels, mutated and / or modified cells comprising enhanced / increased carbon yield of produced heterologous protein (carbon yield efficiency), etc. As outlined herein, certain embodiments of the present disclosure relate to mutated and / or modified Bacillus strains comprising a variant ilvE gene, which mutated and / or modified Bacillus strains are particularly useful for enhanced production of a protein of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to the fields of bacteriology, microbiology, genetics, molecular biology, enzymology, industrial protein production, etc. Certain embodiments of the present disclosure relate to Bacillus sp. strains comprising an enhanced protein production phenotype, compositions and methods for constructing such recombinant Bacillus sp. strains, etc.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 380,706, filed October 24, 2022, which is incorporated herein by reference in its entirety.

[0003] Sequence Listing Reference The contents of the electronic submission of the sequence listing text file entitled "NB41976-WO-PCT_SequenceListing.xml" was created on October 2, 2023, is 37KB in size, and is incorporated herein by reference in its entirety. [Background technology]

[0004] Gram-positive bacteria, such as Bacillus subtilis, Bacillus licheniformis, and Bacillus amyloliquefaciens, are frequently used as microbial factories to produce industrially relevant proteins due to their excellent fermentation properties and high yields (e.g., up to 25 grams per liter of culture; Van Dijl and Hecker, 2013). For example, Bacillus sp. host cells are well known for their production of enzymes (e.g., amylase, cellulase, mannanase, pectate lyase, protease, pullulanase, etc.) required for food, textile, laundry applications, medical device cleaning, and pharmaceutical industries. These non-pathogenic Gram-positive bacteria produce proteins that are completely free of harmful by-products (e.g., lipopolysaccharides (LPS), also known as endotoxins), and have been awarded a "Qualified Presumption of Safety" (QPS) rating by the European Food Safety Authority (EFSA), and many of their products have been awarded a "Generally Recognized As Safe" (GRAS) rating by the US Food and Drug Administration (Olempska-Beer et al., 2006; Earl et al., 2008; Caspers et al., 2010).

[0005] Thus, the production of proteins (e.g., enzymes, antibodies, receptors, etc.) by microbial host cells is of particular interest in the field of biotechnology. Similarly, the optimization of Bacillus host cells to produce and secrete one or more proteins of interest is particularly relevant in industrial biotechnology settings, where even small improvements in protein yield can be crucial for industrially large-scale protein production. For example, the expression of many heterologous proteins can remain challenging and unpredictable, such as in terms of yield. As described below, the present disclosure relates to a highly desirable and unmet need for obtaining and constructing Bacillus sp. cells (e.g., protein-producing hosts) with enhanced protein production capabilities. Summary of the Invention [Means for solving the problem]

[0006] As generally described herein, certain embodiments of the present disclosure relate to, inter alia, variant ilvE genes, variant ilvE gene 5′-untranslated region (5′-UTR) sequences, mutant Bacillus strains comprising variant ilvE gene sequences, recombinant (genetically modified) Bacillus strains comprising variant ilvE gene sequences, mutant and / or recombinant Bacillus strains comprising variant ilvE gene sequences and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains comprising variant ilvE gene sequences, expression cassettes encoding proteins of interest, methods and compositions for culturing recombinant Bacillus strains comprising variant ilvE gene sequences to enhance production of proteins of interest, etc. More specifically, as described below, the novel mutant and / or recombinant Bacillus cells of the present disclosure are particularly useful for producing proteins of interest when cultured under suitable conditions.

[0007] Accordingly, certain embodiments of the present disclosure relate to a variant ilvE gene comprising a single nucleotide polymorphism (SNP) mutation in the 5'-untranslated region (5'-UTR) of the ilvE gene. In one or more embodiments, the variant ilvE gene of the present disclosure encodes a functional IlvE protein. In certain other embodiments, the present disclosure provides a synthetic ilvE gene construct comprising a heterologous promoter sequence operably linked, in a 5' to 3' direction, to a variant ilvE 5'-UTR sequence operably linked to an ilvE gene coding sequence (CDS) encoding a functional IlvE protein. In other embodiments, the present disclosure relates to a mutant Bacillus subtilis strain comprising a variant ilvE gene with a SNP in the 5'-untranslated region (5'-UTR) of the ilvE gene. In one or more other embodiments, the mutant and / or recombinant B. subtilis cells of the present disclosure produce one or more proteins of interest. In other related embodiments, mutant and / or recombinant B. subtilis cells producing one or more proteins of interest comprise an enhanced carbon yield phenotype (when the mutant and control cells are fermented under suitable conditions) compared to a control B. subtilis producing the same one or more proteins of interest and containing a wild-type ilvE gene. In one or more specific related embodiments, the present disclosure provides mutant and / or engineered Bacillus cells (strains) comprising an enhanced protein productivity phenotype, mutant and / or engineered cells comprising enhanced / increased ilvE messenger RNA (mRNA) levels, mutant and / or engineered cells comprising enhanced / increased carbon yield (carbon yield efficiency) of produced heterologous protein, and the like.

[0008] In other embodiments, the disclosure relates to methods for increasing ilvE messenger RNA (mRNA) levels in recombinant B. subtilis cells, generally comprising obtaining a parent B. subtilis cell having a wild-type (WT) ilvE gene, replacing the WT ilvE gene with a variant ilvE gene, wherein the variant ilvE gene comprises a SNP in the 5'-untranslated region (5'-UTR) of the ilvE gene, and fermenting the parent and recombinant cells under suitable conditions for at least about 16 hours, wherein the recombinant cells comprise increased levels of ilvE mRNA compared to the parent cell. In other embodiments, the disclosure relates to a method for increasing the carbon yield of a heterologous protein produced in a recombinant B. subtilis cell, comprising obtaining or constructing a parent B. subtilis cell that produces a heterologous protein of interest (POI) and contains a WT ilvE gene; replacing the WT ilvE gene with a variant ilvE gene that contains a SNP mutation in the 5'-untranslated region (5'-UTR); and fermenting the parent and recombinant cells for at least about 16 hours under conditions suitable for production of the POI, wherein the recombinant cells comprise an increased carbon yield efficiency of the POI produced compared to the parent cell. [Brief explanation of the drawings]

[0009] [Figure 1]The DNA sequences of the wild-type (WT) B. subtilis ilvE 5'-UTR (FIG. 1A) and mutant B. subtilis ilvE 5'-UTR (FIG. 1B) are shown. For example, as shown in FIG. 1A, the WT ilvE 5'-UTR sequence contains a cytosine (C) at nucleotide position 73, and as shown in FIG. 1B, the mutant ilvE 5'-UTR contains a thymine (T) at nucleotide position 73 (where the C and T nucleotides at that position are shown in bold and double-underlined in FIG. 1). Thus, as shown in FIG. 1, the WT B. subtilis ilvE 5'-UTR sequence contains SEQ ID NO: 17 (FIG. 1A), and the mutant ilvE 5'-UTR sequence contains SEQ ID NO: 18 (FIG. 1B). [Figure 2] 1 shows a schematic map depicting nucleotide position 73 (+73) of the mutant ilvE 5'-UTR sequence (SEQ ID NO: 18). In particular, the WT (SEQ ID NO: 17) and mutant (SEQ ID NO: 18) ilvE 5'-UTR sequences shown in FIG. 1 are numbered from 5' to 3', with nucleotide position 1 (+1) being the position of the first nucleotide of the 5' untranslated region, identified as the transcription start site. In particular, as shown in FIG. 2, the C>T mutation at position 73 in the mutant ilvE 5'-UTR is located near CodY-binding motif 2 and putative binding motif 5. [Figure 3] The nucleic acid sequences of the wild-type (WT) ilvE promoter (FIG. 3A; SEQ ID NO: 16), the WT ilvE 5'-UTR (FIG. 3B; SEQ ID NO: 17), the WT ilvE gene coding sequence (FIG. 3C; SEQ ID NO: 14), and the WT ilvE gene (FIG. 3D; SEQ ID NO: 25) are shown. As shown in FIG. 3D, the WT ilvE gene (SEQ ID NO: 25) includes, from 5' to 3', the WT ilvE promoter (nucleotides in italics; SEQ ID NO: 16), the WT ilvE 5'-UTR (nucleotides in bold; SEQ ID NO: 17), and the WT ilvE gene CDS (nucleotides underlined; SEQ ID NO: 14). Similarly, FIG. 3 shows the amino acid sequence of the native (mature) IlvE protein (FIG. 3E; SEQ ID NO: 15) encoded by the WT ilvE gene CDS (SEQ ID NO: 14). [Figure 4]Figure 4 shows data from real-time qPCR (RT qPCR) analysis of Bacillus strain A (containing the mutant ilvE 5'-UTR) relative to Bacillus strain B (containing the WT ilvE 5'-UTR), as described in Example 2 below. Specifically, the data shown in Figure 4 show the results of RT qPCR in a time-course experiment, in which the 2-ΔΔC method (Livak and Schmittgen, 2001) was used to calculate the log fold change between the housekeeping ftsY gene and the ilvE gene. As shown in Figure 4, the bars and values ​​represent the fold change of ilvE mRNA in B. subtilis strain A (containing the mutant ilvE 5'-UTR) relative to the isogenic B. subtilis strain B (containing the WT ilvE 5'-UTR) at 16, 24, and 32 hours of fermentation. DETAILED DESCRIPTION OF THE INVENTION

[0010] A brief description of biological sequences SEQ ID NO: 1 is a nucleotide (DNA) sequence containing the aprE 5'-UTR sequence of wild-type B. subtilis. SEQ ID NO:2 is the wild-type DNA sequence encoding the native B. subtilis aprE signal sequence. SEQ ID NO:3 is the amino acid sequence of the native B. subtilis aprE signal sequence encoded by SEQ ID NO:2. SEQ ID NO: 4 is the DNA sequence encoding the B. clausii GG36 Pro region sequence. SEQ ID NO:5 is the amino acid sequence of the native B. clausii GG36 Pro region sequence encoded by SEQ ID NO:4. SEQ ID NO: 6 is the wild-type DNA sequence encoding the native B. clausii protease (Eraser11). SEQ ID NO:7 is the amino acid sequence of the native B. clausii protease (Eraser11) encoded by SEQ ID NO:6. SEQ ID NO: 8 is a DNA sequence containing the B. amyloliquefaciens BPN' terminator sequence. SEQ ID NO: 9 is a DNA sequence containing the B. subtilis 5'skfA flanking region (FR) sequence. SEQ ID NO: 10 is a DNA sequence containing the B. subtilis 3'skfA FR sequence. SEQ ID NO: 11 is a DNA sequence containing the B. subtilis 5'aprE FR sequence. SEQ ID NO: 12 is the DNA sequence containing the wild-type B. subtilis alrA gene. SEQ ID NO: 13 is a DNA sequence containing the B. subtilis 3'aprE FR sequence. SEQ ID NO: 14 is the DNA sequence containing the wild-type B. subtilis IlvE gene coding sequence (CDS). SEQ ID NO:15 is the amino acid sequence of the native B. subtilis IlvE protein encoded by SEQ ID NO:14. SEQ ID NO: 16 is the DNA sequence containing the IlvE promoter of wild-type B. subtilis. SEQ ID NO: 17 is a DNA sequence containing the IlvE 5'-UTR sequence of wild-type B. subtilis. SEQ ID NO: 18 is a DNA sequence containing a mutated B. subtilis IlvE 5'-UTR sequence. SEQ ID NO: 19 is the B. subtilis IlvE-forward (FW) primer (DNA) sequence. SEQ ID NO: 20 is the B. subtilis IlvE-reverse (RV) primer (DNA) sequence. SEQ ID NO: 21 is a synthetic DNA probe designated "IlvE-BBQ." SEQ ID NO: 22 is the B. subtilis ftsY-forward (FW) primer (DNA) sequence. SEQ ID NO: 23 is the B. subtilis ftsY-reverse (RV) primer (DNA) sequence DNA. SEQ ID NO: 24 is a synthetic DNA probe designated "ftsY-BBQ." SEQ ID NO: 25 is a DNA sequence containing the wild-type (WT) B. subtilis IlvE gene comprising (in the 5' to 3' direction) the WTilvE promoter (SEQ ID NO: 16) operably linked to the WTilvE 5'-UTR (SEQ ID NO: 17), which is operably linked to the WTilvE gene CDS (SEQ ID NO: 14). SEQ ID NO:26 is the DNA sequence of the native B. subtilis Hbs promoter region sequence. SEQ ID NO: 27 is a synthetic DNA construct comprising the upstream (5') Hbs promoter operably linked to the wild-type IlvE gene. SEQ ID NO: 28 is a synthetic DNA construct comprising the upstream (5') Hbs promoter operably linked to a variant IlvE gene.

[0011] As described herein, certain embodiments of the present disclosure relate to compositions and methods for enhancing protein production in mutated / recombinant Bacillus sp. (host) cells / strains. More specifically, as described below and further illustrated in the Examples below, the recombinant Bacillus sp. cells of the present disclosure are particularly useful for enhancing the production of proteins of interest when cultured under suitable conditions. Accordingly, certain embodiments of the present disclosure relate to, inter alia, mutated Bacillus strains comprising variant ilvE gene sequences, recombinant (genetically modified) Bacillus strains comprising variant ilvE gene sequences, mutated and / or recombinant Bacillus strains comprising variant ilvE gene sequences and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains comprising variant ilvE gene sequences, expression cassettes encoding proteins of interest, methods and compositions for culturing recombinant Bacillus strains comprising variant ilvE gene sequences to enhance production of proteins of interest, etc.

[0012] I. Definition In view of the disclosed recombinant (modified) cells and the methods thereof described herein, the following terms and phrases are defined. Terms not defined herein shall be accorded the meaning commonly used in the art.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the compositions and methods of the present invention belong. Although any methods and materials similar or equivalent to those described herein can be used to practice or test the compositions and methods of the present invention, representative methods and materials are described below. All publications and patents cited herein are incorporated herein by reference in their entirety.

[0014] It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a prelude to the use of exclusive terminology such as "solely," "only," "excluding," "not including," or the use of any "negative" limitation or qualification in connection with the recitation of claim elements.

[0015] It will be apparent to those skilled in the art upon reading this disclosure that each of the individual embodiments depicted and illustrated herein has distinct components and features that may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the compositions and methods described herein. Any described method can be carried out in the order of events recited or in any other order that is logically possible.

[0016] As used herein, the terms "recombinant" or "non-naturally occurring" refer to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic change or that has been modified by the introduction of a heterologous nucleic acid molecule, or to a cell (e.g., a microbial cell) that has been altered so that expression of a heterologous or endogenous nucleic acid molecule or gene can be controlled. Recombinant also refers to a cell that is derived from or is the progeny of a non-naturally occurring cell that has one or more such modifications. Genetic modifications include, for example, modifications that introduce an expressible nucleic acid molecule that encodes a protein, or the addition, deletion, substitution, or other functional change of other nucleic acid molecules in the genetic material of a cell. For example, recombinant cells may express genes or other nucleic acid molecules that are not found in the same or homologous form within native (wild-type) cells (e.g., fusion or chimeric proteins), or may provide an altered expression pattern of an endogenous gene, such as overexpression, underexpression, minimal expression, or no expression at all. "Recombination," "recombining," or producing a "recombinant" nucleic acid is generally the assembly of two or more nucleic acid fragments, which assembly gives rise to a chimeric gene.

[0017] As used herein, the terms "Gram-positive bacteria," "Gram-positive cells," "Gram-positive bacterial strains," and / or "Gram-positive bacterial cells" have the same meaning as used in the art. For example, Gram-positive bacterial cells include all strains of the phyla Actinobacteria and Firmicutes. In certain embodiments, such Gram-positive bacteria are from the classes Bacilli, Clostridia, and Mollicutes.

[0018] As used herein, the term "Bacillus" refers to all species within the genus "Bacillus" known to those skilled in the art, including, but not limited to, B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophilus, B. erythrocytes, B. erythrocytes, B. cholerae, B. ulcers, B. ulcerative colitis ... Examples of Bacillus species include B. alkalophilus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis. It is recognized that the genus Bacillus continues to undergo taxonomic reorganization. Thus, the genus is intended to include organisms such as reclassified species, for example, but not limited to, "B. stearothermophilus," now referred to as "Geobacillus stearothermophilus."

[0019] As used herein, the "wild-type B. subtilis ilvE promoter" sequence (abbreviated as "WT ilvE pro") comprises the nucleotide sequence set forth in SEQ ID NO: 16, as shown in Figure 3A.

[0020] As used herein, the "wild-type B. subtilis ilvE 5'-untranslated region" sequence (abbreviated as "WT ilvE 5'-UTR") comprises the nucleotide sequence set forth in SEQ ID NO: 17, as shown in Figure 3B.

[0021] As used herein, the "wild-type B. subtilis ilvE gene coding sequence (abbreviated "gene CDS, CDS or ORF")" comprises the nucleotides set forth in SEQ ID NO: 14, as shown in Figure 3C.

[0022] As used herein, the wild-type ilvE gene comprises the nucleotide sequence set forth in SEQ ID NO: 25, as shown in Figure 3D.

[0023] As used herein, the "native B. subtilis IlvE protein" encoded by the B. subtilis gene CDS comprises the amino acid sequence set forth in SEQ ID NO: 15, as shown in Figure 3E.

[0024] As used herein, a "mutant B. subtilis ilvE 5'-untranslated region" sequence (abbreviated "mutant ilvE 5'-UTR") comprises the nucleotide sequence set forth in SEQ ID NO: 18, as shown in Figure 1. In particular, as shown in Figure 1, the mutant ilvE 5'-UTR sequence (SEQ ID NO: 18; Figure 1B) comprises an unexpected single nucleotide polymorphism (SNP) at nucleotide position 73 compared to the wild-type B. subtilis ilvE 5'-UTR (SEQ ID NO: 17; Figure 1A). For example, as shown in Figure 1, the wild-type B. subtilis ilvE 5'-UTR sequence comprises a cytosine (C) at nucleotide position 73 (73C; SEQ ID NO: 17), while the mutant ilvE 5'-UTR SNP sequence comprises a thymine (T) at nucleotide position 73 (73T; SEQ ID NO: 18). As shown in Figure 1, the WT ilvE 5'-UTR sequence (Figure 1A) is shown with a double-underlined cytosine (C) at nucleotide position 73, and the mutant ilvE 5'-UTR sequence (Figure 1B) is shown with a double-underlined thymine (T) at nucleotide position 73. As shown in the schematic diagram in Figure 2, nucleotide position 73 of the ilvE 5'-UTR sequence is numbered from the beginning (+1) of the transcription start site, and position 73 may alternatively be referred to as position +73.

[0025] As used herein, phrases such as "B. subtilis P2 promoter" and / or "operably linked to a P2 promoter" refer specifically to the B. subtilis P2 promoter sequence as described and illustrated in WO 2020 / 112609, which is incorporated herein by reference in its entirety. More specifically, the B. subtilis P2 promoter is set forth in WO 2020 / 112609 as SEQ ID NO: 40.

[0026] As used herein, "host cell" refers to a cell that has the ability to act as a host or expression vehicle for a newly introduced DNA sequence. Thus, in certain embodiments of the present disclosure, the host cell is a Gram-positive cell, a Bacillus sp. cell, or an E. coli cell.

[0027] As used herein, the phrases "altered Bacillus cell" and / or "daughter Bacillus cell" refer to a recombinant Bacillus cell that contains at least one genetic modification that is not present in the parent cell from which the modified cell is derived. In certain embodiments, an "unmodified" Bacillus cell may be referred to as a "control cell," particularly when compared to or in relation to an "altered" Bacillus cell.

[0028] As used herein, when comparing the expression and / or production of a protein of interest (POI) in an "unmodified" (parent or control) cell to the expression and / or production of the same POI in an "modified" (daughter) cell, it will be understood that the "modified" and "unmodified" cells are grown / cultured / fermented under identical conditions (i.e., identical conditions of medium, temperature, pH, etc.). In certain embodiments, the increased amount of POI may be an endogenous Bacillus POI (e.g., a native protease, a native amylase, etc.) or a heterologous POI (e.g., a recombinant protease, a recombinant amylase, etc.) expressed in a recombinant Bacillus cell of the present disclosure.

[0029] As used herein, "increased" protein production or "increased" protein production refers to an increased amount of a produced protein (e.g., a protein of interest). The protein may be produced inside the host cell or secreted (or transported) into the culture medium. In certain embodiments, the protein of interest is produced (secreted) into the culture medium. Increased protein production can be detected, for example, as a higher maximum level of protein or enzyme activity (e.g., protease activity, amylase activity, pullulanase activity, cellulase activity, etc.) or as total extracellular protein produced compared to the parent cell.

[0030] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) the introduction, substitution, or removal of one or more nucleotides in a gene (or its ORF), or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of a gene or its ORF; (b) gene disruption; (c) gene conversion; (d) gene deletion; (e) gene downregulation; (f) directed mutagenesis; and / or (g) random mutagenesis of any one or more genes disclosed herein.

[0031] As used herein, the term "expression" refers to the transcription and stable accumulation of sense (mRNA) or antisense RNA derived from a nucleic acid molecule of the present disclosure. Expression can also refer to the translation of mRNA into a polypeptide. Thus, the term "expression" includes any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0032] As used herein, "nucleic acid" refers to nucleotide or polynucleotide sequences, and fragments or portions thereof, and to DNA, cDNA, and RNA of genomic or synthetic origin, which may be double-stranded or single-stranded, whether representing the sense or antisense strand. It will be understood that, as a result of the degeneracy of the genetic code, many nucleotide sequences can encode a given protein.

[0033] The polynucleotides (or nucleic acid molecules) described herein are understood to include "genes," "vectors," and "plasmids."

[0034] Thus, the term "gene" refers to a polynucleotide that encodes a specific sequence of amino acids, including all or part of a protein's coding sequence, and may include regulatory (non-transcribed) DNA sequences, such as promoter sequences, that determine the conditions under which the gene is expressed. The transcribed region of a gene may include introns, untranslated regions (UTRs), including 5'-untranslated regions (UTRs) and 3'-UTRs, and coding sequences (CDSs).

[0035] As used herein, the term "coding sequence" (CDS) refers to a nucleotide sequence that directly specifies the amino acid sequence of its (encoded) protein product. The boundaries of the coding sequence are generally determined by an open reading frame (hereinafter "ORF"), which usually begins with the ATG start codon. Coding sequences typically include DNA, cDNA, and recombinant nucleotide sequences.

[0036] As used herein, the term "promoter" refers to a nucleic acid sequence capable of controlling the expression of a coding sequence or functional RNA. Generally, the coding sequence is located 3' (downstream) of the promoter sequence. Promoters may be derived entirely from a native gene, be composed of different elements from different naturally occurring promoters, or include synthetic nucleic acid segments. Those skilled in the art will appreciate that different promoters can direct the expression of a gene in different cell types, at different developmental stages, or in response to different environmental or physiological conditions. Promoters that most frequently cause gene expression in most cell types are generally referred to as "constitutive promoters." Furthermore, it is recognized that because the exact boundaries of regulatory sequences in most cases are not completely defined, DNA fragments of different lengths may have the same promoter activity.

[0037] As used herein, the term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is affected by the other. For example, a promoter is operably linked to a coding sequence (e.g., ORF) when it is capable of affecting the expression of the coding sequence (i.e., when the coding sequence is under the transcriptional control of the promoter). A coding sequence can be operably linked to a regulatory sequence in either a sense or antisense orientation.

[0038] A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, DNA encoding a secretory leader (i.e., signal peptide) is operably linked to DNA for a polypeptide if it is expressed as a preprotein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of that coding sequence; or a ribosome binding site is operably linked to a coding sequence if it is positioned so as to promote translation. Generally, "operably linked" means that the DNA sequences being linked are contiguous, and, in the case of a secretory leader, contiguous and in reading phase. Enhancers, however, need not be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide adaptors or linkers are used in accordance with conventional practice.

[0039] As used herein, "a functional promoter sequence that controls the expression of a gene of interest (or its open reading frame) linked to a protein-coding sequence of the gene of interest" refers to a promoter sequence that controls the transcription and translation of a coding sequence in Bacillus. For example, in certain embodiments, the present disclosure is directed to a polynucleotide comprising a 5' promoter (or a 5' promoter region or a tandem 5' promoter, etc.), wherein the promoter region is operably linked to a nucleic acid sequence (e.g., an ORF) that encodes a protein.

[0040] As used herein, "suitable regulatory sequences" refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence that influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences can include promoters, translation leader sequences, RNA processing sites, effector binding sites, and stem-loop structures.

[0041] As used herein, the term "introducing," when used in phrases such as "introducing at least one polynucleotide open reading frame (ORF), or gene thereof, or vector thereof, into a bacterial cell" or "introducing into a Bacillus cell," includes methods known in the art for introducing polynucleotides into cells, including, but not limited to, protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation, and the like.

[0042] As used herein, "transformed" or "transformation" means that a cell has been transformed by the use of recombinant DNA techniques. Transformation generally occurs by inserting one or more nucleotide sequences (e.g., polynucleotides, ORFs, or genes) into a cell. The inserted nucleotide sequences may be heterologous nucleotide sequences (i.e., sequences that are not naturally present in the cell being transformed). Transformation generally refers to the introduction of exogenous DNA into a host cell such that the DNA is maintained as a chromosomal integrant or a self-replicating extrachromosomal vector.

[0043] As used herein, "transforming DNA," "transforming sequence," and "DNA construct" refer to DNA used to introduce a sequence into a host cell or organism. Transforming DNA is DNA used to introduce a sequence into a host cell or organism. This DNA can be generated in vitro by PCR or any other suitable technique. In some embodiments, the transforming DNA includes the incoming sequence, while in other embodiments, the transforming DNA further includes the incoming sequence flanked by homology boxes. In yet other embodiments, the transforming DNA includes other non-homologous sequences (i.e., stuffer sequences or flanking sequences) added to the ends. The ends can be closed such that the transforming DNA forms a closed circle, such as, for example, for insertion into a vector.

[0044] As used herein, "gene disruption" or "gene disruption" are used interchangeably and refer broadly to any genetic modification that substantially prevents a host cell from producing a functional gene product (e.g., a protein). Thus, as used herein, gene disruption includes, but is not limited to, frameshift mutations, premature stop codons (i.e., so that a functional protein is not produced), substitutions that eliminate or reduce the activity of internal deletions of proteins (so that a functional protein is not produced), insertions that disrupt coding sequences, mutations that remove the operable link between the native promoter and the open reading frame required for transcription, and the like.

[0045] As used herein, "incoming sequence" refers to a DNA sequence that is introduced into a Bacillus sp. chromosome. In some embodiments, the incoming sequence is part of a DNA construct. In other embodiments, the incoming sequence encodes one or more proteins of interest. In some embodiments, the incoming sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it may be a homologous or heterologous sequence). In some embodiments, the incoming sequence encodes one or more proteins of interest, genes, and / or mutant or modified genes. In alternative embodiments, the incoming sequence encodes a functional wild-type gene or operon, a functional mutant gene or operon, or a non-functional gene or operon. In some embodiments, a non-functional sequence can be inserted into a gene to disrupt gene function. In another embodiment, the incoming sequence comprises a selectable marker. In yet another embodiment, the incoming sequence comprises two homology boxes.

[0046] As used herein, a "homology box" refers to a nucleic acid sequence that is homologous to a sequence within a Bacillus chromosome. More specifically, a homology box is an upstream or downstream region that shares about 80-100% sequence identity, about 90-100% sequence identity, or about 95-100% sequence identity with the coding region immediately flanking a gene or portion of a gene to be deleted, disrupted, inactivated, downregulated, etc., according to the present invention. These sequences direct where a DNA construct will integrate within the Bacillus chromosome and direct which portion of the Bacillus chromosome will be replaced by the incoming sequence. While not intended to limit the present disclosure, a homology box can include from about 1 base pair (bp) to 200 kilobases (kb). Preferably, the homology box comprises approximately 1 bp to 10.0 kb, 1 bp to 5.0 kb, 1 bp to 2.5 kb, 1 bp to 1.0 kb, and 0.25 kb to 2.5 kb. The homology box may also comprise approximately 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb, and 0.1 kb. In some embodiments, the 5' and 3' ends of the selectable marker are flanked by homology boxes, which comprise nucleic acid sequences that immediately flank the coding region of the gene.

[0047] As used herein, the term "nucleotide sequence encoding a selectable marker" refers to a nucleotide sequence that is expressible in a host cell and in which expression of the selectable marker confers on cells containing the expressed gene the ability to grow in the presence of a corresponding selection agent or in the absence of an essential nutrient.

[0048] As used herein, the terms "selectable marker" and "selection marker" refer to a nucleic acid (e.g., a gene) that can be expressed in a host cell, allowing for easy selection of those hosts that contain the vector. Examples of such selectable markers include, but are not limited to, antimicrobial agents. Thus, the term "selectable marker" refers to a gene that indicates that a host cell has taken up incoming DNA of interest or that some other reaction has occurred. Generally, a selectable marker is a gene that confers antimicrobial resistance or a metabolic advantage to a host cell, allowing cells containing foreign DNA to be distinguished from cells that have not received the foreign sequence during transformation.

[0049] A "present selectable marker" is a marker located on the chromosome of the microorganism being transformed. The present selectable marker encodes a different gene than the selectable marker on the transforming DNA construct. Selectable markers are well known to those skilled in the art. As noted above, markers can be antimicrobial resistance markers (e.g., amplicons). R , phleo R , spec R , kan R ,ery R , tet R , cmp R and neo R In some embodiments, the present invention provides a chloramphenicol resistance gene (e.g., the gene present on pC194 and the resistance gene present in the genome of Bacillus licheniformis). This resistance gene is particularly useful in the present invention and in embodiments involving chromosomal integration cassettes and chromosomal amplification of the integration plasmid. Other markers useful according to the present invention include, but are not limited to, auxotrophic markers such as serine, lysine, tryptophan, and the like, and detectable markers such as β-galactosidase.

[0050] As defined herein, a host cell "genome," bacterial (host) cell "genome," or Bacillus sp. (host) cell "genome" includes chromosomal genes and extrachromosomal genes.

[0051] As used herein, the terms "plasmid," "vector," and "cassette" refer to extrachromosomal elements that often carry genes that are not generally part of the cell's central metabolism and usually have the form of circular double-stranded DNA molecules. Such elements can be linear or circular, single- or double-stranded, DNA or RNA autonomously replicating sequences, genome-integrating sequences, phage, or nucleotide sequences from any source in which multiple nucleotide sequences have been joined or recombined into a unique structure that can introduce into a cell a promoter fragment and DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences.

[0052] As used herein, the term "plasmid" refers to a circular double-stranded (ds) DNA construct that is used as a cloning vector and forms an extrachromosomal, self-replicating genetic element in many bacteria and some eukaryotes. In some embodiments, the plasmid is integrated into the genome of the host cell. In some embodiments, the plasmid is present in the parent cell and is lost in daughter cells.

[0053] As used herein, "transformation cassette" refers to a specific vector containing a gene (or its ORF) and having elements in addition to the foreign gene that facilitate transformation of a particular host cell.

[0054] As used herein, the term "vector" refers to any nucleic acid that can replicate (multiply) within a cell and carry new genes or DNA segments into the cell. Thus, the term refers to a nucleic acid construct designed for transport between various host cells. Vectors include viruses, bacteriophages, proviruses, plasmids, phagemids, transposons, which are "episomal" (i.e., capable of autonomous replication or integration into the chromosomes of the host organism), as well as artificial chromosomes such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), and PLACs (plant artificial chromosomes).

[0055] An "expression vector" refers to a vector capable of incorporating and expressing heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and known to those skilled in the art. The selection of an appropriate expression vector is within the knowledge of one skilled in the art.

[0056] As used herein, the terms "expression cassette" and "expression vector" refer to nucleic acid constructs produced recombinantly or synthetically with a set of specific nucleic acid elements that allow for transcription of a specific nucleic acid in a target cell (i.e., they are vectors or vector elements as described above). Recombinant expression cassettes can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter. In some embodiments, the DNA construct also includes a set of specific nucleic acid elements that allow for transcription of a specific nucleic acid in a target cell. In certain embodiments, the DNA construct of the present disclosure includes a selectable marker and an inactivated chromosomal segment or gene segment or DNA segment, as defined herein.

[0057] As used herein, a "targeting vector" is a vector that contains a polynucleotide sequence homologous to a region in a host cell chromosome into which the targeting vector is transformed and is capable of driving homologous recombination at that region. For example, targeting vectors are used to introduce mutations into a host cell chromosome by homologous recombination. In some embodiments, the targeting vector contains other non-homologous sequences (i.e., stuffer sequences or flanking sequences), for example, added to the ends. The ends can be closed such that the targeting vector forms a closed circle, e.g., by insertion into a vector. For example, in certain embodiments, parent B. licheniformis (host) cells are modified (e.g., transformed) by introducing one or more "targeting vectors" into them.

[0058] As used herein, the term "protein of interest" or "POI" refers to an expression that is desired in modified B. licheniformis (daughter) host cells, where the POI is expressed at an increased level (i.e., compared to the "unmodified" (parent) cells). Thus, as used herein, a POI can be an enzyme, a substrate-binding protein, a surfactant protein, a structural protein, a receptor protein, etc. In certain embodiments, the modified cells of the present disclosure produce an increased amount of a heterologous or endogenous protein of interest compared to the parent cell. In certain embodiments, the increased amount of a protein of interest produced by the modified cells of the present disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or more than a 5.0% increase compared to the parent cell.

[0059] Similarly, as defined herein, "gene of interest" or "GOI" refers to a nucleic acid sequence (e.g., polynucleotide, gene, or open reading frame) that encodes a POI. A "gene of interest" that encodes a "protein of interest" can be a naturally occurring gene, a mutated gene, or a synthetic gene.

[0060] As used herein, the terms "polypeptide" and "protein" are used interchangeably and refer to polymers of any length comprising amino acid residues linked by peptide bonds. Conventional one-letter or three-letter codes for amino acid residues are used herein. Polypeptides can be linear or branched, can contain modified amino acids, and can be interrupted by non-amino acids. The term polypeptide also encompasses amino acid polymers that are modified naturally or by intervention, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within this definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids), as well as other modifications known in the art.

[0061] In certain embodiments, the genes of the present disclosure are those encoding enzymes (e.g., acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysate, lyase), endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.

[0062] As used herein, a "variant" polypeptide refers to a polypeptide that is derived from a parent (or reference) polypeptide by one or more amino acid substitutions, additions, or deletions, typically by recombinant DNA techniques. A variant polypeptide may differ from the parent polypeptide by a small number of amino acid residues and may be defined by the level of primary amino acid sequence homology / identity with the parent (reference) polypeptide.

[0063] Preferably, a variant polypeptide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity with a parent (reference) polypeptide sequence. As used herein, a "variant" polynucleotide refers to a polynucleotide that encodes a variant polypeptide, which "variant polynucleotide" has a particular degree of sequence homology / identity with a parent polynucleotide or hybridizes to a parent polynucleotide (or its complement) under stringent hybridization conditions. Preferably, the variant polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% nucleotide sequence identity with the parent (reference) polynucleotide sequence.

[0064] As used herein, the term "mutation" refers to any change or alteration in a nucleic acid sequence. There are several types of mutations, including point mutations, deletion mutations, silent mutations, frameshift mutations, splicing mutations, etc. Mutations can be made specifically (e.g., by site-directed mutagenesis) or randomly (e.g., by chemical agents, repair minus passaging through bacterial strains).

[0065] As used herein, the term "substitution" in reference to a polypeptide or sequence thereof means the replacement (ie, substitution) of one amino acid with another.

[0066] As defined herein, "endogenous gene" refers to a gene that is present in its natural location in the genome of an organism.

[0067] As defined herein, a "heterologous" gene, "non-endogenous" gene, or "foreign" gene refers to a gene (or ORF) that is not normally found in the host organism, but that has been introduced into the host organism by gene transfer. As used herein, the term "foreign" gene includes a native gene (or ORF) inserted into a non-native organism and / or a chimeric gene inserted into a native or non-native organism.

[0068] As defined herein, a "heterologous regulatory sequence" refers to a gene expression control sequence (e.g., a promoter or enhancer) that does not naturally function to regulate (control) the expression of a gene of interest. Generally, heterologous nucleic acid sequences are not endogenous (natural) to the cell or part of the genome in which they are present, but have been added to the cell by infection, transfection, transformation, microinjection, electroporation, etc. A "heterologous" nucleic acid construct can contain regulatory sequence / DNA coding (ORF) sequence combinations that are the same as or different from regulatory sequence / DNA coding sequence combinations found in the native host cell.

[0069] As used herein, the terms "signal sequence" and "signal peptide" refer to a sequence of amino acid residues that may be involved in the secretion or direct export of a mature protein or a precursor form of a protein. A signal sequence is generally located at the N-terminus of the precursor or mature protein sequence. A signal sequence may be endogenous or foreign. A signal sequence is usually not present in the mature protein. A signal sequence is generally cleaved from a protein by a signal peptidase after the protein has been exported.

[0070] The term "derived from" encompasses the terms "originating from," "obtained from," "available from," and "made from," and generally indicates that one particular material or composition finds its origin in, or has characteristics that can be described with reference to, another particular material or composition.

[0071] As used herein, the term "homology" refers to homologous polynucleotides or homologous polypeptides. When two or more polynucleotides or two or more polypeptides are homologous, this means that the homologous polynucleotides or polypeptides have a "degree of identity" of at least 60%, more preferably at least 70%, even more preferably at least 85%, even more preferably at least 90%, more preferably at least 95%, and most preferably at least 98%. The degree of homology between sequences can be determined using any suitable method known in the art (see, for example, Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; programs such as GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package (Genetics Computer Group, Madison, WI); and Devereux et al., 1984). For purposes of the present invention, the degree of identity between two amino acid sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970) as implemented in the Needle program of the EMBOSS package (Rice et al., 2000), preferably version 3.0.0 or later. Optional parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. The output of the Needle label "longest identity" (obtained using the nobrief option) is used as the percent identity and is calculated as follows: (Identical residues × 100) / (length of alignment − total number of gaps in alignment)

[0072] As used herein, the term "percent identity" refers to the level of nucleic acid or amino acid sequence identity between nucleic acid sequences encoding polypeptides or between the amino acid sequences of polypeptides when aligned using a sequence alignment program.

[0073] As used herein, "specific productivity" is the total amount of protein produced per cell per unit time over a given period of time.

[0074] As defined herein, the terms "purified," "isolated," or "enriched" mean that a biomolecule (e.g., a polypeptide or polynucleotide) has been altered from its native state by separation from some or all of the naturally occurring components with which it is naturally associated. Such isolation or purification can be accomplished by separation techniques known in the art, such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulfate precipitation or other protein salting-out, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis, or gradient separation to remove whole cells, cell debris, impurities, extraneous proteins, or enzymes that are not desired in the final composition. Purified or isolated biomolecule compositions can then be supplemented with components that confer additional benefits, such as activators, anti-inhibitors, desirable ions, pH-adjusting compounds, or other enzymes or chemicals.

[0075] As used herein, "flanking sequence" refers to any sequence upstream or downstream of the sequence under consideration (e.g., in gene ABC, gene B is flanked by gene sequences A and C). In certain embodiments, the incoming sequence is flanked by homology boxes on both sides. In other embodiments, the incoming sequence and homology box comprise a unit flanked by stuffer sequences on both sides. In some embodiments, flanking sequences are present on only one side (3' or 5'), but in preferred embodiments, they are present on both sides of the flanking sequence. The sequence of each homology box is homologous to a sequence within the Bacillus chromosome. These sequences direct where in the Bacillus chromosome the novel construct will be integrated and which portion of the Bacillus chromosome will be replaced by the incoming sequence. In other embodiments, the 5' and 3' ends of the selectable marker are flanked by polynucleotide sequences comprising a portion of an inactivating chromosomal segment. In some embodiments, a flanking sequence is present on only one side (3' or 5'), while in other embodiments, it is present on both sides of the sequence it is flanking.

[0076] II. Mutant Bacillus strains with enhanced protein production and carbon yield phenotypes As described generally herein and further illustrated in the Examples below, Applicants have identified a mutant B. subtilis strain (designated "CZ437") that has an enhanced protein production phenotype. In particular, Applicants performed next-generation sequencing (NGS) on the mutant CZ437 strain to further characterize the observed enhanced protein productivity phenotype and identified an unexpected single nucleotide polymorphism (SNP) mutation in the wild-type (WT) B. subtilis ilvE 5'-UTR sequence (SEQ ID NO: 17). For example, the DNA sequences of the WT ilvE 5'-UTR (SEQ ID NO: 17) and the mutant ilvE 5'-UTR (SEQ ID NO: 18) are shown in Figures 1A and 1B, respectively, where the WT ilvE 5'-UTR sequence contains a cytosine (C) at nucleotide position 73 (SEQ ID NO: 17), and the mutant ilvE 5'-UTR SNP sequence contains a thymine (T) at nucleotide position 73 (SEQ ID NO: 18).

[0077] Without wishing to be bound by any particular theory, mechanism, or mode of operation, the applicants believe that the unexpected SNP mutation (i.e., SNP C→T) identified herein in the ilvE 5'-UTR may affect ilvE messenger RNA (mRNA) stability. For example, the ilvE (ybgE) gene is known to encode a branched-chain amino acid aminotransferase that transaminates branched-chain amino acids and ketoglutarate. Berger et al. (2003) demonstrated another function of IlvE in the methionine regeneration pathway by converting ketomethiobutyrate (KMTB) to methionine. For this reaction, IvlE aminotransferase can use leucine, isoleucine, valine, phenylalanine, and tyrosine as amino donors, whereas the B. subtilis homologous enzyme YkrV uses only glutamine as the amino donor. As generally described in Mader et al. (2004), the ilvE gene is negatively regulated by the global transcriptional regulator Cody, which controls ilvE transcription by binding to its transcription leader (5'-UTR) sequence, acting as a roadblock for RNA polymerase, and repressing ilvE in the presence of casamino acids or free amino acids. As shown in Figure 2, a SNP mutation in the ilvE 5'-UTR is located 55 nucleotides (bp) upstream (5') of the translation start site, which may affect ilvE mRNA stability. Notably, as shown in Figure 2, the C>T mutation in the ilvE 5'-UTR is located near CodY-binding motif 2, a putative weak binding motif 5. As proposed here, weak codY-binding motif 5 could potentially mask codY from binding to motif 2, thereby causing derepression of ilvE transcription.

[0078] Therefore, as described herein and further illustrated in the Examples below, Applicants designed, constructed, and screened recombinant Bacillus strains expressing a reporter protein (GG36) to further evaluate the enhanced protein production phenotype identified in the mutant B. subtilis strain CZ437. More specifically, as described in Example 1, two GG36 reporter protein expression cassettes were constructed and introduced into B. subtilis strain A (containing a mutant ilvE 5′-UTR (SEQ ID NO: 18)) and isogenic B. subtilis strain B (containing a WT ilvE 5′-UTR (SEQ ID NO: 17)), and strains A and B were fermented under the same conditions in a large-scale (approximately 14 L) fermentor using standard fermentation conditions. As shown in Table 1 (Example 1), the relative improvement in carbon yield of mutant B. subtilis reporter strain A is significantly enhanced compared to the isogenic reporter strain B.

[0079] Similarly, time-course samples from B. subtilis strains A and B (expressing the GG36 reporter protein) were used for real-time quantitative PCR (RT qPCR) analysis, with samples taken after 8, 16, 24, and 32 hours of fermentation and total RNA extracted, as described in Example 2. For example, the fold change in ilvE mRNA for B. subtilis strain A (mutated ilvE 5'-UTR) versus isogenic B. subtilis strain B (WT ilvE 5'-UTR) after 8, 16, 24, and 32 hours of fermentation is shown in Figure 4. More specifically, as shown in Figure 4, the amount of ilvE mRNA at 16, 24, and 32 hours of fermentation was significantly increased by 2.91-fold, 1.93-fold, and 1.79-fold, respectively, in B. subtilis strain A compared to the isogenic B. subtilis strain B.

[0080] Accordingly, certain embodiments of the present disclosure relate to mutated Bacillus strains comprising variant ilvE gene sequences, recombinant Bacillus strains comprising variant ilvE gene sequences, mutated / recombinant strains comprising variant ilvE gene sequences and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains comprising variant ilvE gene sequences, expression cassettes encoding proteins of interest, methods and compositions for culturing recombinant Bacillus strains comprising variant ilvE gene sequences to enhance production of proteins of interest, etc.

[0081] III. Recombinant Polynucleotides and Molecular Biology As generally described above, certain embodiments of the present disclosure relate to, inter alia, variant ilvE genes, mutated Bacillus strains comprising variant ilvE genes, recombinant (genetically modified) Bacillus strains comprising variant ilvE gene sequences, mutated and / or recombinant Bacillus strains comprising variant ilvE gene sequences and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains comprising variant ilvE gene sequences, expression cassettes encoding proteins of interest, methods and compositions for culturing recombinant Bacillus strains comprising variant ilvE gene sequences to enhance production of proteins of interest, and the like.

[0082] Thus, in one or more embodiments, the present disclosure provides recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant (genetically modified) Gram-positive bacterial cells / strains expressing a protein of interest, etc. In certain one or more embodiments, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells to enhance production of a protein of interest.

[0083] In certain embodiments, polynucleotide constructs of the present disclosure are referred to as expression cassettes (or expression constructs), and an expression cassette comprises at least an upstream (5') promoter sequence operably linked in a 5' to 3' direction and in operable combination to a downstream (3') gene coding sequence CDS. In certain embodiments, the expression cassette encodes one or more proteins of interest (e.g., 5'-[promoter sequence]-[gene coding sequence]-3'; abbreviated as 5'-[Pro]-[gene CDS]-3').

[0084] In other embodiments, the present disclosure provides an ilvE expression cassette. In certain embodiments, the ilvE expression cassette comprises, in a 5' to 3' direction and in operable combination, at least an upstream (5') promoter sequence operably linked to a variant ilvE 5'-untranslated region (5'-UTR) sequence operably linked to a wild-type ilvE gene CDS (abbreviated as 5'-[Pro]-[ilvE*5'-UTR)]-[WT ilvE CDS]-3'. As abbreviated above, the variant (mutated) ilvE* gene 5'-UTR sequence is presented / denoted with an asterisk (*) to distinguish it from the wild-type ilvE gene 5'-UTR sequence (i.e., SEQ ID NO: 17).

[0085] In certain other embodiments, the expression cassette may include one or more DNA sequence elements, such as, but not limited to, a DNA sequence element encoding a protein / peptide signal (secretory) sequence, a DNA sequence element encoding propeptide (pro-region) amino acid residues, a DNA sequence element comprising a transcription terminator sequence, a DNA sequence element comprising a 5'-UTR, a 3'-UTR, etc.

[0086] Thus, one or more of the nucleic acid sequences described herein can be produced by using any suitable synthesis, manipulation, and / or isolation technique, or a combination thereof. For example, one or more of the polynucleotides described herein can be produced using standard nucleic acid synthesis techniques, such as solid-phase synthesis techniques, which are well known to those of skill in the art. In such techniques, generally, fragments of up to 50 or more nucleobases are synthesized and then joined (e.g., by enzymatic or chemical ligation methods) to form substantially the required contiguous nucleic acid sequence. Synthesis of one or more polynucleotides described herein can also be facilitated by chemical synthesis using any suitable method known in the art, including, but not limited to, the classical phosphoramidite method or methods commonly practiced in automated synthesis. One or more polynucleotides described herein can also be generated using an automated DNA synthesizer. Customized nucleic acids can also be obtained from a variety of commercial sources (e.g., ATUM (DNA 2.0), Newark, CA, USA; Life Tech (GeneArt), Carlsbad, CA, USA; GenScript, Ontario, Canada; Base Clear BV, Leiden, Netherlands; Integrated DNA Technologies, Skokie, IL, USA; Ginkgo Bioworks (Gen9), Boston, MA, USA; and Twist Bioscience, San Diego, CA, USA). San Francisco, CA, USA. Other techniques for synthesizing nucleic acids and the associated principles are described and known in the art.

[0087] Recombinant DNA techniques useful for modifying nucleic acids are well known in the art, such as restriction endonuclease digestion, ligation, reverse transcription and cDNA production, and polymerase chain reaction (e.g., PCR). One or more polynucleotides described herein can also be obtained by screening a cDNA library with one or more oligonucleotide probes capable of hybridizing to or PCR amplifying a polynucleotide encoding one or more variants described herein. Methods for screening and isolating cDNA clones and PCR amplification methods are well known to those of skill in the art and are described in standard references known to those of skill in the art. One or more polynucleotides described herein can be obtained by altering a naturally occurring polynucleotide backbone (e.g., encoding one or more variant pro-region sequences described herein) by, for example, known mutagenesis methods (e.g., site-directed mutagenesis, site-saturation mutagenesis, and in vitro recombination). A variety of methods suitable for generating modified polynucleotides described herein that encode one or more variants described herein are known in the art, including, but not limited to, site-saturation mutagenesis, systematic mutagenesis, insertional mutagenesis, deletion mutagenesis, random mutagenesis, site-directed mutagenesis and directed evolution, as well as various other recombinant methods.

[0088] As outlined above and further illustrated in the Examples below, certain embodiments of the present disclosure relate to recombinant (engineered) Gram-positive cells capable of producing heterologous proteins of interest. Accordingly, certain embodiments relate to methods for constructing such recombinant Gram-positive cells with improved protein production capabilities. In certain embodiments, one or more expression cassettes encoding one or more proteins of interest are introduced into a Gram-positive cell of the present disclosure. In exemplary embodiments, the cassettes are integrated into the genome of the cell. Accordingly, certain embodiments relate to nucleic acid molecules, polynucleotides (e.g., vectors, plasmids, expression cassettes), regulatory elements, and the like, suitable for use in constructing recombinant (engineered) Gram-positive host cells.

[0089] Thus, as presented in the Examples and outlined herein, recombinant cells of the present disclosure can be constructed by those skilled in the art using standard and routine recombinant DNA and molecular cloning techniques well known in the art. Methods of genetic modification include, but are not limited to, (a) introduction, substitution, or removal of one or more nucleotides in a gene, or introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of a gene, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-directed mutagenesis, and / or (g) random mutagenesis.

[0090] In certain embodiments, modified cells of the present disclosure may be constructed by reducing or eliminating expression of a gene using methods well known in the art, such as insertion, disruption, substitution, or deletion. The portion of a gene to be modified or inactivated may be, for example, a coding region or a regulatory element required for expression of the coding region.

[0091] An example of such a regulatory or control sequence may be a promoter sequence or a functional portion thereof (i.e., a portion sufficient to affect the expression of a nucleic acid sequence). Other control sequences for modification include, but are not limited to, a leader sequence, a propeptide sequence, a signal sequence, a transcription terminator, a transcription activator, and the like.

[0092] In certain other embodiments, modified cells are constructed by gene deletion, which eliminates or reduces gene expression. Gene deletion techniques allow for the partial or complete removal of genes, thereby eliminating their expression or expressing non-functional (or reduced activity) protein products. In such methods, gene deletion can be achieved by homologous recombination using a plasmid constructed to contain adjacent 5' and 3' regions flanking the gene. The flanking 5' and 3' regions can be introduced into cells, for example, on a temperature-sensitive plasmid associated with a second selectable marker at a permissive temperature that allows the plasmid to establish in the cells. Cells are then shifted to a non-permissive temperature to select for cells with the plasmid integrated into the chromosome at one of the homologous flanking regions. Selection for plasmid integration is influenced by selection for the second selectable marker. After integration, recombination events at the second homologous flanking region are stimulated by shifting the cells to a permissive temperature for several generations without selection. Cells are plated to obtain single colonies, which are then tested for the loss of both selectable markers. Thus, one skilled in the art can readily identify nucleotide regions within the coding sequence of a gene and / or the non-coding sequence of a gene that are suitable for complete or partial deletion.

[0093] In other embodiments, modified cells are constructed by introducing, substituting, or removing one or more nucleotides within a gene or a regulatory element required for its transcription or translation. For example, nucleotides can be inserted or removed to introduce a stop codon, remove a start codon, or cause a frameshift in the open reading frame. Such modifications can be achieved by site-directed mutagenesis or PCR-generated mutagenesis, according to methods known in the art. Thus, in certain embodiments, the genes of the present disclosure are inactivated by complete or partial deletion.

[0094] In another embodiment, modified cells are constructed by a gene conversion process. For example, in gene conversion methods, a nucleic acid sequence corresponding to a gene is mutated in vitro to generate a defective nucleic acid sequence, which is then transformed into a parent cell to generate the defective gene. The defective nucleic acid sequence replaces the endogenous gene through homologous recombination. It may be desirable for the defective gene or gene fragment to also encode a marker that can be used to select for transformants containing the defective gene. For example, the defective gene can be introduced into a non-replicating or temperature-sensitive plasmid associated with a selectable marker. Selection for plasmid integration is affected by selecting for that marker under conditions that do not allow replication of the plasmid. Selection for a second recombination event resulting in gene replacement is affected by examining colonies for the loss of the selectable marker and the acquisition of a mutated gene. Alternatively, the defective nucleic acid sequence can contain an insertion, substitution, or deletion of one or more nucleotides of the gene, as described below.

[0095] In other embodiments, modified cells are constructed using established antisense technology, using a nucleotide sequence complementary to the nucleic acid sequence of a gene. More specifically, the expression of a gene by a Gram-positive cell can be reduced (downregulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which can be transcribed in the cell and hybridize to the mRNA produced in the cell. Thus, under conditions where the complementary antisense nucleotide sequence can hybridize to the mRNA, the amount of translated protein is reduced or eliminated. Such antisense methods include, but are not limited to, RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), antisense oligonucleotides, etc., all of which are well known to those skilled in the art.

[0096] In other embodiments, modified cells are generated / constructed by CRISPR-Cas9 editing. For example, genes encoding proteins of interest may be disrupted (or deleted or downregulated) by nucleic acid-guided endonucleases that find their target DNA by binding to a guide RNA (e.g., Cas9) and Cpf1 or a guide DNA (e.g., NgAgo), which recruit the endonucleases to target sequences on the DNA, where the endonucleases can generate single- or double-strand breaks in the DNA. This targeted DNA break can then serve as a substrate for DNA repair and recombine with the editing template provided to disrupt or delete the gene. For example, a gene encoding a nucleic acid-guided endonuclease (in this case, Cas9 from S. pyogenes) or a codon-optimized gene encoding a Cas9 nuclease is operably linked to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells, thereby generating a Gram-positive cell Cas9 expression cassette. Similarly, one or more target sites unique to a gene of interest can be easily identified by one skilled in the art. For example, to construct a DNA construct encoding a gRNA directed to a target site within a gene of interest, a variable targeting domain (VT) would contain the nucleotides of the target site 5' to the (PAM) protospacer adjacent motif (TGG), fused to DNA encoding the Cas9 endonuclease recognition domain (CER) for S. pyogenes Cas9. DNA encoding the gRNA is generated by combining the DNA encoding the VT domain with the DNA encoding the CER domain. Thus, a Gram-positive expression cassette for the gRNA is generated by operably linking the DNA encoding the gRNA to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells.

[0097] In certain embodiments, the DNA break induced by endonuclease is repaired / replaced using the incoming sequence.For example, to precisely repair the DNA break generated by the above-mentioned Cas9 expression cassette and gRNA expression cassette, a nucleotide editing template is provided so that the DNA repair mechanism of the cell can use the editing template.For example, about 500bp of the 5' side of the targeting gene can be fused to about 500bp of the 3' side of the targeting gene to generate an editing template, and this template is used by the mechanism of the Gram-positive host to repair the DNA break generated by RGEN.

[0098] The Cas9 expression cassette, gRNA expression cassette, and editing template can be co-delivered into filamentous fungal cells using a variety of methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence). Transformed cells are screened by PCR amplification of the target locus using forward and reverse primers. These primers can amplify the wild-type locus or the modified locus edited by RGEN. These fragments are then sequenced using sequencing primers to identify edited colonies.

[0099] In yet other embodiments, modified cells are constructed by random or directed mutagenesis using methods well known in the art, including, but not limited to, chemical mutagenesis and transposition. Genetic modification can be performed by mutagenizing parent cells and screening for mutant cells in which expression of the gene is reduced or eliminated. Mutagenesis can be directed or random, for example, by using suitable physical or chemical mutagenizing agents, by using suitable oligonucleotides, or by subjecting the DNA sequence to PCR-generated mutagenesis. Furthermore, mutagenesis can be performed by using any combination of these mutagenesis methods.

[0100] Examples of physical or chemical mutagenic agents suitable for the purposes of the present invention include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl-N'-nitrosoguanidine (NTG), O-methylhydroxylamine, nitrous acid, ethyl methanesulfonate (EMS), sodium bisulfite, formic acid, and nucleotide analogs. When using such agents, mutagenesis is generally carried out by incubating parent cells to be mutagenized under appropriate conditions in the presence of the mutagen of choice, and selecting mutant cells that exhibit reduced or no gene expression.

[0101] WO 2003 / 083125 discloses methods for modifying Gram-positive (Bacillus) cells, such as creating Bacillus deletion strains and DNA constructs using PCR fusion to bypass E. coli. WO 2002 / 14490 discloses methods for modifying Bacillus cells, including (1) construction and transformation of an integration plasmid (pComK), (2) random mutagenesis of coding, signal, and propeptide sequences, (3) homologous recombination, (4) increasing transformation efficiency by adding non-homologous flanks to the transforming DNA, (5) optimizing double-crossover integration, (6) site-directed mutagenesis, and (7) markerless deletion.

[0102] Those skilled in the art are aware of suitable methods for introducing polynucleotide sequences into bacterial cells (e.g., Gram-negative cells, Gram-positive cells). Indeed, methods such as transformation, including protoplast transformation and aggregation, transduction, and protoplast fusion, are known and suitable for use in the present disclosure. Transformation methods are particularly preferred for introducing the DNA constructs of the present disclosure into host cells.

[0103] In addition to commonly used methods, in some embodiments, host cells are transformed directly (i.e., no intermediate cells are used to amplify or otherwise process the DNA construct before introduction into the host cell). Introduction of the DNA construct into the host cell includes those physical and chemical methods known in the art for introducing DNA into a host cell without insertion into a plasmid or vector. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, liposomes, and the like. In additional embodiments, the DNA construct is co-transformed with a plasmid without being inserted into the plasmid. In further embodiments, the selectable marker is deleted or substantially excised from the modified Bacillus strain by methods known in the art. In some embodiments, degradation of the vector from the host chromosome leaves flanking regions within the chromosome while removing the unique chromosomal region.

[0104] Promoters and promoter sequence regions, coding sequences (CDS), open reading frames (ORF) and / or variant sequences thereof for use in expressing genes in Gram-positive cells are generally known to those skilled in the art. The promoter sequences of the present disclosure are generally selected so that they function in Gram-positive cells. For example, promoters useful for driving gene expression in Bacillus cells include, but are not limited to, the B. subtilis alkaline protease (aprE) promoter, the B. subtilis α-amylase promoter (amyE), the B. licheniformis α-amylase promoter (amyL), the B. amyloliquefaciens α-amylase promoter, the neutral protease (nprE) promoter from B. subtilis, a mutant aprE promoter, or any other promoter from B. licheniformis or other related Bacilli. Methods for screening and generating promoter libraries with different activities (promoter strengths) in Bacillus cells are described in WO 2002 / 14490.

[0105] IV. FERMENTATION OF BACILLUS CELLS TO PRODUCE PROTEINS As outlined above, certain embodiments relate to compositions and methods for constructing and obtaining Gram-positive cells that express / produce one or more proteins of interest. Accordingly, certain other embodiments of the present disclosure relate to methods for producing proteins of interest in Gram-positive cells by fermenting the cells in a suitable medium. Fermentation methods well known in the art may be applied to ferment the Gram-positive cells of the present disclosure.

[0106] In some embodiments, cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system in which the composition of the medium is set at the beginning of the fermentation and remains unchanged throughout the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. In this method, fermentation occurs without adding any components to the system. Batch fermentation is generally considered "batch" with respect to the addition of a carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes constantly until the fermentation is stopped. In a typical batch culture, cells progress through a static lag phase to a high-growth logarithmic phase and may eventually progress to a stationary phase where growth rate decreases or stops. If untreated, cells in the stationary phase eventually die. Generally, cells in the logarithmic phase are responsible for the majority of product production.

[0107] A suitable variation on the standard batch system is the "fed-batch" fermentation system. In this variation of the typical batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cellular metabolism and when a limited amount of substrate is desired in the medium. In fed-batch systems, measurement of the actual substrate concentration is difficult and therefore is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of waste gases such as CO2. Batch and fed-batch fermentation are common and known in the art.

[0108] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density, with cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon or nitrogen source, is maintained at a fixed ratio, while all other parameters are adjustable. In other systems, multiple factors affecting growth can be continuously varied while the cell concentration, measured by medium turbidity, remains constant. Continuous systems attempt to maintain steady-state growth conditions. Therefore, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes and techniques for maximizing product formation rates are well known in the art of industrial microbiology.

[0109] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, such as separating the host cells from the medium by centrifugation or filtration, or, if necessary, disrupting the cells and removing the supernatant from cell debris and debris. Generally, after clarification, the protein component of the supernatant or filtrate is precipitated with a salt, such as ammonium sulfate. The precipitated protein can then be solubilized and purified by various chromatographic methods, such as ion exchange chromatography, gel filtration, etc.

[0110] In some embodiments, cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system in which the composition of the medium is set at the beginning of the fermentation and remains unchanged throughout the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. In this method, fermentation occurs without adding any components to the system. Batch fermentation is generally considered "batch" with respect to the addition of a carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes constantly until the fermentation is stopped. In a typical batch culture, cells progress through a static lag phase to a high-growth logarithmic phase and may eventually progress to a stationary phase where growth rate decreases or stops. If untreated, cells in the stationary phase eventually die. Generally, cells in the logarithmic phase are responsible for the majority of product production.

[0111] A suitable variation on the standard batch system is the "fed-batch" fermentation system. In this variation of the typical batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cellular metabolism and when a limited amount of substrate is desired in the medium. In fed-batch systems, measurement of the actual substrate concentration is difficult and therefore is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of waste gases such as CO2. Batch and fed-batch fermentation are common and known in the art.

[0112] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density, with cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon or nitrogen source, is maintained at a fixed ratio, while all other parameters are adjustable. In other systems, multiple factors affecting growth can be continuously varied while the cell concentration, measured by medium turbidity, remains constant. Continuous systems attempt to maintain steady-state growth conditions. Therefore, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes and techniques for maximizing product formation rates are well known in the art of industrial microbiology.

[0113] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, such as separating the host cells from the medium by centrifugation or filtration, or, if necessary, disrupting the cells and removing the supernatant from cell debris and debris. Generally, after clarification, the protein component of the supernatant or filtrate is precipitated with a salt, such as ammonium sulfate. The precipitated protein can then be solubilized and purified by various chromatographic methods, such as ion exchange chromatography, gel filtration, etc.

[0114] V. Target Protein The protein of interest (POI) of the present disclosure can be any endogenous or heterologous protein, or a variant of such a POI. The protein may contain one or more disulfide bridges or may be a protein whose functional form is monomeric or multimeric, i.e., a protein having a quaternary structure and composed of multiple identical (homologous) or non-identical (heterologous) subunits, wherein the POI or variant POI is preferably a protein with a desired property.

[0115] For example, in certain embodiments, a mutant or modified (recombinant) Gram-positive cell of the present disclosure produces at least about 0.1% more, at least about 0.5% more, at least about 1% more, at least about 5% more, at least about 6% more, at least about 7% more, at least about 8% more, at least about 9% more, or at least about 10% or more POI compared to its unmodified (parental or control) cell.

[0116] In certain embodiments, the mutant or engineered Gram-positive cells of the present disclosure exhibit increased specific productivity (Qp) of the POI compared to control cells. For example, detecting specific productivity (Qp) is a suitable method for assessing protein production. Specific productivity (Qp) can be determined using the following equation: "Qp = gP / gDCW·hr" (where "gP" is grams of protein produced in the tank, "gDCW" is grams of dry cell weight (DCW) in the tank, and "hr" is the fermentation time (hours) from the time of inoculation, which includes the production time and growth time).

[0117] Thus, in certain embodiments, a mutant or modified Gram-positive cell of the present disclosure comprises an increase in specific productivity (Qp) of at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to its unmodified (parental or control) cell.

[0118] In certain other embodiments, the mutant or modified Gram-positive cells comprise enhanced / increased levels of ilvE messenger RNA (mRNA) compared to control cell levels of ilvE. Suitable methods for detecting and analyzing such mRNA are generally known to those of skill in the art and include, but are not limited to, real-time quantitative PCR (RT-qPCR) analysis, RNA sequencing, etc. In certain embodiments, the mutant or modified Gram-positive cells exhibit at least about a 0.1% increase, at least about a 0.5% increase, at least about a 1.0% increase, at least about a 5% increase, or about a 10% increase in ilvE mRNA compared to unmodified (parental / control) cell levels of ilvE mRNA.

[0119] In certain other embodiments, the mutated or engineered Gram-positive cell comprises a phenotype of enhanced / increased carbon yield when expressing / producing one or more proteins of interest. In certain embodiments, the enhanced / increased carbon yield (i.e., when expressing / producing one or more proteins of interest) may be referred to as enhanced / increased carbon yield efficiency. For example, product formation by Gram-positive bacterial cells is a biological conversion process in which chemical nutrients supplied to the bacterial cells during fermentation are converted into metabolic products.

[0120] In certain other embodiments, the variant, modified or mutant Gram-positive cells exhibit increased total protein yield compared to the (unmodified / control) parent strain, where total protein yield is defined as the amount (g) of protein of interest produced per total carbohydrate equivalent of the batch and carbohydrate fed. Thus, as used herein, total protein yield (g / g) may be calculated using the following equation: "Yf=Tp / Tc" (where "Yf" is the total protein yield (g / g), "Tp" is the total amount of protein of interest produced during fermentation (g), and "Tc" is the total carbohydrate equivalents (g) of carbohydrates fed during batch and fermentation (bioreactor) runs.) In certain embodiments, the increase in total protein yield of the engineered strain (i.e., the increase compared to the control strain) is at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to the unmodified (parental) cell.

[0121] Total protein yield can also be described as carbon conversion efficiency / carbon yield, e.g., as the percentage (%) of batch and fed carbon incorporated into total protein of interest. Thus, in certain embodiments, the variant Bacillus strain comprises an increased carbon conversion efficiency (e.g., an increased percentage (%) of batch and fed carbon incorporated into total protein) compared to the (control) parent strain. In certain embodiments, the increase in carbon conversion efficiency of the engineered strain (i.e., the increase compared to the control strain) is at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to the unmodified (parent) cell.

[0122] Enhanced carbon yield, enhanced carbon yield efficiency, etc. may be assessed / determined using routine methods / techniques known to those of skill in the art. In one or more embodiments, fermentation of the engineered cells comprising enhanced carbon yield is carried out under conditions suitable for the production of a protein of interest, and the enhanced carbon yield is the result of more efficient incorporation of nutrients in the fermentation medium into the protein product, indicating an enhanced protein productivity or yield coefficient (Y).

[0123] In certain embodiments, the POI or variant POI thereof is selected from the group consisting of acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysate, and the like. lyase), endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.

[0124] Thus, in certain embodiments, the POI or variant POI thereof is an enzyme selected from the Enzyme Code (EC) EC1, EC2, EC3, EC4, EC5 or EC6.

[0125] A variety of assays for detecting and measuring the activity of intracellularly and extracellularly expressed proteins are known to those of skill in the art.

[0126] VI. Illustrative Embodiments Non-limiting embodiments of the compositions and methods disclosed herein are as follows. 1. A variant ilvE gene containing a mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene. 2. The variant ilvE gene of embodiment 1, which encodes a functional IlvE protein. 3. The variant ilvE gene of embodiment 1, wherein the mutation is a single nucleotide polymorphism (SNP) mutation in the 5'-UTR of the ilvE gene. 4. The variant ilvE gene of embodiment 4, wherein the mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5'-UTR are numbered by correspondence with the wild-type (WT) ilvE 5'-UTR sequence of SEQ ID NO: 17. 5. The variant ilvE gene of embodiment 1, comprising at least about 90%, 91%, 92%, 93%, 94% 95%, 96%, 97%, 98%, 99 or 100% sequence identity to the wild-type B. subtilis ilvE gene of SEQ ID NO: 25. 6. The variant ilvE gene of embodiment 1, wherein the ilvE 5'-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 18 and a thymine (T) at nucleotide position 73. 7. The variant ilvE gene of embodiment 2, wherein the functional IlvE protein comprises at least about 95% sequence identity with the native IlvE protein of SEQ ID NO: 15. 8. The variant ilvE gene of embodiment 1, wherein the WT ilvE gene promoter is replaced with a heterologous promoter. 9. A synthetic ilvE gene construct comprising, in the 5' to 3' direction, a wild-type ilvE gene promoter or a heterologous promoter sequence operably linked to a mutant ilvE 5'-UTR sequence operably linked to an ilvE gene CDS encoding a functional IlvE protein. 10. The genetic construct of embodiment 9, wherein the mutated ilvE 5'-UTR comprises a single nucleotide polymorphism (SNP) mutation in the 5'-UTR of the ilvE gene. 11. The genetic construct of embodiment 10, wherein the mutated ilvE 5'-UTR sequence comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 18 and a thymine (T) at nucleotide position 73. 12. A mutant Bacillus subtilis cell containing a variant ilvE gene containing a mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene. 13. The mutant cell of embodiment 12, comprising a single nucleotide polymorphism (SNP) in the 5'-UTR of the ilvE gene, wherein the mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, and the nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the wild-type (WT) ilvE 5'-UTR sequence of SEQ ID NO: 17. 14. The mutant cell of embodiment 12, which produces one or more proteins of interest. 15. The one or more proteins of interest may be acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysate (glucanase), or a combination thereof. lyase), endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, lectin, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof. 16. The mutant cell of embodiment 12, wherein the variant ilvE gene comprises at least about 90%, 91%, 92%, 93%, 94% 95%, 96%, 97%, 98%, 99, or 100% sequence identity to the wild-type B. subtilis ilvE gene of SEQ ID NO: 25. 17. The mutant cell of embodiment 12, wherein the ilvE 5'-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 18 and a thymine (T) at nucleotide position 73. 18. The mutant cell of embodiment 12, which encodes a functional IlvE protein. 19. The mutant cell of embodiment 12, wherein the WT ilvE gene promoter is replaced with a heterologous promoter. 20. The mutant cell of embodiment 19, wherein the heterologous promoter overexpresses the ilvE gene compared to the WT ilvE promoter when fermented under the same conditions. 21. The mutant cell of embodiment 14, wherein the mutant cell produces the same one or more proteins of interest and comprises an enhanced carbon yield phenotype compared to a control cell comprising a wild-type ilvE gene, when the mutant cell and the control cell are fermented under the same conditions for the production of the one or more proteins of interest. 22. The mutant cell of embodiment 12, comprising increased ilvE messenger RNA (mRNA) levels compared to control cells comprising a WT ilvE gene, when the mutant and control cells are fermented under the same conditions. 23. The mutant cell of embodiment 22, comprising increased ilvE mRNA levels relative to control cells at about 16 hours of fermentation. 24. The mutant cell of embodiment 22, comprising increased ilvE mRNA levels relative to control cells at about 24 hours of fermentation. 25. The mutant cell of embodiment 22, comprising increased ilvE mRNA levels relative to control cells at about 32 hours of fermentation. 26. A genetically modified Bacillus subtilis cell derived from a parent cell containing a wild-type (WT) ilvE gene, the modified cell comprising a variant ilvE gene comprising a mutation in the 5'-UTR sequence of the ilvE gene. 27. The modified cell of embodiment 26, comprising a single nucleotide polymorphism (SNP) mutation in the 5'-UTR sequence of the ilvE gene. 28. The modified cell of embodiment 27, wherein the SNP mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the 5'-UTR are numbered according to their correspondence with the WT ilvE 5'-UTR sequence of SEQ ID NO: 17. 29. The modified cell of embodiment 26, which produces one or more proteins of interest. 30. The one or more target proteins are selected from the group consisting of acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysate, and the like. lyase), endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, lectin, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof. 31. The mutant cell of embodiment 26, wherein the variant ilvE gene comprises at least about 90%, 91%, 92%, 93%, 94% 95%, 96%, 97%, 98%, 99, or 100% sequence identity to the wild-type B. subtilis ilvE gene of SEQ ID NO: 25. 32. The modified cell of embodiment 26, wherein the ilvE 5'-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 18 and a thymine (T) at nucleotide position 73. 33. The modified cell of embodiment 26, encoding a functional IlvE protein. 34. The modified cell of embodiment 26, wherein the WT ilvE gene promoter is replaced with a heterologous promoter. 35. The modified cell of embodiment 34, wherein the heterologous promoter overexpresses the ilvE gene compared to the WT ilvE promoter when fermented under the same conditions. 36. The modified cell of embodiment 29, wherein the modified cell produces the same one or more proteins of interest and comprises an enhanced carbon yield phenotype compared to a control cell comprising a wild-type ilvE gene, when the modified cell and the control cell are fermented under the same conditions for the production of the one or more proteins of interest. 37. The modified cell of embodiment 26, comprising increased ilvE messenger RNA (mRNA) levels compared to control cells comprising a WT ilvE gene, when the modified cell and the control cell are fermented under suitable conditions. 38. The modified cell of embodiment 37, comprising increased ilvE mRNA levels compared to control cells at about 16 hours of fermentation. 39. The modified cell of embodiment 37, comprising increased ilvE mRNA levels compared to control cells at about 24 hours of fermentation. 40. The modified cell of embodiment 37, comprising increased ilvE mRNA levels relative to control cells at about 32 hours of fermentation. 41. A method for increasing ilvE messenger RNA (mRNA) levels in recombinant Bacillus subtilis cells, comprising: (a) obtaining parent B. subtilis cells containing a wild-type (WT) ilvE gene and replacing the WT ilvE gene with a variant ilvE gene containing a mutation in the 5'-untranslated region (5'-UTR) of the ilvE gene; and (b) fermenting the parent and modified cells under the same conditions for at least about 16 hours, wherein the modified cells contain increased levels of ilvE mRNA compared to the parent cells. 42. The method of embodiment 41, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5'-UTR sequence. 43. The method of embodiment 42, wherein the SNP mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the wild-type (WT) ilvE 5'-UTR sequence of SEQ ID NO: 17. 44. A method for increasing ilvE messenger RNA (mRNA) levels in modified Bacillus subtilis cells, comprising: (a) obtaining a parent B. subtilis containing a wild-type (WT) ilvE gene and mutating the 5'-untranslated region (5'-UTR) of the WT ilvE gene to obtain modified B. subtilis cells containing a mutation in the 5'-untranslated region (5'-UTR) of the ilvE gene; and (b) fermenting the parent and modified cells under the same conditions for at least about 16 hours, wherein the modified cells contain increased levels of ilvE mRNA compared to the parent cells. 45. The method of embodiment 44, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5'-UTR sequence. 46. ​​The method of embodiment 45, wherein the SNP mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the WT ilvE 5'-UTR sequence of SEQ ID NO: 17. 47. A method for increasing ilvE messenger RNA (mRNA) levels in modified Bacillus subtilis cells, comprising: (a) obtaining a parent B. subtilis containing a wild-type (WT) ilvE gene and mutating the 5'-untranslated region (5'-UTR) of the WT ilvE gene to obtain modified B. subtilis cells containing a variant ilvE gene; and (b) fermenting the parent and modified cells under the same conditions for at least about 16 hours, wherein the modified cells contain increased levels of ilvE mRNA compared to the parent cells. 48. The method of embodiment 47, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5'-UTR sequence. 49. The method of embodiment 48, wherein the SNP mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the WT ilvE 5'-UTR sequence of SEQ ID NO: 17. 50. A method for increasing the carbon yield of a heterologous protein produced in an engineered Bacillus subtilis cell, comprising: (a) obtaining or constructing a parent B. subtilis cell that produces a heterologous protein of interest (POI) and replacing the wild-type (WT) ilvE gene with a variant ilvE gene; and (b) fermenting the parent and engineered cells under the same conditions suitable for production of the POI for at least about 16 hours, wherein the engineered cells comprise an increased carbon yield efficiency of the POI produced compared to the parent cell. 51. The method of embodiment 50, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5'-UTR sequence. 52. The method of embodiment 51, wherein the SNP mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the WT ilvE 5'-UTR sequence of SEQ ID NO: 17. 53. A method for increasing the carbon yield of a heterologous protein expressed / produced in an engineered Bacillus subtilis cell, comprising: (a) obtaining a parent B. subtilis containing a wild-type (WT) ilvE gene and mutating the 5'-untranslated region (5'-UTR) of the WT ilvE gene to obtain an engineered B. subtilis cell containing a variant ilvE gene; and (b) fermenting the parent and engineered cells under the same conditions for at least about 16 hours, wherein the engineered cells comprise an increased carbon yield efficiency of the POI produced compared to the parent cells. 54. The method of embodiment 53, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5'-UTR sequence. 55. The method of embodiment 54, wherein the SNP mutation in the 5'-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the WT ilvE 5'-UTR sequence of SEQ ID NO: 17. 56. The method of any one of embodiments 41-55, wherein the variant ilvE gene comprises at least about 90%, 91%, 92%, 93%, 94% 95%, 96%, 97%, 98%, 99, or 100% sequence identity to the wild-type B. subtilis ilvE gene of SEQ ID NO: 25. 57. The method of any one of embodiments 41-55, wherein the mutated ilvE 5'-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 18 and a thymine (T) at nucleotide position 73. 58. The method of any one of embodiments 41 to 55, wherein the WT ilvE gene promoter is replaced with a heterologous promoter. 59. The method of embodiment 58, wherein the heterologous promoter overexpresses the ilvE gene compared to the WT ilvE promoter when fermented under the same conditions. 60. Heterologous POIs include acetylesterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysate (glucanase), and lyase), endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, lectin, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectinomerase 56. The method of any one of embodiments 41, 44, 47, 50 and 55, wherein the enzyme is selected from the group consisting of acetyl esterase, pectolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamno-galacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase and combinations thereof. 61. The method of any one of embodiments 41-55, wherein the variant ilvE gene encodes a functional IlvE protein. 62. The method of embodiment 41 or embodiment 44, wherein the modified cells comprise increased levels of ilvE messenger RNA (mRNA) compared to control cells when fermented under the same conditions for at least about 16 hours. 63. The method of embodiment 41 or embodiment 44, wherein the modified cells comprise increased ilvE mRNA levels compared to control cells when fermented under the same conditions for at least about 24 hours. 64. The method of embodiment 41 or embodiment 44, wherein the modified cells comprise increased levels of ilvE messenger RNA (mRNA) compared to control cells when fermented under the same conditions for at least about 32 hours. 65. The method of embodiment 50 or embodiment 53, wherein the modified cells comprise an increased carbon yield efficiency compared to the parent cells when fermented under the same conditions for at least about 16 hours. 66. The method of embodiment 50 or embodiment 53, wherein the modified cells comprise an increased carbon yield efficiency compared to the parent cells when fermented under the same conditions for at least about 24 hours. 67. The method of embodiment 50 or embodiment 53, wherein the modified cells comprise an increased carbon yield efficiency compared to the parent cells when fermented under the same conditions for at least about 32 hours. [Example]

[0127] Certain embodiments of the present disclosure may be further understood in light of the following examples, which should not be construed as limiting. Modifications to materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989).

[0128] Example 1 Construction of Bacillus protease reporter strains As outlined above, Applicants have identified mutant B. subtilis cells (strain CZ437) with an enhanced protein production phenotype. In particular, Applicants performed next-generation sequencing (NGS) on the mutant CZ437 strain to further characterize the observed enhanced protein productivity phenotype and identified an unexpected SNP in the B. subtilis ilvE 5'-UTR sequence. For example, the DNA sequences of the WT ilvE 5'-UTR (SEQ ID NO: 17) and mutant ilvE 5'-UTR (SEQ ID NO: 18) are shown in Figures 1A and 1B, respectively. As shown in Figure 1, the WT ilvE 5'-UTR sequence (SEQ ID NO: 17; Figure 1A) contains a cytosine (C) at nucleotide position 73 ( + 73C), the mutated ilvE 5′-UTR sequence (SEQ ID NO: 18; Figure 1B) contains a thymine (T) at nucleotide position 73 ( + 73T). In this example, Applicants constructed recombinant B. subtilis strains expressing a heterologous reporter (GG36) protein to assess the association of the mutant ilvE 5'-UTR sequence (SEQ ID NO: 18) with the enhanced protein production phenotype identified in the mutant B. subtilis strain CZ437. In particular, these DNA fragments described herein were assembled using standard molecular biology techniques and used as templates to develop linear DNA expression cassettes for integration into the B. subtilis strains described herein.

[0129] A. Construction of reporter protein expression cassettes. The construction of reporter protein cassettes was carried out as follows: the first (1) containing the 5'skfA flanking region (FR) sequence of B. subtilis (5'skfA FR; SEQ ID NO: 9) stThe .) DNA fragment was operably linked to an expression cassette comprising the upstream (5') B. subtilis P2 promoter operably linked to a DNA sequence comprising a wild-type B. subtilis aprE 5'-untranslated region (5'-UTR; SEQ ID NO: 1), operably linked to a DNA sequence encoding a wild-type B. subtilis aprE signal sequence (SEQ ID NO: 2), operably linked to a DNA sequence encoding a variant B. lentus propeptide sequence (SEQ ID NO: 4), operably linked to a DNA sequence encoding a mature (GG36) subtilisin reporter (SEQ ID NO: 6), operably linked to a 3' skfH FR sequence (3' skfH FR; SEQ ID NO: 10).

[0130] A second (2) gene containing the 5' yhfN flanking region (FR) sequence located in the chromosomal region of the B. subtilis 5' aprE flanking region (FR) sequence (5' aprE FR; SEQ ID NO: 11). nd The GG36 subtilisin reporter expression cassette was further ligated to the B. subtilis alanine racemase (alrA) gene (SEQ ID NO: 12) and 3' aprE FR sequence (3' aprE FR; SEQ ID NO: 13).

[0131] Mutagenesis of the B. ilvE transcription leader sequence As briefly described above, a cytosine (C) to thymine (T) mutation at position 73 of the ilvE 5'-UTR was introduced into the genome of B. subtilis using random strain mutagenesis. st and 2 nd The cassette was integrated into a B. subtilis strain containing a C to T (SNP) mutation at position 73 in the ilvE 5′-UTR (reporter strain A; 73T) and an isogenic B. subtilis strain containing the wild-type ilvE 5′-UTR (reporter strain B; 73C).

[0132] More specifically, B. subtilis strain A (mutated ilvE 5'-UTR; SEQ ID NO: 18) and isogenic B. subtilis strain B (WT ilvE 5'-UTR; SEQ ID NO: 17) were fermented in large-scale (approximately 14 L) fermentors using standard fermentation conditions. As shown in Table 1 below, the relative improvement in carbon yield of mutant B. subtilis reporter strain A is significantly enhanced compared to the isogenic B. subtilis reporter strain.

[0133] [Table 1]

[0134] C. Replacement of the wild-type ilvE promoter with the heterologous Hbs promoter Using conventional molecular biology techniques, two cassettes for expression of the ilvE amino acid aminotransferase were constructed. More specifically, the ilvE expression cassette contains an upstream (5') wild-type (WT) B. subtilis hbs promoter (Phbs) region sequence (SEQ ID NO:26) operably linked to a downstream (3') DNA sequence containing either the WT ilvE transcription leader (WT 5'-UTR; SEQ ID NO:17) or a mutant ilvE transcription leader (mutant 5'-UTR; SEQ ID NO:18), operably linked to a downstream (3') WT ilvE gene CDS (SEQ ID NO:14). Thus, the hbs promoter region (SEQ ID NO:26) drives expression of both cassettes. The DNA sequence of the WT ilvE gene cassette is set forth in SEQ ID NO:27, and the sequence of the mutant ilvE gene cassette is set forth in SEQ ID NO:28. More specifically, the cassette (SEQ ID NO: 27 or SEQ ID NO: 28) was integrated into the spoIIIAA genomic locus of a parent B. subtilis strain containing two GG36 reporter protein cassettes.

[0135] B. subtilis strains overexpressing the ilvE gene under the control of the hbs promoter and containing either the WT ilvE 5'-UTR (strain CZ477) or the mutant ilvE 5'-UTR (strain CZ488), as outlined below in Table 2, were fermented in large-scale (approximately 14 L) bioreactors under standard fermentation conditions and compared with strain CZ450 containing the WT ilvE promoter and the WT ilvE 5'-UTR. As shown in Table 2, both strains overexpressing the ilvE gene showed increased carbon efficiency compared to the strain with the WT ilvE promoter.

[0136] [Table 2]

[0137] Thus, in one or more embodiments of the present disclosure, as shown in this Example, a highly expressed heterologous promoter region sequence (e.g., the hbs promoter) may be used to overexpress a variant ilvE gene (or an ilvE gene expression construct thereof) comprising a WT ilvE 5'-UTR sequence operably linked to a downstream WT ilvE gene CDS (or a variant ilvE gene CDS thereof encoding a functional ilvE protein) and / or to overexpress a variant ilvE gene (or an ilvE gene expression construct thereof) comprising a mutated ilvE 5'-UTR sequence operably linked to a downstream WT ilvE gene CDS (or a variant ilvE gene CDS thereof encoding a functional ilvE protein). Accordingly, such methods, genetic elements, expression constructs, engineered cells comprising an enhanced protein production phenotype, engineered cells comprising an enhanced carbon yield phenotype, etc., when cultured under suitable conditions, are particularly useful for the expression / production of proteins of interest.

[0138] Example 2 Real-time quantitative PCR RNA analysis In this example, time-course samples from B. subtilis reporter strains A (mutated ilvE 5'-UTR) and B (WT ilvE 5'-UTR) were used for real-time quantitative PCR (RT-qPCR) analysis. For example, samples were collected after 8, 16, 24, and 32 hours of fermentation, and total RNA was extracted. Specifically, the extracted RNA samples were treated with DNase-I to remove genomic DNA from the samples, and then cDNA was synthesized using the Transcriptor First Strand cDNA Synthesis Kit (Roche). Subsequently, 1000-fold diluted cDNA from each sample was used as a template for qPCR, and the ftsY gene was used as a housekeeping gene for data normalization. For example, sequence-specific ilvE forward (SEQ ID NO: 19) and reverse (SEQ ID NO: 20) primers and an ilvE probe (SEQ ID NO: 21) were used to amplify sequences within the ilvE gene. Similarly, sequence-specific ftsY forward (SEQ ID NO: 22) and reverse (SEQ ID NO: 23) primers and an ftsY probe (SEQ ID NO: 24) were used to amplify sequences within the ftsY gene.

[0139] In particular, the RT qPCR time course experimental data is shown in Figure 4, and 2ΔΔC TThe log fold change between the housekeeping ftsY gene and the ilvE gene was calculated using the method (Livak and Schmittgen, 2001). As shown in Figure 4, the bars and values ​​represent the fold change of ilvE mRNA in B. subtilis strain A (mutated ilvE 5'-UTR) relative to the isogenic B. subtilis strain B (WT ilvE 5'-UTR). Notably, except for the initial time point after 8 hours of fermentation (i.e., the amount of ilvE mRNA changed by only 0.84-fold at this time point), the amount of ilvE mRNA was significantly increased by 2.91, 1.93, and 1.79-fold at 16, 24, and 32 hours of fermentation in B. subtilis strain A (mutated ilvE 5'-UTR) compared to isogenic B. subtilis strain B (WT ilvE 5'-UTR), respectively (Figure 4). Based on the above, the results presented and described herein demonstrate that the C-to-T SNP identified herein (SEQ ID NO: 18) significantly increases the level of ilvE mRNA.

[0140] References International Publication No. 2002 / 14490 Brochure International Publication No. 2003 / 083125 Brochure International Publication No. 2020 / 112609 Brochure Ausubel et al., “Current Protocols in Molecular Biology”, published by Greene Publishing Assoc. and Wiley-Interscience (1987). Belitsky and Sonenshein, “Roadblock repression of transcription by Bacillus subtilis CodY”, J.Mol.Biol.26;411(4):729-743, 2011. Berger et al.,“Methionine Regeneration and Aminotransferases in Bacillus subtilis,Bacillus cereus,and Bacillus anthracis”,J.Bacteriol.185(8):2418-2431,2003. Caspers et al.,“Improvement of Sec-dependent secretion of a heterologous model protein in Bacillus subtilis by saturation mutagenesis of the N-domain of the AmyE signal peptide”,Appl.Microbiol.Biotechnol.,86(6):1877-1885,2010. Earl et al.,“Ecology and genomics of Bacillus subtilis”,Trends in Microbiology.,16(6):269-275,2008. Livak and Schmittgen,“Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2 - ΔΔC T Method”,Methods.Vol.25,pp.402-408,2001. Maeder et al.,“Transcriptional organization and posttranscriptional regulation of the Bacillus subtilis branched-chain amino acid biosynthesis genes”,J Bacteriol.186(8):2240-2252,2004. Olempska-Beer et al.,“Food-processing enzymes from recombinant microorganisms--a review”’ Regul.Toxicol.Pharmacol.,45(2):144-158,2006. Sambrook et al.,“Molecular Cloning:A Laboratory Manual” Cold Spring Harbor Laboratory:Cold Spring Harbor,N.Y.(1989),(2001) and (2012). Van Dijl and Hecker,“Bacillus subtilis:from soil bacterium to super-secreting cell factory”,Microbial Cell Factories,12(3),2013.

Claims

1. A mutated ilvE 5'-untranslated region (5'-UTR) nucleic acid sequence comprising SEQ ID NO:

18.

2. 1. A variant ilvE gene comprising a single nucleotide polymorphism (SNP) mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene, wherein the ilvE gene comprises at least 90% identity with a wild-type B. subtilis ilvE gene of SEQ ID NO: 2 and encodes a native IlvE protein, and the SNP in the 5′-UTR of the ilvE gene is a cytosine (C) to thymine (T) mutation at position 73 of the ilvE 5′-UTR sequence shown as SEQ ID NO:

18.

3. A synthetic ilvE gene comprising a wild-type (WT) ilvE gene promoter or a heterologous gene promoter operably linked to a mutated ilvE 5'-UTR sequence comprising SEQ ID NO: 18 operably linked to a wild-type (WT) ilvE gene CDS.

4. 1. A mutant Bacillus subtilis cell comprising a single nucleotide polymorphism (SNP) in the 5'-untranslated region (5'-UTR) of the ilvE gene, wherein the SNP is a cytosine (C) to thymine (T) mutation at position 73, wherein nucleotide positions of the ilvE 5'-UTR are numbered according to their correspondence with the wild-type (WT) ilvE 5'-UTR of SEQ ID NO:

17.

5. 1. A genetically modified Bacillus subtilis cell derived from a parent cell containing a wild-type (WT) ilvE gene, the modified cell comprising an introduced synthetic ilvE gene comprising a heterologous gene promoter operably linked to a mutated ilvE 5′-UTR sequence comprising SEQ ID NO: 18 operably linked to the WT ilvE gene CDS.

6. The mutant cell of claim 4, which produces one or more proteins of interest.

7. The modified cell of claim 5 , which produces one or more proteins of interest.

8. The modified cell of claim 5 , wherein the heterologous gene promoter overexpresses the ilvE gene.

9. 1. A method for increasing ilvE messenger RNA (mRNA) levels in an engineered Bacillus subtilis cell, comprising: (a) obtaining a parent B. subtilis cell containing a wild-type (WT) ilvE gene and replacing the WT ilvE gene with a synthetic ilvE gene comprising a heterologous gene promoter operably linked to a mutated ilvE 5'-UTR sequence comprising SEQ ID NO: 18 operably linked to a WT ilvE gene CDS; (b) fermenting the parental and modified cells under the same conditions for at least about 16 hours; Including, The modified cell comprises an increased level of ilvE mRNA compared to the parent cell.

10. 10. The method of claim 9, wherein the synthetic ilvE gene comprises at least 90% identity with the ilvE gene of SEQ ID NO: 25 and comprises a thymine (T) at nucleotide position 73 of the ilvE 5'-UTR.

11. 1. A method for increasing the carbon yield of a heterologous protein produced in an engineered Bacillus subtilis cell, comprising: (a) obtaining or constructing a parent B. subtilis cell that produces a heterologous protein of interest (POI), and introducing into said cell a synthetic ilvE gene comprising a heterologous gene promoter operably linked to a mutated ilvE 5′-UTR sequence comprising SEQ ID NO: 18 operably linked to a WT ilvE gene CDS; (b) fermenting the parental and modified cells for at least about 16 hours under the same conditions suitable for producing the POI; Including, The method, wherein the modified cell comprises an increased carbon yield efficiency of the POI produced compared to the parent cell.

12. 12. The method of claim 11, wherein the synthetic ilvE gene comprises at least 90% identity with the ilvE gene of SEQ ID NO: 25 and comprises a thymine (T) at nucleotide position 73 of the ilvE 5'-UTR.