Compositions and methods for enhancing protein production in bacillus cells
By introducing SNP mutations into the ilvE gene 5'-UTR of Bacillus species, the variant ilvE gene was constructed, which solved the problem of low protein yield in the prior art and achieved a significant improvement in protein productivity in the industrial biotechnology environment.
Patent Information
- Application Number
- CN202380073333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-24
- Filing Date
- 2023-10-13
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively improve the protein yield of Bacillus species in industrial biotechnology environments, especially in heterologous protein expression.
The variant ilvE gene was constructed by introducing single nucleotide polymorphism (SNP) mutations into the 5'-untranslated region (5'-UTR) of the ilvE gene, which was used to construct recombinant Bacillus strains to enhance the productivity of the protein of interest.
It significantly improves the carbon yield and productivity of the target protein in Bacillus cells, enhances the ilvE messenger RNA level, and thus improves the expression efficiency of heterologous proteins.
Smart Images

Figure CN120187846A_ABST
Abstract
Description
Field of the Invention
[0001] The present disclosure generally relates to the fields of bacteriology, microbiology, genetics, molecular biology, enzymology, industrial protein production, and the like. Certain embodiments of the present disclosure relate to Bacillus sp. strains having an enhanced protein productivity phenotype, compositions and methods for constructing recombinant Bacillus sp. strains, and the like.
[0002] Cross - Reference to Related Applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 380,706, filed on October 24, 2022, which is hereby incorporated by reference in its entirety.
[0004] Incorporation by Reference of Sequence Listing
[0005] The content of the electronically - submitted sequence listing of the text file named "NB41976 - WO - PCT_SequenceListing.xml", created on October 2, 2023, and having a size of 37 KB, is hereby incorporated by reference in its entirety. Background of the Invention
[0006] Gram - positive bacteria such as Bacillus subtilis, Bacillus licheniformis, Bacillus amyloliquefaciens, etc. are often used as microbial factories for producing industrially relevant proteins due to their excellent fermentation characteristics and high yields (e.g., up to 25 grams per liter of culture; Van Dijl and Hecker, 2013). For example, Bacillus sp. host cells are well - known for producing enzymes (e.g., amylase, cellulase, mannanase, pectate lyase, protease, pullulanase, etc.) required for the food, textile, laundry detergent, medical device cleaning, pharmaceutical industries, and the like. Since these non - pathogenic Gram - positive bacteria produce proteins that are completely free of toxic by - products (e.g., lipopolysaccharide; LPS, also known as endotoxin), they have obtained the "Qualified Presumption of Safety" (QPS) status from the European Food Safety Authority (EFSA), and many of their products have obtained the "Generally Recognized As Safe" (GRAS) status from the U.S. Food and Drug Administration (Olempska - Beer et al., 2006; Earl et al., 2008; Caspers et al., 2010).
[0007] Thus, the production of proteins (e.g., enzymes, antibodies, receptors, etc.) via microbial host cells is of particular significance in the field of biotechnology. Similarly, the optimization of Bacillus host cells for the production and secretion of one or more proteins of interest is highly relevant, especially in the context of industrial biotechnology, where minor improvements in protein productivity, etc., can be significant when the protein is produced in large industrial quantities. For example, the expression of many heterologous proteins can still be challenging and unpredictable in terms of productivity, etc. As described herein, the present disclosure relates to a highly desired and unmet need for obtaining and constructing Bacillus species cells (e.g., protein-producing hosts) with enhanced protein production capabilities. Summary of the Invention
[0008] As generally described herein, certain embodiments of the present disclosure particularly relate to variant ilvE genes, variant ilvE gene 5'-untranslated region (5'-UTR) sequences, mutant Bacillus strains comprising variant ilvE gene sequences, recombinant (genetically modified) Bacillus strains comprising variant ilvE gene sequences, mutant and / or recombinant Bacillus strains comprising variant ilvE gene sequences and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains comprising variant ilvE gene sequences, expression cassettes encoding proteins of interest, methods and compositions for culturing recombinant Bacillus strains comprising variant ilvE gene sequences for enhanced production of proteins of interest, etc. More particularly, as described herein, the novel mutant and / or recombinant Bacillus cells of the present disclosure are particularly useful for the production of proteins of interest when cultured under suitable conditions.
[0009] Accordingly, certain embodiments of the present disclosure relate to variant ilvE genes that contain single nucleotide polymorphism (SNP) mutations in the 5′-untranslated region (5′-UTR) of the ilvE gene. In one or more embodiments, the variant ilvE genes of the present disclosure encode functional IlvE proteins. In certain other embodiments, the present disclosure provides synthetic ilvE gene constructs that contain, in the 5′ to 3′ direction, a heterologous promoter sequence operably linked to a mutant ilvE 5′-UTR sequence, which is operably linked to an ilvE gene coding sequence (CDS) that encodes a functional IlvE protein. In other embodiments, the present disclosure relates to mutant Bacillus subtilis strains that contain variant ilvE genes having SNPs in the 5′-untranslated region (5′-UTR) of the ilvE gene. In one or more other embodiments, the mutant and / or recombinant Bacillus subtilis cells of the present disclosure produce one or more proteins of interest. In other related embodiments, when the mutant and control cells are fermented under suitable conditions, the mutant and / or recombinant Bacillus subtilis cells that produce one or more proteins of interest have an enhanced carbon yield phenotype relative to control Bacillus subtilis cells that produce the same one or more proteins of interest and contain the wild-type ilvE gene. In certain related one or more embodiments, the present disclosure provides mutant and / or modified Bacillus cells (strains) that contain an enhanced protein productivity phenotype, mutant and / or modified cells that contain enhanced / increased ilvE messenger RNA (mRNA) levels, mutant and / or modified cells that have an enhanced / increased carbon yield (carbon yield efficiency) of the heterologous protein produced, etc.
[0010] In other embodiments, the present disclosure relates to methods for increasing the level of ilvE messenger RNA (mRNA) in recombinant Bacillus subtilis cells, which generally comprise obtaining parental Bacillus subtilis cells having a wild-type (WT) ilvE gene and replacing the WT ilvE gene with a variant ilvE gene, wherein the variant ilvE gene comprises an SNP mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene, and fermenting the parental cells and the recombinant cells under suitable conditions for at least about sixteen hours, wherein the recombinant cells comprise an increased level of ilvE mRNA compared to the parental cells. In other embodiments, the present disclosure relates to methods for increasing the carbon yield of a heterologous protein produced in recombinant Bacillus subtilis cells, which comprise obtaining or constructing parental Bacillus subtilis cells that produce a heterologous protein of interest (POI) and comprise a WT ilvE gene, and replacing the WT ilvE gene with a variant ilvE gene that comprises an SNP mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene, and fermenting the parental cells and the recombinant cells under suitable conditions for producing the POI for at least about sixteen hours, wherein the recombinant cells have an increased carbon yield efficiency of the POI produced compared to the parental cells. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 Shows the wild-type (WT) Bacillus subtilis ilvE 5′-UTR ( Figure 1 A) and the mutant Bacillus subtilis ilvE 5′-UTR ( Figure 1 B) DNA sequences. For example, as Figure 1 shown in A, the WT ilvE 5′-UTR sequence contains a cytosine (C) at nucleotide position 73, and the mutant ilvE 5′-UTR, as Figure 1 shown in B, contains a thymine (T) at nucleotide position 73, where the and nucleotides at this position are presented in bold and double-underlined in Figure 1 . Thus, as Figure 1 presented, the WT Bacillus subtilis ilvE 5′-UTR sequence contains SEQ ID NO:17 ( Figure 1 A), and the mutant ilvE 5′-UTR sequence contains SEQ ID NO:18 ( Figure 1 B).
[0012] Figure 2 Presents a schematic diagram showing nucleotide position 73 ( + 73) of the mutant ilvE 5′-UTR sequence (SEQ ID NO:18). In particular, Figure 1The WT (SEQ ID NO:17) and mutant (SEQ ID NO:18) ilvE 5′-UTR sequences presented are numbered in the 5′ to 3′ direction, where nucleotide position 1 ( + 1) is the first nucleotide position of the 5′ untranslated region identified as the transcription start site. In particular, as Figure 2 shown, the C>T mutation at position 73 in the mutant ilvE 5′-UTR is near CodY binding motif 2 and putative binding motif 5.
[0013] Figure 3 shows the nucleic acid sequences of the wild-type (WT) ilvE promoter ( Figure 3 A; SEQ ID NO:16), WT ilvE 5′-UTR ( Figure 3 B; SEQ ID NO:17), WT ilvE gene coding sequence ( Figure 3 C; SEQ ID NO:14) and WT ilvE gene ( Figure 3 D; SEQ IDNO:25). As presented in the 5′ to 3′ direction in Figure 3 D, the WT ilvE gene (SEQ ID NO:25) contains the WT ilvE promoter (italicized nucleotides; SEQ ID NO:16), WT ilvE 5′-UTR (bold nucleotides; SEQ ID NO:17) and the nucleotides of the WT ilvE gene CDS ( Underlined ; SEQ ID NO:14). Similarly, Figure 3 shows the amino acid sequence of the native (mature) IlvE protein encoded by the WT ilvE gene CDS (SEQ ID NO:14) ( Figure 3 E; SEQ ID NO:15).
[0014] Figure 4 Presents data from a real-time qPCR (RT qPCR) analysis of Bacillus strain A (containing the mutant ilvE 5′-UTR) relative to Bacillus strain B (containing the WT ilvE 5′-UTR), as described in Example 2 below. In particular, Figure 4 the data presented shows the results of RT qPCR in a time-course experiment, where the 2-ΔΔC T method (Livak and Schmittgen, 2001) was used to calculate the log fold change between the housekeeping ftsY gene and the ilvE gene. As Figure 4As shown, the bars and values represent the fold change in ilvE mRNA of Bacillus subtilis strain A (mutated ilvE 5′-UTR) compared to isogenic Bacillus subtilis strain B (WT ilvE 5′-UTR) at 16, 24, and 32 hour fermentation time points.
[0015] Biological Sequence Description
[0016] SEQ ID NO:1 is a nucleotide (DNA) sequence containing the wild-type Bacillus subtilis aprE 5′-UTR sequence
[0017] SEQ ID NO:2 is the wild-type DNA sequence encoding the native Bacillus subtilis aprE signal sequence.
[0018] SEQ ID NO:3 is the amino acid sequence of the native Bacillus subtilis aprE signal sequence encoded by SEQ ID NO:2.
[0019] SEQ ID NO:4 is the DNA sequence encoding the native Bacillus clausii GG36 Pro region sequence.
[0020] SEQ ID NO:5 is the amino acid sequence of the native Bacillus clausii GG36 Pro region sequence encoded by SEQ ID NO:4.
[0021] SEQ ID NO:6 is the wild-type DNA sequence encoding the native Bacillus clausii protease (Eraser11).
[0022] SEQ ID NO:7 is the amino acid sequence of the native Bacillus clausii protease (Eraser11) encoded by SEQ ID NO:6.
[0023] SEQ ID NO:8 is a DNA sequence containing the Bacillus amyloliquefaciens BPN′ terminator sequence.
[0024] SEQ ID NO:9 is a DNA sequence containing the Bacillus subtilis 5′skfA flanking region (FR) sequence.
[0025] SEQ ID NO:10 is a DNA sequence containing the Bacillus subtilis 3′skfA FR sequence.
[0026] SEQ ID NO:11 is a DNA sequence containing the Bacillus subtilis 5′aprE FR sequence.
[0027] SEQ ID NO:12 is a DNA sequence containing the wild-type Bacillus subtilis alrA gene.
[0028] SEQ ID NO:13 is a DNA sequence containing the 3′ aprE FR sequence of Bacillus subtilis.
[0029] SEQ ID NO:14 is a DNA sequence containing the coding sequence (CDS) of the wild-type Bacillus subtilis IlvE gene.
[0030] SEQ ID NO:15 is the amino acid sequence of the native Bacillus subtilis IlvE protein encoded by SEQ ID NO:14.
[0031] SEQ ID NO:16 is a DNA sequence containing the wild-type Bacillus subtilis IlvE promoter.
[0032] SEQ ID NO:17 is a DNA sequence containing the wild-type Bacillus subtilis IlvE 5′-UTR sequence.
[0033] SEQ ID NO:18 is a DNA sequence containing the mutant Bacillus subtilis IlvE 5′-UTR sequence.
[0034] SEQ ID NO:19 is the Bacillus subtilis IlvE-forward (FW) primer (DNA) sequence.
[0035] SEQ ID NO:20 is the Bacillus subtilis IlvE-reverse (RV) primer (DNA) sequence.
[0036] SEQ ID NO:21 is a synthetic DNA probe named “IlvE-BBQ”.
[0037] SEQ ID NO:22 is the Bacillus subtilis ftsY-forward (FW) primer (DNA) sequence.
[0038] SEQ ID NO:23 is the Bacillus subtilis ftsY-reverse (RV) primer (DNA) sequence DNA.
[0039] SEQ ID NO:24 is a synthetic DNA probe named “ftsY-BBQ”.
[0040] SEQ ID NO:25 is a DNA sequence containing the wild-type (WT) Bacillus subtilis IlvE gene, which contains (in the 5′ to 3′ direction) the WT ilvE promoter (SEQ ID NO:16) operably linked to the WT ilvE 5′-UTR (SEQ ID NO:17), and the WT ilvE 5′-UTR is operably linked to the WT ilvE gene CDS (SEQ ID NO:14).
[0041] SEQ ID NO:26 is the DNA sequence of the natural Bacillus subtilis Hbs promoter region sequence.
[0042] SEQ ID NO:27 is a synthetic DNA construct that comprises the upstream (5′) Hbs promoter operably linked to the wild-type IlvE gene.
[0043] SEQ ID NO:28 is a synthetic DNA construct that comprises the upstream (5′) Hbs promoter operably linked to the variant IlvE gene. Detailed Description
[0044] As described herein, certain embodiments of the present disclosure relate to compositions and methods for enhancing protein production in mutant / recombinant Bacillus species (host) cells / strains. More particularly, as set forth below and further described in the examples below, the recombinant Bacillus cells of the present disclosure are particularly useful for enhancing the production of a protein of interest when cultured under suitable conditions. Accordingly, certain embodiments of the present disclosure particularly provide mutant Bacillus strains comprising a variant ilvE gene sequence, recombinant (genetically modified) Bacillus strains comprising a variant ilvE gene sequence, mutant and / or recombinant Bacillus strains comprising a variant ilvE gene sequence and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains comprising a variant ilvE gene sequence, expression cassettes encoding a protein of interest, methods and compositions for culturing recombinant Bacillus strains comprising a variant ilvE gene sequence for enhancing the production of a protein of interest, and the like.
[0045] I. Definitions
[0046] In view of the recombinant (modified) cells and methods of the present disclosure described herein, the following terms and phrases are defined. Terms not defined herein should conform to their ordinary meaning as used in the art.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the compositions and methods of the present invention pertain. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the compositions and methods of the present invention, representative illustrative methods and materials are now described. All publications and patents cited herein are incorporated by reference in their entirety.
[0048] It should be further noted that claims can be drafted to exclude any optional elements. Accordingly, this statement is intended to serve as a prerequisite (or condition) for the use of exclusive terms such as "alone", "only", "exclude", "do not include", etc. in connection with the recitation of claim elements or for the use of "negative" limitations.
[0049] Upon reading this disclosure, it will be apparent to those skilled in the art that each of the individual embodiments described and illustrated herein has discrete components and features that can be readily separated from or combined with the features of any one of several other embodiments without departing from the scope or spirit of the compositions and methods of the invention described herein. Any recited method can be carried out in the order of the recited events or in any other order that is logically possible.
[0050] As used herein, the term "recombinant" or "non-natural" refers to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic alteration or has been modified by the introduction of a heterologous nucleic acid molecule, or to a cell (e.g., a microbial cell) that has been altered such that the expression of a heterologous or endogenous nucleic acid molecule or gene can be controlled. Recombinant also refers to a cell that is derived from a non-natural cell or to the progeny of a non-natural cell that has one or more such modifications. Genetic alterations include, for example, modifications that introduce an expressible nucleic acid molecule encoding a protein, or the addition, deletion, substitution, or other functional alteration of other nucleic acid molecules in the genetic material of the cell. For example, a recombinant cell can express the same or a homologous form of a gene or other nucleic acid molecule that is not found in a native (wild-type) cell (e.g., a fusion protein or a chimeric protein), or can provide an altered pattern of endogenous gene expression, such as overexpression, underexpression, minimal expression, or no expression at all. "Recombination" or generating a "recombined" nucleic acid generally is the assembly of two or more nucleic acid fragments, wherein the assembly results in a chimeric gene.
[0051] As used herein, the phrases "Gram-positive bacterium", "Gram-positive cell", "Gram-positive bacterial strain", and / or "Gram-positive bacterial cell" have the same meaning as used in the art. For example, Gram-positive bacterial cells include all strains of the phyla Actinobacteria and Firmicutes. In certain embodiments, such Gram-positive bacteria belong to the classes Bacilli, Clostridia, and Mollicutes.
[0052] As used herein, "Bacillus" includes all species within the genus "Bacillus" known to those of skill in the art, including but not limited to Bacillus subtilis, Bacillus licheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus clausii, Bacillus halodurans, Bacillus megaterium, Bacillus coagulans, Bacillus circulans, Bacillus lautus, and Bacillus thuringiensis. It should be recognized that the genus Bacillus is constantly undergoing taxonomic reorganization. Thus, the genus is intended to include species that have been reclassified, including but not limited to organisms such as Bacillus stearothermophilus (which is now named "Geobacillus stearothermophilus").
[0053] As used herein, the "wild-type Bacillus subtilis ilvE promoter" sequence (abbreviated as "WT ilvE pro") comprises the nucleotide sequence shown in SEQ ID NO:16, as Figure 3 shown in A.
[0054] As used herein, the "wild-type Bacillus subtilis ilvE 5'-untranslated region" sequence (abbreviated as "WT ilvE5'-UTR") comprises the nucleotide sequence shown in SEQ ID NO:17, as Figure 3 shown in B.
[0055] As used herein, the "wild-type Bacillus subtilis ilvE gene coding sequence (abbreviated as "gene CDS, CDS, or ORF") comprises the nucleotides shown in SEQ ID NO:14, as Figure 3 shown in C.
[0056] As used herein, the wild-type ilvE gene comprises the nucleotides shown in SEQ ID NO:25, as Figure 3 shown in D.
[0057] As used herein, the "native Bacillus subtilis IlvE protein" encoded by the WT Bacillus subtilis gene CDS comprises the amino acid sequence shown in SEQ ID NO:15, as Figure 3 shown in E.
[0058] As used herein, the "mutant Bacillus subtilis ilvE 5'-untranslated region" sequence (abbreviated as "mutant ilvE 5'-UTR") comprises the nucleotide sequence shown in SEQ ID NO:18, as Figure 1 shown. In particular, as Figure 1 shown, compared to the WT Bacillus subtilis ilvE 5'-UTR (SEQ ID NO:17; Figure 1 A), the mutant ilvE 5'-UTR sequence (SEQ ID NO:18; Figure 1 B) contains an unexpected single nucleotide polymorphism (SNP) at nucleotide position 73. For example, as Figure 1 presented, the WT Bacillus subtilis ilvE 5'-UTR sequence contains cytosine (C) at nucleotide position 73 (73C; SEQ ID NO:17), and the mutant ilvE 5'-UTR SNP sequence contains thymine (T) at nucleotide position 73 (73T; SEQ ID NO:18). As Figure 1 presented, the WT ilvE 5'-UTR sequence ( Figure 1 A) shows cytosine at nucleotide position 73 ( C ) double-underlined, and the mutant ilvE 5'-UTR sequence ( Figure 1 B) shows thymine at nucleotide position 73 ( T ) double-underlined. As generally shown in the Figure 2 schematic diagram, nucleotide position 73 of the ilvE 5'-UTR sequence is numbered from the start (+1) of the transcription start site, where position 73 can alternatively be referred to as position +73.
[0059] As used herein, phrases such as "Bacillus subtilis P2 promoter" and / or "operably linked to the P2 promoter" specifically refer to the Bacillus subtilis P2 promoter sequence set forth and described in PCT Publication No. WO 2020 / 112609 (incorporated herein by reference in its entirety). More particularly, the Bacillus subtilis P2 promoter is listed as SEQ ID NO:40 in PCT Publication No. WO 2020 / 112609.
[0060] As used herein, a "host cell" refers to a cell having the ability to serve as a host or expression vehicle for newly introduced DNA sequences. Thus, in certain embodiments of the present disclosure, the host cell is a Gram-positive cell, a Bacillus species, or an Escherichia coli (E. coli) cell.
[0061] As used herein, the phrases "modified Bacillus cell" and / or "Bacillus daughter cell" refer to a recombinant Bacillus cell comprising at least one genetic modification that is not present in the parental cell from which the modified cell is derived. In certain embodiments, an "unmodified" Bacillus cell may be referred to as a "control cell", particularly when compared to or relative to a modified Bacillus cell.
[0062] As used herein, when comparing the expression and / or production of a protein of interest (POI) in an "unmodified" (parental or control) cell to the expression and / or production of the same POI in a "modified" (daughter) cell, it is understood that the "unmodified" and "modified" cells are grown / cultured / fermented under the same conditions (e.g., the same conditions such as medium, temperature, pH, etc.). In certain embodiments, the increased amount of POI may be an endogenous POI (e.g., a native protease, a native amylase, etc.) or a heterologous POI (e.g., a recombinant protease, a recombinant amylase, etc.) expressed in the recombinant Bacillus cells of the present disclosure.
[0063] As used herein, "increasing" protein production or "increased" protein production means an increase in the amount of protein produced (e.g., the protein of interest). The protein may be produced within the host cell or secreted (or transported) into the culture medium. In certain embodiments, the protein of interest is secreted into the culture medium. Increased protein production, compared to the parental cell, may be detected as, for example, a higher maximum level of protein or enzyme activity (e.g., protease activity, amylase activity, pullulanase activity, cellulase activity, etc.) or the total extracellular protein produced.
[0064] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) introducing, substituting, or removing one or more nucleotides in a gene (or its ORF), or introducing, substituting, or removing one or more nucleotides in a regulatory element required for transcription or translation of a gene or its ORF, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-directed mutagenesis of any one or more of the genes disclosed herein, and / or (g) random mutagenesis.
[0065] As used herein, the term "expression" refers to the transcription and stable accumulation of sense (mRNA) or antisense RNA derived from a nucleic acid molecule of the present disclosure. Expression may also refer to the translation of mRNA into a polypeptide. Thus, the term "expression" includes any steps involved in polypeptide production, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, secretion, etc.
[0066] As used herein, "nucleic acid" refers to nucleotide or polynucleotide sequences, and fragments or portions thereof, as well as DNA, cDNA, and RNA of genomic or synthetic origin, which may be double-stranded or single-stranded, whether representing the sense or antisense strand. It should be understood that due to the degeneracy of the genetic code, multiple nucleotide sequences can encode a given protein.
[0067] It should be understood that the polynucleotides (or nucleic acid molecules) described herein include "genes", "vectors", and "plasmids".
[0068] Accordingly, the term "gene" refers to a polynucleotide encoding a specific sequence of amino acids, which includes all or part of the protein-coding sequence and may include regulatory (non-transcribed) DNA sequences, such as promoter sequences, which determine, for example, the conditions under which the gene is expressed. The transcribed region of a gene may include untranslated regions (UTRs) (including introns, 5'-untranslated region (UTR), and 3'-UTR) as well as the coding sequence (CDS).
[0069] As used herein, the term "coding sequence" (CDS) refers to a nucleotide sequence that directly specifies the amino acid sequence of its (encoded) protein product. The boundaries of the coding sequence are typically determined by a reading frame (hereinafter, "ORF") that usually begins with an ATG start codon. Coding sequences typically include DNA, cDNA, and recombinant nucleotide sequences.
[0070] As used herein, the term "promoter" refers to a nucleic acid sequence capable of controlling the expression of a coding sequence or functional RNA. Typically, the coding sequence is located 3' (downstream) of the promoter sequence. A promoter may be derived entirely from a native gene, or may be composed of different elements derived from different promoters found in nature, or may even contain synthetic nucleic acid segments. Those skilled in the art will understand that different promoters can direct the expression of a gene in different cell types, or at different developmental stages, or in response to different environmental or physiological conditions. A promoter that causes a gene to be expressed in most cell types most of the time is typically referred to as a "constitutive promoter". It should be further recognized that, in most cases, since the exact boundaries of regulatory sequences have not been fully defined, DNA fragments of different lengths can have the same promoter activity.
[0071] As used herein, the term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment such that the function of one is affected by the other. For example, a promoter is operably linked to a coding sequence (e.g., an ORF) when the expression of the coding sequence can be achieved (i.e., the coding sequence is under the transcriptional control of the promoter). The coding sequence can be operably linked to a regulatory sequence in the sense or antisense orientation.
[0072] A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, if DNA encoding a secretory leader sequence (i.e., a signal peptide) is expressed as a preprotein that participates in the secretion of a polypeptide, then the DNA encoding the secretory leader sequence (i.e., the signal peptide) is operably linked to the DNA of the polypeptide; if a promoter or enhancer affects the transcription of a coding sequence, then the promoter or enhancer is operably linked to the sequence; or if a ribosome binding site is positioned to facilitate translation, then the ribosome binding site is operably linked to the coding sequence. Generally, "operably linked" means that the DNA sequences being linked are contiguous, and in the case of a secretory leader sequence, are contiguous and in the reading phase. However, an enhancer does not have to be contiguous. Ligation is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide linkers or adaptors are used according to conventional practice.
[0073] As used herein, a "functional promoter sequence that controls the expression of a gene of interest (or its open reading frame) and is linked to the protein-coding sequence of the gene of interest" refers to a promoter sequence that controls the transcription and translation of a coding sequence in Bacillus. For example, in certain embodiments, the present disclosure relates to a polynucleotide comprising a 5' promoter (or 5' promoter region, or tandem 5' promoters, etc.), wherein the promoter region is operably linked to a nucleic acid sequence encoding a protein (e.g., an ORF).
[0074] As used herein, a "suitable regulatory sequence" refers to a nucleotide sequence that is located upstream (5' non-coding sequence), within, or downstream (3' non-coding sequence) of a coding sequence and affects the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences can include promoters, translational leader sequences, RNA processing sites, effector binding sites, and stem-loop structures.
[0075] As used herein, as used in phrases such as "introducing into a bacterial cell" or "introducing into a Bacillus cell" at least one polynucleotide open reading frame (ORF), or its gene, or its vector, the term "introducing" includes methods known in the art for introducing polynucleotides into cells, including but not limited to protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation, etc.
[0076] As used herein, "transformed" or "transformation" means the transformation of a cell by the use of recombinant DNA techniques. Transformation typically occurs by inserting one or more nucleotide sequences (e.g., polynucleotide, ORF, or gene) into the cell. The inserted nucleotide sequence can be a heterologous nucleotide sequence (i.e., a sequence that is not naturally present in the cell to be transformed). Thus, transformation generally refers to the introduction of foreign DNA into a host cell such that the DNA remains as a chromosomal integrant or a self-replicating extrachromosomal vector.
[0077] As used herein, "transforming DNA", "transformation sequence", and "DNA construct" refer to DNA used to introduce a sequence into a host cell or organism. Transforming DNA is the DNA used to introduce a sequence into a host cell or organism. The DNA can be generated in vitro by PCR or any other suitable technique. In some embodiments, the transforming DNA contains an input sequence, and in other embodiments, it further contains an input sequence flanked by homology boxes. In still other embodiments, the transforming DNA contains other non-homologous sequences (i.e., filler sequences or flanks) added to the ends. The ends can be closed such that the transforming DNA forms a closed loop, such as when inserted into a vector.
[0078] As used herein, "disruption of a gene" or "gene disruption" are used interchangeably and broadly refer to any genetic modification that substantially prevents a host cell from producing a functional gene product (e.g., a protein). Thus, as used herein, gene disruption includes, but is not limited to, frameshift mutations, premature stop codons (i.e., such that a functional protein is not produced), substitutions that eliminate or reduce the activity of an internal deletion of a protein (such that a functional protein is not produced), insertions that disrupt the coding sequence, mutations that remove the operable linkage between the native promoter required for transcription and the reading frame, etc.
[0079] As used herein, "input sequence" refers to a DNA sequence introduced into the chromosome of a Bacillus species. In some embodiments, the input sequence is part of a DNA construct. In other embodiments, the input sequence encodes one or more proteins of interest. In some embodiments, the input sequence contains a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it can be a homologous or heterologous sequence). In some embodiments, the input sequence encodes one or more proteins of interest, genes, and / or mutated or modified genes. In alternative embodiments, the input sequence encodes a functional wild-type gene or operon, a functional mutated gene or operon, or a non-functional gene or operon. In some embodiments, a non-functional sequence can be inserted into a gene to disrupt its function. In another embodiment, the input sequence includes a selectable marker. In additional embodiments, the input sequence includes two homology boxes.
[0080] As used herein, "homologous box" refers to a nucleic acid sequence that is homologous to a sequence in the Bacillus chromosome. More particularly, according to the present invention, the homologous box is an upstream or downstream region that has a sequence identity of between about 80% and 100%, between about 90% and 100%, or between about 95% and 100% with the directly flanking coding regions of the gene or a portion of the gene to be deleted, disrupted, inactivated, downregulated, etc. These sequences direct where the DNA construct is integrated in the Bacillus chromosome and which portion of the Bacillus chromosome is replaced by the input sequence. While not intended to limit the disclosure, the homologous box can include between about 1 base pair (bp) and 200 kilobases (kb). Preferably, the homologous box includes between about 1 bp and 10.0 kb; between 1 bp and 5.0 kb; between 1 bp and 2.5 kb; between 1 bp and 1.0 kb; and between 0.25 kb and 2.5 kb. The homologous box can also include about 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb, and 0.1 kb. In some embodiments, the 5' and 3' ends of the selectable marker are flanked by homologous boxes, where the homologous box contains a nucleic acid sequence that is closely flanked by the coding region of the gene.
[0081] As used herein, the term "nucleotide sequence encoding a selectable marker" refers to a nucleotide sequence that is capable of being expressed in a host cell and where the expression of the selectable marker confers upon the cell containing the expressed gene the ability to grow in the presence of the corresponding selective reagent or in the absence of an essential nutrient.
[0082] As used herein, the terms "selectable marker" and "selection marker" refer to a nucleic acid (e.g., a gene) that is capable of being expressed in a host cell, which allows for the easy selection of those hosts that contain the vector. Examples of such selectable markers include, but are not limited to, antimicrobial agents. Thus, the term "selectable marker" refers to a gene that provides an indication that the host cell has taken up the input DNA of interest or that some other reaction has occurred. Typically, the selectable marker is a gene that confers antimicrobial resistance or a metabolic advantage to the host cell to allow the differentiation of cells containing foreign DNA from cells that have not received any foreign sequences during transformation.
[0083] A "residing selectable marker" is a marker that is located on the chromosome of the microorganism to be transformed. The residing selectable marker encodes a gene that is different from the selectable marker on the transforming DNA construct. Selection markers are well known to those of skill in the art. As indicated above, the marker can be an antimicrobial resistance marker (e.g., amp R , phleo R , specR , kan R , ery R , tet R , cmp R , and neo R ). In some embodiments, the present invention provides chloramphenicol resistance genes (e.g., the gene present on pC194, and the resistance gene present in the genome of Bacillus licheniformis). Such resistance genes are particularly useful in the present invention and in embodiments involving chromosomal integration of cassettes and chromosomal amplification of integrative plasmids. Other markers useful according to the present invention include, but are not limited to, auxotrophic markers such as serine, lysine, tryptophan; and detection markers such as β-galactosidase.
[0084] As defined herein, the "genome" of a host cell, the "genome" of a bacterial (host) cell, or the "genome" of a Bacillus species (host) cell includes chromosomal and extrachromosomal genes.
[0085] As used herein, the terms "plasmid", "vector", and "cassette" refer to extrachromosomal elements that typically carry genes that are not part of the central metabolism of the cell and are usually in the form of circular double-stranded DNA molecules. Such elements can be linear or circular self-replicating sequences, genomic integration sequences, phages, or nucleotide sequences of single-stranded or double-stranded DNA or RNA derived from any source, wherein multiple nucleotide sequences have been ligated or recombined into a single construct that is capable of introducing a promoter fragment and a DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences, into a cell.
[0086] As used herein, the term "plasmid" refers to a circular double-stranded (ds) DNA construct that serves as a cloning vector and forms an extrachromosomal self-replicating genetic element in many bacteria and some eukaryotes. In some embodiments, the plasmid is incorporated into the genome of the host cell. In some embodiments, the plasmid is present in the parental cell and is lost in the daughter cells.
[0087] As used herein, a "transformation cassette" refers to a specific vector that contains a gene (or its ORF) and, in addition to the foreign gene, has elements that facilitate the transformation of a specific host cell.
[0088] As used herein, the term "vector" refers to any nucleic acid that can replicate (propagate) in a cell and can carry a new gene or DNA segment into the cell. Thus, the term refers to a nucleic acid construct designed for transfer between different host cells. Vectors include viruses, bacteriophages, proviruses, plasmids, phagemids, transposons, and artificial chromosomes such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), PLACs (plant artificial chromosomes), etc. that are "episomes" (i.e., they replicate autonomously or can integrate into the chromosome of the host organism).
[0089] An "expression vector" is a vector that has the ability to incorporate and express heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and are known to those of skill in the art. The selection of an appropriate expression vector is within the knowledge of those of skill in the art.
[0090] As used herein, the terms "expression cassette" and "expression vector" refer to a nucleic acid construct that is recombinantly or synthetically produced and has a series of designated nucleic acid elements (i.e., these are vectors or vector elements as described above) that permit transcription of a specific nucleic acid in a target cell. A recombinant expression cassette can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes (among other sequences) the nucleic acid sequence to be transcribed and a promoter. In some embodiments, the DNA construct also includes a series of designated nucleic acid elements that permit transcription of a specific nucleic acid in a target cell. In certain embodiments, the DNA constructs of the present disclosure contain selectable markers and inactivated chromosomes, or genes, or DNA segments as defined herein.
[0091] As used herein, a "targeting vector" is a vector that includes a polynucleotide sequence that is homologous to a region in the chromosome of the host cell into which the targeting vector is transformed and that can drive homologous recombination at that region. For example, a targeting vector can be used to introduce a mutation into the chromosome of a host cell by homologous recombination. In some embodiments, the targeting vector contains, for example, additional non-homologous sequences (i.e., filler sequences or flanking sequences) added to the ends. The ends can be closed such that the targeting vector forms a closed loop, such as in an insertion vector. For example, in certain embodiments, a (host) cell is modified (e.g., transformed) by introducing one or more "targeting vectors" into a parental Bacillus licheniformis cell.
[0092] As used herein, the term "protein of interest" or "POI" refers to a polypeptide of interest that is desired to be expressed in a modified Bacillus licheniformis (sub) host cell, wherein the POI is preferably expressed at an increased level (i.e., relative to an "unmodified" (parent) cell). Thus, as used herein, the POI can be an enzyme, a substrate-binding protein, a surfactant protein, a structural protein, a receptor protein, etc. In certain embodiments, the modified cells of the present disclosure produce an increased amount of a heterologous or endogenous protein of interest relative to the parent cell. In specific embodiments, the increased amount of the protein of interest produced by the modified cells of the present disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or more than a 5.0% increase relative to the parent cell.
[0093] Similarly, as defined herein, "gene of interest" or "GOI" refers to a nucleic acid sequence (e.g., polynucleotide, gene, or ORF) encoding a POI. The "gene of interest" encoding the "protein of interest" can be a naturally occurring gene, a mutated gene, or a synthetic gene.
[0094] As used herein, the terms "polypeptide" and "protein" are used interchangeably and refer to any length polymer of amino acid residues joined by peptide bonds. The conventional one (1)-letter or three (3)-letter codes for amino acid residues are used herein. The polypeptide can be linear or branched, it can contain modified amino acids, and it can be interrupted by non-amino acids. The term polypeptide also encompasses amino acid polymers that have been modified either naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within this definition are, for example, polypeptides containing one or more amino acid analogs (including, for example, non-natural amino acids, etc.) and other modifications known in the art.
[0095] In certain embodiments, the genes of the present disclosure encode proteins for commercially relevant industrial purposes, such as enzymes (e.g., acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof).
[0096] As used herein, a "variant" polypeptide refers to a polypeptide typically derived from a parental (or reference) polypeptide by substitution, addition, or deletion of one or more amino acids through recombinant DNA techniques. Variant polypeptides may differ from the parental polypeptide by a small number of amino acid residues and can be defined by the level of their amino acid sequence homology / identity to the parental (reference) polypeptide.
[0097] Preferably, the variant polypeptide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% amino acid sequence identity to the parental (reference) polypeptide sequence. As used herein, a "variant" polynucleotide refers to a polynucleotide encoding a variant polypeptide, wherein the "variant polynucleotide" has a specified degree of sequence homology / identity to the parental polynucleotide or hybridizes to the parental polynucleotide (or its complementary sequence) under stringent hybridization conditions. Preferably, the variant polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% nucleotide sequence identity to the parental (reference) polynucleotide sequence.
[0098] As used herein, "mutation" refers to any change or alteration in a nucleic acid sequence. There are several types of mutations, including point mutations, deletion mutations, silent mutations, frameshift mutations, splicing mutations, and the like. Mutations can be made specifically (e.g., via site-directed mutagenesis) or randomly (e.g., via chemical agents, by passage of a repair-minus bacterial strain).
[0099] As used herein, in the context of a polypeptide or its sequence, the term "substitution" means that one amino acid is replaced by another amino acid (i.e., substituted).
[0100] As defined herein, an "endogenous gene" is a gene located in its natural position in the genome of an organism.
[0101] As defined herein, a "heterologous" gene, "non-endogenous" gene, or "exogenous" gene is a gene (or ORF) that is not normally found in a host organism but is introduced into the host organism by gene transfer. As used herein, the term one or more "exogenous" genes includes natural genes (or ORFs) inserted into a non-natural organism and / or chimeric genes inserted into a natural or non-natural organism.
[0102] As defined herein, a "heterologous control sequence" is a gene expression control sequence (e.g., a promoter or enhancer) that does not function in nature to regulate the expression of a target gene. Typically, a heterologous nucleic acid sequence is not endogenous (natural) to the part of the cell or genome in which it is present and has been added to the cell by infection, transfection, transformation, microinjection, electroporation, or the like. A "heterologous" nucleic acid construct can contain a control sequence / DNA coding sequence combination that is the same as or different from the control sequence / DNA coding (ORF) sequence combination found in a natural host cell.
[0103] As used herein, the terms "signal sequence" and "signal peptide" refer to a sequence of amino acid residues that can participate in the secretion or directed transport of a mature protein or a precursor form of a protein. Typically, the signal sequence is located at the N-terminus of the precursor or mature protein sequence. The signal sequence can be endogenous or exogenous. The signal sequence is generally not present in the mature protein. Typically, after protein transport, the signal sequence is cleaved from the protein by signal peptidase.
[0104] The term "derived from" encompasses the terms "originating from", "obtained from", "obtainable from", and "produced from", and generally indicates that a specified material or composition finds its origin in another specified material or composition, or has characteristics that can be described with reference to another specified material or composition.
[0105] As used herein, the term "homology" pertains to homologous polynucleotides or polypeptides. If two or more polynucleotides or two or more polypeptides are homologous, this means that the homologous polynucleotides or polypeptides have a "degree of identity" of at least 60%, more preferably at least 70%, even more preferably at least 85%, still more preferably at least 90%, more preferably at least 95%, and most preferably at least 98%. The degree of homology between sequences can be determined using any suitable method known in the art (see, e.g., Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; programs in the Wisconsin Genetics Software Package (Genetics Computer Group, Madison, WI) such as GAP, BESTFIT, FASTA, and TFASTA; and Devereux et al., 1984). For the purposes of the present invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970) implemented as the Needle program in the EMBOSS package (Rice et al., 2000) (preferably version 3.0.0 or later) is used to determine the degree of identity between two amino acid sequences. The optional parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EBLOSUM62 (the EMBOSS version of BLOSUM62) substitution matrix. The output of Needle labeled "longest identity" (obtained using the non-simplified option) is used as the percentage identity and is calculated as follows:
[0106] (Identical residues × 100) / (Alignment length - Total number of gaps in the alignment)
[0107] As used herein, the term "percent identity (%)" refers to the level of nucleic acid or amino acid sequence identity between a nucleic acid sequence encoding a polypeptide or the amino acid sequence of a polypeptide when aligned using a sequence alignment program.
[0108] As used herein, "specific productivity" is the total amount of protein produced per cell per time over a given period.
[0109] As defined herein, the terms "purified", "isolated", or "enriched" mean that a biomolecule (e.g., a polypeptide or polynucleotide) is altered from its native state by separating it from some or all of the naturally-occurring components with which it is associated in nature. Such separation or purification can be accomplished by separation techniques well-known in the art such as ion-exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulfate precipitation or other protein salt precipitation, centrifugation, size-exclusion chromatography, filtration, microfiltration, gel electrophoresis, or gradient separation to remove unwanted whole cells, cell debris, impurities, foreign proteins, or enzymes from the final composition. Components that provide additional benefits, such as activators, anti-inhibitors, desired ions, pH-controlling compounds, or other enzymes or chemicals, can then be further added to the purified or isolated biomolecule composition.
[0110] As used herein, "flanking sequence" refers to any sequence that is upstream or downstream of the sequence being discussed (e.g., for gene A-B-C, gene B is flanked by the A and C gene sequences). In certain embodiments, the input sequence is flanked by homeoboxes on each side. In another embodiment, the input sequence and the homeoboxes are contained within a unit that is flanked by filler sequences on each side. In some embodiments, the flanking sequence is present only on one side (3' or 5'), but in preferred embodiments, it is on each side of the sequence being flanked. The sequence of each homeobox is homologous to a sequence in the Bacillus chromosome. These sequences direct the integration position of the new construct in the Bacillus chromosome and which part of the Bacillus chromosome will be replaced by the input sequence. In other embodiments, the 5' and 3' ends of the selectable marker are flanked by polynucleotide sequences that comprise portions of inactivated chromosomal segments. In some embodiments, the flanking sequence is present only on one side (3' or 5'), while in other embodiments, it is present on each side of the sequence being flanked.
[0111] II. Mutant Bacillus Strains with Enhanced Protein Production and Carbon Yield Phenotypes
[0112] As generally set forth herein, and further described in the examples below, the applicant has identified a mutant Bacillus subtilis strain (designated "CZ437") with an enhanced protein production phenotype. In particular, the applicant performed next-generation sequencing (NGS) on the mutant CZ437 strain to further characterize the observed enhanced protein productivity phenotype, in which an unexpected single nucleotide polymorphism (SNP) mutation in the wild-type (WT) Bacillus subtilis ilvE 5'-UTR sequence (SEQ ID NO:17) was identified. For example, the DNA sequences of the WT ilvE 5'-UTR (SEQ ID NO:17) and the mutant ilvE 5'-UTR (SEQ ID NO:18) are presented in Figure 1 A andFigure 1 in B, wherein the WT ilvE 5′-UTR sequence contains cytosine (C) at nucleotide position 73 (SEQ ID NO:17), and the mutant ilvE 5′-UTR SNP sequence contains thymine (T) at nucleotide position 73 (SEQ ID NO:18).
[0113] Without wishing to be bound by any particular theory, mechanism, or mode of operation, Applicants contemplate herein that the unexpected SNP mutation identified in the ilvE 5′-UTR (i.e., SNP C→T) may affect ilvE messenger RNA (mRNA) stability. For example, the ilvE (ybgE) gene is known to encode a branched-chain amino acid transaminase that induces the transamination of branched-chain amino acids and α-ketoglutarate. Berger et al. (2003) demonstrated another function of IlvE in the methionine regeneration pathway by converting ketomethylthiolbutyrate (KMTB) to methionine. For this reaction, the IvlE aminotransferase can use leucine, isoleucine, valine, phenylalanine, and tyrosine as amino donors, while the Bacillus subtilis homolog YkrV uses only glutamine as an amino donor. As generally described in Mader et al. (2004), the ilvE gene is negatively regulated by the global transcriptional regulator CodY, where CodY controls the transcription of ilvE by binding to its transcriptional leader sequence (5′-UTR) sequence and acting as a roadblock to RNA polymerase, resulting in repression of ilvE in the presence of casein amino acids or free amino acids. As Figure 2 shown / annotated, the SNP mutation in the ilvE 5′-UTR occurs fifty-five (55) nucleotides (bp) upstream (5′) of the translation start site, which may affect ilvE mRNA stability. In particular, as Figure 2 presented, the C>T mutation in the ilvE 5′-UTR is located near CodY binding motif 2 and at putative weak binding motif 5. As expected herein, the weak codY binding motif 5 may potentially mask codY from binding to motif 2, resulting in derepression of ilvE transcription.
[0114] Accordingly, as described herein and further described in the examples below, the Applicant designed, constructed, and screened recombinant Bacillus strains expressing a reporter protein (GG36) to further evaluate the enhanced protein production phenotype identified in the mutant Bacillus subtilis CZ437 strain. More particularly, as described in Example 1, two (2) GG36 reporter protein expression cassettes were constructed and introduced into Bacillus subtilis strain A (containing the mutant ilvE 5′-UTR (SEQ ID NO:18)) and the isogenic Bacillus subtilis strain B (containing the WT ilvE 5′-UTR (SEQ ID NO:17)), where strains A and B were fermented in a large-scale (approx. 14L) fermenter under the same conditions using standard fermentation conditions. As shown in Table 1 (Example 1), the relative improvement in carbon productivity of the mutant Bacillus subtilis reporter strain A was significantly enhanced compared to the isogenic reporter strain B.
[0115] Similarly, as described in Example 2, time-course samples from Bacillus subtilis strains A and B (expressing the GG36 reporter protein) were used for real-time quantitative PCR (RT qPCR) analysis, where samples were collected at 8, 16, 24, and 32 hours of fermentation and total RNA was extracted. For example, at 8, 16, 24, and 32 hours of fermentation, the fold change in ilvE mRNA of Bacillus subtilis strain A (mutant ilvE 5′-UTR) compared to the isogenic Bacillus subtilis strain B (WT ilvE 5′-UTR) was as Figure 4 shown. More particularly, as Figure 4 presented, the amount of ilvE mRNA in Bacillus subtilis strain A was significantly increased by 2.91, 1.93, and 1.79-fold at the 16, 24, and 32-hour fermentation time points, respectively, compared to the isogenic Bacillus subtilis strain B.
[0116] Accordingly, certain embodiments of the present disclosure relate to mutant Bacillus strains containing a variant ilvE gene sequence, recombinant Bacillus strains containing a variant ilvE gene sequence, mutant / recombinant strains containing a variant ilvE gene sequence and expressing / producing one or more proteins of interest, methods and compositions for constructing recombinant Bacillus strains containing a variant ilvE gene sequence, expression cassettes encoding proteins of interest, methods and compositions for culturing recombinant Bacillus strains containing a variant ilvE gene sequence for enhancing the production of proteins of interest, etc.
[0117] III. Recombinant Polynucleotides and Molecular Biology
[0118] As generally described above, certain embodiments of the present disclosure particularly relate to variant ilvE genes, mutant Bacillus strains comprising variant ilvE genes, recombinant (genetically modified) Bacillus strains comprising variant ilvE genes, mutant and / or recombinant Bacillus strains comprising variant ilvE genes and expressing / producing one or more target proteins, methods and compositions for constructing recombinant Bacillus strains comprising variant ilvE gene sequences, expression cassettes encoding target proteins, methods and compositions for culturing recombinant Bacillus strains comprising variant ilvE gene sequences for enhancing the production of target proteins, and the like.
[0119] Thus, in one or more embodiments, the present disclosure provides recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant (genetically modified) Gram-positive bacterial cells / strains expressing target proteins, and the like. In certain one or more embodiments, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells for enhancing the production of target proteins.
[0120] In certain embodiments, the polynucleotide constructs of the present disclosure are referred to as expression cassettes (or expression constructs), wherein the expression cassette comprises, in the 5' to 3' direction and in an operable combination, at least one upstream (5') promoter sequence operably linked to a downstream (3') gene coding sequence CDS. In certain embodiments, the expression cassette encodes one or more target proteins (e.g., 5'-[promoter sequence]-[gene coding sequence]-3'; abbreviated as 5'-[pro]-[gene CDS]-3').
[0121] In other embodiments, the present disclosure provides ilvE expression cassettes. In certain embodiments, the ilvE expression cassette comprises, in the 5' to 3' direction and in an operable combination, at least one upstream (5') promoter sequence operably linked to a variant ilvE 5'-untranslated region (5'-UTR) sequence, which variant ilvE 5'-untranslated region (5'-UTR) sequence is operably linked to a wild-type ilvE gene CDS (abbreviated as 5'-[pro]-[ilvE* 5'-UTR]-[WT ilvE CDS]-3'). As abbreviated above, the variant (mutant) ilvE* gene 5'-UTR sequence is presented / shown with an asterisk (*) to distinguish it from the wild-type ilvE gene 5'-UTR sequence (i.e., SEQ ID NO:17).
[0122] In certain other embodiments, the expression cassette can comprise one or more DNA sequence elements, including but not limited to DNA sequence elements encoding protein / peptide signal (secretion) sequences, DNA sequence elements encoding propeptide (proregion) amino acid residues, DNA sequence elements comprising transcription terminator sequences, DNA sequence elements comprising 5′-UTR, 3′-UTR, and the like.
[0123] Accordingly, one or more of the nucleic acid sequences described herein can be produced by using any suitable synthetic, manipulative, and / or isolation techniques, or combinations thereof. For example, one or more of the polynucleotides described herein can be produced using standard nucleic acid synthesis techniques well known to those skilled in the art, such as solid-phase synthesis techniques. In such techniques, typically fragments of up to fifty (50) or more nucleotide bases are synthesized and then ligated (e.g., by enzymatic or chemical ligation methods) to substantially form any desired continuous nucleic acid sequence. The synthesis of one or more of the polynucleotides described herein can also be facilitated by any suitable method known in the art, including but not limited to chemical synthesis using classical phosphoramidite methods or methods typically implemented in automated synthesis methods. One or more of the polynucleotides described herein can also be produced using an automated DNA synthesizer. Custom nucleic acids can be ordered from a variety of commercial sources (e.g., ATUM (DNA 2.0), Newark, CA, USA; Life Tech (GeneArt), Carlsbad, CA, USA; GenScript, Ontario, Canada; BaseClear B.V., Leiden, Netherlands; Integrated DNA Technologies, Skokie, IL, USA; Ginkgo Bioworks (Gen9), Boston, MA, USA; and Twist Bioscience, San Francisco, CA, USA). Other techniques and related principles for synthesizing nucleic acids are described and known in the art.
[0124] Recombinant DNA techniques for modifying nucleic acids are well known in the art, such as, for example, restriction endonuclease digestion, ligation, reverse transcription and cDNA production, and polymerase chain reaction (e.g., PCR). One or more of the polynucleotides described herein can also be obtained by screening a cDNA library using one or more oligonucleotide probes that can hybridize to or PCR amplify polynucleotides encoding one or more of the variants described herein. Procedures for screening and isolating cDNA clones, as well as PCR amplification procedures, are well known to those skilled in the art and are described in standard references known to those skilled in the art. One or more of the polynucleotides described herein can be obtained by altering a naturally occurring polynucleotide backbone (e.g., encoding one or more of the precursor region sequences of the variants described herein) by, for example, known mutagenesis procedures (e.g., site-directed mutagenesis, site saturation mutagenesis, and in vitro recombination). A variety of methods suitable for generating modified polynucleotides described herein encoding one or more of the variants described herein are known in the art and include, but are not limited to, for example, site saturation mutagenesis, scanning mutagenesis, insertional mutagenesis, deletion mutagenesis, random mutagenesis, site-directed mutagenesis and directed evolution, and various other recombination methods.
[0125] As generally set forth above, and further described in the examples below, certain embodiments of the present disclosure relate to recombinant (modified) Gram-positive cells capable of producing a heterologous target protein. Accordingly, certain embodiments relate to methods for constructing such recombinant Gram-positive cells with increased protein production capabilities. In certain embodiments, one or more expression cassettes encoding one or more target proteins are introduced into the Gram-positive cells of the present disclosure. In an exemplary embodiment, these cassettes are integrated into the genome of the cell. Accordingly, certain embodiments relate to nucleic acid molecules, polynucleotides (e.g., vectors, plasmids, expression cassettes), regulatory elements, etc. suitable for constructing recombinant (modified) Gram-positive host cells.
[0126] Thus, as presented in the examples and generally described herein, the recombinant cells of the present disclosure can be constructed by those skilled in the art using standard and conventional recombinant DNA and molecular cloning techniques well known in the art. Methods for genetic modification include, but are not limited to, (a) introducing, substituting, or removing one or more nucleotides in a gene, or introducing, substituting, or removing one or more nucleotides in a regulatory element required for transcription or translation of a gene, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-specific mutagenesis, and / or (g) random mutagenesis.
[0127] In certain embodiments, modified cells of the present disclosure can be constructed by reducing or eliminating the expression of a gene by using methods well-known in the art (e.g., insertion, disruption, substitution, or deletion). The portion of the gene to be modified or inactivated can be, for example, the coding region or regulatory elements required for the expression of the coding region.
[0128] Examples of such regulatory or control sequences can be a promoter sequence or a functional portion thereof (i.e., a portion sufficient to affect the expression of a nucleic acid sequence). Other control sequences for modification include, but are not limited to, leader sequences, propeptide sequences, signal sequences, transcription terminators, transcription activators, and the like.
[0129] In certain other embodiments, modified cells are constructed by gene deletion to eliminate or reduce the expression of a gene. Gene deletion techniques enable the partial or complete removal of one or more genes, thereby eliminating their expression or expressing non-functional (or reduced-activity) protein products. In such methods, the deletion of a gene can be accomplished by homologous recombination using a plasmid that has been constructed to continuously contain the 5' and 3' regions flanking the gene. The contiguous 5' and 3' regions can be introduced into the cell, for example, on a temperature-sensitive plasmid, in combination with a second selectable marker at the permissive temperature to allow the plasmid to establish in the cell. The cell is then transferred to the non-permissive temperature to select cells that have integrated the plasmid into one of the chromosomal homologous flanking regions. The selection of plasmid integration is effected by selecting the second selectable marker. After integration, the recombination event at the second homologous flanking region is stimulated by transferring the cells to the permissive temperature for several generations without selection. The cells are plated to obtain single colonies, and the colonies are examined for the loss of both selectable markers. Thus, those skilled in the art can readily identify nucleotide regions (suitable for complete or partial deletion) in the coding sequence of the gene and / or the non-coding sequence of the gene.
[0130] In other embodiments, modified cells are constructed by introducing, substituting, or removing one or more nucleotides in the gene or regulatory elements required for its transcription or translation. For example, nucleotides can be inserted or removed so as to cause the introduction of a stop codon, the removal of a start codon, or a frameshift of the reading frame. Such modifications can be accomplished by site-directed mutagenesis or mutagenesis generated by PCR according to methods known in the art. Thus, in certain embodiments, the genes of the present disclosure are inactivated by complete or partial deletion.
[0131] In another embodiment, modified cells are constructed by a gene conversion process. For example, in a gene conversion method, a nucleic acid sequence corresponding to one or more genes is mutagenized in vitro to produce a defective nucleic acid sequence, which is then transformed into a parental cell to produce a defective gene. By homologous recombination, the defective nucleic acid sequence replaces the endogenous gene. It may be desirable that the defective gene or gene fragment also encodes a marker that can be used to select for transformants containing the defective gene. For example, the defective gene can be associated with a selectable marker and introduced on a non-replicating or temperature-sensitive plasmid. Selection for plasmid integration is effected by selecting for the marker under conditions that do not permit plasmid replication. Selection for the second recombination event leading to gene replacement is effected by examining the colonies for loss of the selectable marker and for acquisition of the mutated gene. Alternatively, the defective nucleic acid sequence can contain an insertion, substitution, or deletion of one or more nucleotides of the gene, as described below.
[0132] In other embodiments, modified cells are constructed using nucleotide sequences complementary to the nucleic acid sequence of a gene by established antisense techniques. More particularly, the expression of a gene in a Gram-positive cell can be reduced (downregulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which can be transcribed in the cell and is capable of hybridizing to the mRNA produced in the cell. Under conditions that permit the complementary antisense nucleotide sequence to hybridize to the mRNA, the amount of translated protein is thus reduced or eliminated. Such antisense methods include, but are not limited to, RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), antisense oligonucleotides, etc., all of which are well known to those skilled in the art.
[0133] In other embodiments, modified cells are generated / constructed via CRISPR-Cas9 editing. For example, a gene encoding a protein of interest can be edited or disrupted (or deleted or downregulated) by means of a nucleic acid-guided endonuclease, which finds its target DNA by binding a guide RNA (e.g., Cas9) and Cpf1 or a guide DNA (e.g., NgAgo), which recruits the endonuclease to the target sequence on the DNA, where the endonuclease can create a single-stranded or double-stranded break in the DNA. This targeted DNA break becomes a substrate for DNA repair and can be recombined with a provided editing template to effect gene disruption or deletion. For example, a gene encoding a nucleic acid-guided endonuclease (for this purpose, Cas9 from Streptococcus pyogenes) or a codon-optimized gene encoding the Cas9 nuclease can be operably linked to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells, thereby generating a Gram-positive cell Cas9 expression cassette. Similarly, one or more target sites specific to the gene of interest can be readily identified by those skilled in the art. For example, to construct a DNA construct encoding a gRNA - directed to a target site within the gene of interest, a variable targeting domain (VT) will contain the nucleotides of the target site that are 5' of the protospacer adjacent motif (PAM) (TGG) and these nucleotides are fused to DNA encoding the Cas9 endonuclease recognition domain (CER) of Streptococcus pyogenes Cas9. The DNA encoding the VT domain and the DNA encoding the CER domain are combined, thereby generating DNA encoding the gRNA. Thus, a Gram-positive expression cassette for the gRNA is generated by operably linking the DNA encoding the gRNA to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells.
[0134] In certain embodiments, DNA breaks induced by endonucleases are repaired / replaced with an input sequence. For example, to precisely repair DNA breaks generated by the above-described Cas9 expression cassette and gRNA expression cassette, a nucleotide editing template is provided such that the cell's DNA repair machinery can utilize the editing template. For example, approximately 500 bp 5' of the target gene can be fused with approximately 500 bp 3' of the target gene to generate an editing template that is used by the Gram-positive host machinery to repair DNA breaks generated by the RGEN.
[0135] Many different methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence) can be used to co - deliver the Cas9 expression cassette, the gRNA expression cassette, and the editing template to filamentous fungal cells. Transformed cells are screened by amplifying the locus of the gene by PCR using forward and reverse primers. These primers can amplify the wild - type locus or the modified locus that has been edited by the RGEN. Then, sequencing primers are used to sequence these fragments to identify the edited colonies.
[0136] In still other embodiments, modified cells are constructed by random or specific mutagenesis using methods well - known in the art, including but not limited to chemical mutagenesis and transposition. Modification of a gene can be carried out by subjecting parental cells to mutagenesis and screening for mutant cells in which gene expression has been reduced or eliminated. Mutagenesis, which can be specific or random, can be carried out, for example, by using a suitable physical or chemical mutagen, using a suitable oligonucleotide, or subjecting a DNA sequence to PCR - generated mutagenesis. In addition, mutagenesis can be carried out by using any combination of these mutagenesis methods.
[0137] Examples of physical or chemical mutagens suitable for the purposes of the present invention include ultraviolet (UV) irradiation, hydroxylamine, N - methyl - N'- nitro - N - nitrosoguanidine (MNNG), N - methyl - N'- nitrosoguanidine (NTG), O - methylhydroxylamine, nitrous acid, ethyl methane sulfonate (EMS), sodium bisulfite, formic acid, and nucleotide analogs. When using such reagents, mutagenesis is typically carried out by incubating the parental cells to be mutagenized in the presence of the selected mutagen under suitable conditions and selecting mutant cells that exhibit reduced or no expression of the gene.
[0138] PCT Publication No. WO 2003 / 083125 discloses methods for modifying Gram - positive (Bacillus) cells, such as using PCR fusion to generate Bacillus deletion strains and DNA constructs to bypass Escherichia coli. PCT Publication No. WO2002 / 14490 discloses methods for modifying Bacillus cells, which include (1) constructing and transforming an integrative plasmid (pComK), (2) randomly mutating coding sequences, signal sequences, and propeptide sequences, (3) homologous recombination, (4) increasing transformation efficiency by adding non - homologous flanks to the transforming DNA, (5) optimizing double - crossover integration, (6) site - directed mutagenesis, and (7) marker - less deletion.
[0139] Those skilled in the art know suitable methods for introducing polynucleotide sequences into bacterial cells (e.g., Gram-negative cells, Gram-positive cells). In fact, methods such as transformation, including protoplast transformation and mid-plate assembly, transduction, and protoplast fusion, are known and suitable for the present disclosure. The transformation method is particularly preferably used to introduce the DNA constructs of the present disclosure into host cells.
[0140] In addition to the common methods, in some embodiments, the host cells are directly transformed (i.e., without using intermediate cells to amplify the DNA construct or otherwise process the DNA construct before introducing it into the host cell). Introducing the DNA construct into the host cell includes those physical and chemical methods known in the art for introducing DNA into the host cell without inserting a plasmid or vector. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, liposomes, etc. In additional embodiments, the DNA construct is co-transformed with a plasmid without inserting into the plasmid. In additional embodiments, a selectable marker is deleted or substantially excised from a modified Bacillus strain by methods known in the art. In some embodiments, the vector is resolved from the host chromosome, leaving the flanking regions on the chromosome while removing the native chromosomal region.
[0141] Promoters and promoter sequence regions for expressing genes, their coding sequences (CDS), open reading frames (ORF), and / or variant sequences in Gram-positive cells are generally known to those skilled in the art. The promoter sequences of the present disclosure are generally selected such that they function in Gram-positive cells. For example, promoters that can be used to drive gene expression in Bacillus cells include, but are not limited to, the Bacillus subtilis alkaline protease (aprE) promoter, the α-amylase promoter (amyE) of Bacillus subtilis, the α-amylase promoter (amyL) of Bacillus licheniformis, the α-amylase promoter of Bacillus amyloliquefaciens, the neutral protease (nprE) promoter from Bacillus subtilis, the mutant aprE promoter, or any other promoter from Bacillus licheniformis or other related Bacillus. Methods for screening and generating a library of promoters with a range of activities (promoter strength) in Bacillus cells are described in published patent application WO 2002 / 14490.
[0142] IV. Fermenting Bacillus cells for protein production
[0143] As generally described above, certain embodiments relate to compositions and methods for constructing and obtaining Gram-positive cells that express / produce one or more target proteins. Accordingly, certain other embodiments of the present disclosure relate to methods for producing a target protein in Gram-positive cells by fermenting the cells in a suitable medium. Fermentation methods well known in the art can be used to ferment the Gram-positive cells of the present disclosure.
[0144] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system where the composition of the medium is set at the start of fermentation and does not change during fermentation. At the start of fermentation, the medium is inoculated with the desired organism. In this method, fermentation occurs without adding any components to the system. Typically, batch fermentation qualifies as "batch" with respect to the addition of carbon source, and attempts are often made to control factors such as pH and oxygen concentration. The metabolite and biomass composition of the batch system changes continuously until fermentation stops. In a typical batch culture, cells can progress through a static lag phase to a high-growth logarithmic phase and finally enter a stationary phase where the growth rate decreases or stops. If untreated, cells in the stationary phase eventually die. Generally, cells in the logarithmic phase are responsible for the bulk production of the product.
[0145] A suitable variation of the standard batch system is the "fed-batch" fermentation system. In this variant of the typical batch system, substrate is added incrementally as fermentation progresses. The fed-batch system is useful when catabolite repression might inhibit the metabolism of the cells and when a limited amount of substrate is desired in the medium. Measurement of the actual substrate concentration in the fed-batch system is difficult and thus it is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of off-gases (such as CO2). Batch and fed-batch fermentations are commonly used and known in the art.
[0146] Continuous fermentation is an open system where a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation typically maintains the culture at a constant high density where the cells are mainly in the logarithmic phase of growth. Continuous fermentation allows for the regulation of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient (such as a carbon or nitrogen source) is maintained at a fixed rate and all other parameters are allowed to be adjusted. In other systems, many factors affecting growth can be continuously changed while the cell concentration, measured by the turbidity of the medium, remains constant. The continuous system strives to maintain steady-state growth conditions. Thus, the cell loss due to the withdrawal of the medium should be balanced with the cell growth rate in the fermentation. Methods for regulating nutrients and growth factors for continuous fermentation processes and techniques for maximizing the rate of product formation are well known in the field of industrial microbiology.
[0147] In certain embodiments, the desired protein expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, which include separating the host cells from the culture medium by centrifugation or filtration, or, if desired, disrupting the cells and removing the supernatant from the cell fractions and debris. Typically, after clarification, the protein fraction of the supernatant or filtrate is precipitated with a salt (e.g., ammonium sulfate). The precipitated protein is then dissolved and can be purified by a variety of chromatographic procedures (e.g., ion exchange chromatography, gel filtration).
[0148] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. A classic batch fermentation is a closed system where the composition of the culture medium is set at the start of the fermentation and does not change during the fermentation. At the start of the fermentation, the culture medium is inoculated with the desired organism. In this method, the fermentation is allowed to occur without adding any components to the system. Typically, batch fermentation qualifies as a "batch" with respect to the addition of a carbon source, and often attempts are made to control factors such as pH and oxygen concentration. The metabolite and biomass composition of the batch system changes continuously until the fermentation stops. In a typical batch culture, the cells can progress through a stationary lag phase to a high-growth logarithmic phase and finally enter a stationary phase where the growth rate decreases or stops. If untreated, the cells in the stationary phase will eventually die. Usually, the cells in the logarithmic phase are responsible for the bulk production of the product.
[0149] A suitable variation of the standard batch system is the "fed-batch" fermentation system. In this variant of the typical batch system, the substrate is added incrementally as the fermentation progresses. The fed-batch system is useful when catabolite repression might inhibit the metabolism of the cells and when a limited amount of substrate is desired in the culture medium. Measurement of the actual substrate concentration in the fed-batch system is difficult and thus it is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of exhaust gases (e.g., CO2). Batch and fed-batch fermentations are commonly used and known in the art.
[0150] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation typically maintains the culture at a constant high density, where the cells are mainly in the logarithmic growth phase. Continuous fermentation allows for the regulation of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient (such as a carbon source or a nitrogen source) is maintained at a fixed rate, and all other parameters are allowed to be adjusted. In other systems, many factors that affect growth can be continuously changed while the cell concentration measured by the turbidity of the medium remains constant. Continuous systems strive to maintain steady-state growth conditions. Therefore, the cell loss caused by the withdrawal of the medium should be balanced with the cell growth rate in the fermentation. Methods for regulating nutrients and growth factors for continuous fermentation processes and techniques for maximizing the product formation rate are well known in the field of industrial microbiology.
[0151] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, which include separating the host cells from the culture medium by centrifugation or filtration, or, if desired, disrupting the cells and removing the supernatant from the cell fractions and debris. Typically, after clarification, the protein component of the supernatant or filtrate is precipitated with a salt (e.g., ammonium sulfate). The precipitated protein is then dissolved and can be purified by various chromatographic procedures (e.g., ion-exchange chromatography, gel filtration).
[0152] V. Protein of interest
[0153] The protein of interest (POI) of the present disclosure can be any endogenous or heterologous protein, and it can be a variant of such a POI. The protein can contain one or more disulfide bridges, or it can be a protein whose functional form is monomeric or polymeric, i.e., the protein has a quaternary structure and is composed of multiple identical (homologous) or different (heterologous) subunits, where the POI or its variant POI is preferably a POI with the desired characteristics.
[0154] For example, in certain embodiments, the mutant or modified (recombinant) Gram-positive cells of the present disclosure produce at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more of the POI compared to their unmodified (parent or control) cells.
[0155] In certain embodiments, the mutant or modified Gram-positive cells of the present disclosure exhibit an increased specific productivity (Qp) of the POI relative to control cells. For example, the measurement of specific productivity (Qp) is a suitable method for evaluating protein production. The specific productivity (Qp) can be determined using the following equation:
[0156] “Qp = gP / gDCW·hr”
[0157] where “gP” is the grams of protein produced in the vessel; “gDCW” is the grams of dry cell weight (DCW) in the vessel; and “hr” is the fermentation time in hours starting from the inoculation time, which includes the production time as well as the growth time.
[0158] Thus, in certain other embodiments, relative to unmodified (parental / control) cells, the mutant or modified Gram-positive cells of the present disclosure contain an increase in specific productivity (Qp) of at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more.
[0159] In certain other embodiments, the mutant or modified Gram-positive cells contain an enhanced / increased ilvE messenger RNA (mRNA) level relative to the control cells' ilvE mRNA level. Suitable methods for such mRNA detection and analysis are generally known to those skilled in the art and include, but are not limited to, real-time quantitative PCR (RT qPCR) analysis, RNA sequencing, etc. In certain embodiments, the ilvE mRNA of the mutant or modified Gram-positive cells is increased by at least about 0.1%, at least about 0.5%, at least about 1.0%, at least about 5% to about 10% relative to the ilvE mRNA level of the unmodified (parental / control) cells.
[0160] In certain other embodiments, the mutant or modified Gram-positive cells contain an enhanced / increased carbon yield phenotype when expressing / producing one or more proteins of interest. In certain embodiments, the enhanced / increased carbon yield (i.e., when expressing / producing one or more proteins of interest) can be referred to as enhanced / increased carbon yield efficiency. For example, product formation in Gram-positive bacterial cells is a biotransformation process in which the chemotrophic nutrients fed to the bacterial cells during fermentation are converted into metabolites.
[0161] In certain other embodiments, variant, modified or mutant Gram-positive cells exhibit increased total protein yield, where total protein yield is defined as the amount of the protein of interest produced (g) per total carbohydrate equivalent of each batch and carbohydrate feed relative to the (unmodified / control) parental strain. Thus, as used herein, the total protein yield (g / g) can be calculated using the following equation:
[0162] "Yf = Tp / Tc"
[0163] where "Yf" is the total protein yield (g / g), "Tp" is the total protein of interest produced (g) during fermentation, and "Tc" is the total carbohydrate equivalent (g) of the batch and carbohydrate feeds during the fermentation (bioreactor) run. In certain embodiments, the increase in total protein yield of the modified strain (i.e., relative to the control strain) is at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to the unmodified (parental) cells.
[0164] The total protein carbon yield can also be described as carbon conversion efficiency / carbon yield, e.g., in the form of the percentage (%) of the carbon of the batch and feed incorporated into the total protein of interest. Thus, in certain embodiments, relative to the (control) parental strain, the variant Bacillus strain comprises increased carbon conversion efficiency (e.g., an increase in the percentage (%) of the carbon of the batch and feed incorporated into the total protein). In certain embodiments, the increase in carbon conversion efficiency of the modified strain (i.e., relative to the control strain) is at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to the unmodified (parental / control) cells.
[0165] Conventional methods / techniques known to those skilled in the art can be used to evaluate / determine enhanced carbon yield, enhanced carbon yield efficiency, etc. In one or more embodiments, the modified cells comprising enhanced carbon yield are fermented under suitable conditions for the production of the protein of interest, where the enhanced carbon yield is the result of more efficient incorporation of the nutrients in the fermentation medium into the protein product, thereby demonstrating enhanced protein productivity or yield coefficient (Y).
[0166] In certain embodiments, the POI or its variant POI is selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0167] Thus, in certain embodiments, the POI or its variant POI is an enzyme selected from Enzyme Commission (EC) numbers EC 1, EC 2, EC 3, EC 4, EC 5, or EC 6.
[0168] There are various assays known to those of ordinary skill in the art for detecting and measuring the activity of proteins expressed intracellularly and extracellularly.
[0169] VI. Exemplary Embodiments
[0170] Non-limiting examples of the compositions and methods disclosed herein are as follows:
[0171] 1. A variant ilvE gene that contains a mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene.
[0172] 2. The variant ilvE gene according to Example 1, which encodes a functional IlvE protein.
[0173] 3. The variant ilvE gene according to Example 1, wherein the mutation is a single nucleotide polymorphism (SNP) mutation in the 5′-UTR of the ilvE gene.
[0174] 4. The variant ilvE gene according to Example 4, wherein the mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, where the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the wild-type (WT) ilvE 5′-UTR sequence of SEQ ID NO:17.
[0175] 5. The variant ilvE gene as described in Example 1, which has at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the wild-type Bacillus subtilis ilvE gene of SEQ ID NO:25.
[0176] 6. The variant ilvE gene as described in Example 1, wherein the ilvE 5′-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO:18 and thymine (T) at nucleotide position 73.
[0177] 7. The variant ilvE gene as described in Example 2, wherein the functional IlvE protein has at least about 95% sequence identity with the native IlvE protein of SEQ ID NO:15.
[0178] 8. The variant ilvE gene as described in Example 1, wherein the WT ilvE gene promoter is replaced by a heterologous promoter.
[0179] 9. A synthetic ilvE gene construct that, in the 5′ to 3′ direction, comprises a wild-type ilvE gene promoter or a heterologous promoter sequence operably linked to a mutant ilvE 5′-UTR sequence, which mutant ilvE 5′-UTR sequence is operably linked to an ilvE gene CDS encoding a functional IlvE protein.
[0180] 10. The gene construct as described in Example 9, wherein the mutant ilvE 5′-UTR comprises a single nucleotide polymorphism (SNP) mutation in the 5′-UTR of the ilvE gene.
[0181] 11. The gene construct as described in Example 10, wherein the mutant ilvE 5′-UTR sequence comprises at least 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO:18 and thymine (T) at nucleotide position 73.
[0182] 12. A mutant Bacillus subtilis cell that comprises a variant ilvE gene comprising a mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene.
[0183] 13. The mutant cell as described in Example 12, which contains a single nucleotide polymorphism (SNP) in the 5′-UTR of the ilvE gene, wherein the mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, and wherein the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the wild-type (WT) ilvE 5′-UTR sequence of SEQ ID NO:17.
[0184] 14. The mutant cell as described in Example 12, which produces one or more target proteins.
[0185] 15. The mutant cell as described in Example 14, wherein the one or more target proteins are selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, lectin, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0186] 16. The mutant cell as described in Example 12, wherein the variant ilvE gene has at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the wild-type Bacillus subtilis ilvE gene of SEQ ID NO:25.
[0187] 17. The mutant cell as described in Example 12, wherein the ilvE 5′-UTR contains at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO:18 and thymine (T) at nucleotide position 73.
[0188] 18. The mutant cell as described in Example 12, which encodes a functional IlvE protein.
[0189] 19. The mutant cell as described in Example 12, wherein the WT ilvE gene promoter is replaced by a heterologous promoter.
[0190] 20. The mutant cell as described in Example 19, wherein when fermented under the same conditions, the heterologous promoter overexpresses the ilvE gene relative to the WT ilvE promoter.
[0191] 21. The mutant cell as described in Example 14, wherein when the mutant cell and the control cell are fermented under the same conditions for producing the one or more target proteins, the mutant cell has an enhanced carbon yield phenotype compared to a control cell that produces the same one or more target proteins and contains the wild-type ilvE gene.
[0192] 22. The mutant cell as described in Example 12, wherein when the mutant cell and the control cell are fermented under the same conditions, the mutant cell has an increased level of ilvE messenger RNA (mRNA) compared to a control cell containing the WT ilvE gene.
[0193] 23. The mutant cell as described in Example 22, which has an increased level of ilvE mRNA at about sixteen (16) hours of fermentation compared to a control cell.
[0194] 24. The mutant cell as described in Example 22, which has an increased level of ilvE mRNA at about twenty-four (24) hours of fermentation compared to a control cell.
[0195] 25. The mutant cell as described in Example 22, which has an increased level of ilvE mRNA at about thirty-two (32) hours of fermentation compared to a control cell.
[0196] 26. A genetically modified Bacillus subtilis cell derived from a parental cell containing a wild-type (WT) ilvE gene, wherein the modified cell contains a variant ilvE gene that contains a mutation in the 5′-UTR sequence of the ilvE gene.
[0197] 27. The modified cell as described in Example 26, which contains a single nucleotide polymorphism (SNP) mutation in the 5′-UTR sequence of the ilvE gene.
[0198] 28. The modified cell as described in Example 27, wherein the SNP mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, where the nucleotide positions of the 5′-UTR are numbered corresponding to the WT ilvE 5′-UTR sequence of SEQ ID NO:17.
[0199] 29. The modified cell as described in Example 26, which produces one or more target proteins.
[0200] 30. The modified cell as described in Example 28, wherein the one or more target proteins are selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, lectin, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0201] 31. The modified cell as described in Example 26, wherein the variant ilvE gene has at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the wild-type Bacillus subtilis ilvE gene of SEQ ID NO: 25.
[0202] 32. The modified cell as described in Example 26, wherein the ilvE 5′-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 18 and a thymine (T) at nucleotide position 73.
[0203] 33. The modified cell as described in Example 26, which encodes a functional IlvE protein.
[0204] 34. The modified cell as described in Example 26, wherein the WT ilvE gene promoter is replaced by a heterologous promoter.
[0205] 35. The modified cell as described in Example 34, wherein when fermented under the same conditions, the heterologous promoter overexpresses the ilvE gene relative to the WT ilvE promoter.
[0206] 36. The modified cell as described in Example 29, when the modified cell and the control cell are fermented under the same conditions for producing the one or more target proteins, the modified cell has an enhanced carbon yield phenotype compared to the control cell that produces the same one or more target proteins and contains the wild-type ilvE gene.
[0207] 37. The modified cell as described in Example 26, when the modified cell and the control cell are fermented under suitable conditions, the modified cell has an increased ilvE messenger RNA (mRNA) level compared to the control cell containing the WT ilvE gene.
[0208] 38. The modified cell as described in Example 37, compared to the control cell, the modified cell has an increased ilvE mRNA level at about sixteen (16) hours of fermentation.
[0209] 39. The modified cell as described in Example 37, compared to the control cell, the modified cell has an increased ilvE mRNA level at about twenty-four (24) hours of fermentation.
[0210] 40. The modified cell as described in Example 37, compared to the control cell, the modified cell has an increased ilvE mRNA level at about thirty-two (32) hours of fermentation.
[0211] 41. A method for increasing the ilvE messenger RNA (mRNA) level in recombinant Bacillus subtilis cells, the method comprising (a) obtaining parental Bacillus subtilis cells containing the wild-type (WT) ilvE gene and replacing the WT ilvE gene with a variant ilvE gene that contains a mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene, and (b) fermenting the parental cells and the modified cells under the same conditions for at least about sixteen (16) hours, wherein the modified cells have an increased level of ilvE mRNA compared to the parental cells.
[0212] 42. The method as described in Example 41, wherein the variant ilvE gene contains a single nucleotide polymorphism (SNP) mutation in the ilvE 5′-UTR sequence.
[0213] 43. The method as described in Example 42, wherein the SNP mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the wild-type (WT) ilvE 5′-UTR sequence of SEQ ID NO:17.
[0214] 44. A method for increasing the ilvE messenger RNA (mRNA) level in a modified Bacillus subtilis cell, the method comprising: (a) obtaining a parental Bacillus subtilis containing the wild-type (WT) ilvE gene and mutating the 5′-untranslated region (5′-UTR) of the WT ilvE gene to obtain a modified Bacillus subtilis cell containing a mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene, and (b) fermenting the parental cells and the modified cells under the same conditions for at least about sixteen (16) hours, wherein the modified cells contain an increased level of ilvE mRNA compared to the parental cells.
[0215] 45. The method according to embodiment 44, wherein the variant ilvE gene contains a single nucleotide polymorphism (SNP) mutation in the ilvE 5′-UTR sequence.
[0216] 46. The method according to embodiment 45, wherein the SNP mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the WT ilvE 5′-UTR sequence of SEQ ID NO:17.
[0217] 47. A method for increasing the ilvE messenger RNA (mRNA) level in a modified Bacillus subtilis cell, the method comprising: (a) obtaining a parental Bacillus subtilis containing the wild-type (WT) ilvE gene and mutating the 5′-untranslated region (5′-UTR) of the WT ilvE gene to obtain a modified Bacillus subtilis cell containing a variant ilvE gene, and (b) fermenting the parental cells and the modified cells under the same conditions for at least about sixteen (16) hours, wherein the modified cells contain an increased level of ilvE mRNA compared to the parental cells.
[0218] 48. The method according to embodiment 47, wherein the variant ilvE gene contains a single nucleotide polymorphism (SNP) mutation in the ilvE 5′-UTR sequence.
[0219] 49. The method according to embodiment 48, wherein the SNP mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the WT ilvE 5′-UTR sequence of SEQ ID NO:17.
[0220] 50. A method for increasing the carbon yield of a heterologous protein produced in a modified Bacillus subtilis cell, the method comprising (a) obtaining or constructing a parental Bacillus subtilis cell that produces a heterologous protein of interest (POI) and replacing the wild-type (WT) ilvE gene with a variant ilvE gene, and (b) fermenting the parental cells and the modified cells for at least about sixteen (16) hours under the same suitable conditions for producing the POI, wherein the carbon yield efficiency of the POI produced by the modified cells is increased compared to the parental cells.
[0221] 51. The method according to embodiment 50, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5′-UTR sequence.
[0222] 52. The method according to embodiment 51, wherein the SNP mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the WT ilvE 5′-UTR sequence of SEQ ID NO:17.
[0223] 53. A method for increasing the carbon yield of a heterologous protein expressed / produced in a modified Bacillus subtilis cell, the method comprising: (a) obtaining a parental Bacillus subtilis comprising a wild-type (WT) ilvE gene and mutating the 5′-untranslated region (5′-UTR) of the WT ilvE gene to obtain a modified Bacillus subtilis cell comprising a variant ilvE gene, and (b) fermenting the parental cells and the modified cells for at least about sixteen (16) hours under the same conditions, wherein the carbon yield efficiency of the POI produced by the modified cells is increased compared to the parental cells.
[0224] 54. The method according to embodiment 53, wherein the variant ilvE gene comprises a single nucleotide polymorphism (SNP) mutation in the ilvE 5′-UTR sequence.
[0225] 55. The method according to embodiment 54, wherein the SNP mutation in the 5′-UTR is a cytosine (C) to thymine (T) mutation at position 73, wherein the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the WT ilvE 5′-UTR sequence of SEQ ID NO:17.
[0226] 56. The method according to any one of Examples 41 - 55, wherein the variant ilvE gene has at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the wild-type Bacillus subtilis ilvE gene of SEQ ID NO:25.
[0227] 57. The method according to any one of Examples 41 - 55, wherein the mutated ilvE 5′-UTR comprises at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO:18 and thymine (T) at nucleotide position 73.
[0228] 58. The method according to any one of Examples 41 - 55, wherein the WT ilvE gene promoter is replaced by a heterologous promoter.
[0229] 59. The method according to Example 58, wherein when fermented under the same conditions, the heterologous promoter overexpresses the ilvE gene relative to the WT ilvE promoter.
[0230] 60. The method according to any one of Examples 41, 44, 47, 50 or 55, wherein the heterologous POI is selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, lectin, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0231] 61. The method according to any one of Examples 41 - 55, wherein the variant ilvE gene encodes a functional IlvE protein.
[0232] 62. The method according to Example 41 or Example 44, wherein when fermented for at least about sixteen (16) hours under the same conditions, the modified cells contain increased levels of ilvE messenger RNA (mRNA) relative to the control cells.
[0233] 63. The method according to Example 41 or Example 44, wherein when fermented for at least about twenty-four (24) hours under the same conditions, the modified cell contains an increased level of ilvE mRNA relative to the control cell.
[0234] 64. The method according to Example 41 or Example 44, wherein when fermented for at least about thirty-two (32) hours under the same conditions, the modified cell contains an increased level of ilvE mRNA relative to the control cell.
[0235] 65. The method according to Example 50 or Example 53, wherein when fermented for at least about sixteen (16) hours under the same conditions, the carbon yield efficiency of the modified cell is increased relative to the parental cell.
[0236] 66. The method according to Example 50 or Example 53, wherein when fermented for at least about twenty-four (24) hours under the same conditions, the carbon yield efficiency of the modified cell is increased relative to the parental cell.
[0237] 67. The method according to Example 50 or Example 53, wherein when fermented for at least about thirty-two (32) hours under the same conditions, the carbon yield efficiency of the modified cell is increased relative to the parental cell.
[0238] Examples
[0239] Certain embodiments of the present disclosure can be further understood from the following examples, which should not be construed as limiting. Modifications to the materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989).
[0240] Example 1
[0241] Construction of a Bacillus protease reporter strain
[0242] As generally described above, the applicant has identified mutant Bacillus subtilis cells (strain CZ437) with an enhanced protein production phenotype. In particular, the applicant performed next-generation sequencing (NGS) on the mutant CZ437 strain to further characterize the observed enhanced protein productivity phenotype, in which an unexpected SNP in the Bacillus subtilis ilvE 5′-UTR sequence was identified. For example, the DNA sequences of the WT ilvE 5′-UTR (SEQ ID NO:17) and the mutant ilvE 5′-UTR (SEQ ID NO:18) are presented in Figure 1 A and Figure 1In B. As Figure 1 shown, the WT ilvE 5′-UTR sequence (SEQ ID NO:17; Figure 1 A) contains cytosine (C) at nucleotide position 73 ( + 73C), and the mutant ilvE 5′-UTR sequence (SEQ ID NO:18; Figure 1 B) contains thymine (T) at nucleotide position 73 ( + 73T). In this example, the applicant constructed a recombinant Bacillus subtilis strain expressing a heterologous reporter (GG36) protein to evaluate the mutant ilvE 5′-UTR sequence (SEQ ID NO:18) associated with the enhanced protein production phenotype identified in the mutant Bacillus subtilis CZ437 strain. In particular, the DNA fragments described herein were assembled using standard molecular biology techniques and used as templates to develop linear DNA expression cassettes for integration into the Bacillus subtilis strains described herein.
[0243] A. Construction of the reporter protein expression cassette.
[0244] The construction of the reporter protein cassette was carried out as follows: A first (1st) DNA fragment (5′skfA FR; SEQ ID NO:9) containing the 5′ skfA flanking region (FR) sequence of Bacillus subtilis was operably linked to an expression cassette that contains an upstream (5′) Bacillus subtilis P2 promoter, which was operably linked to a DNA sequence containing the wild-type Bacillus subtilis aprE 5′-untranslated region (5′-UTR; SEQ ID NO:1), which was operably linked to a DNA sequence encoding the wild-type Bacillus subtilis aprE signal sequence (SEQ ID NO:2), which was operably linked to a DNA sequence encoding a variant Bacillus lentus propeptide sequence (SEQ ID NO:4), which was operably linked to a DNA sequence encoding the mature (GG36) subtilisin reporter gene (SEQ ID NO:6), which was operably linked to the BPN′ terminator sequence (SEQ ID NO:8), which was operably linked to the 3′ skfH FR sequence (3′ skfH FR; SEQ ID NO:10).
[0245] A second (2nd) DNA fragment comprising a 5′ yhfN flanking region (FR) sequence in a chromosomal region containing a 5′ aprE flanking region (FR) sequence (5′ aprE FR; SEQ ID NO:11) of Bacillus subtilis is operably linked to an expression cassette comprising an upstream (5′) Bacillus subtilis P2 promoter operably linked to a DNA sequence (SEQ ID NO:1) containing the wild-type Bacillus subtilis aprE 5′-UTR, which DNA sequence (SEQ ID NO:1) is operably linked to a DNA (SEQ ID NO:2) encoding the wild-type Bacillus subtilis aprE signal sequence, which DNA (SEQ ID NO:2) is operably linked to a DNA sequence (SEQ ID NO:4) encoding a variant Bacillus lentus propeptide sequence, which DNA sequence (SEQ ID NO:4) is operably linked to a DNA sequence (SEQ ID NO:6) encoding a mature (GG36) subtilisin reporter gene, which DNA sequence (SEQ ID NO:6) is operably linked to a BPN′ terminator (SEQ ID NO:8). The GG36 subtilisin reporter gene expression cassette is further linked to a Bacillus subtilis alanine racemase (alrA) gene (SEQ ID NO:12) and a 3′ aprE FR sequence (3′ aprE FR; SEQ ID NO:13).
[0246] Mutagenesis of the B. ilvE transcriptional leader sequence
[0247] As briefly described above, a cytosine (C) to thymine (T) mutation at position 73 of the ilvE 5′-UTR was introduced into the genome of Bacillus subtilis using random strain mutagenesis. Specifically, the above-described 1st and 2nd cassettes were integrated into a Bacillus subtilis strain (reporter strain A; 73T) containing a position 73C to T (SNP) mutation in the ilvE 5′-UTR and an isogenic Bacillus subtilis strain (reporter strain B; 73C) containing the wild-type ilvE 5′-UTR.
[0248] More specifically, Bacillus subtilis strain A (mutant ilvE 5′-UTR; SEQ ID NO:18) and isogenic Bacillus subtilis strain B (WT ilvE 5′-UTR; SEQ ID NO:17) were fermented in a large-scale (approx. 14 L) fermenter using standard fermentation conditions. As presented in Table 1 below, the relative improvement in carbon yield of mutant Bacillus subtilis reporter strain A was significantly enhanced compared to isogenic Bacillus subtilis reporter strain B.
[0249] Table 1
[0250] Relative improvement in carbon yield of Bacillus subtilis reporter strain A containing mutant ilvE 5′-UTR
[0251]
[0252] C. Replace the wild-type ilvE promoter with a heterologous Hbs promoter
[0253] Two (2) expression cassettes expressing the ilvE amino acid transaminase were constructed using conventional molecular biology techniques. More particularly, the ilvE expression cassette contains an upstream (5′) wild-type (WT) Bacillus subtilis hbs promoter (Phbs) region sequence (SEQ ID NO:26) operably linked to a downstream (3′) DNA sequence, which downstream (3′) DNA sequence contains a WT ilvE transcriptional leader sequence (WT 5′-UTR; SEQ ID NO:17) or a mutant ilvE transcriptional leader sequence (mutant 5′-UTR; SEQ ID NO:18) operably linked to a downstream (3′) WT ilvE gene CDS (SEQ ID NO:14). Thus, the hbs promoter region (SEQ ID NO:26) drives the expression of the two cassettes, where the DNA sequence of the WT ilvE gene cassette is shown in SEQ ID NO:27, and the sequence of the mutant ilvE gene cassette is shown on SEQ ID NO:28. More particularly, the cassette (SEQ ID NO:27 or SEQ ID NO:28) was integrated into the spoIIIAA genomic locus of the parental Bacillus subtilis strain containing two GG36 reporter protein cassettes.
[0254] As generally shown in Table 2 below, under standard fermentation conditions, Bacillus subtilis strains overexpressing the ilvE gene under the control of the hbs promoter and containing the WT ilvE 5′-UTR (strain CZ477) or the mutant ilvE 5′-UTR (strain CZ488) were fermented in a large-scale (approx. 14L) bioreactor and compared to strain CZ450 containing the WT ilvE promoter and the WT ilvE 5′-UTR. As presented in Table 2, the two strains with overexpression of the ilvE gene showed an increase in carbon efficiency compared to the strain with the WT ilvE promoter.
[0255] Table 2
[0256] Relative improvement in carbon yield of Bacillus subtilis reporter strains overexpressing the ilvE gene using the hbs promoter
[0257]
[0258] Thus, in one or more embodiments of the present disclosure, as shown in this example, highly expressed heterologous promoter region sequences (e.g., hbs promoter, etc.) can be used to overexpress a variant ilvE gene (or its ilvE gene expression construct) comprising a WT ilvE 5′-UTR sequence operably linked to a downstream WT ilvE gene CDS (or a variant ilvE gene CDS encoding a functional ilvE protein), and / or to overexpress a variant ilvE gene (or its ilvE gene expression construct) comprising a mutated ilvE 5′-UTR sequence operably linked to a downstream WT ilvE gene CDS (or a variant ilvE gene CDS encoding a functional ilvE protein). Thus, when cultured under suitable conditions, such methods, genetic elements, expression constructs, modified cells comprising an enhanced protein production phenotype, modified cells comprising an enhanced carbon yield phenotype, etc. are particularly beneficial in terms of the expression / production of the protein of interest.
[0259] Example 2
[0260] Real-time quantitative PCR RNA analysis
[0261] In this example, time-course samples from Bacillus subtilis reporter strains A (mutated ilvE 5′-UTR) and B (WT ilvE 5′-UTR) were used for real-time quantitative PCR (RT qPCR) analysis. For example, samples were collected at 8, 16, 24, and 32 hours of fermentation and total RNA extraction was performed. Specifically, the extracted RNA samples were treated with DNase-I to remove genomic DNA from the samples, and then cDNA was synthesized using a Transcriptor First Strand cDNA Synthesis Kit (Roche). Subsequently, 1000-fold diluted cDNA from each sample was used as a template for qPCR, and the ftsY gene was used as a housekeeping gene for data normalization. For example, sequence-specific ilvE forward (SEQ ID NO:19) and reverse (SEQ ID NO:20) primers, and an ilvE probe (SEQ ID NO:21) were used to amplify a sequence within the ilvE gene. Similarly, sequence-specific ftsY forward (SEQ ID NO:22) and reverse (SEQ ID NO:23) primers, and an ftsY probe (SEQ ID NO:24) were used to amplify a sequence within the ftsY gene.
[0262] Specifically, the RT qPCR time-course experimental data are presented in Figure 4 where the 2 - ΔΔC T method (Livak and Schmittgen, 2001) was used to calculate the log fold change between the housekeeping ftsY gene and the ilvE gene. AsFigure 4 As shown, the bars and values represent the fold change in ilvE mRNA of Bacillus subtilis strain A (mutant ilvE 5′-UTR) compared to isogenic Bacillus subtilis strain B (WT ilvE 5′-UTR). In particular, except for the first time point at 8 hours of fermentation (i.e., where the amount of ilvE mRNA differed by only 0.84-fold), compared to isogenic Bacillus subtilis strain B (WT ilvE 5′-UTR), at the 16, 24, and 32-hour fermentation time points, the amount of ilvE mRNA in Bacillus subtilis strain A (mutant ilvE 5′-UTR) was significantly increased by 2.91, 1.93, and 1.79-fold, respectively ( Figure 4 ). Based on the foregoing, the results presented and described herein demonstrate that the C to T SNP (SEQ ID NO:18) identified herein significantly increases the level of ilvE mRNA.
[0263] References
[0264] PCT Publication No. WO 2002 / 14490
[0265] PCT Publication No. WO 2003 / 083125
[0266] PCT Publication No. WO 2020 / 112609
[0267] Ausubel et al., “Current Protocols in Molecular Biology”, published by Greene Publishing Assoc. and Wiley-Interscience (1987).
[0268] Belitsky and Sonenshein, “Roadblock repression of transcription by Bacillus subtilis CodY”, J. Mol. Biol. 26; 411(4):729–743, 2011.
[0269] Berger et al., “Methionine Regeneration and Aminotransferases in Bacillus subtilis, Bacillus cereus, and Bacillus anthracis”, J. Bacteriol. 185(8):2418-2431, 2003.
[0270] Caspers et al., “Improvement of Sec-dependent secretion of a heterologous model protein in Bacillus subtilis by saturation mutagenesis of the N-domain of the AmyE signal peptide”, Appl. Microbiol. Biotechnol., 86(6):1877-1885, 2010.
[0271] Earl et al., “Ecology and genomics of Bacillus subtilis”, Trends in Microbiology., 16(6):269-275, 2008.
[0272] Livak and Schmittgen, “Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2 - ΔΔC T Method”, Methods. Vol.25, pp.402–408, 2001.
[0273] et al., “Transcriptional organization and posttranscriptional regulation of the Bacillus subtilis branched-chain amino acid biosynthesis genes”, J Bacteriol.186(8):2240-2252, 2004.
[0274] Olempska-Beer et al., “Food-processing enzymes from recombinant microorganisms--a review”’ Regul. Toxicol. Pharmacol., 45(2):144-158, 2006.
[0275] Sambrook et al., “Molecular Cloning: A Laboratory Manual” Cold Spring Harbor Laboratory: Cold Spring Harbor, N.Y. (1989), (2001) and (2012).
[0276] Van Dijl and Hecker, “Bacillus subtilis: from soil bacterium to super-secreting cell factory”, Microbial Cell Factories, 12(3), 2013.
Claims
1. A mutant ilvE 5′-untranslated region (5′-UTR) nucleic acid sequence, which comprises SEQ ID NO:
18.
2. A variant ilvE gene, which comprises a single nucleotide polymorphism (SNP) mutation in the 5′-untranslated region (5′-UTR) of the ilvE gene, wherein the ilvE gene has at least 90% identity with the wild-type Bacillus subtilis (B. subtilis) ilvE gene of SEQ ID NO:2 and encodes a native IlvE protein, and the SNP in the 5′-UTR of the ilvE gene is a cytosine (C) to thymine (T) mutation at position 73 of the ilvE 5′-UTR sequence as shown in SEQ ID NO:
18.
3. A synthetic ilvE gene, which comprises a wild-type (WT) ilvE gene promoter or a heterologous gene promoter operably linked to a mutant ilvE 5′-UTR sequence comprising SEQ ID NO:18, and the mutant ilvE 5′-UTR sequence is operably linked to a wild-type (WT) ilvE gene CDS.
4. A mutant Bacillus subtilis cell, which comprises a single nucleotide polymorphism (SNP) in the 5′-untranslated region (5′-UTR) of the ilvE gene, wherein the SNP is a cytosine (C) to thymine (T) mutation at position 73, and the nucleotide positions of the ilvE 5′-UTR are numbered corresponding to the wild-type (WT) ilvE 5′-UTR of SEQ ID NO:
17.
5. A genetically modified Bacillus subtilis cell derived from a parental cell comprising a wild-type (WT) ilvE gene, wherein the modified cell comprises an introduced synthetic ilvE gene, and the introduced synthetic ilvE gene comprises a heterologous gene promoter operably linked to a mutant ilvE 5′-UTR sequence comprising SEQ ID NO:18, and the mutant ilvE 5′-UTR sequence is operably linked to a WT ilvE gene CDS.
6. The mutant cell according to claim 5, which produces one or more target proteins.
7. The modified cell according to claim 6, which produces one or more target proteins.
8. The modified cell according to claim 6, wherein the heterologous gene promoter overexpresses the ilvE gene.
9. A method for increasing the level of ilvE messenger RNA (mRNA) in a modified Bacillus subtilis cell, the method comprising: (a) Obtain parental Bacillus subtilis cells containing the wild-type (WT) ilvE gene and replace the WT ilvE gene with a synthetic ilvE gene, the synthetic ilvE gene comprising a heterologous gene promoter operably linked to a mutant ilvE 5′-UTR sequence comprising SEQ ID NO:18, the mutant ilvE 5′-UTR sequence operably linked to the WT ilvE gene CDS, and (b) Ferment the parental cells and the modified cells under the same conditions for at least about sixteen (16) hours, wherein the modified cells contain an increased level of ilvE mRNA compared to the parental cells.
10. The method according to claim 10, wherein the synthesized ilvE gene has at least 90% identity with the ilvE gene of SEQ ID NO: 25 and contains thymine (T) at nucleotide position 73 of the ilvE 5′-UTR.
11. A method for increasing the carbon yield of a heterologous protein produced in a modified Bacillus subtilis cell, the method comprising: (a) Obtain or construct parental Bacillus subtilis cells that produce a heterologous protein of interest (POI) and introduce into the cells a synthetic ilvE gene, the synthetic ilvE gene comprising a heterologous gene promoter operably linked to a mutant ilvE 5′-UTR sequence comprising SEQ ID NO:18, the mutant ilvE 5′-UTR sequence operably linked to the WT ilvE gene CDS, and (b) Ferment the parental cells and the modified cells under the same suitable conditions for producing the POI for at least about sixteen (16) hours, wherein the carbon productivity efficiency of the POI produced by the modified cells is increased compared to the parental cells.
12. The method according to claim 12, wherein the synthesized ilvE gene has at least 90% identity with the ilvE gene of SEQ ID NO: 25 and contains thymine (T) at nucleotide position 73 of the ilvE 5′-UTR.
Citation Information
Patent Citations
Bacillus transformation, transformants and mutant libraries
WO2002014490A2
Ehanced protein expression in bacillus
WO2003083125A1
Novel promoter sequences and methods thereof for enhanced protein production in bacillus cells
WO2020112609A1