Novel promoter and 5' untranslated region mutations that enhance protein production in Gram-positive cells
Novel promoter and 5' UTR sequences with specific mutations improve protein production in Bacillus strains by increasing expression levels, addressing the need for enhanced production in industrial biotechnology.
Patent Information
- Application Number
- JP2025512641
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-09-01
- Publication Date
- 2025-09-04
AI Technical Summary
There is a need for improved promoter and 5' untranslated region (UTR) nucleic acid sequences to enhance protein expression and production in Gram-positive bacterial host cells, such as Bacillus species, which are commonly used in industrial biotechnology for producing proteins like enzymes and antibodies, as existing promoters do not adequately increase production levels.
Introduction of novel promoter and 5' UTR nucleic acid sequences, including specific mutations, into recombinant polynucleotides and expression cassettes, which are operably linked with protein signal and pro-region sequences, to enhance protein production in Gram-positive bacterial cells.
The novel sequences result in increased protein production, particularly for enzymes and antibodies, leading to cost and time savings in industrial biotechnology by enhancing expression levels in engineered Bacillus strains.
Smart Images

Figure 2025529135000008 
Figure 2025529135000009 
Figure 2025529135000010
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to the fields of microbial host cells, molecular biology, protein engineering, fermentation, protein production, etc. Particular aspects of the disclosure relate to novel promoter and 5' untranslated region nucleic acid (DNA) sequences.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 374,450, filed September 02, 2022, the entire disclosure of which is incorporated herein by reference.
[0003] Sequence Listing Reference The contents of the electronic submission of the Sequence Listing text file entitled "NB41861-WO-PCT_SequenceListing.xml", created on August 30, 2023, is 53KB in size, and is incorporated herein by reference in its entirety. [Background technology]
[0004] Genetic engineering has facilitated various improvements in host microorganisms used as industrial bioreactors or cell factories. For example, Gram-positive bacterial host cells can produce and secrete many useful proteins and metabolites. The most common Bacillus species used in industry are B. licheniformis, B. amyloliquefaciens, and B. subtilis. These Bacillus strains have generally recognized as safe (GRAS) status and are therefore natural candidates for producing proteins used in the food and pharmaceutical industries. For example, important enzymes produced include α-amylase, neutral protease, alkaline (or serine) protease, etc. However, despite advances in knowledge about protein production in Bacillus host cells, there remains a need for methods and compositions for improving the expression / production of these proteins by microorganisms.
[0005] Recombinant production of a protein of interest (POI) encoded by a gene (or ORF) of interest is typically accomplished by constructing an expression vector suitable for use in a desired host cell, in which a nucleic acid encoding the desired POI is placed under the expression control of a promoter. Thus, the expression vector is introduced into the host cell by various techniques (e.g., by transformation), and production of the desired protein product is achieved by culturing the transformed host cell under conditions suitable for expression and production of the protein product. For example, Bacillus sp. promoters (and their associated elements) have been described for recombinant expression of functional polypeptides (e.g., Kim et al., 2008 and U.S. Pat. No. 4,559,300). While numerous promoters are known, there remains a need in the art for novel promoter (nucleic acid) sequences that can improve expression of heterologous nucleic acids encoding proteins of interest. For example, in the technical field of industrial biotechnology, even relatively small increases in expression / production levels of industrially important proteins (e.g., enzymes, antibodies, receptors, etc.) can result in significant cost, energy, and time savings for the recombinant protein produced. Summary of the Invention [Means for solving the problem]
[0006] As outlined herein, the present disclosure provides, inter alia, compositions and methods for producing proteins of interest in Gram-positive bacterial (host) cells. Certain embodiments relate to novel promoter and 5'-UTR nucleic acid (DNA) sequences, recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.) comprising the novel promoter / 5'-UTR sequences, recombinant polynucleotides comprising the novel promoter / 5'-UTR sequences operably linked to a DNA sequence encoding a protein signal (secretion) sequence and / or a DNA sequence encoding a pro-region sequence, operably linked to a DNA sequence encoding a protein of interest, and the like.
[0007] Accordingly, certain embodiments of the present disclosure provide a variant nucleic acid sequence comprising at least one mutation set forth in any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant nucleic acid sequence are numbered according to SEQ ID NO:1. In certain other embodiments, the variant nucleic acid comprises the nucleotide sequence of any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant sequence are numbered according to SEQ ID NO:1. In certain other embodiments, the variant nucleic acid sequence comprises at least about 97.5% identity to SEQ ID NO:1. In related aspects, the variant nucleic acid sequences of the present disclosure may be referred to as variant promoter / 5'-untranslated region (5-UTR) sequences.
[0008] Certain other embodiments relate to polynucleotides (DNA) comprising the variant nucleic acid sequences of the present disclosure. Accordingly, certain embodiments provide polynucleotides comprising the variant nucleic acid sequences of the present disclosure operably linked to a downstream (3') nucleic acid sequence encoding a protein of interest.
[0009] In certain other embodiments or aspects, a polynucleotide comprising a variant nucleic acid sequence of the present disclosure is operably linked to a downstream (3') nucleic acid sequence encoding a pro-region sequence operably linked to a downstream nucleic acid sequence encoding a mature protein of interest (POI). In certain other embodiments or aspects, a polynucleotide comprising a variant nucleic acid sequence of the present disclosure is operably linked to a downstream (3') nucleic acid sequence encoding a protein signal (secretory) sequence (SS) operably linked to a downstream nucleic acid sequence encoding a mature protein of interest (POI). In certain other embodiments or aspects, a polynucleotide comprising a variant nucleic acid sequence of the present disclosure is operably linked to a downstream (3') nucleic acid sequence encoding a protein signal (secretory) sequence (SS) operably linked to a downstream nucleic acid sequence encoding a pro-region (PRO) sequence operably linked to a downstream nucleic acid sequence encoding a mature protein of interest (POI).
[0010] In certain related embodiments, the POI is selected from the group consisting of an enzyme, an antibody, a receptor protein, a lectin, and a regulatory protein.
[0011] Other embodiments provide expression cassettes comprising the polynucleotides of the present disclosure. In related embodiments, a Gram-positive bacterial (host) cell comprises one or more transfer cassettes of the present disclosure. In certain embodiments, the Gram-positive host cell is a Bacillus sp. cell. In related embodiments, the Bacillus sp. (host) cell is selected from the group consisting of B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophilus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis.
[0012] Certain other embodiments of the present disclosure provide methods for producing a protein of interest (POI) in Gram-positive bacterial cells, the method comprising: (a) introducing into the Gram-positive cells an expression cassette comprising a variant nucleic acid sequence of any one of SEQ ID NOs: 8 through 46 operably linked to a downstream (3') nucleic acid sequence encoding the protein of interest (POI); and (b) culturing the engineered cells under conditions suitable for producing the POI. In certain preferred embodiments of the method, the engineered cells produce an increased amount of the POI compared to (vis-a-vis) a control Gram-positive cell comprising an introduced expression cassette comprising a reference nucleic acid sequence of SEQ ID NO: 1 operably linked to a downstream (3') nucleic acid sequence encoding the same POI, wherein the engineered and control cells are cultured under the same conditions. In certain other embodiments of the method, the engineered cells produce an increased amount of the POI compared to the control cell after at least about 72 hours of culture. In still other embodiments of the method, the protein of interest (POI) is selected from the group consisting of an enzyme, an antibody, a receptor protein, a lectin, and a regulatory protein.In certain embodiments of the method, the enzyme is selected from the group consisting of acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, insulin In certain other embodiments of the method, the Gram-positive bacterial cell is selected from the group consisting of: ATPase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectin degrading enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof. In certain other embodiments of the method, the Gram-positive bacterial cell is a Bacillus sp. cell. [Brief explanation of the drawings]
[0013] [Figure 1] The nucleotide sequence of the DNA sequence of the variant rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1) is shown. More specifically, as shown in Figure 1, the variant (reference) rrnI-P2 promoter / 5'-UTR region sequence comprises nucleotide positions 1-149 of SEQ ID NO: 1, where the "UP," "-35," "-10," and "Shine-Dalgarno" sequence elements are shown in bold.
[0014] [Figure 2]Figure 2A shows a DNA sequence alignment of the reference variant rrnI-P2 promoter / 5'-UTR (SEQ ID NO: 1) and specific SEL variant promoter / 5'-UTR region sequences of the present disclosure. More specifically, as shown in Figures 2A and 2B, the nucleotide positions of the SEL variant sequences described in this example are aligned with the reference rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1; nucleotide positions 1-149). As shown in Figure 2A, nucleotide positions 1-83 of the reference rrnI-P2 promoter region contain the UP, -35, and -10 elements, shown in gray shading. As shown in Figure 2B, nucleotide positions 83-149 of the reference rrnI-P2 promoter / 5'-UTR region contain the Shine-Dalgarno element, shown in gray shading. As shown in Figure 2A, the modified nucleotide position in the rrnI-P2 promoter region is indicated by the black-shaded nucleotide residue (e.g., Figure 2A, UTR-00664; TGA).
[0015] [Figure 3] Figure 3A shows a DNA sequence alignment of the reference variant rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1) and specific SEL variant promoter / 5'-UTR region sequences of the present disclosure. More specifically, as shown in Figures 3A and 3B, the nucleotide positions of the SEL variant sequences described in this example are aligned with the reference rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1; nucleotide positions 1-149). As shown in Figure 3A, nucleotide positions 1-83 of the reference rrnI-P2 promoter / 5'-UTR region contain the UP element, -35, and -10 elements, shown in gray shading. As shown in Figure 3B, nucleotide positions 83-149 of the reference rrnI-P2 promoter / 5'-UTR region contain the Shine-Dalgarno element, shown in gray shading. As shown in Figure 3A / 3B, the altered nucleotide positions in the rrnI-P2 promoter region are indicated by black shaded nucleotide residues (e.g., Figure 3A, UTR-00798; A). DETAILED DESCRIPTION OF THE INVENTION
[0016] A brief description of biological sequences SEQ ID NO: 1 is the nucleic acid (DNA) sequence of the variant (reference) rrnI-P2 promoter / 5'-UTR region.
[0017] SEQ ID NO:2 is the amino acid sequence of the wild-type Bacillus gibsonii subtilisin designated "BG46."
[0018] SEQ ID NO: 3 is the amino acid sequence of a variant B. gibsonii BG46 subtilisin designated "BG46_variant."
[0019] SEQ ID NO: 4 is the DNA sequence encoding the AprE protein signal sequence of wild-type B. subtilis.
[0020] SEQ ID NO: 5 is the DNA sequence encoding the wild-type B. lentus pro region sequence.
[0021] SEQ ID NO: 6 is the DNA sequence of the wild-type B. amyloliquefaciens BPN' terminator.
[0022] SEQ ID NO:7 is the DNA sequence of the kanamycin (kan) gene expression cassette.
[0023] SEQ ID NO: 8 is the DNA sequence of variant UTR-00664.
[0024] SEQ ID NO: 9 is the DNA sequence of variant UTR-00692.
[0025] SEQ ID NO: 10 is the DNA sequence of variant UTR-00330.
[0026] SEQ ID NO: 11 is the DNA sequence of variant UTR-00411.
[0027] SEQ ID NO: 12 is the DNA sequence of variant UTR-00325.
[0028] SEQ ID NO: 13 is the DNA sequence of variant UTR-00730.
[0029] SEQ ID NO: 14 is the DNA sequence of variant UTR-00348.
[0030] SEQ ID NO: 15 is the DNA sequence of variant UTR-00738.
[0031] SEQ ID NO: 16 is the DNA sequence of variant UTR-00788.
[0032] SEQ ID NO: 17 is the DNA sequence of variant UTR-00792.
[0033] SEQ ID NO: 18 is the DNA sequence of variant UTR-00800.
[0034] SEQ ID NO: 19 is the DNA sequence of variant UTR-01018.
[0035] SEQ ID NO: 20 is the DNA sequence of variant UTR-01112.
[0036] SEQ ID NO: 21 is the DNA sequence of variant UTR-00037.
[0037] SEQ ID NO: 22 is the DNA sequence of variant UTR-00039.
[0038] SEQ ID NO: 23 is the DNA sequence of variant UTR-00661.
[0039] SEQ ID NO: 24 is the DNA sequence of variant UTR-00891.
[0040] SEQ ID NO: 25 is the DNA sequence of variant UTR-00084.
[0041] SEQ ID NO: 26 is the DNA sequence of variant UTR-00362.
[0042] SEQ ID NO: 27 is the DNA sequence of variant UTR-00424.
[0043] SEQ ID NO: 28 is the DNA sequence of variant UTR-00643.
[0044] SEQ ID NO: 29 is the DNA sequence of variant UTR-00645.
[0045] SEQ ID NO: 30 is the DNA sequence of variant UTR-00741.
[0046] SEQ ID NO: 31 is the DNA sequence of variant UTR-00798.
[0047] SEQ ID NO: 32 is the DNA sequence of variant UTR-00960.
[0048] SEQ ID NO: 33 is the DNA sequence of variant UTR-01223.
[0049] SEQ ID NO: 34 is the DNA sequence of variant UTR-00656.
[0050] SEQ ID NO: 35 is the DNA sequence of variant UTR-00657.
[0051] SEQ ID NO: 36 is the DNA sequence of variant UTR-00030.
[0052] SEQ ID NO: 37 is the DNA sequence of variant UTR-01092.
[0053] SEQ ID NO: 38 is the DNA sequence of variant UTR-00721.
[0054] SEQ ID NO: 39 is the DNA sequence of variant UTR-00651.
[0055] SEQ ID NO: 40 is the DNA sequence of variant UTR-00301.
[0056] SEQ ID NO: 41 is the DNA sequence of variant UTR-00187.
[0057] SEQ ID NO: 42 is the DNA sequence of variant UTR-00035.
[0058] SEQ ID NO: 43 is the DNA sequence of variant UTR-00005.
[0059] SEQ ID NO: 44 is the DNA sequence of variant UTR-00863.
[0060] SEQ ID NO: 45 is the DNA sequence of variant UTR-00711.
[0061] SEQ ID NO: 46 is the DNA sequence of variant UTR-00752.
[0062] SEQ ID NO: 47 is the DNA sequence of the 5'aprE gene FR.
[0063] SEQ ID NO: 48 is the DNA sequence of the 3'aprE gene FR.
[0064] As briefly described above and in more detail herein below, the present disclosure provides, inter alia, novel promoter and 5'-UTR nucleic acid (DNA) sequences; recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.) comprising the novel promoter / 5'-UTR sequences; recombinant polynucleotides comprising the novel promoter / 5'-UTR sequences operably linked to a DNA sequence encoding a protein of interest operably linked to a DNA sequence encoding a protein signal (secretion) sequence and / or a DNA sequence encoding a pro-region sequence;
[0065] In certain aspects, the disclosure provides recombinant Gram-positive bacterial strains that express one or more introduced polynucleotides encoding a protein of interest. In certain other aspects, the disclosure provides compositions and methods for designing / constructing recombinant Gram-positive bacterial strains that express one or more introduced novel polynucleotide constructs encoding a protein of interest, compositions and methods for culturing recombinant strains that express a protein of interest, compositions and methods for enhancing production of a protein of interest, etc.
[0066] I. Definition With respect to the recombinant polynucleotides, recombinant (engineered) strains, and methods thereof described herein, the following terms and phrases are defined. Terms not defined herein should be given the meaning commonly used in the art.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the compositions and methods of the present invention apply. Although any methods and materials similar or equivalent to those described herein can also be used to practice or test the compositions and methods of the present invention, exemplary methods and materials are described below. All publications and patents cited herein are incorporated herein by reference in their entirety.
[0068] It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a prelude to the use of exclusive terminology such as "solely," "only," "excluding," "not including," or the use of any "negative" limitation or qualification in connection with the recitation of claim elements.
[0069] It will be apparent to those skilled in the art upon reading this disclosure that each of the individual embodiments described and illustrated herein has distinct components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the compositions and methods described herein. Any described method may be carried out in the order of events recited or in any other order which is logically possible.
[0070] As used herein, the terms "Gram-positive bacteria," "Gram-positive cells," "Gram-positive bacterial strains," and / or "Gram-positive bacterial cells" have the same meaning as used in the art. For example, Gram-positive bacterial cells include all strains of Actinobacteria and Firmicutes. In certain embodiments, such Gram-positive bacteria are of the classes Bacilli, Clostridia, and Mollicutes.
[0071] As used herein, the term "Bacillus" includes all species within the genus "Bacillus" known to those skilled in the art, such as B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkal ... Examples of Bacillus species include, but are not limited to, B. lus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis. It is recognized that the genus Bacillus continues to undergo taxonomic reorganization. Thus, the genus is intended to include organisms such as reclassified species, for example, but not limited to, "B. stearothermophilus," which is now referred to as "Geobacillus stearothermophilus."
[0072] As used herein, the terms "recombinant" or "non-naturally occurring" refer to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic change or that has been modified by the introduction of a heterologous nucleic acid molecule, or to a cell (e.g., a microbial cell) that has been modified so that expression of a heterologous or endogenous nucleic acid molecule or gene can be controlled. Recombinant also refers to a cell that is derived from, or is the progeny of, a non-naturally occurring cell that has one or more such modifications. Genetic modifications include, for example, modifications that introduce an expressible nucleic acid molecule encoding a protein, or the addition, deletion, substitution, or other functional change of other nucleic acid molecules in the genetic material of a cell. For example, recombinant cells may express genes or other nucleic acid molecules not found in the same or homologous form in native (wild-type) cells (e.g., fusion or chimeric proteins), or may provide an altered expression pattern of an endogenous gene, such as overexpressed, underexpressed, minimally expressed, or not expressed at all. "Recombination," "recombining," or producing a "recombinant" nucleic acid generally refers to the assembly of two or more nucleic acid fragments, which assembly gives rise to a chimeric gene.
[0073] The term "derived" includes the terms "originating from," "obtained from," "obtainable from," and "made from," and generally indicates that one particular material or composition finds its origin in another material or composition, or has characteristics that can be described with reference to that other particular material or composition. For example, the recombinant Gram-positive bacterial cells of the present disclosure can be derived / obtained from any known Gram-positive bacterial strain.
[0074] As used herein, "nucleic acid" refers to DNA, cDNA, and RNA of genomic or synthetic origin, which may be double-stranded or single-stranded, including nucleotide or polynucleotide sequences and fragments or portions thereof, and whether representing the sense or antisense strand. It will be understood that, as a result of the degeneracy of the genetic code, a large number of nucleotide sequences can encode a given protein.
[0075] The polynucleotides (or nucleic acid molecules) described herein are understood to include "genes," "vectors," and "plasmids."
[0076] Thus, the term "gene" refers to a polynucleotide that encodes a specific sequence of amino acids, including all or part of a protein's coding sequence, and may include regulatory (non-transcribed) DNA sequences, such as promoter sequences, that determine the conditions under which the gene is expressed. The transcribed region of a gene may include untranslated regions (UTRs), including 5'-untranslated regions (UTRs) and 3'-UTRs, as well as the coding sequence.
[0077] As used herein, "endogenous gene" refers to a gene that is present in its natural location in the genome of an organism.
[0078] As used herein, a "heterologous" gene, "non-endogenous" gene, or "foreign" gene refers to a gene that is not normally found in the host organism, but that has been introduced into the host organism by gene transfer. The term "foreign" gene includes a native gene inserted into a non-native organism and / or a chimeric gene inserted into a natural or non-native organism.
[0079] As used herein, "heterologous regulatory sequence" refers to a gene expression control sequence (e.g., promoter, enhancer, terminator, etc.) that does not naturally function to regulate (control) the expression of a gene of interest. Generally, heterologous nucleic acids are not endogenous (natural) to the cell or part of the genome in which they are present, but have been added to the cell by infection, transfection, transformation, transduction, microinjection, electroporation, etc. A "heterologous" nucleic acid construct can contain regulatory sequence / DNA coding (ORF) sequence combinations that are the same as or different from regulatory sequence / DNA coding sequence combinations found in the native host cell.
[0080] As used herein, the term "expression" refers to the transcription and stable accumulation of sense (mRNA) or antisense RNA derived from a nucleic acid molecule of the present disclosure. Expression can also refer to the translation of mRNA into a polypeptide. Thus, the term "expression" includes any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.
[0081] As used herein, the term "coding sequence" (CDS) refers to a nucleotide sequence that directly specifies the amino acid sequence of its (encoded) protein product. The boundaries of the coding sequence are generally determined by an open reading frame (hereinafter "ORF"), which usually begins with the ATG start codon. Coding sequences typically include DNA, cDNA, and recombinant nucleotide sequences.
[0082] As used herein, the terms "promoter," "promoter element," "promoter sequence," and the like refer to a nucleic acid (DNA) sequence capable of controlling transcription of a gene coding sequence (CDS) into messenger RNA (mRNA) when the promoter region sequence is located upstream (5') and operably linked to a downstream (3') gene CDS. As generally understood by those skilled in the art, a promoter generally provides a site for specific binding and transcription initiation by RNA polymerase. In certain aspects, the term "promoter" refers to the minimal portion of a promoter nucleic acid sequence necessary for transcription initiation (i.e., including the RNA polymerase binding site). For example, a promoter generally includes a "-10" (consensus sequence) element and a "-35" (consensus sequence) element, which are upstream (5') of the gene CDS to be translated and relative to the +1 transcription start site (TSS). The core promoter -10 and -35 elements are commonly referred to in the art as the "TATAAT" (Pribnow box) consensus region and the "TTGACA" consensus region, respectively. The core promoter (-10 and -35) regions are generally spaced apart (ie, spaced apart) by about 15-20 intervening base pairs (nucleotides), as shown in FIG.
[0083] Promoters may be derived entirely from native genes, or may be composed of various elements from various naturally occurring promoters, or may even include synthetic nucleic acid segments. It is understood by those skilled in the art that various promoters may direct the expression of genes in various cell types, at various developmental stages, or in response to various environmental or physiological conditions. Promoters may be constitutive, inducible, variable, hybrid, synthetic, tandem, and the like. The promoter that most frequently drives gene expression in the majority of cell types is generally referred to as a "constitutive promoter." Furthermore, because the exact boundaries of regulatory sequences are in most cases not completely defined, it has been recognized that DNA fragments of various lengths may have the same promoter activity.
[0084] In certain embodiments, an upstream (5') promoter sequence (pro) operably linked to a downstream DNA sequence encoding a protein signal sequence (SS) operably linked to a downstream DNA sequence encoding a pro region sequence (PRO) operably linked to a downstream (3') DNA sequence encoding a mature protein of interest (ORF) may be depicted diagrammatically as 5'-[pro]-[SS]-[PRO]-[ORF]-3'.
[0085] As used herein, a "functional promoter sequence" that controls expression of a gene of interest linked to a protein-coding sequence of the gene of interest refers to a promoter sequence that controls the transcription and translation of the coding sequence in a desired Gram-positive host cell. For example, in certain embodiments, the present disclosure provides a polynucleotide comprising an upstream (5') promoter (or 5' promoter region or tandem 5' promoter, etc.) that functions in a Gram-positive cell, wherein the functional promoter region is operably linked to a nucleic acid sequence encoding a protein of interest.
[0086] As used herein, the term "precursor protein" refers to an inactive form of a protein. In certain embodiments, a full-length protein is synthesized as a prosequence, a precursor of the mature protein form (abbreviated "preprotein"). In other embodiments, a full-length protein is synthesized as a signal peptide sequence, a prosequence, and a precursor of the mature protein form (abbreviated "preproprotein"). For example, the presequence usually acts as a signal peptide for transport, and the prosequence is generally essential for correct folding of the associated (mature) protein.
[0087] As used herein, the term "mature protein" refers to the active form of a protein, as opposed to the inactive precursor (full-length) protein.
[0088] As used herein, the terms "signal sequence," "secretion signal," and "signal peptide" may be used interchangeably and refer to a sequence of amino acid residues that may be involved in the secretion or direct transport of a precursor protein. A signal (pre) sequence is generally cleaved from the precursor protein by a signal peptidase during translocation. A signal (pre) sequence is generally located at the N-terminus of the mature protein sequence or at the N-terminus of the pro-region (pro) sequence when the signal (pre) sequence and pro-region (pro) sequence are used in operable combination upstream (5') of the mature POI sequence.
[0089] As used herein, phrases such as "variant rrnI-P2 promoter / 5'-UTR region" and / or "reference rrnI-P2 promoter / 5'-UTR region" specifically refer to the DNA sequence of the variant B. subtilis rrnI-P2 promoter / 5'-UTR region set forth in SEQ ID NO: 1 (see, e.g., Figure 1). Thus, in certain aspects, the variant rrnI-P2 promoter / 5'-UTR region sequence (SEQ ID NO: 1) may be referred to as a reference sequence, or control sequence, particularly when compared to one or more SEL variant promoter / 5'-UTR region sequences of the present disclosure. For example, as shown in Figures 2 and 3, the reference rrnI-P2 promoter / 5'-UTR region sequence (nucleotide positions 1-149) has been aligned with certain site evaluation library (SEL) variant promoter / 5'-UTR region sequences of the present disclosure. More specifically, the SEL variant promoter / 5'-UTR region sequence (see, e.g., Table 2; SEQ ID NOS: 8-48) was aligned with the reference rrnIp2 promoter / 5'-UTR region sequence, where nucleotide positions 1-83 of the reference promoter / 5'-UTR region sequence and the SEL variant are shown in Figures 2A and 3A, respectively, and nucleotide positions 84-149 of the reference promoter / 5'-UTR region sequence and the same SEL variant are shown in Figures 2B and 3B, respectively. As annotated in Figure 1B, DNA sequence elements of the reference promoter / 5'-UTR region sequence (SEQ ID NO: 1) include (in the 5' to 3' direction) the upstream (UP) element, the -35 element, the -10 element, and the Shine-Dalgarno (SD) element, which are indicated by bolded nucleotides.
[0090] In certain aspects, a promoter comprises nucleotides upstream (5') of the promoter, and these upstream (5') nucleotides are referred to herein as "upstream promoter elements" (abbreviated "UP elements" or "UP sequences"). Thus, as used herein, "UP sequence" refers to an "A+T"-rich (nucleic acid) sequence region located upstream of the -35 core promoter element. The UP sequence can also be described as a nucleic acid sequence region located upstream of the -35 core promoter element that directly interacts with the C-terminal domain of the α-subunit of RNA polymerase. Thus, in certain embodiments, a promoter comprises one (or more) UP sequences located upstream of the promoter and operably linked to the promoter.
[0091] As used herein, the phrase "transcription start site" (abbreviated "TIS") generally refers to the base pairs at which transcription begins. By convention, transcription start sites (TIS) in the DNA sequence of a transcription unit are assigned numbers with nucleotide positions extending in the direction of transcription (i.e., 3'; downstream) being assigned positive "(+)" numbers, and nucleotide positions extending in the opposite direction (i.e., 5'; upstream) being assigned negative "(-)" numbers.
[0092] As used herein, the phrases "translation start site" (abbreviated "tss") and "translation start site (tss) codon" may be used interchangeably and refer to the three-nucleotide translation start site (tss) codon. For example, prokaryotic "tss codons" include, but are not limited to, "AUG," "GUG," "UGG," etc.
[0093] As used herein, the term "Shine-Dalgarno" sequence (abbreviated "SD" sequence) refers to a messenger RNA (mRNA) ribosome binding site, generally located approximately 8 nucleotides upstream (5') of a start codon (e.g., AUG). As understood in the art, the SD sequence helps recruit ribosomes to mRNA to initiate protein synthesis by aligning the ribosome with the start codon (e.g., AUG), where transfer RNA (t-RNA) can add amino acids in a sequence determined by the codon and move downstream (3') from the translation start site (TSS).
[0094] As used herein, the terms "pro sequence," "pro-sequence," and "pro region sequence" may be used interchangeably and may be abbreviated as "PRO" sequence, "Pro" sequence, "pro" sequence, etc. As used herein, the term pro sequence has the same meaning as understood in the art. For example, the B. subtilis alkaline serine protease "subtilisin" is initially produced as a preprosubtilisin, which consists of a signal (pre) sequence for protein secretion, followed by a 77-amino acid pro region (pro) sequence, followed by an amino acid sequence encoding mature subtilisin (e.g., preprosubtilisin). Pro sequences act as intramolecular chaperones (e.g., directly catalyzing protein folding reactions) and are often essential for the correct folding of the associated (mature) protein. Similarly, pro sequences may be required for both the folding and intracellular transport (or secretion) of the mature protein of interest, suggesting that these two functions are closely related. In certain aspects, the pro region sequence of the present disclosure comprises an amino acid sequence derived from the wild-type (WT, see) B. lentus pro region sequence of SEQ ID NO:5.
[0095] As used herein, the phrase "polynucleotide encoding a full-length protein" refers to a DNA sequence encoding a "precursor" protein, and the phrase "polynucleotide encoding a mature protein" refers to a DNA sequence encoding a "mature" protein, as defined herein. In certain embodiments, a polynucleotide encoding a precursor protein comprises at least an upstream (5') DNA sequence encoding a pro-region amino acid sequence operably linked to a downstream (3') DNA sequence (e.g., open reading frame, ORF) encoding the amino acid sequence of the mature protein of interest (POI). In other embodiments, a polynucleotide encoding a precursor protein comprises at least an upstream (5') DNA sequence encoding a protein signal sequence operably linked to a downstream (3') DNA sequence encoding a pro-region amino acid sequence operably linked to a downstream (3') ORF encoding the amino acid sequence of the mature POI.
[0096] As used herein, the term "untranslated region" may be abbreviated as "UTR."
[0097] As used herein, the terms "5' prime (5') untranslated region," "5' untranslated region," and / or "5' transcript leader" can be used interchangeably and can be abbreviated as "5'-UTR." As generally understood in the art, the 5'-UTR is known as the region of a messenger RNA (mRNA) that is immediately upstream (5') of the start codon.
[0098] A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, DNA encoding a secretory leader (i.e., signal sequence) is operably linked to DNA encoding a polypeptide if it is expressed as a preprotein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence (CDS, ORF) if it affects the transcription of the sequence; or a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation. Generally, "operably linked" means that the DNA sequences being linked are contiguous, and, in the case of a secretory leader, contiguous and in reading phase. Enhancers, however, need not be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide adapters or linkers are used in accordance with conventional practice. Thus, the term operably linked generally refers to the association (juxtaposition) of nucleic acid sequences on a single nucleic acid fragment so that the function of one is affected by the function of the other. For example, a promoter (pro) is operably linked to a gene coding sequence (gene CDS) if it controls the transcription of the gene CDS.
[0099] As used herein, DNA encoding the "wild-type B. subtilis aprE signal peptide sequence" may be abbreviated as "aprE SS," which comprises the nucleotide sequence of SEQ ID NO:4.
[0100] As used herein, DNA encoding the "wild-type B. lentus pro-peptide region DNA sequence," which may be abbreviated as "PRO sequence," "PRO region," or "PRO," comprises the nucleotide sequence of SEQ ID NO:5.
[0101] As used herein, the "wild-type B. amyloliquefaciens BPN' terminator (BPN' term)", which may be abbreviated as "term", comprises the nucleotide sequence of SEQ ID NO:6.
[0102] As used herein, the wild-type Bacillus gibsonii (BG46) subtilisin comprising the amino acid sequence of SEQ ID NO:2 is abbreviated as "WT BG46."
[0103] As used herein, a variant Bacillus gibsonii (BG46) subtilisin comprising the amino acid sequence of SEQ ID NO:3 is abbreviated as "BG46 variant," which BG46 variant is derived from the wild-type B. gibsonii (BG46) subtilisin reporter protein (SEQ ID NO:2).
[0104] As used herein, "upstream (5') aprE flanking region (FR) sequence" includes SEQ ID NO:49.
[0105] As used herein, the "downstream (3') aprE flanking region (FR) sequence" comprises SEQ ID NO: 50 and contains a "kanamycin (Kan) gene expression cassette" for selection (SEQ ID NO: 7).
[0106] As used herein, exemplary proteases may be referred to as "reporter proteins." In one or more specific embodiments of the present disclosure, exemplary reporter proteins are expressed / produced by one or more recombinant (modified) cells of the present disclosure. In certain embodiments, reporter proteins include, but are not limited to, native and variant Bacillus sp. subtilisins.
[0107] As used herein, the term "subtilisin" refers to any member of the S8 serine protease family described in MEROPS - The Peptidase Database (Rawlings et al., 2006). The term subtilisin includes the wide variety of Bacillus subtilisins that have been identified and sequenced, such as subtilisin 168, subtilisin BPN', subtilisin Carlsberg, and mutant (variant) proteases derived therefrom.
[0108] In certain one or more embodiments, exemplary subtilisin reporters include, but are not limited to, naturally occurring B. clausii subtilisin and functional variants thereof, naturally occurring B. gibsonii subtilisin and functional variants thereof, naturally occurring B. lentus subtilisin and functional variants thereof, naturally occurring B. licheniformis subtilisin (AprL) and functional variants thereof, naturally occurring B. subtilis subtilisin (AprE) and functional variants thereof, naturally occurring B. amyloliquefaciens subtilisin (BPN') and functional variants thereof, and the like. In certain aspects, exemplary B. clausii, B. gibsonii, and / or B. lentus subtilisin reporters may be referred to as alkaline proteases. For example, alkaline subtilisins generally have an isoelectric point (pI) of about 9.5, while B. licheniformis, B. subtilis, and B. amyloliquefaciens subtilisins have a pI of about 6.5.
[0109] In certain embodiments, the present disclosure relates to one or more variant subtilisins derived from a parent (natural) subtilisin sequence, such as a natural B. subtilis subtilisin (e.g., 168), a natural B. amyloliquefaciens (e.g., BPN'), a natural B. licheniformis subtilisin (e.g., Carlsberg), a natural B. lentus subtilisin (e.g., 309), a natural B. alcalophilus subtilisin (e.g., PB92), etc. One of ordinary skill in the art can readily design, construct, screen, and identify functional subtilisin variants using routine methods known in the art. In particular, WO 2010 / 056634, WO 2011 / 130222, WO 2015 / 089447, WO 2016 / 202839, WO 2017 / 207762, and WO 2023 / 114936 (each of which is incorporated by reference in its entirety) describe suitable methods and compositions for constructing functional subtilisin variants derived from naturally occurring B. clausii subtilisin, functional subtilisin variants derived from naturally occurring B. amyloliquefaciens subtilisin, functional subtilisin variants derived from naturally occurring B. gibsonii subtilisin, and the like.
[0110] As used herein, "suitable regulatory sequences" refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence that influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences can include promoters, transcription leader sequences, RNA processing sites, effector binding sites, and stem-loop structures.
[0111] As used herein, "host cell" refers to a cell that has the ability to act as a host or expression vehicle for a newly introduced DNA sequence. Thus, in certain embodiments of the present disclosure, the host cell is a Gram-positive cell (e.g., Bacillus sp.) and / or a Gram-negative cell (e.g., E. coli).
[0112] As used herein, an "modified cell" refers to a recombinant cell that contains at least one genetic modification that is not present in the parent, reference, or control cell from which the modified cell is derived.
[0113] As used herein, when comparing the expression and / or production of a protein of interest (POI) in a recombinant (modified) cell with the expression and / or production of the same POI in an unmodified (control) cell, it will be understood that the modified and unmodified cells are grown / cultured / fermented under identical conditions (e.g., identical conditions of medium, temperature, pH, etc.).
[0114] As used herein, "increased amount," when used in phrases such as "the recombinant cells 'express / produce' an increased amount of a protein of interest compared to unmodified (control) cells," specifically refers to the "increased amount" of a protein of interest (POI) expressed / produced by the recombinant cells, but this "increased amount" is always compared to unmodified (control) cells expressing / producing the same POI, and the modified and unmodified cells are grown / cultured / fermented under the same conditions.
[0115] As used herein, "increasing" protein production or "increased" protein production refers to an increased amount of a protein (e.g., a protein of interest) produced. The protein may be produced inside the host cell or secreted (or exported) into the culture medium. In certain embodiments, the protein of interest is produced (secreted) into the culture medium. Increased protein production can be detected, for example, as a higher maximum level of protein or enzyme activity (e.g., amylase activity) or total extracellular protein produced compared to the parent host cell.
[0116] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) the introduction, substitution, or removal of one or more nucleotides in a gene (or its ORF), or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of a gene or its ORF; (b) gene disruption; (c) gene conversion; (d) gene deletion; (e) gene downregulation; (f) directed mutagenesis; and / or (g) random mutagenesis of any one or more genes disclosed herein.
[0117] As used herein, the term "introducing," when used in phrases such as "introducing a gene, polynucleotide, open reading frame (ORF), gene coding sequence, vector, expression cassette, etc. into a Gram-positive bacterial cell," includes methods known in the art for introducing polynucleotides (DNA) into cells, including, but not limited to, protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation, etc.
[0118] As used herein, "transformed" or "transformation" refers to a cell that has been transformed through the use of recombinant DNA technology. Transformation generally occurs by inserting one or more nucleotide sequences (e.g., polynucleotides, ORFs, or genes) into a cell. The inserted nucleotide sequences may be heterologous nucleotide sequences (i.e., sequences that do not naturally occur in the cell being transformed). Thus, transformation generally refers to the introduction of exogenous DNA into a host cell such that the DNA is maintained as a chromosomal integrant or a self-replicating extrachromosomal vector.
[0119] As used herein, "transforming DNA," "transforming sequence," and "DNA construct" refer to DNA used to introduce a sequence into a host cell or organism. Transforming DNA is DNA used to introduce a sequence into a host cell or organism. This DNA can be generated in vitro by PCR or any other suitable technique. In some embodiments, the transforming DNA includes the incoming sequence, while in other embodiments, the transforming DNA further includes the incoming sequence flanked by homology boxes. In yet other embodiments, the transforming DNA includes other non-homologous sequences (i.e., stuffer sequences or flanking sequences) added to the ends. The ends can be closed such that the transforming DNA forms a closed circle, such as, for example, for insertion into a vector.
[0120] As used herein, "gene disruption" or "gene disruption" are used interchangeably and refer broadly to any genetic modification that substantially prevents a host cell from producing a functional gene product (e.g., a protein). Thus, as used herein, gene disruption includes, but is not limited to, frameshift mutations, premature stop codons (i.e., so that a functional protein is not produced), substitutions or internal deletions that eliminate or reduce the activity of a protein (so that a functional protein is not produced), insertions that disrupt the coding sequence, mutations that remove the operable link between the native promoter and the open reading frame required for transcription, and the like.
[0121] As used herein, "incoming sequence" refers to a DNA sequence that is introduced into a bacterial cell chromosome. In some embodiments, the incoming sequence is part of a DNA construct. In other embodiments, the incoming sequence encodes one or more proteins of interest. In some embodiments, the incoming sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it may be a homologous or heterologous sequence). In some embodiments, the incoming sequence encodes one or more proteins of interest, genes, and / or mutant or modified genes. In alternative embodiments, the incoming sequence encodes a functional wild-type gene or operon, a functional mutant gene or operon, or a non-functional gene or operon. In some embodiments, a non-functional sequence can be inserted into a gene to disrupt the function of the gene. In another embodiment, the incoming sequence comprises a selectable marker. In yet another embodiment, the incoming sequence comprises two homology boxes.
[0122] As used herein, a "homology box" refers to a nucleic acid sequence that is homologous to a sequence within a bacterial cell chromosome. More specifically, a homology box is an upstream or downstream region that shares about 80-100% sequence identity, about 90-100% sequence identity, or about 95-100% sequence identity with the coding region immediately flanking a gene or portion of a gene to be deleted, disrupted, inactivated, downregulated, etc., according to the present invention. These sequences direct where a DNA construct will integrate within the bacterial cell chromosome and which portion of the chromosome will be replaced by the incoming sequence. While not intended to limit the present disclosure, a homology box can comprise from about 1 base pair (bp) to 200 kilobases (kb). Preferably, the homology box comprises from about 1 bp to 10.0 kb; 1 bp to 5.0 kb; 1 bp to 2.5 kb; 1 bp to 1.0 kb; and 0.25 kb to 2.5 kb. The homology box may also comprise about 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb, and 0.1 kb. In some embodiments, the 5' and 3' ends of the selectable marker are flanked by homology boxes, which comprise nucleic acid sequences that immediately flank the coding region of the gene.
[0123] As used herein, host cell "genome," bacterial (host) cell "genome," or Bacillus sp. (host) cell "genome" includes chromosomal genes and extrachromosomal genes.
[0124] As used herein, the terms "plasmid," "vector," and "cassette" refer to an extrachromosomal element that typically carries genes that are not part of the cell's central metabolism and is usually in the form of a circular, double-stranded DNA molecule. Such elements can be linear or circular, single- or double-stranded, DNA or RNA autonomously replicating sequences, genome-integrating sequences, phage, or nucleotide sequences from any source in which multiple nucleotide sequences have been joined or recombined into a unique structure that can introduce into a cell a promoter fragment and DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences.
[0125] As used herein, the term "plasmid" refers to a circular double-stranded (ds) DNA construct that is used as a cloning vector and forms an extrachromosomal, self-replicating genetic element in many bacteria and some eukaryotes. In some embodiments, the plasmid becomes integrated into the genome of the host cell. In some embodiments, the plasmid is present in the parent cell and is lost in the daughter cells.
[0126] As used herein, "transformation cassette" refers to a specific vector containing a gene (or its ORF) and having elements in addition to the foreign gene that facilitate transformation of a particular host cell.
[0127] As used herein, the term "vector" refers to any nucleic acid that can replicate (multiply) within a cell and carry new genes or DNA segments into the cell. Thus, the term refers to a nucleic acid construct designed for transport between various host cells. Vectors include viruses, bacteriophages, proviruses, plasmids, phagemids, transposons, which are "episomal" (i.e., capable of autonomous replication or integration into the chromosomes of the host organism), as well as artificial chromosomes such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), and PLACs (plant artificial chromosomes).
[0128] An "expression vector" refers to a vector capable of incorporating and expressing heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and known to those skilled in the art. Selection of an appropriate expression vector is within the knowledge of one of ordinary skill in the art.
[0129] As used herein, the terms "expression cassette" and "expression vector" refer to nucleic acid constructs (i.e., they are vectors or vector elements as described above) that are recombinantly or synthetically produced with a set of specific nucleic acid elements that allow for transcription of a specific nucleic acid in a target cell. Recombinant expression cassettes can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter. In some embodiments, the DNA construct also includes a set of specific nucleic acid elements that allow for transcription of a specific nucleic acid in a target cell. In certain embodiments, the DNA construct of the present disclosure includes a selectable marker and an inactivated chromosomal segment or gene segment or DNA segment, as defined herein.
[0130] As used herein, a "targeting vector" is a vector that contains a polynucleotide sequence homologous to a region in a host cell chromosome into which the targeting vector is transformed and is capable of driving homologous recombination at that region. For example, targeting vectors are used to introduce mutations into a host cell chromosome by homologous recombination. In some embodiments, the targeting vector contains other non-homologous sequences (i.e., stuffer sequences or flanking sequences), for example, added to the ends. The ends can be closed such that the targeting vector forms a closed circle, e.g., by insertion into a vector. For example, in certain embodiments, parent B. licheniformis (host) cells are modified (e.g., transformed) by introducing one or more "targeting vectors" into them.
[0131] As used herein, the term "protein of interest" or "POI" refers to a polypeptide of interest desired to be expressed in an engineered (recombinant) Gram-positive host cell, where the POI is preferably expressed at increased levels (i.e., compared to "unmodified" (parent or control) cells). Thus, as used herein, a POI can be an enzyme, substrate-binding protein, surfactant protein, structural protein, receptor protein, etc. In certain embodiments, the engineered cells of the present disclosure produce an increased amount of a heterologous protein of interest compared to control cells. In certain embodiments, the increased amount of protein of interest produced by the engineered cells of the present disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or more than a 5.0% increase compared to control cells.
[0132] Similarly, as defined herein, "gene of interest" or "GOI" refers to a nucleic acid sequence (e.g., polynucleotide, gene, or ORF) that encodes a POI. A "gene of interest" that encodes a "protein of interest" may be a naturally occurring gene, a mutated gene, or a synthetic gene.
[0133] As used herein, the terms "polypeptide" and "protein" are used interchangeably and refer to polymers of any length comprising amino acid residues linked by peptide bonds. Conventional one-letter or three-letter codes for amino acid residues are used herein. Polypeptides can be linear or branched, can contain modified amino acids, and can be interrupted by non-amino acids. The term polypeptide also encompasses amino acid polymers that are modified, either naturally or by intervention, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within this definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids), as well as other modifications known in the art.
[0134] In certain embodiments, the genes of the present disclosure encode enzymes (e.g., acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysate, lysase), endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0135] As used herein, a "variant" polypeptide refers to a polypeptide derived from a parent (or reference) polypeptide by one or more amino acid substitutions, additions, or deletions, generally by recombinant DNA techniques. A variant polypeptide may differ from the parent polypeptide by a small number of amino acid residues and may be defined by the level of primary amino acid sequence homology / identity with the parent (reference) polypeptide. Preferably, a variant polypeptide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity with the parent (reference) polypeptide sequence.
[0136] As used herein, a "variant" polynucleotide refers to a polynucleotide that has a particular degree of sequence homology / identity with a parent polynucleotide or that hybridizes to a parent polynucleotide (or its complement) under stringent hybridization conditions. Preferably, a variant polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% nucleotide sequence identity with the parent (reference) polynucleotide sequence.
[0137] As used herein, the term "mutation" refers to any change or alteration in a nucleic acid sequence. There are several types of mutations, including point mutations, deletion mutations, silent mutations, frameshift mutations, splicing mutations, etc. Mutations can be made specifically (e.g., by site-directed mutagenesis) or randomly (e.g., by chemical agents, repair minus passaging through bacterial strains).
[0138] As used herein, the term "substitution" in reference to a polypeptide or sequence thereof means the replacement (ie, substitution) of one amino acid with another.
[0139] As used herein, the term "homology" refers to homologous polynucleotides or homologous polypeptides. When two or more polynucleotides or two or more polypeptides are homologous, this means that the homologous polynucleotides or polypeptides have a "degree of identity" of at least 50%, at least 60%, more preferably at least 70%, even more preferably at least 85%, even more preferably at least 90%, more preferably at least 95%, and most preferably at least 98%.
[0140] The degree of homology between sequences can be determined using any suitable method known in the art (see, for example, Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; programs such as GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package (Genetics Computer Group, Madison, WI); and Devereux et al., 1984). For purposes of the present invention, the degree of identity between two amino acid sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970) as implemented in the Needle program in the EMBOSS package (Rice et al., 2000), preferably version 3.0.0 or later. Optional parameters used may include a gap open penalty of 10, a gap extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. The output of Needle labeled "longest identity" (obtained using the nobrief option) is used as the percent identity, calculated as follows: (Identical residues × 100) / (alignment length – total number of gaps in the alignment)
[0141] For purposes of the present invention, the degree of identity between two deoxyribonucleotide sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, supra) as implemented in the Needle program in the EMBOSS package (Rice et al., 2000, supra), preferably version 3.0.0 or later. Optional parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EDNAFULL (EMBOSS version in NCBI NUC4.4) substitution matrix. The output of Needle labeled "Longest Identity" (obtained using the -nobrief option) is used as the percent identity, calculated as follows: (identical deoxyribonucleotides × 100) / (length of alignment – total number of gaps in alignment)
[0142] As used herein, the phrases "substantially similar" and "substantially identical," in the context of at least two nucleic acids or polypeptides, typically mean that the polynucleotide or polypeptide comprises a sequence having at least about 40% identity, at least about 50% identity, at least about 60% identity, at least about 70% identity, at least about 75% identity, at least about 80% identity, at least about 85% identity, at least about 90% identity, at least about 91% identity, at least about 92% identity, at least about 93% identity, at least about 94% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, or even at least about 99% identity or greater, compared to a reference (i.e., wild-type) sequence. Sequence identity can be determined using known programs such as BLAST, ALIGN, and CLUSTAL, using standard parameters.
[0143] As used herein, the term "percent identity" refers to the level of nucleic acid or amino acid sequence identity between nucleic acid sequences encoding polypeptides or between the amino acid sequences of polypeptides when aligned using a sequence alignment program.
[0144] As used herein, the term "specific productivity" refers to the total amount of protein produced per cell per unit time over a given period of time.
[0145] As used herein, the terms "purified," "isolated," or "enriched" mean that a biomolecule (e.g., a polypeptide or polynucleotide) has been altered from its native state by separation from some or all of the naturally occurring components with which it is naturally associated. Such isolation or purification can be carried out by separation techniques known in the art, such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulfate precipitation or other protein salting-out, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis, or gradient separation to remove unwanted whole cells, cell debris, impurities, extraneous proteins, or enzymes in the final composition. Purified or isolated biomolecule compositions can then be supplemented with components that confer additional benefits, such as activators, anti-inhibitors, desirable ions, pH-adjusting compounds, or other enzymes or chemicals.
[0146] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) the introduction, substitution, or removal of one or more nucleotides in a gene (or its ORF) or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of a gene or its ORF; (b) gene disruption; (c) gene conversion; (d) gene deletion; (e) gene downregulation; (f) directed mutagenesis; and / or (g) random mutagenesis of any one or more genes disclosed herein.
[0147] II. Novel promoter region variant sequences Certain Gram-positive bacterial promoters suitable for expression of proteins of interest have been described (see, e.g., Kim et al., 2008; U.S. Pat. No. 4,559,300; WO 2013 / 086219; and WO 2017 / 152169). As presented herein and described in the Examples below, Applicant designed and constructed a site evaluation library (SEL) to test / screen genetic modifications (mutations) to enhance recombinant protein productivity.
[0148] Specifically, genetic modifications were made to the promoter and 5'-UTR regions of expression constructs encoding exemplary reporter proteins, and the SEL constructs were introduced and expressed in recombinant Bacillus sp. cells. More specifically, as outlined in the Examples, mature subtilisin reporter proteins were expressed in B. subtilis under the control of variant promoter / 5'-UTR region sequences (see Table 2 and SEQ ID NOS: 8-46) constructed by site-scanning mutagenesis libraries (see Table 1), or under the control of the reference rrnI-P2 promoter / 5'-UTR region of SEQ ID NO: 1 (Example 1).
[0149] Recombinant strains expressing subtilisin reporter proteins under the control of the mutant promoter / 5'-UTR region sequences (Table 2) were compared to a reference strain expressing the same subtilisin reporter under the control of the (reference) rrnI-P2 promoter / 5'-UTR region sequence (SEQ ID NO: 1), as described in Example 2. Similarly, performance index (PI) values of recombinant strains showing increased protein productivity after 72 hours of growth compared to the control strain, as described in Example 3 (Table 3).
[0150] Specifically, as described in Example 3, approximately 50% of the constructed SEL variants contained mutations upstream (5') of the UP element, while approximately 22% of the SEL variants contained mutations in the UP element, compared to the reference rrnI-P2 promoter / 5'-UTR region of SEQ ID NO: 1 (see Figures 2A and 3A). Approximately 22% of the SEL variants contained mutations around the Shine-Dalgarno (SD) region, compared to the reference rrnI-P2 / 5'-UTR region of SEQ ID NO: 1 (see Figures 2B and 3B). In contrast, approximately 72% of SEL variants that contained mutations in the −35 and / or −10 promoter elements had performance index (PI) values of 0.1 to 0.5 after 24, 48, and 72 h of growth, and another 21% of SEL variants that contained mutations in the SD region had PI values of 0.1 to 0.5 after 24, 48, and 72 h of growth (data not shown).
[0151] Thus, as generally described above, certain aspects of the present disclosure relate to the novel mutant promoter / 5'-UTR region nucleic acids described herein. Accordingly, certain embodiments relate to such novel mutant promoter / 5'-UTR region nucleic acid (DNA) sequences suitable for expression of a gene coding sequence (CDS) of a protein of interest.
[0152] III. Recombinant Polynucleotides and Molecular Biology As generally described above, certain embodiments of the present disclosure relate to novel mutant promoter / 5'-UTR region sequences suitable for expression of a gene CDS encoding a protein of interest. In related aspects, the present disclosure provides recombinant polynucleotides comprising one or more mutant promoter / 5'-UTR region nucleic acid (DNA) sequences. Accordingly, certain embodiments relate to recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant Gram-positive bacterial cells / strains expressing a protein of interest, etc. In certain aspects, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells (strains) to enhance production of a protein of interest. In certain aspects, the polynucleotide constructs of the present disclosure are referred to as expression cassettes, which comprise, in a 5' to 3' direction and in operable combination, at least an upstream (5') promoter / 5'-UTR region DNA sequence linked to a downstream (3') gene CDS encoding a mature protein of interest (POI).
[0153] In certain embodiments, the expression cassette comprises a variant promoter / 5'-UTR region of the present disclosure operably linked to a downstream gene CDS encoding a mature POI. In certain other embodiments, the expression cassette of the present disclosure comprises one or more DNA sequence elements, including, but not limited to, a DNA sequence element encoding a protein / peptide signal (secretory) sequence (SS), a DNA sequence element encoding propeptide (proregion) amino acid residues (PRO), a DNA sequence element comprising a transcription terminator sequence (term), a DNA sequence element comprising a 5'-UTR, a 3'-UTR, etc.
[0154] As generally described above, certain embodiments of the present disclosure relate to novel variant pro-region sequences. In related aspects, the present disclosure provides recombinant polynucleotides comprising one or more variant pro-region nucleic acid (DNA) sequences. Accordingly, certain embodiments relate to recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant Gram-positive bacterial cells / strains expressing a protein of interest, etc. In certain aspects, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells (strains) to enhance production of a protein of interest. In certain aspects, the polynucleotide constructs of the present disclosure are referred to as expression cassettes, which comprise, in a 5' to 3' direction and in operative combination, at least an upstream (5') pro-region DNA sequence linked to a downstream (3') gene CDS encoding a mature protein of interest (POI).
[0155] For example, one or more nucleic acid sequences described herein can be produced by using any suitable synthesis, manipulation, and / or isolation technique, or a combination thereof. For example, one or more polynucleotides described herein can be produced using standard nucleic acid synthesis techniques, such as solid-phase synthesis techniques well known to those of skill in the art. In such techniques, fragments, typically consisting of up to 50 or more nucleotide bases, are synthesized and then joined (e.g., by enzymatic or chemical ligation methods) to form essentially any desired contiguous nucleic acid sequence. Synthesis of one or more polynucleotides described herein can be facilitated by any suitable method known in the art, including, but not limited to, chemical synthesis using the classical phosphoramidite method (e.g., Beaucage and Caruthers, 1981) or the method described by Matthes et al. (1984), as commonly practiced in automated synthesis methods. One or more polynucleotides described herein can also be produced using an automated DNA synthesizer. Customized nucleic acids can be ordered from a variety of commercial sources (e.g., ATUM (DNA 2.0), Newark, CA, USA; Life Tech (GeneArt), Carlsbad, CA, USA; GenScript, Ontario, Canada; Base Clear BV, Leiden, Netherlands; Integrated DNA Technologies, Skokie, IL, USA; Ginkgo Bioworks (Gen9), Boston, MA, USA; and Twist Bioscience, San Francisco, CA, USA). Other techniques for synthesizing nucleic acids and related principles are described and known in the art.
[0156] Recombinant DNA techniques useful for modifying nucleic acids are well known in the art, such as restriction endonuclease digestion, ligation, reverse transcription and cDNA production, and polymerase chain reaction (e.g., PCR). One or more polynucleotides described herein can also be obtained by screening a cDNA library with one or more oligonucleotide probes that hybridize to or PCR amplify a polynucleotide encoding one or more variants described herein. Methods for screening and isolating cDNA clones and PCR amplification methods are well known to those of skill in the art and are described in standard references known to those of skill in the art. One or more polynucleotides described herein can be obtained by modifying a native polynucleotide backbone (e.g., encoding one or more variant pro-region sequences described herein) by, for example, known mutagenesis procedures (e.g., site-directed mutagenesis, site-saturation mutagenesis, and in vitro recombination). A variety of methods suitable for generating modified polynucleotides described herein that encode one or more variants described herein are known in the art, including, but not limited to, site-saturation mutagenesis, systematic mutagenesis, insertional mutagenesis, deletion mutagenesis, random mutagenesis, site-directed mutagenesis, and directed evolution, as well as various other recombinant methods.
[0157] As generally described above and explained in more detail in the Examples below, certain embodiments of the present disclosure relate to recombinant (modified) Gram-positive cells capable of producing increased amounts of a heterologous protein of interest. Accordingly, certain embodiments relate to methods for constructing such recombinant Gram-positive cells with increased protein production capacity. In certain embodiments, one or more expression cassettes encoding the protein of interest are introduced into the Gram-positive cells of the present disclosure. In exemplary embodiments, the cassettes are integrated into the genome of the cell. Accordingly, certain embodiments relate to nucleic acid molecules, polynucleotides (e.g., vectors, plasmids, expression cassettes), regulatory elements, and the like, suitable for use in constructing recombinant (modified) Gram-positive host cells.
[0158] Thus, as presented in the Examples and generally described herein, recombinant cells of the present disclosure can be constructed by one of skill in the art using standard and routine recombinant DNA and molecular cloning techniques well known in the art. Methods of genetic modification include, but are not limited to, (a) introduction, substitution, or removal of one or more nucleotides in a gene or introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of a gene, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-directed mutagenesis, and / or (g) random mutagenesis.
[0159] In certain embodiments, modified cells of the present disclosure can be constructed by reducing or eliminating expression of a gene using methods well known in the art, such as insertion, disruption, substitution, or deletion. The portion of a gene to be modified or inactivated can be, for example, a coding region or a regulatory element required for expression of the coding region.
[0160] An example of such a regulatory or control sequence may be a promoter sequence or a functional portion thereof (i.e., a portion sufficient to affect the expression of a nucleic acid sequence). Other control sequences for modification include, but are not limited to, a leader sequence, a propeptide sequence, a signal sequence, a transcription terminator, a transcription activator, and the like.
[0161] In certain other embodiments, modified cells are constructed by gene deletion, which eliminates or reduces gene expression. Gene deletion techniques allow for the partial or complete removal of genes, thereby eliminating their expression or expressing non-functional (or reduced activity) protein products. In such methods, gene deletion can be accomplished by homologous recombination using a plasmid constructed to contain adjacent 5' and 3' regions flanking the gene. The flanking 5' and 3' regions can be introduced into cells, for example, on a temperature-sensitive plasmid associated with a second selectable marker at a permissive temperature that allows the plasmid to establish in the cells. Cells are then shifted to a non-permissive temperature to select for cells with the plasmid integrated into the chromosome at one of the homologous flanking regions. Selection for plasmid integration is influenced by selection for the second selectable marker. After integration, recombination events at the second homologous flanking region are stimulated by shifting the cells to a permissive temperature for several generations without selection. Cells are plated to obtain single colonies, which are then tested for the loss of both selectable markers. Thus, one skilled in the art can readily identify nucleotide regions within the coding sequence of a gene and / or the non-coding sequence of a gene that are suitable for complete or partial deletion.
[0162] In other embodiments, modified cells are constructed by introducing, substituting, or removing one or more nucleotides in a gene or regulatory element required for its transcription or translation. For example, nucleotides can be inserted or removed to introduce a stop codon, remove a start codon, or cause a frameshift in the open reading frame. Such modifications can be achieved by site-directed mutagenesis or PCR-generated mutagenesis, according to methods known in the art. Thus, in certain embodiments, the genes of the present disclosure are inactivated by complete or partial deletion.
[0163] In another embodiment, modified cells are constructed by a gene conversion process. For example, in gene conversion methods, a nucleic acid sequence corresponding to a gene is mutated in vitro to generate a defective nucleic acid sequence, which is then transformed into a parent cell to generate the defective gene. The defective nucleic acid sequence replaces the endogenous gene through homologous recombination. It may be desirable for the defective gene or gene fragment to also encode a marker that can be used to select for transformants containing the defective gene. For example, the defective gene can be introduced on a non-replicating or temperature-sensitive plasmid associated with a selectable marker. Selection for plasmid integration is affected by selecting for that marker under conditions that do not allow replication of the plasmid. Selection for a second recombination event resulting in gene replacement is affected by examining colonies for the loss of the selectable marker and the acquisition of a mutated gene. Alternatively, the defective nucleic acid sequence can contain an insertion, substitution, or deletion of one or more nucleotides of the gene, as described below.
[0164] In other embodiments, modified cells are constructed using established antisense technology, using a nucleotide sequence complementary to the nucleic acid sequence of a gene. More specifically, the expression of a gene by a Gram-positive cell can be reduced (downregulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which can be transcribed in the cell and hybridize to the mRNA produced in the cell. Thus, under conditions where the complementary antisense nucleotide sequence can hybridize to the mRNA, the amount of translated protein is reduced or eliminated. Such antisense methods include, but are not limited to, RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), antisense oligonucleotides, etc., all of which are well known to those skilled in the art.
[0165] In other embodiments, modified cells are produced / constructed by CRISPR-Cas9 editing. For example, genes encoding proteins of interest may be edited or disrupted (or deleted or downregulated) by a nucleic acid-guided endonuclease that finds its target DNA by binding to a guide RNA (e.g., Cas9) that recruits the endonuclease to a target sequence on the DNA and either Cpf1 or a guide DNA (e.g., NgAgo), where the endonuclease can generate single- or double-strand breaks within the DNA. This targeted DNA break can then serve as a substrate for DNA repair and recombine with the provided editing template to disrupt or delete the gene. For example, a gene encoding a nucleic acid-guided endonuclease (in this case, S. pyogenes-derived Cas9) or a codon-optimized gene encoding a Cas9 nuclease is operably linked to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells, thereby generating a Gram-positive cell Cas9 expression cassette. Similarly, one or more target sites unique to a gene of interest can be easily identified by those skilled in the art. For example, to construct a DNA construct encoding a gRNA directed to a target site within a gene of interest, the variable targeting domain (VT) would contain the nucleotides of the target site 5' to the (PAM) protospacer adjacent motif (TGG), whose nucleotides are fused to DNA encoding the Cas9 endonuclease recognition domain (CER) for S. pyogenes Cas9. The combination of the DNA encoding the VT domain and the DNA encoding the CER domain generates DNA encoding the gRNA. Thus, a Gram-positive expression cassette for the gRNA is generated by operably linking the DNA encoding the gRNA to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells.
[0166] In certain embodiments, the DNA break induced by endonuclease is repaired / replaced using the incoming sequence.For example, to precisely repair the DNA break generated by the above-mentioned Cas9 expression cassette and gRNA expression cassette, a nucleotide editing template is provided so that the DNA repair mechanism of the cell can use the editing template.For example, about 500bp of the 5' side of the targeting gene can be fused to about 500bp of the 3' side of the targeting gene to generate an editing template, and this template is used by the mechanism of the Gram-positive host to repair the DNA break generated by RGEN.
[0167] The Cas9 expression cassette, gRNA expression cassette, and editing template can be co-delivered into filamentous fungal cells using a number of different methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence). Transformed cells are screened by PCR amplification of the target locus by amplifying the locus using forward and reverse primers. These primers can amplify the wild-type locus or the modified locus edited by RGEN. These fragments are then sequenced using sequencing primers to identify edited colonies.
[0168] In yet other embodiments, modified cells are constructed by random or directed mutagenesis using methods well known in the art, including, but not limited to, chemical mutagenesis and transposition. Genetic modification can be performed by subjecting parent cells to mutagenesis and screening for mutant cells in which expression of the gene is reduced or eliminated. Mutagenesis can be directed or random, for example, by using suitable physical or chemical mutagenizing agents, by using suitable oligonucleotides, or by subjecting DNA sequences to PCR-generated mutagenesis. Furthermore, mutagenesis can be performed by using any combination of these mutagenesis methods.
[0169] Examples of physical or chemical mutagenic agents suitable for the purposes of the present invention include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl-N'-nitrosoguanidine (NTG), O-methylhydroxylamine, nitrous acid, ethyl methanesulfonate (EMS), sodium bisulfite, formic acid, and nucleotide analogs. When using such agents, mutagenesis is generally carried out by incubating parent cells to be mutagenized under suitable conditions in the presence of the mutagen of choice, and selecting mutant cells that show reduced or no expression of the gene.
[0170] WO 2003 / 083125 discloses methods for modifying Gram-positive (Bacillus) cells, such as creating Bacillus deletion strains and DNA constructs using PCR fusion to bypass E. coli. WO 2002 / 14490 discloses methods for modifying Bacillus cells, including (1) construction and transformation of an integrating plasmid (pComK), (2) random mutagenesis of coding, signal, and propeptide sequences, (3) homologous recombination, (4) increasing transformation efficiency by adding non-homologous flanks to the transforming DNA, (5) optimizing double-crossover integration, (6) site-directed mutagenesis, and (7) markerless deletion.
[0171] Those skilled in the art are well aware of suitable methods for introducing polynucleotide sequences into bacterial cells (e.g., gram-negative cells, gram-positive cells). Indeed, methods such as transformation, including protoplast transformation and aggregation, transduction, and protoplast fusion, are known and suitable for use in the present disclosure. Transformation methods are particularly preferred for introducing the DNA constructs of the present disclosure into host cells.
[0172] In addition to commonly used methods, in some embodiments, host cells are directly transformed (i.e., no intermediate cells are used to amplify or otherwise process the DNA construct prior to introduction into the host cell). Introduction of the DNA construct into the host cell includes those physical and chemical methods known in the art for introducing DNA into a host cell without insertion into a plasmid or vector. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, liposomes, and the like. In additional embodiments, the DNA construct is co-transformed with a plasmid without being inserted into the plasmid. In further embodiments, the selectable marker is deleted or substantially excised from the modified Bacillus strain by methods known in the art. In some embodiments, separation of the vector from the host chromosome leaves flanking regions within the chromosome while removing the unique chromosomal region.
[0173] Promoters and promoter sequence regions, their coding sequences (CDS), open reading frames (ORFs) and / or variant sequences for use in expressing genes in Gram-positive cells are generally known to those skilled in the art. The promoter sequences of the present disclosure are generally selected so that they function in Gram-positive cells. For example, promoters useful for driving gene expression in Bacillus cells include, but are not limited to, the B. subtilis alkaline protease (aprE) promoter, the B. subtilis α-amylase promoter (amyE), the B. licheniformis α-amylase promoter (amyL), the B. amyloliquefaciens α-amylase promoter, the B. subtilis neutral protease (nprE) promoter, a mutant aprE promoter, or any other promoter from B. licheniformis or other related Bacillus species. Methods for screening and generating promoter libraries with varying activities (promoter strengths) in Bacillus cells are described in WO 2002 / 14490.
[0174] IV. Fermentation of Gram-Positive Cells to Produce Proteins As broadly described above, certain embodiments relate to compositions and methods for constructing and obtaining Gram-positive cells with an increased protein production phenotype. Accordingly, certain embodiments relate to methods for producing a protein of interest in Gram-positive cells by fermenting the cells in an appropriate medium. Fermentation methods known in the art may be applied to ferment the Gram-positive cells of the present disclosure.
[0175] In some embodiments, cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system in which the composition of the medium is set at the beginning of the fermentation and remains unchanged during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. This method allows fermentation to occur without adding any components to the system. Generally, batch fermentation is considered "batch" with respect to the addition of a carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes continuously until the fermentation is stopped. In a typical batch culture, cells progress through a static lag phase to a high-growth logarithmic phase and may eventually progress to a stationary phase where growth rate decreases or stops. If untreated, cells in the stationary phase eventually die. Generally, cells in the logarithmic phase are responsible for the majority of product production.
[0176] A suitable variation on the standard batch system is the "fed-batch fermentation" system. In this variation of the typical batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cellular metabolism and when a limited amount of substrate is desired in the medium. In fed-batch systems, the actual substrate concentration is difficult to measure and is therefore estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of waste gases such as CO2. Batch and fed-batch fermentations are common and known in the art.
[0177] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density, with cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon or nitrogen source, is maintained at a fixed ratio, while all other parameters are adjustable. In other systems, multiple factors affecting growth can be continuously varied while the cell concentration, measured by medium turbidity, remains constant. Continuous systems attempt to maintain steady-state growth conditions. Therefore, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes and techniques for maximizing product formation rates are well known in the art of industrial microbiology.
[0178] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, such as separating the host cells from the medium by centrifugation or filtration, or, optionally, disrupting the cells and removing the supernatant from cell debris and debris. Generally, after clarification, the protein component of the supernatant or filtrate is precipitated with a salt, such as ammonium sulfate. The precipitated protein can then be solubilized and purified by various chromatographic methods, such as ion exchange chromatography, gel filtration, etc.
[0179] In some embodiments, cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system in which the composition of the medium is set at the beginning of the fermentation and remains unchanged during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. This method allows fermentation to occur without adding any components to the system. Generally, batch fermentation is considered "batch" with respect to the addition of a carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes continuously until the fermentation is stopped. In a typical batch culture, cells progress through a static lag phase to a high-growth logarithmic phase and may eventually progress to a stationary phase where growth rate decreases or stops. If untreated, cells in the stationary phase eventually die. Generally, cells in the logarithmic phase are responsible for the majority of product production.
[0180] A suitable variation on the standard batch system is the "fed-batch fermentation" system. In this variation of the typical batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cellular metabolism and when a limited amount of substrate is desired in the medium. In fed-batch systems, measurement of the actual substrate concentration is difficult and therefore is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of waste gases such as CO2. Batch and fed-batch fermentation are common and known in the art.
[0181] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density, with cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon or nitrogen source, is maintained at a fixed ratio, while all other parameters are adjustable. In other systems, multiple factors affecting growth can be continuously varied while the cell concentration, measured by medium turbidity, remains constant. Continuous systems attempt to maintain steady-state growth conditions. Therefore, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes and techniques for maximizing product formation rates are well known in the art of industrial microbiology.
[0182] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, such as separating the host cells from the medium by centrifugation or filtration, or, optionally, disrupting the cells and removing the supernatant from cell debris and debris. Generally, after clarification, the protein component of the supernatant or filtrate is precipitated with a salt, such as ammonium sulfate. The precipitated protein can then be solubilized and purified by various chromatographic methods, such as ion exchange chromatography, gel filtration, etc.
[0183] V. Protein of Interest The protein of interest (POI) of the present disclosure can be any endogenous or heterologous protein, or a variant of such a POI. The protein may contain one or more disulfide bridges or may be a protein whose functional form is monomeric or multimeric, i.e., the protein has a quaternary structure and is composed of multiple identical (homologous) or non-identical (heterologous) subunits, where the POI or variant POI is preferably a protein with a desired property.
[0184] For example, in certain embodiments, the modified Gram-positive cells of the present disclosure produce at least about 0.1% more, at least about 0.5% more, at least about 1% more, at least about 5% more, at least about 6% more, at least about 7% more, at least about 8% more, at least about 9% more, or at least about 10% or more POI compared to the unmodified (parent or control) cells.
[0185] In certain embodiments, the engineered Gram-positive cells of the present disclosure exhibit increased specific productivity (Qp) of a POI compared to control cells. For example, detecting specific productivity (Qp) is a suitable method for assessing protein production. Specific productivity (Qp) can be determined using the following equation: "Qp = gP / gDCW·hr" where "gP" is the grams of protein produced in the tank, "gDCW" is the grams of dry cell weight (DCW) in the tank, and "hr" is the fermentation time (hours) from the time of inoculation, which includes the production time and growth time.
[0186] Thus, in certain other embodiments, the modified Gram-positive cells of the present disclosure comprise an increase in specific productivity (Qp) of at least about 0.1%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to the unmodified (parent) cell.
[0187] In certain embodiments, the POI or variant POI thereof is selected from the group consisting of acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, and the like. The enzyme is selected from the group consisting of: oxidase, invertase, isomerase, laccase, ligase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0188] Thus, in certain embodiments, the POI or variant POI thereof is an enzyme selected from the EC numbers EC1, EC2, EC3, EC4, EC5, or EC6.
[0189] There are a variety of assays known to those of skill in the art for detecting and measuring the activity of intracellularly and extracellularly expressed proteins.
[0190] VI. Illustrative Embodiments Non-limiting embodiments of the compositions and methods disclosed herein are as follows.
[0191] 1. A variant nucleic acid (DNA) sequence comprising at least one mutation set forth in any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant nucleic acid sequence are numbered according to SEQ ID NO:1.
[0192] 2. A variant nucleic acid (DNA) sequence comprising any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant sequence are numbered according to SEQ ID NO: 1.
[0193] 3. The variant nucleic acid of embodiment 1 or 2, comprising at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:1.
[0194] 4. The variant nucleic acid of any one of embodiments 1 to 3, further defined as a variant rrnI-P2 promoter and 5' untranslated region (5'-UTR) sequence (variant rrnI-P2 / 5'-UTR sequence).
[0195] 5. A polynucleotide comprising a variant nucleic acid sequence according to any one of embodiments 1 to 4.
[0196] 6. A polynucleotide comprising the variant nucleic acid sequence of any one of embodiments 1 to 4 operably linked to a downstream (3') nucleic acid sequence encoding a protein of interest (POI).
[0197] 7. A polynucleotide comprising the variant nucleic acid sequence of any one of embodiments 1-4 operably linked to a downstream (3') nucleic acid sequence encoding a pro region (PRO), which is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).
[0198] 8. A polynucleotide comprising the variant nucleic acid sequence of any one of embodiments 1-4 operably linked to a downstream (3') nucleic acid sequence encoding a protein signal (secretory) sequence (SS), which is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).
[0199] 9. A polynucleotide comprising the variant nucleic acid sequence of any one of embodiments 1-4 operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI), operably linked to a downstream nucleic acid sequence encoding a proregion (PRO) sequence, operably linked to a downstream (3') nucleic acid sequence encoding a protein signal (secretion) sequence (SS).
[0200] 10. The polynucleotide of any one of embodiments 5 to 9, further comprising a downstream (3') terminator sequence operably linked to the nucleic acid encoding the POI protein.
[0201] 11. The polynucleotide according to any one of embodiments 5 to 9, wherein the POI is selected from the group consisting of an enzyme, an antibody, a receptor protein, a lectin, and a regulatory protein.
[0202] 12. POIs include acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, and laccase. 10. The polynucleotide of any one of embodiments 5 to 9, wherein the polynucleotide is selected from the group consisting of: lipases, lyases, mannosidases, oxidases, oxidoreductases, pectate lyases, pectin acetylesterases, pectin depolymerases, pectin methylesterases, pectinolytic enzymes, perhydrolases, polyol oxidases, peroxidases, phenol oxidases, phytases, polygalacturonases, proteases, peptidases, rhamnogalacturonases, ribonucleases, transferases, transport proteins, transglutaminases, xylanases, hexose oxidases, and combinations thereof.
[0203] 13. The polynucleotide of embodiment 12, wherein the enzyme is a protease.
[0204] 14. The polynucleotide of embodiment 13, wherein the enzyme is subtilisin.
[0205] 15. The polynucleotide of embodiment 14, wherein the subtilisin comprises at least about 80% to 100% identity to SEQ ID NO:2 or SEQ ID NO:3.
[0206] 16. The polynucleotide of embodiment 7 or embodiment 9, wherein the pro region (PRO) sequence comprises at least about 80% to 100% identity to SEQ ID NO:5.
[0207] 17. The polynucleotide of embodiment 8 or embodiment 9, wherein the DNA encoding the protein signal (secretory) sequence (SS) comprises at least about 80% to 100% identity to SEQ ID NO:4.
[0208] 18. The polynucleotide of embodiment 10, wherein the DNA encoding the terminator sequence has at least about 80% to 100% identity to SEQ ID NO: 6.
[0209] 19. An expression cassette comprising a polynucleotide according to any one of embodiments 5 to 18.
[0210] 20. A Gram-positive bacterial cell comprising the transfer cassette of embodiment 19.
[0211] 21. A Gram-positive cell according to embodiment 20, wherein the cassette is integrated into the genome of the cell.
[0212] 22. The Gram-positive cell of embodiment 20, wherein the cell is a Bacillus sp. cell.
[0213] 23. The Gram-positive cell of embodiment 22, wherein the cell is a Bacillus sp. cell selected from the group consisting of B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophilus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis.
[0214] 25. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising: (a) introducing into the Gram-positive bacterial cell a polynucleotide, the polynucleotide comprising an upstream (5') variant promoter and a 5'-untranslated region (5-UTR) nucleic acid sequence, wherein the upstream (5') variant promoter and the 5'-untranslated region (5-UTR) nucleic acid sequence (i) comprises at least one mutation set forth in any one of SEQ ID NO: 8 through SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO: 1, or (ii) comprises any one of SEQ ID NO: 8 through SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO: 1, operably linked to a downstream (3') open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
[0215] 26. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising: (a) introducing into the Gram-positive bacterial cell a polynucleotide, the polynucleotide comprising an upstream (5') variant promoter and a 5'-untranslated region (5-UTR) nucleic acid sequence, wherein the upstream (5') variant promoter and the 5'-untranslated region (5-UTR) nucleic acid sequence (i) comprises at least one mutation set forth in any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO:1, or (ii) comprises any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO:1, operably linked to a downstream (3') nucleic acid sequence encoding a pro region (PRO) that is operably linked to a downstream open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
[0216] 27. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising: (a) introducing into the Gram-positive bacterial cell a polynucleotide, the polynucleotide comprising an upstream (5') variant promoter and a 5'-untranslated region (5-UTR) nucleic acid sequence, wherein the upstream (5') variant promoter and the 5'-untranslated region (5-UTR) nucleic acid sequence (i) comprises at least one mutation set forth in any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO:1, or (ii) comprises any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO:1, operably linked to a downstream (3') nucleic acid sequence encoding a protein signal (secretion) sequence (SS), which is operably linked to a downstream open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
[0217] 28. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, comprising: (a) introducing into the Gram-positive bacterial cell a polynucleotide, wherein the polynucleotide comprises an upstream (5') variant promoter and a 5'-untranslated region (5-UTR) nucleic acid sequence, wherein the upstream (5') variant promoter and the 5'-untranslated region (5-UTR) nucleic acid sequence (i) comprises at least one mutation set forth in any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO:1. or (ii) introducing a variant promoter / 5'-UTR sequence comprising any one of SEQ ID NO:8 through SEQ ID NO:46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO:1, operably linked to a downstream (3') nucleic acid encoding a protein signal (secretion) sequence (SS) operably linked to a downstream nucleic acid sequence encoding a pro region (pro) sequence operably linked to a downstream open reading frame (ORF) encoding a protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
[0218] 29. The method of any one of embodiments 25-28, wherein the variant promoter / 5'-UTR sequence comprises at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:1.
[0219] 30. The method of any one of embodiments 25 to 29, further comprising a downstream (3') terminator sequence operably linked to the ORF encoding the POI.
[0220] 31. The method of any one of embodiments 25 to 29, wherein the modified cells produce an increased amount of the POI compared to control cells cultured under the same conditions.
[0221] 32. The method of any one of embodiments 25 to 29, wherein the modified cells produce an increased amount of POI compared to control cells after about 72 hours of culture.
[0222] 33. The method according to any one of embodiments 25 to 32, wherein the POI is selected from the group consisting of enzymes, antibodies, receptor proteins, lectins, and regulatory proteins.
[0223] 34. POIs include acetylesterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase 33. The method of any one of embodiments 25 to 32, wherein the hydroxylase is selected from the group consisting of: hydroxylase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0224] 35. The method of embodiment 34, wherein the enzyme is a protease.
[0225] 36. The method of embodiment 35, wherein the protease is subtilisin.
[0226] 37. The method of embodiment 36, wherein the subtilisin comprises at least about 80% to 100% identity to SEQ ID NO: 2 or SEQ ID NO: 3.
[0227] 38. The method of embodiment 26 or embodiment 28, wherein the pro region (PRO) sequence comprises at least about 80% to 100% identity to SEQ ID NO:5.
[0228] 39. The method of embodiment 27 or embodiment 28, wherein the DNA encoding the protein signal (secretory) sequence (SS) comprises at least about 80% to 100% identity to SEQ ID NO: 4.
[0229] 40. The method of embodiment 30, wherein the DNA encoding the terminator sequence comprises at least about 80% to 100% identity to SEQ ID NO: 6.
[0230] 41. The method of any one of embodiments 25 to 29, wherein the polynucleotide is integrated into the genome of the cell.
[0231] 42. The method of any one of embodiments 25 to 29, wherein the Gram-positive cell is a Bacillus sp. cell.
[0232] 43. The method of embodiment 42, wherein the Bacillus sp. cell is selected from the group consisting of B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophilus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis. [Example]
[0233] Certain aspects of the present disclosure may be further understood in light of the following examples, which should not be construed as limiting. Variations in materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989).
[0234] Example 1 Evaluation of promoter / 5' untranslated region sequences for insertion and deletion sites In this example, a variant B. gibsonii (BG46 variant) subtilisin (SEQ ID NO: 3) was used as a reporter to monitor protein expression as described herein. More specifically, a DNA fragment containing the upstream (5') aprE gene flanking region (5' aprE gene FR; SEQ ID NO: 49) was operably linked to a DNA sequence encoding the mature BG46 variant reporter protein (SEQ ID NO: 3), operably linked to a B. amyloliquefaciens BPN' terminator DNA sequence (SEQ ID NO: 6), operably linked to a DNA sequence encoding the B. lentus pro-peptide sequence (SEQ ID NO: 5), and operably linked to a DNA sequence encoding the B. lentus pro-peptide sequence (SEQ ID NO: 5). The polynucleotide construct (expression cassette) comprises an upstream (5') B. subtilis variant (control) rrnI-P2 promoter / 5'-UTR region DNA sequence (SEQ ID NO: 1), operably linked to a DNA sequence encoding the E signal sequence (SEQ ID NO: 4), which in turn is operably linked to downstream (3') aprE gene flanking region (3' aprE gene FR; SEQ ID NO: 50) sequences containing a downstream kanamycin (kan) gene expression cassette. More specifically, the DNA fragments were assembled using standard molecular biology techniques and then used as templates to generate linear DNA expression cassettes containing one or more of the promoter region SEL variants described herein.
[0235] Indel site evaluation library
[0236] As outlined herein, 75 insertion-deletion (In-Del) site evaluation libraries (SELs) were performed on the reference (control) rrnI-P2 promoter / 5'-UTR region sequence (SEQ ID NO: 1). SEL variant promoter / 5'-UTR region sequences were designed and constructed by Twist Bioscience HQ (South San Francisco) as outlined in Table 1 and generated as 4.4 kb fragments. More specifically, linear DNA of the expression cassette (Table 2) was used to transform competent B. subtilis cells, and the transformation mixture was plated on LA plates containing 1.8 ppm kanamycin and incubated overnight at 37°C. Single colonies were picked and grown in Luria broth at 37°C under antibiotic selection.
[0237] DNA sequence analysis was performed to determine the sequences of unique (variant) promoter / 5'-UTR regions cherry-picked into 96-well microtiter plates (MTPs). For example, as shown in Figure 1, the reference rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1) contains 149 nucleotides, with nucleotide positions numbered 1 to 149 from 5' to 3'. Specifically, as shown in Figure 1, the reference promoter / 5'-UTR region contains nucleotide positions 1 to 149 of SEQ ID NO: 1, and SEL generated a library of 75 unique promoter / 5'-UTR regions by altering (modifying) two adjacent nucleotide positions of SEQ ID NO: 1.
[0238] For example, Table 1, set forth below, shows 31 possible variants of the first site of the reference promoter / 5'-UTR region (i.e., adjacent nucleotide positions 1 (guanine, "G") and 2 (cytosine, "C") of SEQ ID NO: 1), with the first column (Table 1) showing the two nucleotide positions ("variations") that are altered compared to the two nucleotides at the same positions of the reference rrnI-P2 p promoter / 5'-UTR region (Table 1, second column, "Results using GC as reference"). [Table 1]
[0239] Similarly, Table 2 below shows the reference rrnI-P2 promoter / 5'-UTR region and 40 mutant promoter / 5'-UTR region sequences identified in SELs (Table 2; UTR-00664, SEQ ID NO:8 to UTR-00752, SEQ ID NO:46). As described in more detail in Example 2, for reporter protein expression experiments, transformed cells were grown in culture medium (MOPs buffer-based concentrated semi-defined medium with urea as the primary nitrogen source, maltodextrin as the primary carbon source, supplemented with 3% soytone to enhance cell growth, and containing an antibiotic selection agent) in 96-well MTPs in a shaking incubator at 32°C, 300 rpm, and 80% humidity for 3 days, then centrifuged and filtered. The clarified culture supernatant was used to assay reporter protease activity to determine productivity levels, with samples taken after 72 hours (Example 3, Table 3). The reporter protease activity assay is described in more detail in Example 2 below. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4]
[0240] Example 2 Protease activity assay The protease activity of the reporter protein (BG46 variant) was determined by measuring the hydrolysis of a synthetic suc-AAPF-pNA peptide substrate. For the AAPF assay, the reagent solution used was: 100 mM Tris, pH 8.6, 10 mM CaCl, 0.005% Tween®-80 (Tris / Ca buffer), and 160 mM suc-AAPF-pNA in DMSO (suc-AAPF-pNA stock solution; Sigma; S-7388). To prepare a working solution, 1 mL of suc-AAPF-pNA stock solution was added to 100 mL of Tris / Ca buffer and mixed. Enzyme samples were added to a microtiter plate (MTP) containing a 1 mg / mL suc-AAPF-pNA working solution, and activity was assayed at 405 nm over 3–5 min at room temperature using a SpectraMax plate reader in kinetic mode. The protease activity was expressed in mOD / min. Specifically, the protease activity of each constructed variant was measured and compared with that of a reference construct (rrnI-P2 promoter / 5'-UTR region; SEQ ID NO: 1) grown in the same plate. The performance index (PI) was calculated by dividing the value of the reference sample by the value of the variant sample. This is shown in Table 3, as described in Example 3.
[0241] Example 3 Modification of promoter and 5' untranslated region to enhance protein productivity As described above, a mature subtilisin reporter protein (BG46 variant) was expressed in B. subtilis under the control of variant promoter / 5'-UTR sequences constructed by site-scanning mutagenesis library (SEL, Example 1), and the reporter protein activity of the constructs was assayed (Example 2). As shown in the DNA sequence alignments in Figures 2 and 3, variant promoter / 5'-UTR sequences that showed increased productivity after 72 hours of growth compared to the reference rrnI-P2 promoter / 5'-UTR sequence (SEQ ID NO: 1) and had a performance index (PI) greater than (>) 1.2 are listed in Table 3 below.
[0242] For example, compared to the reference rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1), approximately 50% of the constructed SEL variants contained mutations upstream of the UP element, while approximately 22% of the SEL variants contained mutations in the UP element (Figures 2A / 3A).Similarly, approximately 22% of the SEL variants contained mutations around the Shine-Dalgarno (SD) region in the 5'-UTR compared to the reference rrnI-P2 promoter / 5'-UTR region (SEQ ID NO: 1) (Figures 2B / 3B).
[0243] In contrast, approximately 72% of SEL variants that contained mutations in the -35 / -10 promoter region had PI values of 0.1 to 0.5 after 72 hours of growth, and another 21% of SEL variants that contained mutations around the SD region in the 5'-UTR had PI values of 0.1 to 0.5 after 72 hours of growth (data not shown). [Table 3-1] [Table 3-2] References U.S. Patent No. 4,559,300 International Publication No. 2002 / 14490 Brochure International Publication No. 2003 / 083125 Brochure International Publication No. 2013 / 086219 Brochure International Publication No. 2017 / 152169 Brochure Ausubel et al., “Current Protocols in Molecular Biology”, published by Greene Publishing Assoc. and Wiley-Interscience (1987). Kim et al.,“Comparison of PaprE, PamyE,and PP43 promoter strength for β-galactosidase and staphylokinase expression in Bacillus subtilis”, Biotechnology and Bioprocess Engineering, 13:313,2008. Sambrook et al.,“Molecular Cloning:A Laboratory Manual”Cold Spring Harbor Laboratory:Cold Spring Harbor,N.Y.(1989),(2001)and(2012).
Claims
1. A variant nucleic acid comprising at least one mutation set forth in any one of SEQ ID NO:8 to SEQ ID NO:46, wherein the nucleotide positions of the variant nucleic acid sequence are numbered according to SEQ ID NO:
1.
2. 2. The variant nucleic acid of claim 1, comprising at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:
1.
3. A polynucleotide comprising the variant nucleic acid sequence of claim 1.
4. The polynucleotide of claim 3 , operably linked to a downstream nucleic acid encoding a protein of interest (POI).
5. The polynucleotide of claim 3 , operably linked to a downstream nucleic acid encoding a pro region, which is operably linked to a downstream nucleic acid encoding a protein of interest (POI).
6. The polynucleotide of claim 3 , operably linked to a downstream nucleic acid encoding a protein signal sequence, which is operably linked to a downstream nucleic acid encoding a protein of interest (POI).
7. The polynucleotide of claim 3 , operably linked to a downstream nucleic acid encoding a protein signal sequence, operably linked to a downstream nucleic acid encoding a pro-region sequence, operably linked to a downstream nucleic acid encoding a protein of interest (POI).
8. The polynucleotide of claim 3 , further comprising a downstream terminator sequence operably linked to the nucleic acid encoding the POI protein.
9. The polynucleotide of claim 3 , wherein the POI is selected from the group consisting of an enzyme, an antibody, a receptor protein, a lectin, and a regulatory protein.
10. The POI may be any of acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, lacca 10. The polynucleotide of claim 9, wherein the enzyme is selected from the group consisting of: oxidases, lipases, lyases, mannosidases, oxidases, oxidoreductases, pectate lyases, pectin acetyl esterases, pectin depolymerases, pectin methyl esterases, pectinolytic enzymes, perhydrolases, polyol oxidases, peroxidases, phenol oxidases, phytases, polygalacturonases, proteases, peptidases, rhamnogalacturonase, ribonucleases, transferases, transport proteins, transglutaminases, xylanases, hexose oxidases, and combinations thereof.
11. The polynucleotide of claim 10, wherein the enzyme is a protease.
12. The polynucleotide of claim 11 , wherein the protease is a subtilisin.
13. An expression cassette comprising the polynucleotide of claim 3.
14. A Gram-positive bacterial cell comprising the transfer cassette of claim 13.
15. The Gram-positive cell of claim 14, wherein the cassette is integrated into the genome of the cell.
16. 1. A method for producing a protein of interest (POI) in a gram-positive bacterial cell, comprising: (a) introducing a polynucleotide into a Gram-positive bacterial cell, wherein the polynucleotide comprises an upstream variant promoter and a 5'-untranslated region (5-UTR) nucleic acid sequence, the upstream variant promoter and the 5'-untranslated region (5-UTR) nucleic acid sequence comprising: (i) comprises at least one mutation set forth in any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of said variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO: 1; or (ii) any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5'-UTR sequence are numbered according to SEQ ID NO: 1; introducing, operably linked to a downstream open reading frame (ORF) encoding a protein of interest (POI); (b) culturing the modified cells under conditions suitable for producing said POI.
17. 17. The method of claim 16, wherein the variant promoter / 5'-UTR sequence comprises at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:
1.
18. 17. The method of claim 16, wherein the modified cells produce an increased amount of the POI compared to control cells cultured under the same conditions, and the control cells comprise an introduced polynucleotide comprising an upstream promoter / 5'-UTR sequence of SEQ ID NO: 1 operably linked to a downstream ORF encoding the same POI as the modified cells.
19. 17. The method of claim 16, wherein the modified cells produce an increased amount of the POI compared to the control cells after at least 72 hours of culture.
20. 20. The method of claim 19, wherein the increased amount of POI is at least 5% compared to the control cells after at least 72 hours of culture.
21. 17. The method of claim 16, wherein the POI is selected from the group consisting of an enzyme, an antibody, a receptor protein, a lectin, and a regulatory protein.
22. 17. The method of claim 16, wherein the Gram-positive cell is a Bacillus sp. cell.