Pro-region mutations that enhance protein production in Gram-positive bacterial cells

JP2025510901A5Pending Publication Date: 2026-03-13DANISCO US INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-03-13

Smart Images

  • Figure 00000050_0000
    Figure 00000050_0000
Patent Text Reader

Abstract

The present disclosure relates generally to recombinant polynucleotides comprising novel pro-region DNA sequences. Certain aspects of the disclosure relate to recombinant Gram-positive bacterial strains comprising one or more introduced polynucleotides comprising novel pro-region DNA sequences operably linked to a DNA sequence encoding a protein of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates generally to the fields of microbial host cells, molecular biology, protein engineering, fermentation, protein production, etc. Certain aspects of the disclosure relate to novel pro region nucleic acid (DNA) sequences, recombinant polynucleotides comprising novel pro region DNA sequences, genetically modified (recombinant) Gram-positive bacterial strains comprising one or more introduced polynucleotides comprising novel pro region DNA sequences operably linked to a nucleic acid (DNA) sequence encoding a protein of interest, and the like.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 326,615, filed April 1, 2022, which is incorporated by reference in its entirety.

[0003] Sequence Listing Reference The contents of the electronic submission of the sequence listing text file entitled "NB41862WOPCT_SequenceListing.xml" was created on March 29, 2023, is 51KB in size, and is incorporated herein by reference in its entirety. [Background technology]

[0004] Gram-positive microorganisms are often used in large-scale industrial fermentation due to their ability to secrete fermentation products into the culture medium. Secreted proteins are transported across the cell membrane and cell wall and then subsequently released into the external medium. For example, large-scale industrial fermentation and secretion of heterologous polypeptides is a technique widely used in the industry, where microbial cells are transformed with a nucleic acid encoding the heterologous polypeptide to be expressed. Despite various advances in protein production methods, there is still a need in the art to provide more efficient methods of protein expression aimed at enhancing the production of proteins of interest that find application for use in various industries. Summary of the Invention [Means for solving the problem]

[0005] As outlined herein, the disclosure provides, inter alia, compositions and methods for producing a protein of interest in a Gram-positive bacterial (host) cell. Particular embodiments relate to novel pro region nucleic acid (DNA) sequences, recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.) comprising the novel pro region DNA sequences, recombinant polynucleotides comprising the novel pro region sequences operably linked to downstream (3') gene coding sequences, recombinant polynucleotides comprising the novel pro region sequences operably linked to upstream (5') DNA sequences encoding a preprotein signal (secretion) sequence, and the like.

[0006] In certain aspects, the disclosure provides recombinant Gram-positive bacterial strains expressing one or more introduced polynucleotides encoding a protein of interest. In certain embodiments, the one or more introduced polynucleotides comprise a novel pro-region DNA sequence operably linked to a DNA sequence encoding a protein of interest (which DNA sequence encoding a protein of interest may comprise an upstream (5') protein signal sequence operably linked thereto). In certain other aspects, the disclosure provides compositions and methods for designing / constructing recombinant Gram-positive bacterial strains expressing one or more introduced novel polynucleotide constructs encoding a protein of interest, compositions and methods for culturing recombinant strains expressing a protein of interest, compositions and methods for enhancing production of a protein of interest, and the like.

[0007] For example, in one or more embodiments, the disclosure provides novel variant pro region sequences set forth in one or more amino acid sequences of SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:14, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, and SEQ ID NO:33, and / or combinations thereof. In one or more other embodiments, the disclosure provides novel variant pro region sequences set forth in Figure 1, Figure 2, or Figure 3, and / or combinations thereof. In yet one or more other embodiments, the disclosure provides novel variant pro region sequences set forth or illustrated in one of Tables 1-5 and / or combinations thereof.

[0008] Thus, one or more particular embodiments relate to variant (mutant) proregion sequences that comprise one or more amino acid substitutions, one or more amino acid insertions, etc. In particular embodiments, the variant proregion sequence comprises an amino acid substitution at position 30, where the amino acid positions of the variant proregion are numbered according to SEQ ID NO: 15. In other embodiments, the variant proregion sequence comprises an amino acid substitution at position 30 and one or more positions selected from 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83, and 84, where the amino acid positions of the variant proregion are numbered according to SEQ ID NO: 15. In one or more other embodiments, the variant pro-region sequence is derived from a parent or reference polypeptide that comprises at least about 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to positions 1-84 of SEQ ID NO: 15. In certain other embodiments, the variant pro-region sequence comprises an amino acid insertion of at least a glycine (G) at position 2 and a lysine (K) at position 3, wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO: 14. In one or more other embodiments, the variant pro-region sequence comprises an amino acid insertion of at least a glycine (G) at position 2 and a lysine (K) at position 3, as well as an amino acid substitution at one or more positions selected from 1, 32, 38, 46, 66, 67, 70 and 73, wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO: 14. In some other embodiments, the variant pro region is derived from a parent or reference polypeptide that comprises at least about 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-86 of SEQ ID NO: 14. In still other embodiments, the variant pro region sequence comprises amino acid insertions of at least a glycine (G) at position 2, a lysine (K) at position 3, and an alanine (A) at position 4, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO:30.In certain aspects, the variant pro region is derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-87 of SEQ ID NO: 30. In one or more other embodiments, the variant pro region sequence comprises amino acid insertions of at least a glycine (G) at position 2, a lysine (K) at position 3, and a serine (S) at position 4, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO:29. In some other embodiments, the variant pro region is derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-87 of SEQ ID NO: 29. In one or more other embodiments, the variant pro region sequence comprises amino acid insertions of at least a glycine (G) at position 2, a lysine (K) at position 3, an alanine (A) at position 4, and an alanine (A) at position 5, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO:31. In one or more specific embodiments or aspects, the variant pro-region is derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-88 of SEQ ID NO: 31. In other embodiments, the variant pro-region sequence comprises a glutamic acid (E) to glycine (G) substitution at position 30 (E30G), a leucine (L) to lysine (K) substitution at position 68 (L68K), and an isoleucine (I) to valine (V) substitution at position 73 (I72V), wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO: 32.In one or more other aspects or embodiments, the variant pro-region sequence comprises a glutamic acid (E) to glycine (G) substitution at position 30 (E30G), a leucine (L) to lysine (K) substitution at position 68 (L68K), and a glutamic acid (E) to isoleucine (I) substitution at position 80 (E80I), wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO: 33. In certain embodiments, the variant pro-region is derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-84 of SEQ ID NO: 15, SEQ ID NO: 32, and / or SEQ ID NO: 33. In other aspects, the variant proregion of the disclosure comprises an amino acid modification set forth in any one of Tables 1-5, Figure 1, Figure 2, Figure 3, SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:14, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, and combinations thereof. In one or more embodiments, a variant proregion comprising an amino acid modification set forth in any one of Tables 1-5, Figure 1, Figure 2, Figure 3, SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:14, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, and combinations thereof is derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to SEQ ID NO:15.

[0009] Accordingly, certain other embodiments provide recombinant nucleic acids encoding the novel variant pro-region sequences of the present disclosure. Accordingly, certain embodiments relate to polynucleotides, etc., that include variant pro-region nucleic acids. In one or more embodiments, the present disclosure provides polynucleotides that include an upstream (5') nucleic acid (sequence) that encodes a variant pro-region of the present disclosure, operably linked to a downstream (3') nucleic acid sequence that encodes a heterologous protein of interest (POI). In certain other aspects, the present disclosure provides polynucleotides that include an upstream (5') nucleic acid that encodes a preprotein signal (secretion) sequence, operably linked to a downstream (3') nucleic acid sequence that encodes a variant pro-region of the present disclosure, operably linked to a downstream nucleic acid sequence that encodes a protein of interest (POI). Certain other embodiments relate to expression cassettes that include an upstream (5') promoter region sequence operably linked to a downstream (3') polynucleotide of the present disclosure.

[0010] In certain other embodiments, the present disclosure relates to recombinant (engineered) Gram-positive bacterial cells / strains comprising one or more introduced polynucleotides or expression cassettes of the present disclosure. Accordingly, other embodiments of the present disclosure relate to methods / processes for producing a heterologous protein of interest in a recombinant Gram-positive cell as described and illustrated herein.

[0011] In some one or more embodiments, the disclosure provides a method of producing a heterologous POI in a Gram-positive bacterial cell, the method comprising introducing into the Gram-positive cell an expression cassette comprising an upstream promoter region operably linked to a downstream nucleic acid sequence encoding a variant pro region comprising an amino acid substitution at one or more positions selected from position 30 and 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83 and 84, wherein the variant pro region nucleic acid (sequence) is operably linked to the downstream nucleic acid sequence encoding the POI, and the amino acid positions of the variant pro region are numbered according to SEQ ID NO: 15, and growing / cultivating / fermenting the modified cell under conditions for production of the POI. In certain other embodiments, the present disclosure provides a method of producing a heterologous POI in a Gram-positive bacterial cell, the method comprising introducing into the Gram-positive cell an expression cassette comprising an upstream promoter region sequence operably linked to a downstream nucleic acid (sequence) encoding a variant pro-region comprising an amino acid insertion of glycine (G) at position 2 and lysine (K) at position 3 and an amino acid substitution at one or more positions selected from 1, 32, 38, 46, 66, 67, 70 and 73, wherein the variant pro-region nucleic acid (sequence) is operably linked to the downstream nucleic acid sequence encoding the POI, and the amino acid positions of the variant pro-region are numbered according to SEQ ID NO: 14, and growing / cultivating / fermenting the modified cell under conditions for production of the POI. In yet other embodiments, the present disclosure relates to a method of producing a heterologous POI in a Gram-positive bacterial cell, the method comprising introducing into the Gram-positive cell an expression cassette of the present disclosure and growing / cultivating / fermenting the modified cell under conditions for production of the POI. In one or more particular embodiments of the method, the cassette comprises a nucleic acid (sequence) encoding a preprotein signal (secretion) sequence operably linked and positioned between the promoter and the variant pro region sequence, hi some other embodiments of the method, the modified cells produce increased amounts of the POI as compared to a control Gram-positive cell.For example, the control Gram-positive cell comprises an introduced expression cassette comprising the same upstream promoter sequence (i.e., the same as the modified cell) operably linked to a downstream nucleic acid encoding a pro region sequence comprising SEQ ID NO: 15, which is operably linked to a downstream nucleic acid sequence encoding the same POI (i.e., the same POI as the modified cell). In certain embodiments of the method, the introduced cassette is integrated into the genome of the cell. In certain embodiments of the method, at least two cassettes are introduced into the cell. Certain other embodiments of the disclosure are described herein and below. [Brief description of the drawings]

[0012] [Figure 1]1 shows the amino acid sequence of a wild-type (reference) B. lentus pro-region and specific SEL variant pro-region sequences of the present disclosure. As shown in FIG. 1, the reference B. lentus pro-region sequence includes the 84 amino acid (residue) positions set forth in SEQ ID NO: 15 (with the glutamic acid (E) residue at position 30 underlined), variant pro-region sequence A includes the 84 amino acid (residue) positions set forth in SEQ ID NO: 9 (with the histidine (H) residue at position 30 in bold), and variant pro-region sequence B includes the 84 amino acid (residue) positions set forth in SEQ ID NO: 11 ( Variant pro-region sequence A includes 86 amino acid (residue) positions set forth in SEQ ID NO: 14 (with glycine (G) and lysine (K) residues inserted at positions 2 and 3, respectively, shown in bold (GK), and the E residue at position 32 is underlined), and variant pro-region sequence B includes 87 amino acid (residue) positions set forth in SEQ ID NO: 29 (with glycine (G), lysine (K) and E residues inserted at positions 2-4, respectively). variant pro-region sequence A comprises 87 amino acid (residue) positions set forth in SEQ ID NO:30 (with glycine (G), lysine (K), and alanine (A) residues inserted at positions 2-4, respectively, shown in bold (GKA)); variant pro-region sequence B comprises 87 amino acid (residue) positions set forth in SEQ ID NO:31 (with glycine (G), lysine (K), alanine (A), and alanine (A) residues inserted at positions 2-5, respectively, shown in bold (GKA)); variant pro-region sequence C comprises 88 amino acid (residue) positions set forth in SEQ ID NO:32 (with glycine (G), lysine (K), alanine (A), and alanine (A) residues inserted at positions 2-5, respectively, shown in bold (GKA)); Variant proregion sequence G contains the 84 amino acid (residue) positions set forth in SEQ ID NO:32 (with the glycine (G) residue at position 30 in bold and the L68K and I72V substitutions underlined), and variant proregion sequence H contains the 84 amino acid (residue) positions set forth in SEQ ID NO:33 (with the glycine (G) residue at position 30 in bold and the L68K and E80I substitutions underlined).

[0013] [Diagram 2]The amino acid sequences of variant pro region sequence A (sequence 09), variant pro region sequence B (sequence 11), variant pro region sequence C (sequence 14), variant pro region sequence D (sequence 29), variant pro region sequence E (sequence 30), variant pro region sequence F (sequence 31), variant pro region sequence G (sequence 32), and variant pro region sequence H (sequence 33) aligned with the wild-type (WT) pro region sequence (sequence 15). As shown in FIG. 2, the WT pro-region contains a glutamic acid (E) residue at position 30 (E, sequence 15), variant pro-region sequence A contains a histidine (H) residue at position 30 (H, sequence 09), variant pro-region sequence B contains a glycine (G) residue at position 30 (G, sequence 11), variant pro-region sequence C contains two amino acid insertions of glycine (G) and lysine (K) (GK insertion, sequence 14), variant pro-region sequence D contains three amino acid insertions of glycine (G), lysine (K) and serine (S) (GKS insertion, sequence 29), and variant pro-region sequence F contains two amino acid insertions of glycine (G), lysine (K) and serine (S) (GKS insertion, sequence 29). variant proregion sequence E contains a three amino acid insertion of glycine (G), lysine (K) and alanine (A) (GKA insertion, sequence 30); variant proregion sequence F contains a four amino acid insertion of glycine (G), lysine (K), alanine (A) and alanine (A) (GKAA insertion, sequence 31); variant proregion sequence G contains a glycine (G) residue at position 30, a lysine (K) residue at position 68 and a valine (V) residue at position 72; variant proregion sequence H contains a glycine (G) residue at position 30, a lysine (K) residue at position 68 and an isoleucine (I) residue at position 80. For example, as shown in the alignment of the WT pro region sequence (SEQ ID NO:15) with variant pro region sequence C (SEQ ID NO:14), two amino acid residues GK are inserted between alanine (A) at position 1 and glutamic acid (E) at position 2 of the WT pro region (SEQ ID NO:15), resulting in the 86 amino acid variant pro region sequence C of SEQ ID NO:14. As shown in FIG. 2, the two hyphens (--) shown in SEQ ID NO:15, SEQ ID NO:09, SEQ ID NO:11, SEQ ID NO:32 and SEQ ID NO:33 indicate a gap of two amino acid positions relative to the variant pro region sequence C shown in SEQ ID NO:14.

[0014] [Diagram 3]Reference (wild-type) B. lentus proregion amino acid sequence (E30, SEQ ID NO: 15), variant proregion sequence A (H30, SEQ ID NO: 09), variant proregion sequence B (G30, SEQ ID NO: 11), variant proregion sequence C (G2K3 insertion, SEQ ID NO: 14), variant proregion sequence D (G2K3S4 insertion, SEQ ID NO: 29), variant proregion sequence E (G2K3A4 insertion, SEQ ID NO: 30), variant proregion sequence F (G2K3A4A5 insertion, SEQ ID NO: 31), variant proregion sequence G (G2K3A4A5 insertion, SEQ ID NO: 32), variant proregion sequence H (G2K3A4A5 insertion, SEQ ID NO: 33), variant proregion sequence I (G2K3A4A5 insertion, SEQ ID NO: 34), variant proregion sequence J (G2K3A4A5 insertion, SEQ ID NO: 35), variant proregion sequence K (G2K3A4A5 insertion, SEQ ID NO: 36), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 37), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 38), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 39), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 40), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 41), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 42), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 43), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 44), variant proregion sequence L (G2K3A4A5 insertion, SEQ ID NO: 45), 15 shows the numbering of amino acid (residue) positions of variant pro region sequence G (with substitutions E30G, L68K, and I72V, SEQ ID NO:32) and variant pro region sequence F (with substitutions E30G, L68K, and E80I, SEQ ID NO:33), where the amino acid positions are indicated in the 5' (N-terminus) to 3' (C-terminus) direction with subscripted numbers (e.g., alanine (A1) at position 1 of SEQ ID NO:15, glutamic acid (E2) at position 2 of SEQ ID NO:15, etc.). For example, the variant pro region sequences of the disclosure may refer to the numbering of the wild type (reference) pro region sequence (SEQ ID NO: 15, E30), the numbering of the variant (reference) pro region sequence A (SEQ ID NO: 9, H30), the numbering of the variant (reference) pro region sequence B (SEQ ID NO: 11, G30), the numbering of the variant (reference) pro region sequence C (SEQ ID NO: 14, GK insertion), the numbering of the variant (reference) pro region sequence D (SEQ ID NO: 29, GKS insertion), the numbering of the variant (reference) pro region sequence E (SEQ ID NO: 20, GKA insertion), the numbering of the variant (reference) pro region sequence F (SEQ ID NO: 31, GKAA insertion), the numbering of the variant (reference) pro region sequence G (SEQ ID NO: 32, L68K / I72V substitution), the numbering of the variant (reference) pro region sequence H (SEQ ID NO: 33, L68K / E80I substitution), or a combination thereof, as depicted in FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] A brief description of biological sequences SEQ ID NO:1 is a nucleic acid (DNA) sequence containing the upstream (5') aprE gene flanking regions, the variant rrnI-P2 promoter of B. subtilis and the 5'-aprE UTR region.

[0016] SEQ ID NO:2 is the amino acid sequence of the wild-type Bacillus gibsonii subtilisin designated "BG46."

[0017] SEQ ID NO:3 is the amino acid sequence of the B. gibsonii BG46 subtilisin, a variant designated "BG46_variant 1."

[0018] SEQ ID NO:4 is a DNA sequence encoding the AprE protein signal sequence of wild-type B. subtilis.

[0019] SEQ ID NO:5 is the DNA sequence encoding variant pro region sequence A (30H, SEQ ID NO:9).

[0020] SEQ ID NO:6 is the DNA sequence of the wild-type B. amyloliquefaciens BPN' terminator.

[0021] SEQ ID NO:7 is the DNA sequence of the kanamycin (kan) gene expression cassette.

[0022] SEQ ID NO:8 is the amino acid sequence of the B. gibsonii BG46 subtilisin, a variant designated "BG46_variant 2."

[0023] SEQ ID NO:9 is the amino acid sequence encoded by the DNA sequence of variant proregion A (SEQ ID NO:5).

[0024] SEQ ID NO:10 is the DNA sequence encoding variant pro region sequence B (30G, SEQ ID NO:11).

[0025] SEQ ID NO:11 is the amino acid sequence encoded by the DNA sequence of variant pro region B (SEQ ID NO:10).

[0026] SEQ ID NO:12 is the amino acid sequence of the B. amyloliquefaciens BPN' pro region sequence.

[0027] SEQ ID NO:13 is the DNA sequence encoding the variant pro region sequence C (GK insertion, SEQ ID NO:14).

[0028] SEQ ID NO:14 is the amino acid sequence encoded by the DNA sequence of variant pro region C (SEQ ID NO:13).

[0029] SEQ ID NO:15 is the amino acid sequence of the wild-type (WT) B. lentus pro region sequence.

[0030] SEQ ID NO:16 is the DNA sequence of the B. licheniformis serA gene 5'FR.

[0031] SEQ ID NO: 17 is a DNA sequence containing the rrnI-p3 promoter region.

[0032] SEQ ID NO:18 is the DNA sequence of the B. subtilis aprE 5'-UTR.

[0033] SEQ ID NO:19 is the DNA sequence of the B. licheniformis amyL terminator.

[0034] SEQ ID NO:20 is the DNA sequence of the B. licheniformis serA gene 3'FR.

[0035] SEQ ID NO:21 is the DNA sequence of the B. licheniformis lysA gene 5'FR.

[0036] SEQ ID NO:22 is the DNA sequence of the B. licheniformis lysA gene 3'FR.

[0037] SEQ ID NO:23 is the DNA sequence encoding variant pro region sequence C (GKA insert, SEQ ID NO:14).

[0038] SEQ ID NO:24 is the amino acid sequence encoded by the DNA sequence of variant pro region C (SEQ ID NO:23).

[0039] SEQ ID NO:25 is the DNA sequence encoding the variant pro region sequence C (GKAA insert, SEQ ID NO:14).

[0040] SEQ ID NO:26 is the amino acid sequence encoded by the DNA sequence of variant pro region C (SEQ ID NO:25).

[0041] SEQ ID NO:27 is the DNA sequence encoding variant pro region sequence C (GKAS insert, SEQ ID NO:14).

[0042] SEQ ID NO:28 is the amino acid sequence encoded by the DNA sequence of variant pro region C (SEQ ID NO:27).

[0043] SEQ ID NO:29 is the amino acid sequence of variant pro region sequence D.

[0044] SEQ ID NO:30 is the amino acid sequence of variant pro region sequence E.

[0045] SEQ ID NO:31 is the amino acid sequence of variant pro region sequence F.

[0046] SEQ ID NO:32 is the amino acid sequence of variant pro region sequence G.

[0047] SEQ ID NO:33 is the amino acid sequence of variant pro region sequence H.

[0048] As briefly described above and explained in detail herein below, the present disclosure provides, inter alia, novel pro region nucleic acid (DNA) sequences, recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.) comprising novel pro region DNA sequences, recombinant polynucleotides comprising novel pro region sequences operably linked to downstream (3') gene coding sequences, recombinant polynucleotides comprising novel pro region sequences operably linked to upstream (5') DNA sequences encoding protein signal (secretion) sequences, etc. In certain aspects, the present disclosure provides recombinant Gram-positive bacterial strains expressing one or more introduced polynucleotides encoding a protein of interest. In certain embodiments, the one or more introduced polynucleotides comprise a novel pro region DNA sequence operably linked to a DNA sequence encoding a protein of interest (which may include an upstream (5') protein signal sequence operably linked thereto). Accordingly, certain other aspects of the present disclosure provide, inter alia, compositions and methods for designing / constructing recombinant Gram-positive bacterial strains that express one or more introduced novel polynucleotide constructs encoding a protein of interest, compositions and methods for culturing recombinant strains expressing a protein of interest, compositions and methods for enhancing production of a protein of interest, and the like.

[0049] I. Definition In view of the disclosed recombinant (engineered) strains and methods thereof described herein, the following terms and phrases are defined. Terms not defined herein should be given the meaning commonly used in the art.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the compositions and methods of the present invention belong. Although any methods and materials similar or equivalent to those described herein can be used to carry out or test the compositions and methods of the present invention, exemplary methods and materials are described below. All publications and patents cited herein are incorporated herein by reference in their entirety.

[0051] It is further noted that the claims may be drafted to exclude optional elements, and thus, this statement is intended to serve as a prelude to the use of exclusive terminology such as "solely," "only," "except," "not including," or the use of any "negative" limitation or qualification in connection with the recitation of claim elements.

[0052] As will be apparent to one of ordinary skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features which may be readily separated or combined with the features of any of the other embodiments without departing from the scope or spirit of the compositions and methods described herein. Any of the methods described may be carried out in the order of events recited or in any other order which is logically possible.

[0053] As used herein, the terms "gram-positive bacteria", "gram-positive cells", "gram-positive bacterial strains" and / or "gram-positive bacterial cells" have the same meaning as used in the art. For example, gram-positive bacterial cells include all strains of Actinobacteria and Firmicutes. In certain embodiments, such gram-positive bacteria are of the classes Bacilli, Clostridia, and Mollicutes.

[0054] As used herein, the genus "Bacillus" includes all species within the genus "Bacillus" known to those of skill in the art, such as B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophii, B. arginini ... Examples of species of Bacillus that may be used include, but are not limited to, B. lus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis. It is recognized that the genus Bacillus continues to undergo taxonomic reorganization. Thus, the genus is intended to include reclassified species, such as, but not limited to, organisms such as "B. stearothermophilus," which is now referred to as "Geobacillus stearothermophilus."

[0055] As used herein, the term "recombinant" or "non-naturally occurring" refers to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic change or has been modified by the introduction of a heterologous nucleic acid molecule, or to a cell (e.g., a microbial cell) that has been modified so that the expression of a heterologous or endogenous nucleic acid molecule or gene can be controlled. Recombinant also refers to a cell that is derived from or is the descendant of a non-naturally occurring cell that has one or more such modifications. Genetic changes include, for example, modifications that introduce an expressible nucleic acid molecule that encodes a protein, the addition, deletion, substitution of a nucleic acid molecule, or other functional changes to the genetic material of the cell. For example, recombinant cells can express genes or other nucleic acid molecules that are not found in the same or homologous form in a native (wild-type) cell (e.g., fusion or chimeric proteins), or can provide an altered expression pattern of an endogenous gene, such as being overexpressed, underexpressed, minimally expressed, or not expressed at all. "Recombination," "recombining," or producing a "recombinant" nucleic acid is generally the assembly of two or more nucleic acid fragments, which assembly gives rise to a chimeric gene.

[0056] The term "derived from" includes the terms "originating from," "obtained from," "available from," and "made from," and generally indicates that one particular material or composition finds its origin in, or has characteristics that are describable with reference to, another particular material or composition. For example, the recombinant Gram-positive bacterial cells of the present disclosure can be derived / obtained from any known Gram-positive bacterial strain.

[0057] As used herein, "nucleic acid" refers to DNA, cDNA and RNA of genomic or synthetic origin, which may be double-stranded or single-stranded, including nucleotide or polynucleotide sequences and fragments or portions thereof, and whether representing the sense or antisense strand. It will be understood that, as a result of the degeneracy of the genetic code, a large number of nucleotide sequences can code for a given protein.

[0058] The polynucleotides (or nucleic acid molecules) described herein are understood to include "genes," "vectors," and "plasmids."

[0059] Thus, the term "gene" refers to a polynucleotide that codes for a particular sequence of amino acids, including all or part of a coding sequence for a protein, and may include regulatory (non-transcribed) DNA sequences, such as promoter sequences, that determine the conditions under which the gene is expressed. The transcribed region of a gene may include untranslated regions (UTRs), including 5'-untranslated regions (UTRs) and 3'-UTRs, as well as the coding sequence.

[0060] As used herein, "endogenous gene" refers to a gene that is present in its natural location in the genome of an organism.

[0061] As used herein, a "heterologous" gene, "non-endogenous" gene or "foreign" gene refers to a gene that is not normally found in the host organism, but that has been introduced into the host organism by gene transfer. The term "foreign" gene includes a native gene inserted into a non-native organism and / or a chimeric gene inserted into a native or non-native organism.

[0062] As used herein, "heterologous control sequences" refers to gene expression control sequences (e.g., promoters, enhancers, terminators, etc.) that do not function in nature to regulate (control) the expression of a gene of interest. Generally, heterologous nucleic acids are not endogenous (natural) to the cell or part of the genome in which they are present, but have been added to the cell by infection, transfection, transformation, transduction, microinjection, electroporation, etc. A "heterologous" nucleic acid construct may contain a control sequence / DNA coding (ORF) sequence combination that is the same or different from the control sequence / DNA coding sequence combination found in the native host cell.

[0063] As used herein, the term "expression" refers to the transcription and stable accumulation of sense (mRNA) or antisense RNA derived from the nucleic acid molecule of the present disclosure. Expression can also refer to the translation of mRNA into a polypeptide. Thus, the term "expression" includes any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0064] As used herein, the term "coding sequence" refers to a nucleotide sequence, which directly specifies the amino acid sequence of its (encoded) protein product. The boundaries of the coding sequence are generally determined by an open reading frame (hereinafter "ORF"), which usually begins with the ATG start codon. Coding sequences typically include DNA, cDNA, and recombinant nucleotide sequences.

[0065] As used herein, the terms "promoter", "promoter element", "promoter sequence" and the like refer to a nucleic acid (DNA) sequence capable of controlling the transcription of a gene coding sequence (CDS) into messenger RNA (mRNA) when the promoter region sequence is placed upstream (5') and operably linked to a downstream (3') gene CDS. As generally understood by those skilled in the art, a promoter generally provides a site for specific binding and transcription initiation by an RNA polymerase. In certain aspects, the term "promoter" refers to the minimal portion of a promoter nucleic acid sequence required for transcription initiation (i.e., includes an RNA polymerase binding site). For example, a promoter generally includes a "-10" (consensus sequence) element and a "-35" (consensus sequence) element that are upstream (5') of the gene CDS to be translated and relative to the +1 transcription start site (TSS). The core promoter -10 and -35 elements are commonly referred to in the art as the "TATAAT" (Pribnow box) consensus region and the "TTGACA" consensus region, respectively. The core promoter (-10 and -35) regions are generally spaced (ie, spaced) apart by about 15-20 intervening base pairs (nucleotides).

[0066] Promoters may be derived entirely from natural genes, or may be composed of various elements derived from various naturally occurring promoters, or may even include synthetic nucleic acid segments. Those skilled in the art will appreciate that various promoters may direct the expression of genes in various cell types, or at various developmental stages, or in response to various environmental or physiological conditions. Promoters may be constitutive, inducible, variable, hybrid, synthetic, tandem, etc. The promoters that most often cause expression of genes in the majority of cell types are generally referred to as "constitutive promoters". Furthermore, it has been recognized that DNA fragments of various lengths may have the same promoter activity, since in most cases the exact boundaries of regulatory sequences are not completely clear. In certain embodiments, an upstream (5') promoter sequence (pro) operably linked to a downstream DNA sequence encoding a protein's signal sequence (SS) operably linked to a downstream DNA sequence encoding a pro region sequence (PRO) operably linked to a downstream (3') DNA sequence encoding a mature protein of interest (ORF) may be depicted diagrammatically as 5'-[pro]-[SS]-[PRO]-[ORF]-3'.

[0067] As used herein, a "functional promoter sequence" that controls expression of a gene of interest linked to a protein coding sequence of the gene of interest refers to a promoter sequence that controls the transcription and translation of the coding sequence in a desired Gram-positive host cell. For example, in certain embodiments, the present disclosure provides a polynucleotide that includes an upstream (5') promoter (or 5' promoter region or tandem 5' promoter, etc.) that functions in a Gram-positive cell, where the functional promoter region is operably linked to a nucleic acid sequence that encodes a protein of interest.

[0068] As used herein, the term "precursor protein" refers to an inactive form of a protein. In certain embodiments, full-length proteins are synthesized as a precursor of a prosequence to the mature protein form (abbreviated as "preprotein"). In other embodiments, full-length proteins are synthesized as a signal peptide sequence, a prosequence, and a precursor of the mature protein form (abbreviated as "pre-proprotein"). For example, the presequence usually acts as a signal peptide for transport, and the prosequence is generally essential for correct folding of the associated (mature) protein.

[0069] As used herein, the term "mature protein" refers to the active form of a protein, as opposed to the inactive precursor (full-length) protein.

[0070] As used herein, the terms "signal sequence", "secretion signal" and "signal peptide" may be used interchangeably and refer to a sequence of amino acid residues that may be involved in the secretion or direct transport of a precursor protein. A signal (pre) sequence is generally cleaved from the precursor protein by a signal peptidase during translocation. A signal (pre) sequence is generally located at the N-terminus of the mature protein sequence or at the N-terminus of the proregion (pro) sequence when the signal (pre) sequence and the proregion (pro) sequence are used in operable combination upstream (5') of the mature POI sequence.

[0071] As used herein, the terms "pro sequence", "pro-sequence" and "pro region sequence" may be used interchangeably and may be abbreviated as "PRO" sequence, "Pro" sequence, "pro" sequence, etc. As used herein, the term pro sequence has the same meaning as understood in the art. For example, the B. subtilis alkaline serine protease "subtilisin" is initially produced as a pre-prosubtilisin, which consists of a signal (pre) sequence for protein secretion, followed by a 77 amino acid pro region (pro) sequence, followed by an amino acid sequence encoding the mature subtilisin (e.g., pre-prosubtilisin). Pro sequences are often essential for the correct folding of the associated (mature) protein, acting as intramolecular chaperones (e.g., directly catalyzing protein folding reactions). Similarly, pro sequences may be required for both folding and intracellular transport (or secretion) of the mature protein of interest, suggesting that these two functions are closely related. In certain aspects, a proregion sequence of the present disclosure comprises an amino acid sequence derived from a wild-type (WT, ref.) B. lentus proregion sequence of SEQ ID NO: 15. In certain embodiments, the proregion sequence is derived from an amino acid sequence that includes homology to SEQ ID NO: 15, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 14, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, and / or SEQ ID NO: 33. Thus, in certain aspects, the amino acid modifications of one or more of the proregion variants described herein are numbered by reference to the proregion amino acid sequence of SEQ ID NO: 15, SEQ ID NO: 9, SEQ ID NO: 11, or SEQ ID NO: 14.

[0072] For example, the amino acid sequence of one or more of the pro-region variants described herein can be aligned with the amino acid sequence of SEQ ID NO: 15 using an alignment algorithm, where each amino acid residue in a given amino acid sequence that aligns (preferably optimally aligns) with an amino acid residue in SEQ ID NO: 15 is conveniently numbered by reference to the position number of the corresponding amino acid residue. Sequence alignment algorithms identify positions where insertions or deletions occur in a subject sequence compared to a query sequence (sometimes also referred to as a "reference sequence"). Sequence alignment with other pro-region amino acid sequences can be determined using amino acid alignment algorithms.

[0073] In certain aspects, references to a "position" of a proregion amino acid sequence can be shown as a single letter amino acid (residue) followed by the position number (e.g., in SEQ ID NO:15, an alanine at position 1 is designated as "A1", a glutamic acid at position 2 is designated as "E2", etc.). In related embodiments, references to variant (mutant) proregion amino acid sequences can be shown as a single letter amino acid (residue), where the amino acid position of the parent (reference) sequence is numbered followed by the mutated amino acid (residue) at the same position (e.g., in SEQ ID NO:15, an alanine (A) at position 1 substituted with a glycine (G) is designated as "A1G", a glutamic acid (E) at position 30 substituted with a histidine (H) is designated as "E30H", etc.). Multiple amino acid residues may also be substituted at the same position in the proregion amino acid sequence, where the amino acid position in the parent (reference) sequence is numbered followed by the mutated amino acid residues at the same position separated by a slash (e.g., in SEQ ID NO:15, alanine (A) at position 1 substituted with glycine (G), glutamic acid (E), or lysine (K) may be indicated as "A1G / E / K").

[0074] In this specification, some ranges are indicated by numerical values ​​preceded by the term "about". In this specification, the term "about" is used to provide literal support for the exact number it precedes and a number close to or approximately the number it precedes. In determining whether a number is close to or approximately a specifically recited number, the close or approximate number not recited may be a number that provides a substantial equivalent to the specifically recited number in the context in which it is stated. For example, the term "about" in relation to a numerical value refers to a range of -10% to +10% of the numerical value, unless the term is clearly defined otherwise in the context.

[0075] As used herein, polynucleotides encoding a "pro" sequence, polynucleotides encoding a "pro region" sequence, and polynucleotides encoding a "pro sequence" may be used interchangeably and refer to a DNA sequence located immediately upstream (5') of and operably linked to a downstream (3') gene coding sequence (CDS) encoding a mature protein of interest (POI). In certain aspects, a polynucleotide encoding a pro region sequence refers to a DNA sequence located between and linking an upstream (5') protein signal sequence (SS) and a downstream (3') gene CDS encoding a mature POI. For example, a pro region DNA sequence (PRO) is operably linked to an upstream (5') DNA sequence (SS) encoding a protein signal sequence and operably linked to a downstream (3') DNA gene CDS encoding a mature POI (e.g., 5'-[SS]-[PRO]-[ORF]-3'; N-[pre]-[pro]-[protein]-C).

[0076] As used herein, a "wild-type (WT) B. lentus proregion" sequence comprises the amino acid sequence set forth in SEQ ID NO: 15. In certain embodiments, the WT (reference) proregion sequence (SEQ ID NO: 15) is also referred to as the "E30 proregion" or "E 30 These sequences may be referred to as "pro sequences."

[0077] As used herein, the variant proregion sequence referred to as "variant proregion A" comprises the amino acid sequence set forth in SEQ ID NO: 9. In certain embodiments, the variant proregion A sequence (SEQ ID NO: 9) is a "H30 proregion" or "H 30 These sequences may be referred to as "pro sequences."

[0078] As used herein, the variant proregion sequence referred to as "variant proregion B" comprises the amino acid sequence set forth in SEQ ID NO: 11. In certain embodiments, the variant proregion B sequence (SEQ ID NO: 11) is a "G30 proregion" or "G 30 These sequences may be referred to as "pro sequences."

[0079] As used herein, the variant pro region sequence referred to as "variant pro region C" comprises the amino acid sequence set forth in SEQ ID NO: 14. In certain embodiments, the variant pro region C sequence (SEQ ID NO: 14) may be referred to as the "GK insert pro region" or the "G2K3 insert pro region."

[0080] As used herein, the variant pro region sequence referred to as "variant pro region D" comprises the amino acid sequence set forth in SEQ ID NO: 29. In certain embodiments, the variant pro region D sequence (SEQ ID NO: 29) may be referred to as the "GKS insert pro region" or the "G2K3S4 insert pro region."

[0081] As used herein, the variant pro region sequence referred to as "variant pro region E" comprises the amino acid sequence set forth in SEQ ID NO: 30. In certain embodiments, the variant pro region E sequence (SEQ ID NO: 30) may be referred to as the "GKA insert pro region" or the "G2K3A4 insert pro region."

[0082] As used herein, the variant pro region sequence referred to as "variant pro region F" comprises the amino acid sequence set forth in SEQ ID NO: 31. In certain embodiments, the variant pro region F sequence (SEQ ID NO: 31) may be referred to as the "GKAA insertion pro region" or the "G2K3A4A5 insertion pro region."

[0083] As used herein, the variant pro region sequence referred to as "variant pro region G" comprises the amino acid sequence set forth in SEQ ID NO: 32. In certain embodiments, the variant pro region G sequence (SEQ ID NO: 32) may be referred to as the L68K / I72V pro region.

[0084] As used herein, the variant pro region sequence referred to as "variant pro region H" comprises the amino acid sequence set forth in SEQ ID NO: 32. In certain embodiments, the variant pro region H sequence (SEQ ID NO: 33) may be referred to as the L68K / E80I pro region.

[0085] As roughly shown in FIG. 3, the numbering of amino acid (residue) positions is based on the WT pro-region amino acid sequence (E 30 , SEQ ID NO: 15), variant pro region sequence A (H 30 , SEQ ID NO: 09), variant pro region sequence B (G 30 , SEQ ID NO:11), variant pro region sequence C (GK insertion, SEQ ID NO:14), variant pro region sequence D (GKS insertion, SEQ ID NO:29), variant pro region sequence E (GKA insertion, SEQ ID NO:30), variant pro region sequence F (GKAA insertion, SEQ ID NO:31), variant pro region sequence G (L68K / I72V, SEQ ID NO:32) and / or variant pro region sequence H (L68K / E80I, SEQ ID NO:33).

[0086] As used herein, the phrase "a polynucleotide encoding a full-length protein" refers to a DNA sequence encoding a "precursor" protein, and the phrase "a polynucleotide encoding a mature protein" refers to a DNA sequence encoding a "mature" protein, as defined herein. In certain embodiments, a polynucleotide encoding a precursor protein comprises at least an upstream (5') DNA sequence encoding a pro-region amino acid sequence operably linked to a downstream (3') DNA sequence (e.g., an open reading frame, ORF) encoding the amino acid sequence of a mature protein of interest (POI). In other embodiments, a polynucleotide encoding a precursor protein comprises at least an upstream (5') DNA sequence encoding a protein signal sequence operably linked to a downstream (3') DNA sequence encoding a pro-region amino acid sequence operably linked to a downstream (3') ORF encoding the amino acid sequence of a mature POI.

[0087] As used herein, the term "untranslated region" may be abbreviated as "UTR."

[0088] As used herein, the phrases "five' end (5') untranslated region", "5' untranslated region" and / or "5' transcript leader" may be used interchangeably and may be abbreviated as "5'-UTR". As generally understood in the art, the 5'-UTR is known as the region of a messenger RNA (mRNA) that is immediately upstream (5') of the start codon.

[0089] A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, DNA encoding a secretory leader (i.e., signal sequence) is operably linked to DNA encoding a polypeptide if it is expressed as a preprotein involved in the secretion of the polypeptide, a promoter or enhancer is operably linked to a coding sequence (CDS, ORF) if it affects the transcription of the sequence, or a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation. Generally, "operably linked" means that the DNA sequences being linked are contiguous, and in the case of a secretory leader, contiguous and in reading phase. Enhancers, however, need not be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide adaptors or linkers are used in accordance with conventional practice. Thus, the term operably linked generally refers to the association (juxtaposition) of nucleic acid sequences on a single nucleic acid fragment such that the function of one is affected by the function of the other. For example, a promoter (pro) is operably linked to a gene coding sequence (gene CDS) if it controls the transcription of the gene CDS (eg, 5'-[pro]-[gene CDS] 3').

[0090] As used herein, the phrase "variant rrnI-P2 promoter and 5'-UTR region" (abbreviated "P2 promoter / 5'-UTR region") refers to the DNA sequence of the variant B. subtilis rrnI-P2 promoter / 5'-aprE UTR region set forth in SEQ ID NO:1.

[0091] As used herein, DNA encoding the "wild-type B. subtilis aprE signal peptide sequence" may be abbreviated as "aprE SS," which comprises the nucleotide sequence of SEQ ID NO:4.

[0092] As used herein, the "wild-type B. amyloliquefaciens BPN' terminator (BPN' term)", which may be abbreviated as "term", comprises the nucleotide sequence of SEQ ID NO:6.

[0093] As used herein, the variant Bacillus gibsonii (BG46) subtilisin comprising the amino acid sequence of SEQ ID NO:3 is abbreviated as "BG46_variant 1."

[0094] As used herein, the variant Bacillus gibsonii (BG46) subtilisin comprising the amino acid sequence of SEQ ID NO:8 is abbreviated as "BG46_variant 2", which is derived from the wild-type B. gibsonii (BG46) subtilisin reporter protein (SEQ ID NO:2).

[0095] As used herein, "suitable regulatory sequences" refer to nucleotide sequences located upstream (5' non-coding sequences), within or downstream (3' non-coding sequences) of a coding sequence that influence the transcription, RNA processing or stability or translation of the associated coding sequence. Regulatory sequences may include promoters, transcription leader sequences, RNA processing sites, effector binding sites and stem-loop structures.

[0096] As used herein, "host cell" refers to a cell that has the ability to act as a host or expression vehicle for a newly introduced DNA sequence. Thus, in certain embodiments of the present disclosure, the host cell is a Gram-positive cell (e.g., Bacillus sp.) and / or a Gram-negative cell (e.g., E. coli).

[0097] As used herein, an "modified cell" refers to a recombinant cell that contains at least one genetic modification that is not present in the parent, reference, or control cell from which the modified cell is derived.

[0098] As used herein, when comparing expression and / or production of a protein of interest (POI) in a recombinant (modified) cell with expression and / or production of the same POI in an unmodified (control) cell, it will be understood that the modified and unmodified cells are grown / cultured / fermented under identical conditions (e.g., identical conditions of medium, temperature, pH, etc.).

[0099] As used herein, "increased amount," when used in phrases such as "the recombinant cells "express / produce" an increased amount of a protein of interest compared to unmodified (control) cells," specifically refers to the "increased amount" of the protein of interest (POI) expressed / produced by the recombinant cells, but this "increased amount" is always compared to unmodified (control) cells expressing / producing the same POI, and the modified and unmodified cells are grown / cultured / fermented under the same conditions.

[0100] As used herein, "increasing" protein production or "increased" protein production means that an increased amount of a protein (e.g., a protein of interest) is produced. The protein can be produced inside the host cell or secreted (or transported) into the culture medium. In certain embodiments, the protein of interest is produced (secreted) into the culture medium. An increase in protein production can be detected, for example, as a higher maximum level of protein or enzyme activity (e.g., amylase activity, etc.) or total extracellular protein produced compared to the parent host cell.

[0101] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) the introduction, substitution or removal of one or more nucleotides in a gene (or its ORF) or the introduction, substitution or removal of one or more nucleotides in a regulatory element required for the transcription or translation of a gene or its ORF; (b) gene disruption; (c) gene conversion; (d) gene deletion; (e) gene downregulation; (f) directed mutagenesis; and / or (g) random mutagenesis of any one or more genes disclosed herein.

[0102] In certain aspects, genetic modification specifically refers to the introduction, substitution, or removal of one or more nucleotides in a nucleic acid (DNA) sequence encoding a pro region (amino acid) sequence of the present disclosure. For example, in certain embodiments, a DNA sequence encoding the native pro region amino acid sequence set forth in SEQ ID NO: 15 is genetically modified as described herein.

[0103] As used herein, the term "introducing", when used in phrases such as "introducing a gene, a polynucleotide, an open reading frame (ORF), a gene coding sequence, a vector, an expression cassette, etc. into a Gram-positive bacterial cell", includes methods known in the art for introducing a polynucleotide (DNA) into a cell, including, but not limited to, protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation, etc.

[0104] As used herein, "transformed" or "transformation" refers to a cell that has been transformed by the use of recombinant DNA technology. Transformation generally occurs by inserting one or more nucleotide sequences (e.g., polynucleotides, ORFs, or genes) into a cell. The inserted nucleotide sequence may be a heterologous nucleotide sequence (i.e., a sequence that does not naturally occur in the cell being transformed). Thus, transformation generally refers to the introduction of exogenous DNA into a host cell such that the DNA is maintained as a chromosomal integrant or a self-replicating extrachromosomal vector.

[0105] As used herein, "transforming DNA," "transforming sequence," and "DNA construct" refer to DNA used to introduce a sequence into a host cell or organism. Transforming DNA is DNA used to introduce a sequence into a host cell or organism. This DNA can be generated in vitro by PCR or any other suitable technique. In some embodiments, the transforming DNA includes the incoming sequence, while in other embodiments, the transforming DNA further includes the incoming sequence flanked by homology boxes. In yet other embodiments, the transforming DNA includes other non-homologous sequences (i.e., stuffer sequences or flanks) added to the ends. The ends can be closed such that the transforming DNA forms a closed circle, such as, for example, by insertion into a vector.

[0106] As used herein, "gene disruption" or "gene disruption" are used interchangeably and refer broadly to any genetic modification that substantially prevents a host cell from producing a functional gene product (e.g., protein). Thus, as used herein, gene disruption includes, but is not limited to, frameshift mutations, premature stop codons (i.e., functional protein is not produced), substitutions that eliminate or reduce the activity of protein internal deletions (functional protein is not produced), insertions that disrupt coding sequences, mutations that remove the operable link between the native promoter and the open reading frame required for transcription, and the like.

[0107] As used herein, "incoming sequence" refers to a DNA sequence that is introduced into a bacterial cell chromosome. In some embodiments, the incoming sequence is part of a DNA construct. In other embodiments, the incoming sequence encodes one or more proteins of interest. In some embodiments, the incoming sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it may be a homologous sequence or a heterologous sequence). In some embodiments, the incoming sequence encodes one or more proteins of interest, genes and / or mutant or modified genes. In alternative embodiments, the incoming sequence encodes a functional wild-type gene or operon, a functional mutant gene or operon, or a non-functional gene or operon. In some embodiments, a non-functional sequence may be inserted into a gene to disrupt the function of the gene. In another embodiment, the incoming sequence comprises a selection marker. In yet another embodiment, the incoming sequence comprises two homology boxes.

[0108] As used herein, a "homology box" refers to a nucleic acid sequence that is homologous to a sequence in a bacterial cell chromosome. More specifically, a homology box is an upstream or downstream region that has about 80-100% sequence identity, about 90-100% sequence identity, or about 95-100% sequence identity with the immediately flanking coding region of a gene or part of a gene to be deleted, disrupted, inactivated, downregulated, etc., according to the present invention. These sequences direct where in the bacterial cell chromosome the DNA construct is integrated and direct what part of the chromosome is replaced by the incoming sequence. Without intending to limit the disclosure, a homology box can comprise from about 1 base pair (bp) to 200 kilobases (kb). Preferably, the homology box comprises from about 1 bp to 10.0 kb; 1 bp to 5.0 kb; 1 bp to 2.5 kb; 1 bp to 1.0 kb, and 0.25 kb to 2.5 kb. The homology box may also comprise about 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb and 0.1 kb. In some embodiments, the 5' and 3' ends of the selectable marker are flanked by homology boxes, which comprise nucleic acid sequences that immediately flank the coding region of the gene.

[0109] As used herein, host cell "genome," bacterial (host) cell "genome," or Bacillus sp. (host) cell "genome" includes chromosomal genes and extrachromosomal genes.

[0110] As used herein, the terms "plasmid," "vector," and "cassette" refer to extrachromosomal elements that often carry genes that are not part of the central metabolism of the cell and are usually in the form of circular double-stranded DNA molecules. Such elements can be linear or circular, single-stranded or double-stranded, autonomously replicating sequences of DNA or RNA, genome-integrating sequences, phages, or nucleotide sequences from any source in which multiple nucleotide sequences have been joined or recombined into a unique structure that can introduce into a cell a promoter fragment and DNA sequence for a selected gene product along with appropriate 3' untranslated sequences.

[0111] As used herein, the term "plasmid" refers to a circular double-stranded (ds) DNA construct that is used as a cloning vector and forms an extrachromosomal, self-replicating genetic element in many bacteria and some eukaryotes. In some embodiments, the plasmid becomes integrated into the genome of the host cell, and in some embodiments, the plasmid is present in the parent cell and lost in the daughter cells.

[0112] As used herein, a "transformation cassette" refers to a specific vector that contains a gene (or its ORF) and has elements in addition to a foreign gene that facilitate transformation of a particular host cell.

[0113] As used herein, the term "vector" refers to any nucleic acid that can replicate (multiply) in a cell and can carry new genes or DNA segments into the cell. Thus, the term refers to a nucleic acid construct designed for transport between various host cells. Vectors include viruses, bacteriophages, proviruses, plasmids, phagemids, transposons, which are "episomal" (i.e., they can replicate autonomously or integrate into the chromosomes of the host organism), as well as artificial chromosomes, such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), and PLACs (plant artificial chromosomes).

[0114] "Expression vector" refers to a vector capable of incorporating and expressing heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and known to those skilled in the art. The selection of an appropriate expression vector is within the knowledge of one of ordinary skill in the art.

[0115] As used herein, the terms "expression cassette" and "expression vector" refer to a nucleic acid construct (i.e., they are vectors or vector elements as described above) that is recombinantly or synthetically produced with a set of specific nucleic acid elements that allow transcription of a specific nucleic acid in a target cell. The recombinant expression cassette can be incorporated into a plasmid, a chromosome, mitochondrial DNA, plastid DNA, a virus, or a nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter. In some embodiments, the DNA construct also includes a set of specific nucleic acid elements that allow transcription of a specific nucleic acid in a target cell. In certain embodiments, the DNA construct of the present disclosure includes a selectable marker and an inactivated chromosomal segment or gene segment or DNA segment as defined herein.

[0116] As used herein, a "targeting vector" is a vector that contains a polynucleotide sequence that is homologous to a region in a host cell chromosome into which the targeting vector is transformed and can drive homologous recombination at that region. For example, a targeting vector is used to introduce a mutation into a host cell chromosome by homologous recombination. In some embodiments, the targeting vector contains other non-homologous sequences, which are added, for example, to the ends (i.e., stuffer sequences or flanking sequences). The ends can be closed, for example, by insertion into a vector, to close the targeting vector into a circle. For example, in certain embodiments, parent B. licheniformis (host) cells are modified (e.g., transformed) by introducing one or more "targeting vectors" into them.

[0117] As used herein, the term "protein of interest" or "POI" refers to a polypeptide of interest desired to be expressed in an engineered (recombinant) Gram-positive host cell, with the POI preferably being expressed at an increased level (i.e., compared to "unmodified" (parent or control) cells). Thus, as used herein, a POI can be an enzyme, a substrate binding protein, a surfactant protein, a structural protein, a receptor protein, and the like. In certain embodiments, the engineered cells of the present disclosure produce an increased amount of a heterologous protein of interest compared to a control cell. In certain embodiments, the increased amount of the protein of interest produced by the engineered cells of the present disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or more than a 5.0% increase compared to a control cell.

[0118] Similarly, as defined herein, "gene of interest" or "GOI" refers to a nucleic acid sequence (e.g., polynucleotide, gene or ORF) that encodes a POI. A "gene of interest" that encodes a "protein of interest" can be a naturally occurring gene, a mutated gene or a synthetic gene.

[0119] As used herein, the terms "polypeptide" and "protein" are used interchangeably and refer to polymers of any length that contain amino acid residues linked by peptide bonds. Conventional one-letter or three-letter codes for amino acid residues are used herein. Polypeptides can be linear or branched, can contain modified amino acids, and can be interrupted by non-amino acids. The term polypeptide also encompasses amino acid polymers that are modified naturally or by intervention, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling moiety. Also included within this definition are, for example, polypeptides that contain one or more analogs of an amino acid, including, for example, unnatural amino acids, and other modifications known in the art.

[0120] In certain embodiments, the genes of the present disclosure are encoding enzymes (e.g., acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, In one embodiment, the present invention encodes a commercially relevant protein of industrial interest such as enzymes, enzymes, isomerases, laccases, lipases, lyases, mannosidases, oxidases, oxidoreductases, pectate lyases, pectin acetyl esterases, pectin depolymerases, pectin methyl esterases, pectinolytic enzymes, perhydrolases, polyol oxidases, peroxidases, phenol oxidases, phytases, polygalacturonases, proteases, peptidases, rhamno-galacturonase, ribonucleases, transferases, transport proteins, transglutaminase, xylanase, hexose oxidase, and combinations thereof.

[0121] As used herein, a "variant" polypeptide generally refers to a polypeptide derived from a parent (or reference) polypeptide by one or more amino acid substitutions, additions, or deletions by recombinant DNA techniques. A variant polypeptide may differ from a parent polypeptide by a small number of amino acid residues and may be defined by the level of primary amino acid sequence homology / identity with the parent (reference) polypeptide.

[0122] In one or more particular embodiments, the variant polypeptide has at least about 40% to about 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to a parent (reference) polypeptide sequence.

[0123] As used herein, a "variant" polynucleotide refers to a polynucleotide that has a particular degree of sequence homology / identity with a parent (or reference) polynucleotide or hybridizes to a parent polynucleotide (or its complement) under stringent hybridization conditions. In one or more particular embodiments, the variant polynucleotides of the present disclosure comprise at least about 40% to at least about 100% nucleotide sequence identity with the parent (reference) polynucleotide sequence. In certain other embodiments, the variant polynucleotides comprise at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% nucleotide sequence identity with the parent (reference) polynucleotide sequence.

[0124] As used herein, the term "mutation" refers to any change or alteration in a nucleic acid sequence. There are several types of mutations, including point mutations, deletion mutations, silent mutations, frameshift mutations, splicing mutations, etc. Mutations can be made specifically (e.g., by site-directed mutagenesis) or randomly (e.g., by chemical agents, repair minus passaging through bacterial strains).

[0125] As used herein, in reference to a polypeptide or sequence thereof, the term "substitution" refers to the replacement (ie, substitution) of one amino acid with another.

[0126] As used herein, the term "homology" refers to homologous polynucleotides or homologous polypeptides. When two or more polynucleotides or two or more polypeptides are homologous, this means that the homologous polynucleotides or polypeptides have a "degree of identity" of at least 60%, more preferably at least 70%, even more preferably at least 85%, even more preferably at least 90%, more preferably at least 95% and most preferably at least 98%. Whether two polynucleotide or polypeptide sequences have a sufficiently high degree of identity to be homologous as defined herein can be conveniently determined by aligning the two sequences using a computer program known in the art, for example, "GAP" provided in the GCG program package (Program Manual for the Wisconsin Package, Version 8, August 1994, Genetics Computer Group, 575 Science Drive, Madison, Wisconsin, USA 53711) (Needleman and Wunsch, (1970). For DNA sequence comparison, GAP is used with the following settings: GAP creation penalty 5.0 and GAP extension penalty 0.3.

[0127] As used herein, the term "percent identity" refers to the level of nucleic acid or amino acid sequence identity between nucleic acid sequences encoding a polypeptide or between the amino acid sequences of a polypeptide when aligned using a sequence alignment program.

[0128] As used herein, the term "specific productivity" refers to the total amount of protein produced per cell per unit time over a given period of time.

[0129] As used herein, the terms "purified," "isolated," or "enriched" mean that a biomolecule (e.g., a polypeptide or polynucleotide) is altered from its native state by separation from some or all of the naturally occurring components with which it is naturally associated. Such isolation or purification can be performed by separation techniques known in the art, such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulfate precipitation or other protein salting out, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis, or separation by gradient to remove unwanted whole cells, cell debris, impurities, extraneous proteins, or enzymes in the final composition. Purified or isolated biomolecule compositions can then be supplemented with components that confer additional benefits, such as activators, anti-inhibitors, desirable ions, pH adjusting compounds, or other enzymes or chemicals.

[0130] II. Mutant Pro-Sequences for Enhanced Protein Production As outlined herein and described in the Examples below, the applicants designed and constructed a site evaluation library (SEL) to test / screen specific genetic modifications (mutations) for enhancing recombinant protein productivity. Specifically, as described in the Examples, recombinant proteins (enzymes) were used as reporters to monitor protein expression in the recombinant Bacillus strains of the present disclosure. Certain pro-region sequences suitable for use in the expression / production of heterologous (recombinant) mature proteins have been described. For example, WO 2008 / 112258 and WO 2010 / 123754 describe pro-sequences of B. clausii alkaline precursor protease and B. lentus subtilisin precursor protease, where certain mutant pro-sequences resulted in enhanced productivity of the mature proteases. WO 1993 / 20214 describes specific AprE signal (pre) and pro region sequences for Gram-negative lipase expression in B. subtilis.

[0131] More specifically, as described in the Examples below, genetic engineering (SEL) was performed on the pro region of an expression construct encoding an exemplary (subtilisin) reporter protein. Specifically, as generally described in the Examples (see Figures 1-3), SEL was performed on the pro region sequence derived from the wild-type (reference) pro sequence (SEQ ID NO: 15) to screen for mutations to enhance productivity of the mature subtilisin reporter protein.

[0132] As described in Example 1, a site evaluation library (SEL) was generated as a 4.4 kb fragment using the amino acid sequence encoding variant pro-region sequence A (30H, SEQ ID NO: 9, FIG. 1) as a template, and competent Bacillus cells were transformed with linear DNA of the expression cassette. As described in Example 1, sequence analysis was performed to determine unique pro-region sequence A variants and / or pro-region sequence B variants from the SEL, and each of the 84 amino acid positions of variant pro-region sequence A (SEQ ID NO: 9) was changed (substituted) to the other 19 natural amino acid residues. Specifically, Table 1 (Example 1) shows the results of reporter protein productivity (performance index (PI) value) after 72 hours compared to a reference construct in which the reporter protein was expressed with the WT pro-sequence (SEQ ID NO: 15). For example, as shown in Table 1, the PI of pro-region sequence A (E30H, SEQ ID NO: 9) was increased by about 50% compared to the PI of the WT pro-sequence (30E; SEQ ID NO: 15). Similarly, as shown in Table 1, mutations at any one of positions 1, 2, 3, 4, 6, 14, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 83, or 84 in combination with amino acid position 30 demonstrated enhanced reporter protein productivity at 72 hours compared to the reference pro region sequence (SEQ ID NO: 15). In addition, approximately 80% of the selected mutant combinations carried an extra positive charge (see Table 1, Δ charge).

[0133] Genetic modifications (SEL) were performed on the proregion of an expression construct encoding an exemplary (subtilisin) reporter protein (BG46_variant 2), as described in more detail in Example 3. More specifically, an amino acid sequence encoding variant proregion sequence B (30G, SEQ ID NO:11, FIG. 1) was used as a template to generate linear DNA expression cassettes containing one or more proregions of the present disclosure. Specifically, the 5' sequence of the 78 amino acid wild-type AprE propeptide sequence (SEQ ID NO:12) was used to replace the first three amino acid residues (positions 1-3, "AGK", SEQ ID NO:12) with the first amino acid residue being alanine (A) to obtain variant proregion sequence C (SEQ ID NO:14, GK insertion, FIG. 2). For example, as shown in Table 1 (Example 1), a mutant of the proregion variant A sequence (SEQ ID NO: 9, Table 1, mutant "E30G") showed higher expression of the reporter protein (PI 2.3) after 72 hours of growth compared to the wild-type proregion (Table 1, PI 1.0) and the proregion variant A sequence (SEQ ID NO: 11, Table 1, mutant "E30H", PI 1.5).

[0134] Example 4 of the present disclosure further describes the preparation of a combinatorial library based on a reference pro-region sequence (SEQ ID NO: 15) into which the amino acids glycine (G) and lysine (K) were inserted ("GK"). More specifically, as described in Example 3, amino acid residues GK were inserted at the N-terminus of the wild-type pro-region sequence (SEQ ID NO: 15) to obtain the amino acid sequence of pro-region variant C set forth in SEQ ID NO: 14, which includes the 86 amino acid positions shown in the alignment in FIG. 2. For example, Table 2 (Example 4) shows combinations of pro-region variants that showed increased productivity (PI) of the BG46_variant 2 reporter protein, with changes in charge (Δ charge) generally ranging from +1 to +4 (Table 2). In certain embodiments, Applicants have found that such positive charges are suitable as combinations in the loop regions of the pro-region sequence (e.g., residue positions 66-73 of Reference SEQ ID NO: 15, etc.).

[0135] As shown in Table 2, the PI values ​​of the SEL proregion variant C sequence (derived from SEQ ID NO: 14) are compared to the PI values ​​of the wild-type (WT) proregion sequence (SEQ ID NO: 15), where the numbering of the amino acid positions (Table 2, column 1) is relative to the WT proregion sequence (i.e., without the "AGK" insertion). Thus, as described in Example 4, the proregion variant C sequence described in Table 2 showed higher expression of the reporter protein (PI values ​​of about 1.9 to 2.3) after 72 hours of growth compared to the wild-type proregion (Table 2, PI 1.0) and the proregion variant B sequence (mutant, "E30G", PI 1.7).

[0136] As with the B. licheniformis strains, BG46_variant 2 subtilisin was used as a reporter protein to monitor expression, as with the B. licheniformis strains, generally as described in Example 5. More specifically, DNA encoding the variant pro-region sequences of the present disclosure was operably linked to a DNA sequence encoding the mature BG46_variant 2 protein, and strains were constructed as generally described in Example 5. For example, the PI values ​​of B. licheniformis strains containing variant pro-region sequences compared to control pro-region sequence B, which contains the E30G mutation (SEQ ID NO: 11), are shown in Table 3, where the variant pro-region sequence showed higher expression of the reporter protein compared to the control pro-region sequence (E30G) after 72 hours of growth.

[0137] Additional N-terminal modifications of variant pro-region sequence C (SEQ ID NO: 14) were performed as described in more detail in Example 6. For example, variant pro-region sequence C (having 86 amino acid positions) was further modified to introduce (insert) an additional amino acid, such as alanine (A) or serine (S) after lysine (K) at position 3. Specifically, variant pro-region sequences D (SEQ ID NO: 29, GKS insertion), E (SEQ ID NO: 30, GKA insertion) and F (GKAA insertion; SEQ ID NO: 31) were constructed as described in Example 6 (and shown in Figures 1-3), and the performance index (PI) values ​​(Table 4) of pro-region variant sequences D, E and F are compared to variant pro-region sequence C (SEQ ID NO: 14, "GK" insertion). As shown in Table 4, the PI index of three N-terminal pro-region mutants (GKA, GKS, GKAA) showed higher expression of reporter protein BG46_variant 2 subtilisin (SEQ ID NO: 8) compared to pro-region sequence C (GK, SEQ ID NO: 14) after 72 hours of growth.

[0138] Similarly, specific C-terminal modifications of the proregion sequences were performed as described in Example 7. Specifically, as shown in Figures 1-3, variant proregion sequence B (SEQ ID NO: 11) was further engineered to replace the leucine (L) residue at position 68 with lysine (K) and the isoleucine (I) at position 72 with valine (V) to generate variant proregion sequence G (SEQ ID NO: 32) or engineered to replace the leucine (L) residue at position 68 with lysine (K) and the glutamic acid (E) at position 72 (80) with isoleucine (I) to generate variant proregion sequence F (SEQ ID NO: 33). As shown in Table 5, the PI values ​​of variant proregion sequences G (SEQ ID NO: 32) and H (SEQ ID NO: 33) were increased compared to the reference (control) variant proregion B sequence (SEQ ID NO: 11).

[0139] Accordingly, certain embodiments of the present disclosure provide, inter alia, recombinant polynucleotides encoding the novel pro region sequences, expression cassettes comprising DNA sequences encoding the novel pro region sequences operably linked to a DNA sequence encoding a mature protein of interest, engineered Gram-positive bacterial cells / strains expressing polynucleotides encoding precursor proteins comprising the novel pro region sequences, etc. In certain aspects, the novel variant pro region sequences of the present disclosure comprise an amino acid sequence derived from the wild-type B. lentus pro region sequence of SEQ ID NO:15.

[0140] As used herein, the phrase "variant pro region" refers to a polypeptide sequence derived from a reference pro region sequence, typically by recombinant DNA techniques, by one or more amino acid substitutions, additions, or deletions. A variant pro region (amino acid) sequence may differ from the reference (parent) pro region sequence by only a small number of amino acid residues and may be defined by the level of primary amino acid sequence homology / identity with the reference (parent) pro region sequence.

[0141] As stated above, the term "identical" in relation to two polynucleotide or polypeptide sequences refers to the nucleotides or amino acids in the two sequences that are the same when aligned for maximum correspondence as measured using sequence comparison or sequence analysis algorithms described below and known in the art. The phrase "percent (%) identity" (abbreviated "PID") refers to polynucleotide (nucleic acid) or polypeptide (amino acid) sequence identity. Percent identity may be determined using standard techniques known in the art. In certain aspects, the percent of amino acid identity shared by sequences of interest may be determined by aligning the sequences and directly comparing the sequence information using an alignment program / algorithm such as, for example, BLAST, MUSCLE, or CLUSTAL. For example, the BLAST algorithm is described in Altschul et al. (1990) and Karlin et al. (1993). Specifically, the percent (%) amino acid sequence identity value is determined by dividing the number of matching identical residues by the total number of residues in the "reference" sequence, including any gaps created by the program for optimal / maximal alignment. The BLAST algorithm refers to a "reference" sequence as the "query" sequence.

[0142] As used herein, "homologous proregion" refers to a proregion that has clear similarities in primary, secondary and / or tertiary structure. Protein homology can refer to similarity in linear amino acid sequence when proteins are aligned. Homology can be determined by amino acid sequence alignment, for example, using programs such as BLAST, MUSCLE or CLUSTAL. Homology searches of protein sequences can be performed using BLASTP and PSI-BLAST of NCBI BLAST with a threshold of 0.001 (E-value cutoff) (see, e.g., Altschul et al., 1997).

[0143] The BLAST program uses several search parameters, most of which are set to default values. The NCBI BLAST algorithm finds the most related sequences in terms of biological similarity, but is not recommended for query sequences of less than 20 residues (Altschul et al., 1997 and Schaffer et al., 2001). Typical default BLAST parameters for nucleic acid sequence searches include: adjacent word threshold=11; E-value cutoff=10; scoring matrix=NUC.3.1 (match=1, mismatch=-3); gap opening=5; and gap extension=2. Typical default BLAST parameters for amino acid sequence searches include: word size=3; E-value cutoff=10; scoring matrix=BLOSUM62; gap opening=11; and gap extension=1. This information can be used to group protein sequences and / or generate phylogenetic trees from them. Amino acid sequences can be entered into programs such as the Vector NTI Advance suite, and guide trees can be created using the Neighbor Joining (NJ) method (Saitou and Nei, 1987). Phylogenetic tree construction can be calculated using Kimura's sequence distance correction and ignoring positions containing gaps. Programs such as AlignX can present calculated distance values ​​in parentheses following the molecule names shown on the phylogenetic tree.

[0144] Understanding the homology between molecules can reveal information about the evolutionary history of the molecule as well as its function; if a newly sequenced protein is homologous to a previously characterized protein, it is a strong indication of the biochemical function of the new protein. Two molecules are said to be homologous if they are derived from a common ancestor. Homologous molecules or homologs can be divided into two classes: paralogs and orthologs. Paralogs are homologs that exist within one species. Paralogs often differ in their detailed biochemical function. Orthologs are homologs that exist in different species and have very similar or identical functions. Protein superfamilies are the largest groups of proteins for which a common ancestor can be inferred. Usually, this common ancestor is based on sequence alignment and mechanistic similarity.

[0145] The CLUSTAL W algorithm is another example of a sequence alignment algorithm (Thompson et al., 1994). Default parameters for the CLUSTAL W algorithm include: Gap opening penalty (=10.0; Gap extension penalty = 0.05; Protein weight matrix = BLOSUM series; DNA weight matrix = IUB; Delay divergence sequence (%) = 40; Gap separation distance = 8; DNA transition weight = 0.50; List of hydrophilic residues = GPSNDQEKR; Use negative matrix = OFF; Toggle residue specific penalty = ON; Toggle hydrophilicity penalty = ON; and Toggle end gap separation penalty = OFF. The CLUSTAL algorithm includes deletions that occur at either end. For example, a variant with a 5 amino acid deletion at either end of a 500 amino acid polypeptide (or within the polypeptide) has a percent sequence identity of 99% (495 / 500 identical residues x 100) to the "reference" polypeptide. Such a variant would be encompassed as having "at least 99% sequence identity" to the polypeptide.

[0146] In certain embodiments, the variant proregion sequence has at least about 40% to about 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to a reference (parent) proregion sequence of the present disclosure. In certain embodiments, the proregion sequence is derived from an amino acid sequence that includes homology to SEQ ID NO:15, SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:14, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, and / or SEQ ID NO:33. Thus, in certain aspects, the amino acid modifications of one or more of the proregion variants described herein are numbered by reference to the proregion amino acid sequence of SEQ ID NO:15, SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:14, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, and / or SEQ ID NO:33.

[0147] In certain aspects, the novel (variant) pro region sequence is derived from a parent (reference) pro region sequence that has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence of SEQ ID NO: 15. In certain embodiments, the variant pro region sequence comprises a mutation at one or more amino acid residue positions selected from the group consisting of amino acid (residue) position 30 and positions 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83 and 84 according to the amino acid (position) numbering of SEQ ID NO: 15.

[0148] In other embodiments, a variant pro region sequence of the disclosure comprises an amino acid insertion of a glycine (G) at position 2 and a lysine (K) at position 3, and an amino acid substitution at one or more positions selected from 1, 32, 38, 46, 66, 67, 70, and 73, where the amino acid positions are numbered according to SEQ ID NO: 14. In certain embodiments, the novel (variant) pro region sequence is derived from a reference pro region sequence that comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence of SEQ ID NO: 14.

[0149] In other embodiments, a variant pro region sequence of the disclosure comprises an amino acid insertion of a glycine (G) at position 2, a lysine (K) at position 3, and a serine (S) at position 4, where the amino acid positions are numbered according to SEQ ID NO: 29. In certain embodiments, the novel (variant) pro region sequence is derived from a reference pro region sequence that comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence of SEQ ID NO: 29.

[0150] In other embodiments, a variant pro region sequence of the disclosure comprises an amino acid insertion of a glycine (G) at position 2, a lysine (K) at position 3, and an alanine (A) at position 4, where the amino acid positions are numbered according to SEQ ID NO: 30. In certain embodiments, the variant pro region sequence is derived from a reference pro region sequence that comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence of SEQ ID NO:30.

[0151] In other embodiments, a variant pro region sequence of the disclosure comprises an amino acid insertion of a glycine (G) at position 2, a lysine (K) at position 3, an alanine (A) at position 4, and an alanine (A) at position 5, where the amino acid positions are numbered according to SEQ ID NO: 31. In certain embodiments, the variant pro region sequence is derived from a reference pro region sequence that comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence of SEQ ID NO:31.

[0152] In certain other embodiments, the variant proregion sequence is derived from a parent (reference) proregion sequence that has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to the amino acid sequence of SEQ ID NO:11 (G30; proregion B sequence).

[0153] In certain embodiments, the variant pro region sequence comprises at least one amino acid substitution at a position selected from 68, 72, and 80, wherein the amino acid positions are numbered according to SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:32, or SEQ ID NO:33.

[0154] In certain other embodiments, the variant proregion sequence comprises a mutation at one or more amino acid residue positions selected from the group consisting of amino acid (residue) position 30 and positions 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83, and 84 according to the amino acid (position) numbering of SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:32, or SEQ ID NO:33.

[0155] In certain embodiments, variant pro region sequences are designed, engineered, constructed, etc., such that the net charge of the pro region sequence is positive (e.g., a net charge of +1) relative to the reference (parent) pro region sequence, such as the reference (parent) sequences of SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, and SEQ ID NO:33. In certain embodiments, variant pro region sequences of the disclosure are designed, engineered, constructed, etc., such that the net charge of the pro region sequence is at least positive one (+1) to about positive four (+4). In one or more other embodiments, variant pro region sequences of the disclosure comprise one or more amino acid modifications set forth in Table 1, Table 2, Table 3, Table 4, and / or Table 5.

[0156] III. Recombinant Polynucleotides and Molecular Biology As broadly described above, certain embodiments of the present disclosure relate to novel variant pro-region sequences. In related aspects, the present disclosure provides recombinant polynucleotides comprising one or more variant pro-region nucleic acid (DNA) sequences. Thus, certain embodiments relate to recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant Gram-positive bacterial cells / strains expressing a protein of interest, etc. In certain aspects, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells (strains) to enhance production of a protein of interest. In certain aspects, the polynucleotide constructs of the present disclosure are referred to as expression cassettes, which comprise at least an upstream (5') pro-region DNA sequence linked in a 5' to 3' direction and in operable combination to a downstream (3') gene CDS encoding a mature protein of interest (POI).

[0157] For example, one or more of the nucleic acid sequences described herein can be produced by using any suitable synthesis, manipulation and / or isolation technique or combination thereof. For example, one or more of the polynucleotides described herein can be produced using standard nucleic acid synthesis techniques, such as solid-phase synthesis techniques well known to those of skill in the art. In such techniques, fragments, typically consisting of up to 50 or more nucleotide bases, are synthesized and then linked (e.g., by enzymatic or chemical ligation methods) to form virtually any desired contiguous nucleic acid sequence. The synthesis of one or more of the polynucleotides described herein can also be facilitated by any suitable method known in the art, including, but not limited to, chemical synthesis using the classical phosphoramidite method (e.g., Beaucage and Caruthers, 1981) or the method described by Matthes et al. (1984), as commonly practiced in automated synthesis methods. One or more of the polynucleotides described herein can also be produced by using an automated DNA synthesizer. Customized nucleic acids can be ordered from a variety of commercial sources (e.g., ATUM (DNA 2.0), Newark, CA, USA; Life Tech (GeneArt), Carlsbad, CA, USA; GenScript, Ontario, Canada; Base Clear BV, Leiden, Netherlands; Integrated DNA Technologies, Skokie, IL, USA; Ginkgo Bioworks (Gen9), Boston, MA, USA; and Twist Bioscience, San Francisco, CA, USA). Other techniques for synthesizing nucleic acids and the associated principles are described and known in the art.

[0158] Recombinant DNA techniques useful for modifying nucleic acids are well known in the art, including, for example, restriction endonuclease digestion, ligation, reverse transcription and cDNA production and polymerase chain reaction (e.g., PCR). One or more polynucleotides described herein can also be obtained by screening a cDNA library with one or more oligonucleotide probes that hybridize to or PCR amplify a polynucleotide encoding one or more variants described herein. Methods for screening and isolating cDNA clones and PCR amplification methods are well known to those of skill in the art and are described in standard references known to those of skill in the art. One or more polynucleotides described herein can be obtained by modifying a native (e.g., encoding one or more variant pro-region sequences described herein) polynucleotide backbone, for example, by known mutagenesis procedures (e.g., site-directed mutagenesis, site-saturation mutagenesis and in vitro recombination). A variety of methods suitable for generating modified polynucleotides as described herein that encode one or more variants as described herein are known in the art, including, but not limited to, site-saturation mutagenesis, systematic mutagenesis, insertional mutagenesis, deletion mutagenesis, random mutagenesis, site-directed mutagenesis and directed evolution, as well as various other recombinant methods.

[0159] As broadly described above and more particularly described in the Examples below, certain embodiments of the present disclosure relate to recombinant (modified) Gram-positive cells capable of producing increased amounts of a heterologous protein of interest. Accordingly, certain embodiments relate to methods of constructing such recombinant Gram-positive cells with increased protein production capacity. In certain embodiments, one or more expression cassettes encoding a protein of interest are introduced into the Gram-positive cells of the present disclosure. In exemplary embodiments, the cassettes are integrated into the genome of the cell. Accordingly, certain embodiments relate to nucleic acid molecules, polynucleotides (e.g., vectors, plasmids, expression cassettes), regulatory elements, etc., suitable for use in constructing recombinant (modified) Gram-positive host cells.

[0160] Thus, as presented in the Examples and generally described herein, recombinant cells of the present disclosure can be constructed by those skilled in the art using standard and routine recombinant DNA and molecular cloning techniques well known in the art. Methods of genetic modification include, but are not limited to, (a) introduction, replacement or removal of one or more nucleotides in a gene or introduction, replacement or removal of one or more nucleotides in a regulatory element required for transcription or translation of a gene, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-directed mutagenesis, and / or (g) random mutagenesis.

[0161] In certain embodiments, modified cells of the present disclosure can be constructed by reducing or eliminating expression of a gene using methods well known in the art, such as insertion, disruption, substitution, or deletion. The portion of a gene to be modified or inactivated can be, for example, a coding region or a regulatory element required to express the coding region.

[0162] An example of such a regulatory or control sequence may be a promoter sequence or a functional portion thereof (i.e., a portion sufficient to affect the expression of a nucleic acid sequence). Other control sequences for modification include, but are not limited to, leader sequences, propeptide sequences, signal sequences, transcription terminators, transcription activators, and the like.

[0163] In certain other embodiments, modified cells are constructed by gene deletion, which eliminates or reduces the expression of a gene. Gene deletion techniques allow for partial or complete removal of genes, thereby eliminating their expression or expressing a non-functional (or reduced activity) protein product. In such methods, deletion of a gene can be accomplished by homologous recombination, using a plasmid that has been constructed to contain adjacent 5' and 3' regions flanking the gene. The adjacent 5' and 3' regions can be introduced into the cell, for example, on a temperature-sensitive plasmid associated with a second selectable marker at a permissive temperature that allows the plasmid to establish in the cell. The cells are then shifted to a non-permissive temperature to select for cells that have the plasmid integrated into the chromosome at one of the homologous flanking regions. Selection for plasmid integration is influenced by selection for the second selectable marker. After integration, recombination events at the second homologous flanking region are stimulated by shifting the cells to the permissive temperature for several generations without selection. The cells are plated to obtain single colonies, which are then tested for the loss of both selectable markers. Thus, one of skill in the art can readily identify nucleotide regions within the coding sequence of a gene and / or the non-coding sequence of a gene that are suitable for complete or partial deletion.

[0164] In other embodiments, modified cells are constructed by introducing, substituting or removing one or more nucleotides in a gene or regulatory element required for its transcription or translation. For example, nucleotides can be inserted or removed to cause the introduction of a stop codon, the removal of a start codon, or a frame shift of an open reading frame. Such modifications can be achieved by site-directed mutagenesis or PCR-generated mutagenesis according to methods known in the art. Thus, in certain embodiments, genes of the present disclosure are inactivated by complete or partial deletion.

[0165] In another embodiment, the modified cell is constructed by a gene conversion process. For example, in gene conversion methods, a nucleic acid sequence corresponding to a gene is mutated in vitro to generate a defective nucleic acid sequence, which is then transformed into a parent cell to generate the defective gene. The defective nucleic acid sequence replaces the endogenous gene by homologous recombination. It may be desirable for the defective gene or gene fragment to also encode a marker that can be used to select transformants containing the defective gene. For example, the defective gene can be introduced on a non-replicating or temperature-sensitive plasmid associated with a selectable marker. Selection of plasmid integration is affected by selecting for that marker under conditions that do not allow replication of the plasmid. Selection of a second recombination event resulting in gene replacement is affected by examining the colony for the loss of the selectable marker and the acquisition of a mutated gene. Alternatively, the defective nucleic acid sequence can contain an insertion, substitution, or deletion of one or more nucleotides of the gene, as described below.

[0166] In other embodiments, modified cells are constructed by established antisense technology using a nucleotide sequence that is complementary to the nucleic acid sequence of a gene. More specifically, the expression of a gene by a Gram-positive cell can be reduced (downregulated) or eliminated by introducing a nucleotide sequence that is complementary to the nucleic acid sequence of this gene, which can be transcribed in the cell and hybridize to the mRNA produced in the cell. Thus, under conditions in which the complementary antisense nucleotide sequence can hybridize to the mRNA, the amount of protein translated is reduced or eliminated. Such antisense methods include, but are not limited to, RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), antisense oligonucleotides, and the like, all of which are well known to those skilled in the art.

[0167] In other embodiments, modified cells are produced / constructed by CRISPR-Cas9 editing. For example, genes encoding proteins of interest can be edited or destroyed (or deleted or downregulated) by nucleic acid-guided endonucleases that find their target DNA by binding to either a guide RNA (e.g., Cas9) and Cpf1 or a guide DNA (e.g., NgAgo) that recruits the endonucleases to the target sequence of the DNA, and the endonucleases can generate single- or double-stranded breaks in the DNA. This targeted DNA break can become a substrate for DNA repair and recombine with the editing template provided to disrupt or delete the gene. For example, a gene encoding a nucleic acid-guided endonuclease (for this purpose, Cas9 from S. pyogenes) or a codon-optimized gene encoding a Cas9 nuclease is operably linked to a promoter active in gram-positive cells and a terminator active in gram-positive cells, thereby generating a gram-positive cell Cas9 expression cassette. Similarly, one or more target sites unique to a gene of interest are easily identified by those skilled in the art. For example, to construct a DNA construct encoding a gRNA directed to a target site in a gene of interest, a variable targeting domain (VT) will include the nucleotides of the target site that are 5' to the (PAM) protospacer adjacent motif (TGG), which nucleotides are fused to DNA encoding the Cas9 endonuclease recognition domain (CER) for S. pyogenes Cas9. The combination of the DNA encoding the VT domain and the DNA encoding the CER domain thereby generates a DNA encoding the gRNA. Thus, a Gram-positive expression cassette for a gRNA is generated by operably linking the DNA encoding the gRNA to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells.

[0168] In certain embodiments, the DNA break induced by the endonuclease is repaired / replaced with the incoming sequence.For example, to precisely repair the DNA break generated by the above-mentioned Cas9 expression cassette and gRNA expression cassette, a nucleotide editing template is provided so that the DNA repair mechanism of the cell can utilize the editing template.For example, about 500 bp of the 5' side of the targeting gene can be fused to about 500 bp of the 3' side of the targeting gene to generate an editing template, and this template is used by the mechanism of the Gram-positive host to repair the DNA break generated by the RGEN.

[0169] The Cas9 expression cassette, the gRNA expression cassette and the editing template can be co-delivered into filamentous fungal cells using a number of different methods (e.g., protoplast fusion, electroporation, natural competence or induced competence). Transformed cells are screened by PCR amplification of the target locus by amplifying the locus with forward and reverse primers. These primers can amplify the wild-type locus or the modified locus that has been edited by RGEN. These fragments are then sequenced using sequencing primers to identify edited colonies.

[0170] In yet other embodiments, modified cells are constructed by random or specific mutagenesis using methods well known in the art, including, but not limited to, chemical mutagenesis and transposition. Genetic modification can be performed by subjecting parent cells to mutagenesis and screening for mutant cells in which expression of the gene is reduced or eliminated. Mutagenesis can be specific or random, for example, by using suitable physical or chemical mutagenizing agents, by using suitable oligonucleotides, or by subjecting the DNA sequence to PCR-generated mutagenesis. Furthermore, mutagenesis can be performed by using any combination of these mutagenesis methods.

[0171] Examples of physical or chemical mutagenic agents suitable for the purposes of the present invention include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl-N'-nitrosoguanidine (NTG), O-methylhydroxylamine, nitrous acid, ethyl methanesulfonate (EMS), sodium bisulfite, formic acid and nucleotide analogs. When using such agents, mutagenesis is generally performed by incubating the parent cells to be mutagenized under suitable conditions in the presence of the mutagenizing agent of choice and selecting mutant cells that show reduced or no expression of the gene.

[0172] WO 2003 / 083125 discloses methods for modifying Gram-positive (Bacillus) cells, such as the creation of Bacillus deletion strains and DNA constructs using PCR fusion to bypass E. coli. WO 2002 / 14490 discloses methods for modifying Bacillus cells, including (1) construction and transformation of an integrative plasmid (pComK), (2) random mutagenesis of coding, signal and propeptide sequences, (3) homologous recombination, (4) increasing transformation efficiency by adding non-homologous flanks to the transforming DNA, (5) optimizing double-crossover integration, (6) site-directed mutagenesis, and (7) markerless deletion.

[0173] Those skilled in the art are well aware of suitable methods for introducing polynucleotide sequences into bacterial cells (e.g., gram-negative cells, gram-positive cells). Indeed, methods such as transformation, including transformation and aggregation of protoplasts, transduction, and fusion of protoplasts, are known and suitable for use in the present disclosure. Transformation methods are particularly preferred for introducing the DNA constructs of the present disclosure into host cells.

[0174] In addition to commonly used methods, in some embodiments, the host cell is directly transformed (i.e., no intermediate cells are used to amplify or otherwise process the DNA construct prior to introduction into the host cell). Introduction of the DNA construct into the host cell also includes physical and chemical methods known in the art for introducing DNA into the host cell without insertion into a plasmid or vector. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, liposomes, and the like. In further embodiments, the DNA construct is co-transformed with a plasmid without being inserted into the plasmid. In further embodiments, the selection marker is deleted or substantially excised from the modified Bacillus strain by methods known in the art. In some embodiments, the splitting of the vector from the host chromosome leaves flanking regions in the chromosome while removing the unique chromosomal region.

[0175] Promoters and promoter sequence regions, their coding sequences (CDS), open reading frames (ORFs) and / or variant sequences for use in expressing genes in Gram-positive cells are generally known to those skilled in the art. The promoter sequences of the present disclosure are generally selected such that they function in Gram-positive cells. For example, promoters useful for driving gene expression in Bacillus cells include, but are not limited to, the B. subtilis alkaline protease (aprE) promoter, the B. subtilis α-amylase promoter (amyE), the B. licheniformis α-amylase promoter (amyL), the B. amyloliquefaciens α-amylase promoter, the native protease (nprE) promoter from B. subtilis, a mutant aprE promoter, or any other promoter from B. licheniformis or other related Bacilli. Methods for screening and generating promoter libraries with different activities (promoter strengths) in Bacillus cells are described in WO 2002 / 14490.

[0176] IV. Fermentation of Gram-Positive Cells to Produce Proteins As broadly described above, certain embodiments relate to compositions and methods for constructing and obtaining Gram-positive cells with an increased protein production phenotype. Accordingly, certain embodiments relate to methods for producing a protein of interest in a Gram-positive cell by fermenting the cells in an appropriate medium. Fermentation methods well known in the art may be applied to ferment the Gram-positive cells of the present disclosure.

[0177] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system, where the composition of the medium is set at the beginning of the fermentation and is not changed during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. In this way, fermentation can be carried out without adding any components to the system. Generally, batch fermentation is considered to be "batch" with respect to the addition of the carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes constantly until the point at which the fermentation is stopped. In a typical batch culture, cells can progress through a static lag phase to a high growth log phase and eventually to a stationary phase where the growth rate is reduced or stopped. If not treated, the cells in the stationary phase will eventually die. Generally, the cells in the log phase are responsible for the majority of the production of the product.

[0178] A suitable variation to the standard batch system is the "fed-batch fermentation" system. In this variation of the general batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit the metabolism of the cells and when a limited amount of substrate in the medium is desired. In fed-batch systems, measurement of the actual substrate concentration is difficult and is therefore estimated based on changes in measurable factors such as pH, dissolved oxygen, and partial pressure of waste gases such as CO2. Batch and fed-batch fermentation are common and known in the art.

[0179] Continuous fermentation is an open system in which a defined fermentation medium is added continuously to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density, with the cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, the limiting nutrient, such as the carbon or nitrogen source, is maintained at a constant ratio, and all other parameters can be adjusted. In other systems, a number of factors that affect growth can be continuously varied while the cell concentration, measured by the turbidity of the medium, is kept constant. Continuous systems attempt to maintain steady-state growth conditions. Thus, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes, as well as techniques for maximizing the rate of product formation, are well known in the art of industrial microbiology.

[0180] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure may be recovered from the culture medium by conventional procedures, such as separating the host cells from the medium by centrifugation or filtration, or, if necessary, disrupting the cells and removing the supernatant from cell debris and cell debris. Generally, after clarification, the protein components of the supernatant or filtrate are precipitated by salt, such as ammonium sulfate. The precipitated protein may then be solubilized and purified by various chromatographic methods, such as ion exchange chromatography, gel filtration, etc.

[0181] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system, where the composition of the medium is set at the beginning of the fermentation and is not changed during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. In this way, fermentation can be carried out without adding any components to the system. Generally, batch fermentation is considered to be "batch" with respect to the addition of the carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes constantly until the point at which the fermentation is stopped. In a typical batch culture, cells can progress through a static lag phase to a high growth log phase and eventually to a stationary phase where the growth rate is reduced or stopped. If not treated, the cells in the stationary phase will eventually die. Generally, the cells in the log phase are responsible for the majority of the production of the product.

[0182] A suitable variation to the standard batch system is the "fed-batch fermentation" system. In this variation of the general batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit the metabolism of the cells and when a limited amount of substrate in the medium is desired. In fed-batch systems, measurement of the actual substrate concentration is difficult and is therefore estimated based on changes in measurable factors such as pH, dissolved oxygen, and partial pressure of waste gases such as CO2. Batch and fed-batch fermentation are common and known in the art.

[0183] Continuous fermentation is an open system in which a defined fermentation medium is added continuously to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the culture at a constant high density, with the cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, the limiting nutrient, such as the carbon or nitrogen source, is maintained at a constant ratio, and all other parameters can be adjusted. In other systems, a number of factors that affect growth can be continuously varied while the cell concentration, measured by the turbidity of the medium, is kept constant. Continuous systems attempt to maintain steady-state growth conditions. Thus, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes and techniques for maximizing the rate of product formation are well known in the art of industrial microbiology.

[0184] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure may be recovered from the culture medium by conventional procedures, such as separating the host cells from the medium by centrifugation or filtration, or, if necessary, disrupting the cells and removing the supernatant from cell debris and cell debris. Generally, after clarification, the protein components of the supernatant or filtrate are precipitated by salt, such as ammonium sulfate. The precipitated protein may then be solubilized and purified by various chromatographic methods, such as ion exchange chromatography, gel filtration, etc.

[0185] V. Protein of Interest The protein of interest (POI) of the present disclosure may be any endogenous or heterologous protein, or a variant of such a POI. The protein may contain one or more disulfide bridges, or may be a protein whose functional form is monomeric or multimeric, i.e., the protein has a quaternary structure and is composed of multiple identical (homologous) or non-identical (heterologous) subunits, and the POI or its variant POI is preferably a protein with a property of interest.

[0186] For example, in certain embodiments, the modified Gram-positive cells of the present disclosure produce at least about 0.1% or more, at least about 0.5% or more, at least about 1% or more, at least about 5% or more, at least about 6% or more, at least about 7% or more, at least about 8% or more, at least about 9% or more, or at least about 10% or more of the POI compared to the unmodified (reference or control) cell.

[0187] In certain embodiments, the modified Gram-positive cells of the present disclosure exhibit increased specific productivity (Qp) of the POI compared to control cells. For example, detection of specific productivity (Qp) is a suitable method for assessing protein production. Specific productivity (Qp) is calculated according to the following formula: "Qp = gP / gDCW·hr" where "gP" is the grams of protein produced in the tank, "gDCW" is the grams of dry cell weight (DCW) in the tank, and "hr" is the fermentation time (hours) from the time of inoculation, which includes the production time and the growth time.

[0188] Thus, in certain other embodiments, the modified Gram-positive cells of the present disclosure comprise an increase in specific productivity (Qp) of at least about 0.1%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more compared to the unmodified (parent) cell.

[0189] In certain embodiments, the POI or variant POI is an acetyl esterase, an aminopeptidase, an amylase, an arabinase, an arabinofuranosidase, a carbonic anhydrase, a carboxypeptidase, a catalase, a cellulase, a chitinase, a chymosin, a cutinase, a deoxyribonuclease, an epimerase, an esterase, an α-galactosidase, a β-galactosidase, an α-glucanase, a glucan lyase, an endo-β-glucanase, a glucoamylase, a glucose oxidase, an α-glucosidase, a β-glucosidase, a glucuronidase, a glycosyl hydrolase, a hemicellulase, a hexose oxidase, a hydrolase, a hemicellul ... The enzyme is selected from the group consisting of: oxidase, invertase, isomerase, laccase, ligase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectin degrading enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamno-galacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.

[0190] Thus, in certain embodiments, the POI or variant POI thereof is an enzyme selected from Enzyme Code (EC) EC1, EC2, EC3, EC4, EC5 or EC6.

[0191] There are a variety of assays known to those of skill in the art for detecting and measuring the activity of intracellularly and extracellularly expressed proteins.

[0192] VI. ILLUSTRATIVE EMBODIMENTS Non-limiting embodiments of the compositions and methods disclosed herein are as follows.

[0193] 1. A variant pro region sequence comprising an amino acid substitution at position 30 and one or more positions selected from 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83, and 84, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO:15.

[0194] 2. The variant proregion of embodiment 1, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-84 of SEQ ID NO:15.

[0195] 3. A variant proregion sequence comprising an amino acid insertion of a glycine (G) at position 2 and a lysine (K) at position 3 and an amino acid substitution at one or more positions selected from 1, 32, 38, 46, 66, 67, 70 and 73, wherein the amino acid positions of the variant proregion are numbered according to SEQ ID NO:14.

[0196] 4. The variant proregion of embodiment 3, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-86 of SEQ ID NO:14.

[0197] 5. A variant pro-region sequence comprising the amino acid insertions of a glycine (G) at position 2, a lysine (K) at position 3, and an alanine (A) at position 4, wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO:30.

[0198] 6. The variant proregion of embodiment 5, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-87 of SEQ ID NO:30.

[0199] 7. A variant proregion sequence comprising the amino acid insertions of a glycine (G) at position 2, a lysine (K) at position 3, and a serine (S) at position 4, wherein the amino acid positions of the variant proregion are numbered according to SEQ ID NO:29.

[0200] 8. The variant proregion of embodiment 7, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-87 of SEQ ID NO:29.

[0201] 9. A variant pro region sequence comprising the amino acid insertions of a glycine (G) at position 2, a lysine (K) at position 3, an alanine (A) at position 4, and an alanine (A) at position 5, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO:31.

[0202] 10. The variant proregion of embodiment 9, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-88 of SEQ ID NO:31.

[0203] 11. A variant proregion sequence comprising a glutamic acid (E) to glycine (G) substitution at position 30 (E30G), a leucine (L) to lysine (K) substitution at position 68 (L68K), and an isoleucine (I) to valine (V) substitution at position 73 (I72V), wherein the amino acid positions of the variant proregion are numbered according to SEQ ID NO:32.

[0204] 12. The variant proregion of embodiment 11, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-84 of SEQ ID NO:15 or SEQ ID NO:32.

[0205] 13. A variant proregion sequence comprising a glutamic acid (E) to glycine (G) substitution at position 30 (E30G), a leucine (L) to lysine (K) substitution at position 68 (L68K), and a glutamic acid (E) to isoleucine (I) substitution at position 80 (E80I), wherein the amino acid positions of the variant proregion are numbered according to SEQ ID NO:33.

[0206] 14. The variant proregion of embodiment 13, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1-84 of SEQ ID NO:15 or SEQ ID NO:33.

[0207] 15. A nucleic acid encoding a variant pro region according to any one of embodiments 1 to 14.

[0208] 16. A polynucleotide comprising the variant pro region nucleic acid of embodiment 15.

[0209] 17. A polynucleotide comprising an upstream (5') nucleic acid encoding a variant pro-region according to any one of embodiments 1 to 14, operably linked to a downstream (3') nucleic acid sequence encoding a protein of interest (POI).

[0210] 18. A polynucleotide comprising an upstream (5') nucleic acid encoding a protein signal (secretory) sequence operably linked to a downstream (3') nucleic acid sequence encoding a variant proregion according to any one of embodiments 1 to 14, which is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).

[0211] 19. An expression cassette comprising an upstream (5') promoter region sequence operably linked to a downstream (3') polynucleotide according to any one of embodiments 16 to 18.

[0212] 20. A Gram-positive host cell comprising an introduced polynucleotide according to any one of embodiments 16 to 18 or an introduced cassette according to embodiment 19.

[0213] 21. The variant pro region of embodiment 1, wherein the one or more amino acid substitutions increase the net charge of the variant pro region sequence compared to the reference pro region of SEQ ID NO: 15.

[0214] 22. The variant pro region of embodiment 3, wherein the one or more amino acid substitutions increase the net charge of the variant pro region sequence compared to the reference pro region of SEQ ID NO: 14.

[0215] 23. The variant proregion of embodiment 21 or 22, wherein the increase in net charge is at least one (+1) to four (+4).

[0216] 24. The variant pro region of embodiment 1, comprising an amino acid modification as set forth in Table 1.

[0217] 25. The variant proregion of embodiment 1, wherein the amino acid substitution at position 30 is glycine (G), histidine (H) or asparagine (N).

[0218] 26. The variant pro region of embodiment 1, wherein the amino acid substitution at position 1 is a threonine (T).

[0219] 27. The variant pro region of embodiment 1, wherein the amino acid substitution at position 2 is alanine (A) or proline (P).

[0220] 28. The variant proregion of embodiment 1, wherein the amino acid substitution at position 3 is alanine (A), cystine (C), glycine (G), arginine (R), serine (S), or valine (V).

[0221] 29. The variant pro region of embodiment 1, wherein the amino acid substitution at position 4 is glutamine (Q).

[0222] 30. The variant pro region of embodiment 1, wherein the amino acid substitution at position 6 is arginine (R).

[0223] 31. The variant pro region of embodiment 1, wherein the amino acid substitution at position 14 is a threonine (T).

[0224] 32. The variant pro region of embodiment 1, wherein the amino acid substitution at position 19 is arginine (R).

[0225] 33. The variant pro region of embodiment 1, wherein the amino acid substitution at position 20 is a lysine (K).

[0226] 34. The variant proregion of embodiment 1, wherein the amino acid substitution at position 23 is lysine (K), asparagine (N), or glutamine (Q).

[0227] 35. The variant pro region of embodiment 1, wherein the amino acid substitution at position 36 is glycine (G).

[0228] 36. The variant pro region of embodiment 1, wherein the amino acid substitution at position 37 is a threonine (T).

[0229] 37. The variant pro region of embodiment 1, wherein the amino acid substitution at position 38 is histidine (H).

[0230] 38. The variant proregion of embodiment 1, wherein the amino acid substitution at position 39 is a lysine (K) or a glutamine (Q).

[0231] 39. The variant pro region of embodiment 1, wherein the amino acid substitution at position 44 is isoleucine (I).

[0232] 40. The variant proregion of embodiment 1, wherein the amino acid substitution at position 49 is threonine (T) or glutamine (Q).

[0233] 41. The variant pro region of embodiment 1, wherein the amino acid substitution at position 50 is arginine (R).

[0234] 42. The variant pro region of embodiment 1, wherein the amino acid substitution at position 64 is a lysine (K).

[0235] 43. The variant pro region of embodiment 1, wherein the amino acid substitution at position 65 is a lysine (K).

[0236] 44. The variant pro region of embodiment 1, wherein the amino acid substitution at position 67 is a serine (S).

[0237] 45. The variant pro region of embodiment 1, wherein the amino acid substitution at position 68 is a lysine (K) or a methionine (M).

[0238] 46. ​​The variant pro region of embodiment 1, wherein the amino acid substitution at position 79 is glutamine (Q).

[0239] 47. The variant pro region of embodiment 1, wherein the amino acid substitution at position 83 is isoleucine (I).

[0240] 48. The variant pro region of embodiment 1, wherein the amino acid substitution at position 84 is a leucine (L).

[0241] 49. The variant proregion of embodiment 3, comprising an amino acid modification as set forth in Table 2, Table 3, Table 4, SEQ ID NO: 14, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, and combinations thereof.

[0242] 50. The variant pro region of embodiment 3, wherein the amino acid substitution at position 1 is a threonine (T).

[0243] 51. The variant pro region of embodiment 3, wherein the amino acid substitution at position 32 is glycine (G).

[0244] The variant pro region of embodiment 3, wherein the amino acid substitution at position 52.38 is a glycine (G).

[0245] 53. The variant pro region of embodiment 3, wherein the amino acid substitution at position 46 is isoleucine (I).

[0246] 54. The variant pro region of embodiment 3, wherein the amino acid substitution at position 66 is a lysine (K).

[0247] 55. The variant pro region of embodiment 3, wherein the amino acid substitution at position 67 is a lysine (K).

[0248] The variant pro region of embodiment 3, wherein the amino acid substitution at position 56.70 is a lysine (K).

[0249] The variant pro region of embodiment 3, wherein the amino acid substitution at position 57.73 is arginine (R).

[0250] 58. A variant proregion comprising an amino acid modification as set forth in any one of Tables 1-5, Figure 1, Figure 2, Figure 3, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 14, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 33, and combinations thereof.

[0251] 59. The variant proregion of embodiment 58, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity to positions 1 to 84 of SEQ ID NO:15.

[0252] 60. A method for producing a heterologous protein of interest (POI) in a Gram-positive bacterial cell, comprising: (a) introducing into the Gram-positive cell an expression cassette comprising an upstream (5') promoter operably linked to a downstream (3') nucleic acid encoding a variant pro-region sequence comprising an amino acid substitution at one or more positions selected from position 30 and 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83 and 84, operably linked to a downstream (3') nucleic acid sequence encoding the POI, wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO: 15; and (b) growing / cultivating / fermenting the modified cell under conditions suitable for production of the POI.

[0253] 61. A method for producing a heterologous protein of interest (POI) in a Gram-positive bacterial cell, comprising: (a) introducing into the Gram-positive cell an expression cassette comprising an upstream (5') promoter sequence operably linked to a downstream (3') nucleic acid encoding a variant pro-region sequence comprising an amino acid insertion of glycine (G) at position 2, lysine (K) at position 3, and an amino acid substitution at one or more positions selected from 1, 32, 38, 46, 66, 67, 70, and 73, operably linked to a downstream (3') nucleic acid sequence encoding the POI, wherein the amino acid positions of the variant pro-region are numbered according to SEQ ID NO:14; and (b) growing / cultivating / fermenting the modified cell under conditions suitable for production of the POI.

[0254] 62. A method for producing a heterologous protein of interest (POI) in a Gram-positive bacterial cell, comprising: (a) introducing into the Gram-positive cell an expression cassette as described in embodiment 19; and (b) growing / cultivating / fermenting the modified cell under conditions suitable for the production of the POI.

[0255] 63. The method of any one of embodiments 60 to 62, wherein the cassette further comprises a nucleic acid encoding a preprotein signal (secretion) sequence operably linked and positioned between the promoter and the variant pro region sequence.

[0256] 64. The method of any one of embodiments 60 to 62, wherein the modified cells produce an increased amount of the POI compared to a control Gram-positive cell having an introduced expression cassette comprising the same upstream (5') promoter sequence operably linked to a downstream (3') nucleic acid encoding a proregion sequence of SEQ ID NO: 15, which is operably linked to a downstream nucleic acid sequence encoding the same POI.

[0257] 65. The method of embodiment 64, wherein the increased amount of POI is at least about 0.1% or more, at least about 0.5% or more, at least about 1% or more, at least about 5% or more, at least about 6% or more, at least about 7% or more, at least about 8% or more, at least about 9% or more, or at least about 10% or more of POI compared to control cells.

[0258] 66. The method of any one of embodiments 60 to 63, wherein the performance index (PI) value of the modified cell is at least about 1.1 compared to a control Gram-positive cell having an introduced expression cassette comprising the same upstream promoter sequence operably linked to a downstream nucleic acid encoding a proregion sequence of SEQ ID NO: 15, which is operably linked to a downstream nucleic acid sequence encoding the same POI as the modified cell, and the modified cell and the control cell are cultured under the same conditions.

[0259] 67. The method according to any one of embodiments 60 to 63, wherein the introduced cassette is integrated into the genome of the cell.

[0260] 68. The method according to any one of embodiments 60 to 63, comprising at least two introduced cassettes.

[0261] 69. The method according to any one of embodiments 60 to 63, wherein the POI is selected from the group consisting of enzymes, antibodies, receptors and regulatory proteins.

[0262] 70. POI is an enzyme that activates the enzymes of acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lysase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laminase, glyceryl oxidase ... 64. The method of any one of embodiments 60 to 63, wherein the enzyme is selected from the group consisting of carboxylases, ligases, lipases, lyases, mannosidases, oxidases, oxidoreductases, pectate lyases, pectin acetyl esterases, pectin depolymerases, pectin methyl esterases, pectin degrading enzymes, perhydrolases, polyol oxidases, peroxidases, phenol oxidases, phytases, polygalacturonases, proteases, peptidases, rhamnogalacturonase, ribonucleases, transferases, transport proteins, transglutaminase, xylanase, hexose oxidase, and combinations thereof.

[0263] 71. The method of embodiment 70, wherein the protease is a native or variant subtilisin.

[0264] 72. The Gram-positive bacterial cell is a Bacillus sp. cell, optionally a Bacillus sp. cells were cultured from B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophilus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, B. thuringiensis and Geobacillus stearothermophilus. The method of any one of embodiments 60 to 71, wherein the Bacillus subtilis is selected from the group consisting of Bacillus subtilis (L. stearothermophilus). EXAMPLES

[0265] Certain aspects of the present disclosure may be further understood in light of the following examples, which should not be construed as limiting. Modifications to materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989).

[0266] Example 1 Modification of the GG36 prosequence by site evaluation libraries In this example, a variant subtilisin protein (BG46_variant 1, SEQ ID NO:2) was used as a reporter to monitor protein expression as described herein. More specifically, a DNA fragment comprising the upstream (5') aprE gene flanking region was operably linked to a polynucleotide construct (e.g., expression cassette) comprising a variant of the DNA sequence of the B. subtilis rrnI-P2 promoter / 5'-aprE UTR region to obtain SEQ ID NO:1 operably linked to a DNA sequence encoding an AprE signal sequence (SEQ ID NO:4) operably linked to a DNA sequence encoding variant pro region sequence A (SEQ ID NO:9) operably linked to a DNA sequence encoding mature BG46_variant 1 (SEQ ID NO:2) operably linked to a B. amyloliquefaciens BPN' terminator DNA sequence (SEQ ID NO:6), which in turn was operably linked to a downstream (3') aprE gene flanking region sequence comprising a kanamycin (kan) gene expression cassette (SEQ ID NO:7). More specifically, the DNA fragments were assembled using standard molecular biology techniques and then used as templates to generate linear DNA expression cassettes containing one or more of the pro region modifications (mutations) described herein.

[0267] Site evaluation library (SEL) of proregion sequence A

[0268] The amino acid sequence encoding variant pro-region sequence A (30H, SEQ ID NO: 9) was used as a template to generate the site evaluation library (SEL), which was generated as a 4.4 kb fragment by Twist Bioscience HQ (South San Francisco). Linear DNA of the expression cassette was used to transform competent B. subtilis cells, and the transformation mixture was plated on LA plates containing 1.8 ppm kanamycin and incubated overnight at 37°C. Single colonies were picked and grown in Luria broth at 37°C under antibiotic selection.

[0269] Sequence analysis was performed to determine unique proregion sequence A variants that were cherry-picked into 96-well microtiter plates (MTPs). For example, the first amino acid position of variant proregion sequence A (SEQ ID NO:9) is alanine ("Ala" or "A"), which can be changed (substituted) to any of the other 19 naturally occurring amino acid residues, the second amino acid position of variant proregion sequence A (SEQ ID NO:9) is glutamic acid ("Glu" or "E"), which can be changed (substituted) to any of the other 19 naturally occurring amino acid residues, and so on until all 84 positions of proregion sequence A (SEQ ID NO:9) had been evaluated (see, e.g., FIG. 1).

[0270] For reporter protein expression experiments, transformed cells were grown in culture medium (enriched semi-defined medium based on MOP buffer) in 96-well MTPs in a shaking incubator at 32° C. for 3 days at 300 rpm and 80% humidity, then centrifuged and filtered. The clarified culture supernatants were used to assay reporter protease activity to determine the level of productivity, and samples were taken after 72 hours. The reporter protease activity assay is described in more detail in Example 2 below.

[0271] Positioning that impacts productivity

[0272] As described below, Table 1 shows the results of relative reporter productivity compared to a reference construct, where the reporter gene was expressed with the native GG36 pro-region sequence (SEQ ID NO: 15) and expression was measured by the activity assay described in Example 2 below. For example, as shown in Table 1, performance index (PI) values ​​were given for samples taken after 72 hours (PI was calculated as described in Example 2), and the amino acid (residue) positions that affected productivity in combination with position 30 are 1, 2, 3, 4, 6, 14, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 83, 84. The 5' amino acid change and the 3' amino acid modification affected productivity, and 80% of the selected combinations had an extra positive charge. [Table 1]

[0273] Example 2 Protease activity assay The protease activity of BG46_variant 1 subtilisin was determined by measuring the hydrolysis of a synthetic suc-AAPF-pNA peptide substrate. For the AAPF assay, the reagent solution used was 100 mM Tris at pH 8.6, 10 mM CalCl2, 0.005% Tween®-80 (Tris / Ca buffer) and 160 mM suc-AAPF-pNA in DMSO (suc-AAPF-pNA stock solution) (Sigma: S-7388). To prepare the working solution, 1 mL of suc-AAPF-pNA stock solution was added to 100 mL of Tris / Ca buffer and mixed. Enzyme samples were added to a microtiter plate (MTP) containing 1 mg / mL suc-AAPF-pNA working solution and activity was assayed at room temperature for 3-5 min at 405 nm using a SpectraMax plate reader in kinetic mode. Protease activity was expressed in mOD / min. The activity of each constructed variant was measured and compared to a reference construct grown on the same plate. The value of the reference sample was divided by the value of the variant sample to obtain the performance index (PI).

[0274] Example 3 Modification of the 5'GG36 pro sequence In this example, BG46_variant 2 subtilisin (SEQ ID NO: 8) was used as a reporter protein to monitor expression as described herein. More specifically, a DNA fragment comprising the upstream (5') aprE gene flanking region was operably linked to a polynucleotide construct comprising a variant of the DNA sequence of the (5') B. subtilis rrnI-P2 promoter / 5'-UTR aprE region, yielding SEQ ID NO:1 operably linked to a DNA sequence encoding an AprE signal peptide sequence (SEQ ID NO:4) operably linked to a DNA sequence encoding variant pro region sequence B (SEQ ID NO:9, 30H) operably linked to a DNA sequence encoding mature BG46_variant 2 subtilisin (SEQ ID NO:8) operably linked to a B. amyloliquefaciens BPN' terminator DNA sequence (SEQ ID NO:6), which in turn was operably linked to a downstream (3') aprE gene flanking region sequence comprising a kanamycin (kan) gene expression cassette (SEQ ID NO:7). More specifically, the DNA fragments were assembled using standard molecular biology techniques and then used as templates to generate linear DNA expression cassettes containing one or more of the promoter region modifications described herein.

[0275] Introduction of 5' pro-region modifications

[0276] Using the 5' (N-terminal) sequence of the 78 amino acid AprE proregion sequence (SEQ ID NO: 12), the first three amino acid residues were exchanged with the first amino acid residue (A) of the variant proregion sequence C (SEQ ID NO: 14) to obtain SEQ ID NO: 13. As shown in Table 1 above, the mutant of the proregion variant A sequence (i.e., Table 1, "E030G"; SEQ ID NO: 11) showed high expression of the reporter. Expression constructs were prepared using standard molecular biology techniques and compared to the reference proregion. The variants were generated as 4.4 kb linear DNA fragments that were used to transform competent B. subtilis cells, and the transformation mixture was plated on LA plates containing 1.8 ppm kanamycin and incubated overnight at 37°C. Single colonies were picked and grown in Luria broth at 37°C under antibiotic selection.

[0277] For reporter protein expression experiments, transformed cells were grown in 96-well MTPs in a shaking incubator in culture medium (enriched semi-defined medium based on MOP buffer) at 32°C for 3 days at 300 rpm and 80% humidity, then centrifuged and filtered. The clarified culture supernatant was used to measure (assay) reporter protease activity to determine the level of productivity, and samples were taken after 72 hours. Reporter protease activity assay was performed as described in Example 2 above, and the protein productivity of variant proregion C was approximately 1.1-fold higher than E030G-GG36 pro (SEQ ID NO: 11).

[0278] Example 4 Further modifications of the GG36 prosequence In this example, BG46_variant 2 subtilisin (SEQ ID NO: 8) was used as a reporter protein to monitor expression as described herein. Construction was performed as described in Example 3 above. More specifically, a combinatorial library was prepared on a reference pro-region sequence (SEQ ID NO: 15) with the amino acids glycine (G) and lysine (K) inserted in the 5' (N-terminal) sequence ("GK") (see, e.g., FIG. 2, SEQ ID NO: 14).

[0279] Combinatorial library for variant proregion C

[0280] A selection was made from the data in Table 1 for the combination of pro-region mutations and expression constructs were prepared using standard molecular biology techniques and compared with the reference pro-region. Linear DNA expression cassettes were used to transform competent B. subtilis cells and the transformation mixtures were plated on LA plates containing 1.8 ppm kanamycin and incubated overnight at 37°C. Single colonies were picked and grown in Luria broth at 37°C under antibiotic selection. For reporter protein expression experiments, transformed cells were grown in culture medium (enriched semi-defined medium based on MOP buffer) in 96-well MTPs in a shaking incubator at 300 rpm and 80% humidity for 3 days at 32°C, which were then centrifuged and filtered. Clarified culture supernatants were used to measure (assay) reporter protease activity to determine the level of productivity and samples were taken after 72 hours. Reporter protease activity assays were performed as described in Example 2 above.

[0281] Combinations in professional areas to improve productivity

[0282] As shown below, Table 2 reveals the sequence results of the combinations of pro-region mutations that showed increased productivity of the expression reporter (SEQ ID NO: 8), with delta (D) charges varying from +1 to +4. More specifically, positive charges as combinations in the loop region (i.e., residue positions 66-73). Performance index (PI) values ​​were assigned for samples taken after 72 hours and calculated as described in Example 2 above. For example, the PI values ​​of the pro-region variants are compared to the WT pro-region sequence (Table 2, SEQ ID NO: 15), and the amino acid position numbering (Table 2, first column) is compared to SEQ ID NO: 14 without the "GK" insertion. As shown in FIG. 3, the variant pro-region sequences can be numbered according to the position numbering of the reference variant pro-region sequence C (SEQ ID NO: 14; comprising 86 amino acid positions) or the wild-type (reference) pro-region sequence (SEQ ID NO: 15; comprising 84 amino acid positions). [Table 2]

[0283] Example 5 Modification of the 5'GG36 prosequence in Bacillus licheniformis In this example, BG46_variant 2 subtilisin (SEQ ID NO: 8) was used as a reporter protein to monitor expression as described herein. The construction was generally performed as follows. A first DNA fragment comprising a (5') serA gene flanking region (5' serA gene FR) which comprises a Bacillus licheniformis serA selectable marker expression cassette (SEQ ID NO: 16), a B. subtilis aprE signal sequence (SEQ ID NO: 4) operably linked to DNA encoding a variant pro region C sequence (SEQ ID NO: 14) operably linked to a DNA sequence encoding a mature BG46_variant 2 subtilisin (SEQ ID NO: 8) operably linked to a B. licheniformis amyL terminator (SEQ ID NO: 19) operably linked to a (3') serA gene flanking region (3' serA gene FR) (SEQ ID NO: 20), It was operably linked to a polynucleotide construct comprising the upstream (5') B. subtilis rrnI-p3 promoter region DNA sequence (SEQ ID NO:17) operably linked to the DNA sequence of the 5' untranslated region (5'-UTR; SEQ ID NO:18).

[0284] A second DNA fragment comprising a (5') lysA gene flanking region (5' lysA gene FR) comprising a B. licheniformis lysA selectable marker expression cassette (SEQ ID NO:21) is operably linked to a B. subtilis aprE signal sequence (SEQ ID NO:4) operably linked to DNA encoding a variant pro region C sequence (SEQ ID NO:14) operably linked to a DNA sequence encoding a mature BG46_variant 2 subtilisin (SEQ ID NO:8) operably linked to a B. licheniformis amyL terminator (SEQ ID NO:10) operably linked to a (3') lysA gene flanking region (3' lysA gene FR; SEQ ID NO:22). It was operably linked to a polynucleotide construct comprising the upstream (5') B. subtilis rrnI-p3 promoter region DNA sequence (SEQ ID NO:17) operably linked to the DNA sequence of the 5'-UTR (SEQ ID NO:18).

[0285] More specifically, these DNA fragments were assembled using standard molecular biology techniques and then used as templates to generate linear DNA expression cassettes containing one or more of the pro region sequence modifications described herein. For example, a B. licheniformis strain containing a variant pro region sequence was constructed by integrating a first and a second DNA fragment (described above) into the genome, the first and second fragments containing the pro region sequence variants described in Table 3 below. As shown in Table 3, production of reporter proteins was determined as previously described using standard methods (see, e.g., WO 2019 / 055261) and normalized to a control pro region sequence containing the E30G variant. [Table 3]

[0286] Example 6 Additional modification of the 5'GG36 prosequence In this example, BG46_variant 2 subtilisin (SEQ ID NO: 8) was used as a reporter protein to monitor expression as described herein. Specifically, the 5' (N-terminal) sequence of the 86 amino acid mutant GG36 prosequence (SEQ ID NO: 14; variant proregion sequence C) was further modified to introduce an additional amino acid such as alanine (A) or serine (S) after the lysine at position 3 (e.g., to potentially facilitate release of the mature protease into the medium). For example, these N-terminal modifications resulted in three variant proregion sequences DF (SEQ ID NOs: 29-31), as shown in Figures 1-3.

[0287] These N-terminally modified GG36 pro-sequences (SEQ ID NO:29, SEQ ID NO:30 and SEQ ID NO:31) were assembled in an expression cassette comprising an upstream (5') aprE gene flanking region operably linked to a polynucleotide construct comprising a variant of the (5') B. subtilis rrnI-P2 promoter / 5'-UTR aprE region DNA sequence to yield SEQ ID NO:1 operably linked to a DNA sequence encoding an AprE signal peptide sequence (SEQ ID NO:4) operably linked to a DNA sequence encoding a variant pro region sequence operably linked to a DNA sequence encoding mature BG46_variant 2 subtilisin (SEQ ID NO:8) operably linked to a B. amyloliquefaciens BPN' terminator DNA sequence (SEQ ID NO:6), which in turn was operably linked to a downstream (3') aprE gene flanking region (3' aprE gene FR) sequence comprising a downstream kanamycin (kan) gene expression cassette (SEQ ID NO:7). The cassettes containing the mutated / variant pro sequences were transformed into Bacillus subtilis cells using standard molecular biology techniques.

[0288] The transformed cells were grown in 96-well MTPs in a shaking incubator in culture medium (enriched semi-defined medium based on MOP buffer) at 32° C. for 3 days at 300 rpm and 80% humidity, which were then centrifuged and filtered. The clarified culture supernatants were used to assay the reporter protease activity to determine the level of productivity, and samples were taken after 72 hours. Specifically, Table 4 shows the results of reporter protein productivity (performance index (PI) values) after 72 hours compared to a reference construct in which the reporter protein was expressed with the WT pro sequence (SEQ ID NO: 14). [Table 4]

[0289] As shown in Table 4 above, the PI values ​​of the pro-region variants are compared to the pro-region sequence variant pro-region sequence C (SEQ ID NO: 14, "G2K3"), where the numbering of the amino acid positions is relative to SEQ ID NO: 14 (i.e., with the "GK" insertion). The PI index of the three N-terminal pro-region variants (AGKA, AGKS, AGGAA) showed higher expression of the reporter protein BG46_variant 2 subtilisin (SEQ ID NO: 8) compared to variant C (AGK, SEQ ID NO: 14) after 72 hours of growth.

[0290] Example 7 C-terminal modification of the GG36 prosequence In this example, GG36 variant proregion sequence B (SEQ ID NO: 11) was further engineered to replace leucine (L) at position 68 with lysine (K), isoleucine (I) at position 72 with valine (V), or glutamate (E) at position 80 with isoleucine (I) to obtain two engineered proregion variant sequences G (SEQ ID NO: 32) and H (SEQ ID NO: 33) (see, e.g., Figures 1-3). Specifically, transformed cells were grown as described above and clarified culture supernatants were used to assay reporter protease activity to determine levels of productivity, with samples taken after 72 hours.

[0291] For example, Table 5 below shows the results of reporter protein productivity (PI value) after 72 hours compared to a reference construct in which the reporter protein was expressed with the variant B pro region sequence (SEQ ID NO: 11). More specifically, as shown in Table 5, the PI values ​​of variant pro region sequences G and H were increased compared to the reference / control variant B pro region sequence, with variant pro region sequence G (containing L68K and I72V mutations) resulting in a PI of 1.08 relative to pro region sequence B (E30G), and variant pro region sequence H (containing L68K and E80I mutations) resulting in a PI of 1.18 relative to pro region sequence B (E30G). [Table 5]

[0292] References International Publication No. 1993 / 20214 Brochure International Publication No. 2002 / 14490 Brochure International Publication No. 2003 / 083125 Brochure International Publication No. 2008 / 112258 Brochure International Publication No. 2010 / 123754 Brochure International Publication No. 2016 / 205710 Brochure Altschul et al., “Basic local alignment search tool”, J. Mol. Biol., 215(13):403-410, 1990. Altschul et al., “Gapped BLAST and PSI BLAST a new generation of protein database search programs”, Nucleic Acids Res, Set 1;25(17):3389-402, 1997. Karlin et al.,“Applications and statistics for multiple high-scoring segments in molecular sequences”,PNAS USA,90(12):5873-5787,1993. Saitou and Nei,“The neighbor-joining method:a new method for reconstructing phylogenetic trees”,Mol.Biol.Evol.,Vol.4,issue 4,pages 406-425,1987. Schaffer et al.,“Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements”,Nucleic Acids Res,29(14):2994-3005,2001. Thompson et al.,“CLUSTAL W:improving the sensitivity of progressive multiple sequence alignment through sequence weighting,position-specific gap penalties and weight matrix choice”,Nucleic Acids Res,22(22):4673-4680,1994. Beaucage and Caruthers,“Deoxynucleoside phosphoramidites - A new class of key intermediates for deoxypolynucleotide synthesis”,Tetrahedron Lett.,22,1859-1862,1981. Matthes et al.,“Simultaneous rapid chemical synthesis of over one hundred oligonucleotides on a microscale”,EMBO J.,3:801-805,1984.

Claims

1. A variant pro region sequence comprising an amino acid substitution at position 30 and one or more positions selected from 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83, and 84, wherein the amino acid positions in the variant pro region are numbered according to Sequence ID No.

15.

2. A variant pro region according to claim 1, derived from a parent or reference polypeptide having at least about 70%, about 75%, about 80%, about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO:

15.

3. A variant pro region sequence comprising amino acid insertions of glycine (G) at position 2 and lysine (K) at position 3, and amino acid substitutions at one or more positions selected from 1, 32, 38, 46, 66, 67, 70, and 73, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO:

14.

4. A variant pro region according to claim 3, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO:

14.

5. A variant pro region sequence comprising a substitution of glutamic acid (E) to glycine (G) at position 30 (E30G), a substitution of leucine (L) to lysine (K) at position 68 (L68K), and a substitution of isoleucine (I) to valine (V) at position 73 (I72V), wherein the amino acid positions in the variant pro region are numbered according to SEQ ID NO: 15; or a variant pro region sequence comprising a substitution of glutamic acid (E) to glycine (G) at position 30 (E30G), a substitution of leucine (L) to lysine (K) at position 68 (L68K), and a substitution of glutamic acid (E) to isoleucine (I) at position 80 (E80I), wherein the amino acid positions in the variant pro region are numbered according to SEQ ID NO:

15.

6. A variant pro region according to claim 5, derived from a parent or reference polypeptide having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity with SEQ ID NO:

15.

7. A polynucleotide comprising a variant pro region nucleic acid according to any one of claims 1 to 6.

8. A polynucleotide comprising an upstream nucleic acid encoding a variant pro region according to any one of claims 1 to 6, which is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).

9. A polynucleotide comprising an upstream nucleic acid encoding a signal sequence, which is operably ligated to a downstream nucleic acid sequence encoding a variant pro region according to any one of claims 1 to 6, which is operably ligated to a downstream nucleic acid sequence encoding a target protein (POI).

10. An expression cassette containing an upstream promoter operably ligated to one of the following polynucleotides (i) to (iii) as a downstream polynucleotide: (i) Polynucleotide comprising a variant pro region nucleic acid according to any one of claims 1 to 6; (ii) A polynucleotide comprising an upstream nucleic acid encoding a variant pro region according to any one of claims 1 to 6, which is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI); (iii) A polynucleotide comprising an upstream nucleic acid encoding a signal sequence, which is operably ligated to a downstream nucleic acid sequence encoding a variant pro region according to any one of claims 1 to 6, which is operably ligated to a downstream nucleic acid sequence encoding a target protein (POI).

11. Gram-positive host cells comprising the introduced cassette according to claim 10.

12. A method for producing a target protein (POI) in Gram-positive bacterial cells, (a) Introducing an expression cassette into Gram-positive cells, comprising an upstream promoter operably linked to a downstream nucleic acid encoding a variant pro region sequence, which includes amino acid substitutions at position 30 and one or more positions selected from 1, 2, 3, 4, 6, 14, 16, 19, 20, 23, 36, 37, 38, 39, 42, 43, 44, 49, 50, 64, 65, 67, 68, 71, 79, 83, and 84, wherein the amino acid positions of the variant pro region are numbered according to Sequence ID No. 15, (b) Culturing the modified cells under the conditions for the production of the POI A method that includes this.

13. A method for producing a target protein (POI) in Gram-positive bacterial cells, (a) Introducing an expression cassette into Gram-positive cells, comprising an upstream promoter sequence operably linked to a downstream nucleic acid encoding a variant pro region sequence containing an amino acid insertion of glycine (G) at position 2 and lysine (K) at position 3, and an amino acid substitution at one or more positions selected from 1, 32, 38, 46, 66, 67, 70, and 73, wherein the amino acid positions of the variant pro region are numbered according to SEQ ID NO: 14, (b) Culturing the modified cells under the conditions for the production of the POI A method that includes this.

14. The method according to claim 12 or 13, wherein the cassette further comprises a nucleic acid encoding a signal sequence operably linked and positioned between the upstream promoter and the downstream variant pro region.

15. The method according to claim 12 or 13, wherein the modified cells produce an increased amount of the POI compared to control Gram-positive cells having an introduced expression cassette comprising the same upstream promoter sequence operably ligated to a downstream nucleic acid encoding the pro region sequence of Sequence ID No. 15, which is operably ligated to a downstream nucleic acid sequence encoding the same POI.

16. The method according to claim 12 or 13, wherein the increased amount of POI is increased by at least about 0.1% compared to control cells.

17. The aforementioned POIs include acetylesterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucanylase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, and isomerase. The method according to claim 12 or 13, selected from the group consisting of enzymes, laccase, ligase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectin acid lyase, pectin acetylesterase, pectin depolymerase, pectin methylesterase, pectin-degrading enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transport protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.