Novel promoters and 5 '-untranslated region mutations that enhance protein production in gram-positive cells
By introducing new promoters and 5'-untranslated region nucleic acid sequences into Gram-positive bacterial cells, the problem of insufficient protein expression/production level of Bacillus host cells in the prior art is solved, and the efficient expression and production of the target protein is achieved, reducing production costs and time.
Patent Information
- Application Number
- CN202380076119.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-09-01
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has limitations in improving the expression/production level of proteins produced in Bacillus host cells, resulting in higher production costs, energy and time for proteins in industrial biotechnology.
New promoters and 5'-untranslated region nucleic acid (DNA) sequences are provided for constructing recombinant polynucleotides such as vectors, plasmids, expression cassettes, etc., and enhance the expression and production of the protein of interest by introducing these sequences into Gram-positive bacterial cells.
By using novel promoters and 5'-UTR sequences, the expression and production level of the target protein in Gram-positive bacterial cells is significantly improved, production costs are reduced, and efficiency is improved.
Smart Images

Figure BDA0005380361440000451 
Figure BDA0005380361440000461 
Figure BDA0005380361440000462
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the fields of microbial host cells, molecular biology, protein engineering, fermentation, protein production, etc. Certain aspects of the present disclosure relate to novel promoter and 5′-untranslated region nucleic acid (DNA) sequences.
[0002] Cross-reference to Related Applications
[0003] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 374,450, filed on September 2, 2022, the disclosure of which is hereby incorporated by reference in its entirety.
[0004] Reference to Sequence Listing
[0005] The content of the electronically submitted sequence listing of the text file named "NB41861-WO-PCT_SequenceListing.xml", created on August 30, 2023, and having a size of 53 KB, is hereby incorporated by reference in its entirety. Background Art
[0006] Genetic engineering has facilitated various improvements to host microorganisms used as industrial bioreactors or cell factories. For example, Gram-positive bacterial host cells can produce and secrete large amounts of useful proteins and metabolites. The most common Bacillus species used in industry are Bacillus licheniformis, Bacillus amyloliquefaciens, and Bacillus subtilis. Due to their generally recognized as safe (GRAS) status, strains of Bacillus species are natural candidates for producing proteins for the food and pharmaceutical industries. For example, important production enzymes include α-amylase, neutral protease, alkaline (or serine) protease, etc. However, despite the progress in the understanding of protein production in Bacillus host cells, there is still a need for methods and compositions for improving the expression / production of these proteins by microorganisms.
[0007] The recombinant production of a protein of interest (POI) encoded by a gene of interest (or ORF) is typically accomplished by constructing an expression vector suitable for the desired host cell, wherein the nucleic acid encoding the desired POI is placed under the expression control of a promoter. Thus, the expression vector is introduced into the host cell by various techniques (e.g., via transformation), and the production of the desired protein product is achieved by culturing the transformed host cell under conditions suitable for the expression and production of the protein product. For example, Bacillus species promoters (and associated elements) for the recombinant expression of functional polypeptides have been described (e.g., Kim et al., 2008 and U.S. Patent No. 4,559,300). Although many promoters are known, there remains a need in the art for novel promoter (nucleic acid) sequences that can improve the expression of heterologous nucleic acids encoding proteins of interest. For example, in the field of industrial biotechnology, even a relatively small increase in the expression / production level of industrially relevant proteins (e.g., enzymes, antibodies, receptors, etc.) translates into significant cost, energy, and time savings for the recombinant proteins produced. Summary of the Invention
[0008] As generally described herein, the present disclosure particularly provides compositions and methods for producing a protein of interest in Gram-positive bacterial (host) cells. Certain embodiments relate to novel promoter and 5′-UTR nucleic acid (DNA) sequences, recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.) comprising the novel promoter / 5′-UTR sequences, recombinant polynucleotides comprising a novel promoter / 5′-UTR sequence operably linked to a DNA sequence encoding a protein signal (secretion) sequence, and / or operably linked to a DNA sequence encoding a pro-region sequence, operably linked to a DNA sequence encoding a protein of interest, etc.
[0009] Accordingly, certain embodiments of the present disclosure provide variant nucleic acid sequences comprising at least one mutation shown in any one of SEQ ID NOs: 8 to 46, wherein the nucleotide positions of these variant nucleic acid sequences are numbered according to SEQ ID NO: 1. In certain other embodiments, the variant nucleic acid comprises the nucleotide sequence of any one of SEQ ID NOs: 8 to 46, wherein the nucleotide positions of these variant sequences are numbered according to SEQ ID NO: 1. In certain other embodiments, the variant nucleic acid sequence has at least about 97.5% identity to SEQ ID NO: 1. In a related aspect, the variant nucleic acid sequences of the present disclosure may be referred to as variant promoter / 5′-untranslated region (5-UTR) sequences.
[0010] Certain other one or more embodiments relate to polynucleotides (DNA) comprising variant nucleic acid sequences of the present disclosure. Accordingly, certain embodiments provide polynucleotides comprising variant nucleic acid sequences of the present disclosure, which are operably linked to a downstream (3′) nucleic acid sequence encoding a protein of interest.
[0011] In certain other embodiments or aspects, a polynucleotide comprising a variant nucleic acid sequence of the present disclosure is operably linked to a downstream (3′) nucleic acid sequence encoding a pro-region sequence, which is operably linked to a downstream nucleic acid sequence encoding a mature protein of interest (POI). In certain other embodiments or aspects, a polynucleotide comprising a variant nucleic acid sequence of the present disclosure is operably linked to a downstream (3′) nucleic acid sequence encoding a protein signal (secretion) sequence (SS), which is operably linked to a downstream nucleic acid sequence encoding a mature protein of interest (POI). In certain other embodiments or aspects, a polynucleotide comprising a variant nucleic acid sequence of the present disclosure is operably linked to a downstream (3′) nucleic acid sequence encoding a protein signal (secretion) sequence (SS), which is operably linked to a downstream nucleic acid sequence encoding a pro-region (PRO) sequence, which is operably linked to a downstream nucleic acid sequence encoding a mature protein of interest (POI).
[0012] In certain related embodiments, the POI is selected from the group consisting of: enzymes, antibodies, receptor proteins, lectins, and regulatory proteins.
[0013] Other embodiments provide expression cassettes comprising the polynucleotides of the present disclosure. In related embodiments, a Gram-positive bacterial (host) cell comprises one or more introduced cassettes of the present disclosure. In certain embodiments, the Gram-positive host cell is a Bacillus species cell. In related embodiments, the Bacillus species (host) cell is selected from the group consisting of: Bacillus subtilis, Bacillus licheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus clausii, Bacillus halodurans, Bacillus megaterium, Bacillus coagulans, Bacillus circulans, Bacillus lautus, and Bacillus thuringiensis.
[0014] Certain other embodiments of the present disclosure provide methods for producing a protein of interest (POI) in a Gram - positive bacterial cell, the methods comprising: (a) introducing into the Gram - positive cell an expression cassette comprising a variant nucleic acid sequence of any one of SEQ ID NOs: 8 to 46, the variant nucleic acid sequence being operably linked to a downstream (3′) nucleic acid sequence encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for the production of the POI. In certain preferred embodiments of these methods, the modified cell produces an increased amount of the POI relative to (compared to) a control Gram - positive cell comprising the introduced expression cassette, the introduced expression cassette comprising the reference nucleic acid sequence of SEQ ID NO: 1 operably linked to a downstream (3′) nucleic acid sequence encoding the same POI, wherein the modified cell and the control cell are cultured under the same conditions. In certain other embodiments of these methods, after culturing for at least about seventy - two (72) hours, the modified cell produces an increased amount of the POI relative to the control cell. In still other embodiments of these methods, the protein of interest (POI) is selected from the group consisting of: enzymes, antibodies, receptor proteins, lectins, and regulatory proteins. In certain embodiments of these methods, the enzymes are selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α - galactosidase, β - galactosidase, α - glucanase, glucan lyase, endo - β - glucanase, glucoamylase, glucose oxidase, α - glucosidase, β - glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof. In certain other embodiments of these methods, the Gram - positive bacterial cell is a Bacillus species cell. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 The nucleotide sequence of the variant rrnI - P2 promoter / 5′ - UTR region DNA sequence (SEQ ID NO: 1) is presented. More particularly, as Figure 1As presented, the variant (reference) rrnI-P2 promoter / 5′-UTR region sequence comprises nucleotide positions 1 to 149 of SEQ ID NO:1, wherein the "UP", "-35", "-10" and "Shine-Dalgarno" sequence elements are indicated in bold text.
[0016] Figure 2 DNA sequence alignments of the reference variant rrnI-P2 promoter / 5′-UTR (SEQ ID NO:1) and certain SEL variant promoter / 5′-UTR region sequences of the present disclosure are presented. More particularly, as Figure 2 A and Figure 2 B show, the nucleotide positions of the SEL variant sequences described in the examples are aligned with the reference rrnI-P2 promoter / 5′-UTR region (SEQ ID NO:1; nucleotide positions 1-149). As Figure 2 A shows, nucleotide positions 1-83 of the reference rrnI-P2 promoter region include the UP element, -35 and -10 elements indicated by gray shading, and as Figure 2 B shows, nucleotide positions 83-149 of the reference rrnI-P2 promoter / 5′-UTR region include the Shine-Dalgarno element indicated by gray shading. As Figure 2 As presented in Figure 2 A, the modified nucleotide positions of the rrnI-P2 promoter region are indicated by nucleotide residues with black shading (e.g.,
[0017] Figure 3 DNA sequence alignments of the reference variant rrnI-P2 promoter / 5′-UTR region (SEQ ID NO:1) and certain SEL variant promoter / 5′-UTR region sequences of the present disclosure are presented. More particularly, as Figure 3 A and Figure 3 B show, the nucleotide positions of the SEL variant sequences described in the examples are aligned with the reference rrnI-P2 promoter / 5′-UTR region (SEQ ID NO:1; nucleotide positions 1-149). As Figure 3 A shows, nucleotide positions 1-83 of the reference rrnI-P2 promoter / 5′-UTR region include the UP element, -35 and -10 elements indicated by gray shading, and as Figure 3 B shows, nucleotide positions 83-149 of the reference rrnI-P2 promoter / 5′-UTR region include the Shine-Dalgarno element indicated by gray shading. As Figure 3 As presented in Figure 3A, UTR-00798; A).
[0018] Biological Sequence Description
[0019] SEQ ID NO:1 is the nucleic acid (DNA) sequence of the variant (reference) rrnI-P2 promoter / 5′-UTR region.
[0020] SEQ ID NO:2 is the amino acid sequence of the wild-type subtilisin from Bacillus gibsonii named "BG46".
[0021] SEQ ID NO:3 is the amino acid sequence of the variant subtilisin from Bacillus gibsonii BG46 named "BG46_variant".
[0022] SEQ ID NO:4 is the DNA sequence encoding the signal sequence of the wild-type Bacillus subtilis AprE protein.
[0023] SEQ ID NO:5 is the DNA sequence encoding the pro-region sequence of the wild-type Bacillus lentus.
[0024] SEQ ID NO:6 is the DNA sequence of the wild-type Bacillus amyloliquefaciens BPN′ terminator.
[0025] SEQ ID NO:7 is the DNA sequence of the kanamycin (kan) gene expression cassette.
[0026] SEQ ID NO:8 is the DNA sequence of the variant UTR-00664.
[0027] SEQ ID NO:9 is the DNA sequence of the variant UTR-00692.
[0028] SEQ ID NO:10 is the DNA sequence of the variant UTR-00330.
[0029] SEQ ID NO:11 is the DNA sequence of the variant UTR-00411.
[0030] SEQ ID NO:12 is the DNA sequence of the variant UTR-00325.
[0031] SEQ ID NO:13 is the DNA sequence of the variant UTR-00730.
[0032] SEQ ID NO:14 is the DNA sequence of the variant UTR-00348.
[0033] SEQ ID NO:15 is the DNA sequence of the variant UTR-00738.
[0034] SEQ ID NO:16 is the DNA sequence of variant UTR-00788.
[0035] SEQ ID NO:17 is the DNA sequence of variant UTR-00792.
[0036] SEQ ID NO:18 is the DNA sequence of variant UTR-00800.
[0037] SEQ ID NO:19 is the DNA sequence of variant UTR-01018.
[0038] SEQ ID NO:20 is the DNA sequence of variant UTR-01112.
[0039] SEQ ID NO:21 is the DNA sequence of variant UTR-00037.
[0040] SEQ ID NO:22 is the DNA sequence of variant UTR-00039.
[0041] SEQ ID NO:23 is the DNA sequence of variant UTR-00661.
[0042] SEQ ID NO:24 is the DNA sequence of variant UTR-00891.
[0043] SEQ ID NO:25 is the DNA sequence of variant UTR-00084.
[0044] SEQ ID NO:26 is the DNA sequence of variant UTR-00362.
[0045] SEQ ID NO:27 is the DNA sequence of variant UTR-00424.
[0046] SEQ ID NO:28 is the DNA sequence of variant UTR-00643.
[0047] SEQ ID NO:29 is the DNA sequence of variant UTR-00645.
[0048] SEQ ID NO:30 is the DNA sequence of variant UTR-00741.
[0049] SEQ ID NO:31 is the DNA sequence of variant UTR-00798.
[0050] SEQ ID NO:32 is the DNA sequence of variant UTR-00960.
[0051] SEQ ID NO:33 is the DNA sequence of variant UTR-01223.
[0052] SEQ ID NO:34 is the DNA sequence of variant UTR-00656.
[0053] SEQ ID NO:35 is the DNA sequence of variant UTR-00657.
[0054] SEQ ID NO:36 is the DNA sequence of variant UTR-00030.
[0055] SEQ ID NO:37 is the DNA sequence of variant UTR-01092.
[0056] SEQ ID NO:38 is the DNA sequence of variant UTR-00721.
[0057] SEQ ID NO:39 is the DNA sequence of variant UTR-00651.
[0058] SEQ ID NO:40 is the DNA sequence of variant UTR-00301.
[0059] SEQ ID NO:41 is the DNA sequence of variant UTR-00187.
[0060] SEQ ID NO:42 is the DNA sequence of variant UTR-00035.
[0061] SEQ ID NO:43 is the DNA sequence of variant UTR-00005.
[0062] SEQ ID NO:44 is the DNA sequence of variant UTR-00863.
[0063] SEQ ID NO:45 is the DNA sequence of variant UTR-00711.
[0064] SEQ ID NO:46 is the DNA sequence of variant UTR-00752.
[0065] SEQ ID NO:47 is the DNA sequence of the 5′ aprE gene FR.
[0066] SEQ ID NO:48 is the DNA sequence of the 3′ aprE gene FR. Detailed implementation manners
[0067] As briefly set forth above and described in detail below, the present disclosure particularly provides novel promoter and 5′-UTR nucleic acid (DNA) sequences, recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.) comprising the novel promoter / 5′-UTR sequences, recombinant polynucleotides of the novel promoter / 5′-UTR sequences operably linked to a DNA sequence encoding a protein signal (secretion) sequence, and / or operably linked to a DNA sequence encoding a proregion sequence, operably linked to a DNA sequence encoding a protein of interest, etc.
[0068] In certain aspects, the present disclosure provides recombinant Gram-positive bacterial strains that express one or more introduced polynucleotides encoding a protein of interest. In certain other aspects, the present disclosure provides compositions and methods for designing / constructing recombinant Gram-positive bacterial strains expressing one or more introduced novel polynucleotide constructs encoding a protein of interest, compositions and methods for culturing recombinant strains expressing a protein of interest, compositions and methods for enhancing the production of a protein of interest, etc.
[0069] I. Definitions
[0070] In view of the recombinant polynucleotides, recombinant (modified) strains, and methods described herein, the following terms and phrases are defined. Terms not defined herein should conform to their ordinary meaning as used in the art.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the compositions and methods of the present invention pertain. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the compositions and methods of the present invention, representative illustrative methods and materials are now described. All publications and patents cited herein are incorporated by reference in their entirety.
[0072] It should be further noted that the claims can be drafted to exclude any optional element. Thus, this statement is intended to serve as antecedent basis (or its condition) for the use of exclusive terms relating to the recitation of claim elements such as "solely", "only", "exclude", "do not include", etc. or the use of "negative" limitations.
[0073] Upon reading the present disclosure, it will be apparent to those skilled in the art that each of the individual embodiments described and illustrated herein has discrete components and features that can be readily separated from or combined with the features of any one of several other embodiments without departing from the scope or spirit of the compositions and methods of the present invention described herein. Any of the recited methods can be carried out in the order of the recited events or in any other order that is logically possible.
[0074] As used herein, the phrases "Gram-positive bacterium," "Gram-positive cell," "Gram-positive bacterial strain," and / or "Gram-positive bacterial cell" have the same meaning as used in the art. For example, Gram-positive bacterial cells include all strains of the phyla Actinobacteria and Firmicutes. In certain embodiments, such Gram-positive bacteria are of the classes Bacilli, Clostridia, and Mollicutes.
[0075] As used herein, "Bacillus" includes all species within the genus "Bacillus" known to those of skill in the art, including but not limited to Bacillus subtilis, Bacillus licheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus clausii, Bacillus halodurans, Bacillus megaterium, Bacillus coagulans, Bacillus circulans, Bacillus lautus, and Bacillus thuringiensis. It should be recognized that the genus Bacillus is constantly undergoing taxonomic reorganization. Accordingly, the genus is intended to include species that have been reclassified, including but not limited to organisms such as Bacillus stearothermophilus (which is now named "Geobacillus stearothermophilus").
[0076] As used herein, the terms "recombinant" or "non-natural" refer to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic alteration or has been modified by the introduction of a heterologous nucleic acid molecule, or to a cell (e.g., a microbial cell) that has been altered such that the expression of a heterologous or endogenous nucleic acid molecule or gene can be controlled. Recombinant also refers to a cell that is derived from a non-natural cell or to the progeny of a non-natural cell that has one or more such modifications. Genetic alterations include, for example, modifications that introduce an expressible nucleic acid molecule encoding a protein, or the addition, deletion, substitution, or other functional alteration of other nucleic acid molecules of the cell's genetic material. For example, a recombinant cell can express the same or a homologous form of a gene or other nucleic acid molecule (e.g., a fusion protein or a chimeric protein) that is not found within a natural (wild-type) cell, or can provide an altered pattern of endogenous gene expression, such as overexpression, underexpression, minimal expression, or no expression at all. "Recombination" or generating a "recombined" nucleic acid generally is the assembly of two or more nucleic acid fragments, where the assembly results in a chimeric gene.
[0077] The term "derived from" encompasses the terms "originated from", "obtained from", "obtainable from", and "produced from", and generally indicates that a specified material or composition finds its origin in another specified material or composition, or has characteristics that can be described with reference to another specified material or composition. For example, the recombinant Gram-positive bacterial cells of the present disclosure can be derived from / obtained from any known Gram-positive bacterial strain.
[0078] As used herein, "nucleic acid" refers to nucleotide or polynucleotide sequences and their fragments or portions, as well as DNA, cDNA, and RNA of genomic or synthetic origin, which can be double-stranded or single-stranded, whether representing the sense or antisense strand. It should be understood that due to the degeneracy of the genetic code, multiple nucleotide sequences can encode a given protein.
[0079] It should be understood that the polynucleotides (or nucleic acid molecules) described herein include "genes", "vectors", and "plasmids".
[0080] Accordingly, the term "gene" refers to a polynucleotide that encodes a specific sequence of amino acids, which includes all or part of the protein-coding sequence and can include regulatory (non-transcribed) DNA sequences, such as a promoter sequence, which determines, for example, the conditions under which the gene is expressed. The transcribed region of a gene can include untranslated regions (UTRs) (including 5'-untranslated region (UTR) and 3'-UTR) as well as the coding sequence.
[0081] As used herein, "endogenous gene" refers to a gene that is located in its natural position in the genome of an organism.
[0082] As used herein, a "heterologous" gene, "non-endogenous" gene or "exogenous" gene refers to a gene that is not normally found in a host organism but is introduced into the host organism by gene transfer. The term one or more "exogenous" genes includes natural genes inserted into a non-natural organism and / or chimeric genes inserted into a natural or non-natural organism.
[0083] As used herein, a "heterologous control sequence" refers to a gene expression control sequence (e.g., a promoter, enhancer, terminator, etc.) that does not function intrinsically to regulate (control) the expression of a gene of interest. Typically, heterologous nucleic acids are not endogenous (natural) to the part of the cell or genome in which they are present and have been added to the cell by infection, transfection, transduction, transformation, microinjection, electroporation, etc. A "heterologous" nucleic acid construct can contain a combination of control sequences / DNA coding (ORF) sequences that is the same as or different from the combination found in a natural host cell.
[0084] As used herein, the term "expression" refers to the transcription and stable accumulation of sense (mRNA) or antisense RNA derived from a nucleic acid molecule of the present disclosure. Expression can also refer to the translation of mRNA into a polypeptide. Thus, the term "expression" includes any step involved in polypeptide production, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, secretion, etc.
[0085] As used herein, the term "coding sequence" (CDS) refers to a nucleotide sequence that directly specifies the amino acid sequence of its (encoded) protein product. The boundaries of a coding sequence are typically determined by a reading frame (hereinafter, "ORF") that usually begins with an ATG start codon. Coding sequences typically include DNA, cDNA, and recombinant nucleotide sequences.
[0086] As used herein, the terms "promoter", "promoter element", "promoter sequence", etc. refer to a nucleic acid (DNA) sequence capable of controlling the transcription of a gene coding sequence (CDS) into messenger RNA (mRNA), where the promoter region sequence is placed upstream (5') and operably linked to the downstream (3') gene CDS. As is commonly understood by those skilled in the art, a promoter typically provides a site for specific binding and initiation of transcription by RNA polymerase. In some aspects, the term "promoter" refers to the minimal portion of the promoter nucleic acid sequence required to initiate transcription (i.e., containing the RNA polymerase binding site). For example, a promoter typically contains a "-10" (consensus sequence) element and a "-35" (consensus sequence) element, which are located upstream (5') and are relative to the +1 transcription start site (TSS) of the gene CDS to be transcribed. The core promoter -10 and -35 elements are commonly referred to in the art as the "TATAAT" (Pribnow box) consensus region and the "TTGACA" consensus region, respectively. The spacing of the core promoter (-10 and -35) regions is typically separated (spaced) by approximately fifteen - twenty (15 - 20) intervening base pairs (nucleotides), as Figure 1 shown.
[0087] A promoter can be entirely derived from a natural gene, or composed of different elements derived from different promoters found in nature, or even contain synthetic nucleic acid segments. Those skilled in the art should understand that different promoters can direct the expression of a gene in different cell types, or at different developmental stages, or in response to different environmental or physiological conditions. A promoter can be a constitutive promoter, an inducible promoter, a tunable promoter, a hybrid promoter, a synthetic promoter, a tandem promoter, etc. A promoter that causes a gene to be expressed in most cell types most of the time is typically referred to as a "constitutive promoter". It should be further recognized that, in most cases, since the exact boundaries of regulatory sequences have not been fully defined, DNA fragments of different lengths can have the same promoter activity.
[0088] In some aspects, an upstream (5') promoter sequence (pro) is operably linked to a downstream DNA sequence encoding a protein signal sequence (SS), which is operably linked to a downstream DNA sequence encoding a pro-region sequence (PRO), which is operably linked to a downstream (3') DNA sequence encoding a mature target protein (ORF), which can be schematically presented as 5'-[pro]-[SS]-[PRO]-[ORF]-3'.
[0089] As used herein, "a functional promoter sequence that controls the expression of a target gene linked to the protein-coding sequence of the target gene" refers to a promoter sequence that controls the transcription and translation of the coding sequence in a desired Gram-positive host cell. For example, in certain embodiments, the present disclosure provides a polynucleotide comprising an upstream (5') promoter (or 5'-promoter region, or tandem 5'-promoter, etc.) that is functional in Gram-positive cells, wherein the functional promoter region is operably linked to a nucleic acid sequence encoding a target protein.
[0090] As used herein, the term "precursor protein" refers to the inactive form of a protein. In some aspects, the full-length protein is synthesized as a precursor in the form of a pre-sequence and a mature protein (abbreviated as "pre-protein"). In other aspects, the full-length protein is synthesized as a precursor in the form of a signal peptide sequence, a pre-sequence, and a mature protein (abbreviated as "pre-pro-protein"). For example, the pre-sequence typically serves as a signal peptide for translocation, and the pre-sequence is typically essential for the correct folding of the associated (mature) protein.
[0091] As used herein, the term "mature protein" refers to the active form of a protein as opposed to the inactive precursor (full-length) protein.
[0092] As used herein, the terms "signal sequence", "secretion signal", and "signal peptide" are used interchangeably and refer to a sequence of amino acid residues that can participate in the secretion or directed translocation of a precursor protein. Typically, the signal (pre) sequence is cleaved from the precursor protein by signal peptidase during the translocation process. The signal (pre) sequence is typically located at the N-terminus of the mature protein sequence, or at the N-terminus of the pro-region sequence, where the signal (pre) sequence and the pro-region sequence are used in an operable combination and are located upstream (5') of the mature POI sequence.
[0093] As used herein, phrases such as "variant rrnI-P2 promoter / 5'-UTR region" and / or "reference rrnI-P2 promoter / 5'-UTR region" specifically refer to the variant Bacillus subtilis rrnI-P2 promoter / 5'-UTR region DNA sequence shown in SEQ ID NO:1 (see, for example, Figure 1 ). Thus, in some aspects, the variant rrnI-P2 promoter / 5'-UTR region sequence (SEQ ID NO:1) can be referred to as a reference sequence or a control sequence, particularly when compared to one or more SEL variant promoter / 5'-UTR region sequences of the present disclosure. For example, as Figure 2 and Figure 3As presented, the reference rrnIp2 promoter / 5′-UTR region sequence (nucleotide positions 1-149) has been aligned with certain site evaluation library (SEL) variant promoter 5′-UTR region sequences of the present disclosure. More particularly, the SEL variant promoter / 5′-UTR region sequences (e.g., see Table 2; SEQ ID NOs: 8-48) are aligned with the reference rrnIp2 promoter / 5′-UTR region sequence, where nucleotide positions 1-83 of the reference promoter / 5′-UTR region sequence and the SEL variants are shown in Figure 2 A and Figure 3 A, and nucleotide positions 84-149 of the reference promoter / 5′-UTR region sequence and the same SEL variants are shown in Figure 2 B and Figure 3 B, respectively. As noted in Figure 1 B, the DNA sequence elements of the reference promoter / 5′-UTR region sequence (SEQ ID NO: 1) include (in the 5′ to 3′ direction) an upstream (UP) element, a -35 element, a -10 element, and a Shine-Dalgarno (SD) element, which are indicated by bold nucleotides.
[0094] In some aspects, a promoter contains nucleotides upstream (5′) of the promoter, where such upstream (5′) nucleotides are referred to herein as “upstream promoter elements” (abbreviated as “UP elements” or “UP sequences”). Thus, as used herein, a “UP sequence” refers to an “A+T”-rich (nucleic acid) sequence region located upstream of the -35 core promoter element. The UP sequence can be further described as a nucleic acid sequence region located upstream of the -35 core promoter element that directly interacts with the C-terminal domain of the α-subunit of RNA polymerase. Thus, in certain embodiments, a promoter contains one (or more) UP sequences located upstream of the promoter and operably linked to the promoter.
[0095] As used herein, the phrase “transcription start site” (abbreviated as “TIS”) generally refers to the base pair at which transcription starts. By convention, the transcription start site (TIS) in the DNA sequence of a transcription unit is numbered, where nucleotide positions extending in the transcription direction (i.e., 3′; downstream) are assigned positive “(+)” numbers, and nucleotide positions extending in the opposite direction (i.e., 5′; upstream) are assigned negative “(-)” numbers.
[0096] As used herein, the phrases “transcription start site” (abbreviated as “tss”) and “transcription start site (tss) codon” can be used interchangeably and refer to the three (3)-nucleotide translation start site (tss) codon. For example, prokaryotic “tss codons” include, but are not limited to, “AUG”, “GUG”, “UGG”, etc.
[0097] As used herein, the term "Shine–Dalgarno" sequence (abbreviated as "SD" sequence) refers to the ribosome binding site of messenger RNA (mRNA), which is typically located approximately 8 nucleotides upstream (5′) of the start codon (e.g., AUG). As understood in the art, the SD sequence helps recruit the ribosome to the mRNA to initiate protein synthesis by aligning the ribosome with the start codon (e.g., AUG), where transfer RNA (t-RNA) can add the amino acid specified by the codon and move downstream (3′) from the translation start site (tss).
[0098] As used herein, the terms "pro sequence", "pro-sequence", and "pro-region sequence" are used interchangeably and abbreviated as "PRO" sequence, "Pro" sequence, "pro" sequence, etc. The term pro sequence as used herein has the same meaning as understood in the art. For example, the Bacillus subtilis alkaline serine protease "subtilisin" is first produced as pre-pro-subtilisin, which consists of a signal (pre) sequence for protein secretion, followed by a pro-region sequence of seventy-seven (77) amino acids, followed by the amino acid sequence encoding mature subtilisin (e.g., pre-pro-subtilisin). The pro sequence acts as an intramolecular chaperone (e.g., directly catalyzing the protein folding reaction) and is typically essential for the correct folding of the relevant (mature) protein. Similarly, the folding and intracellular trafficking (or secretion) of the mature protein of interest may both require the pro sequence, indicating that these two functions are closely related. In certain aspects, the pro-region sequence of the present disclosure comprises an amino acid sequence derived from the wild-type (WT, reference) pro-region sequence of Bacillus lentus SEQ ID NO:5.
[0099] As used herein, the phrase "polynucleotide encoding a full-length protein" refers to a DNA sequence encoding a "precursor" protein; and the phrase "polynucleotide encoding a mature protein" refers to a DNA sequence encoding a "mature" protein as defined herein. In some aspects, a polynucleotide encoding a precursor protein comprises at least an upstream (5') DNA sequence encoding a pre-region amino acid sequence, which is operably linked to a downstream (3') DNA sequence encoding an amino acid sequence of a mature protein of interest (POI) (e.g., open reading frame, ORF). In other aspects, a polynucleotide encoding a precursor protein comprises at least an upstream (5') DNA sequence encoding a protein signal sequence, which is operably linked to a downstream (3') DNA sequence encoding a pre-region amino acid sequence, which is operably linked to a downstream (3') ORF encoding an amino acid sequence of a mature POI.
[0100] As used herein, the term "untranslated region" may be abbreviated as "UTR".
[0101] As used herein, the phrases "5'-untranslated region", "5'-UTR", and / or "5'-transcriptional leader sequence" may be used interchangeably and abbreviated as "5'-UTR". As is commonly understood in the art, the 5'-UTR is known to be the region of messenger RNA (mRNA) directly upstream (5') of the start codon.
[0102] A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, if a DNA encoding a secretory leader sequence (i.e., signal sequence) is expressed as a pre-protein involved in polypeptide secretion, then the DNA encoding the secretory leader sequence is operably linked to the DNA encoding the polypeptide; if a promoter or enhancer affects the transcription of a coding sequence, then the promoter or enhancer is operably linked to the sequence (CDS, ORF); or if a ribosome binding site is positioned to facilitate translation, then the ribosome binding site is operably linked to the coding sequence. Generally, "operably linked" means that the DNA sequences being linked are contiguous and, in the case of a secretory leader sequence, are contiguous and in the reading phase. However, an enhancer need not be contiguous. Ligation is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide linkers or adaptors are used according to conventional practice. Thus, the term operably linked generally refers to the association (juxtaposition) of nucleic acid sequences on a single nucleic acid fragment such that the function of one is affected by the other. For example, if a promoter (pro) controls the transcription of a gene coding sequence (gene CDS), then the promoter is operably linked to the gene CDS.
[0103] As used herein, DNA encoding the "wild-type Bacillus subtilis aprE signal peptide sequence" may be abbreviated as "aprESS" and comprises the nucleotide sequence of SEQ ID NO:4.
[0104] As used herein, DNA encoding the "wild-type Bacillus lentus propeptide region DNA sequence" may be abbreviated as "PRO sequence", "PRO region" or "PRO" and comprises the nucleotide sequence of SEQ ID NO:5.
[0105] As used herein, "wild-type Bacillus amyloliquefaciens BPN′ terminator (BPN′term)" may be abbreviated as "term" and comprises the nucleotide sequence of SEQ ID NO:6.
[0106] As used herein, the wild-type Bacillus gibsonii (BG46) subtilisin comprising the amino acid sequence of SEQ ID NO:2 is abbreviated as "WT BG46".
[0107] As used herein, the variant Bacillus gibsonii (BG46) subtilisin comprising the amino acid sequence of SEQ ID NO:3 is abbreviated as "BG46 variant", which is derived from the wild-type Bacillus gibsonii (BG46) subtilisin reporter protein (SEQ ID NO:2).
[0108] As used herein, the "upstream (5′) aprE flanking region (FR) sequence" comprises SEQ ID NO:49.
[0109] As used herein, the "downstream (3′) aprE flanking region (FR) sequence" comprises SEQ ID NO:50 and includes the "kanamycin (Kan) gene expression cassette" (SEQ ID NO:7) for selection.
[0110] As used herein, an exemplary protease may be referred to as a "reporter protein". In certain one or more embodiments of the present disclosure, the exemplary reporter protein is expressed / produced by one or more recombinant (modified) cells of the present disclosure. In certain embodiments, the reporter protein includes, but is not limited to, native and variant Bacillus species subtilisins.
[0111] As used herein, the term "subtilisin" refers to any member of the S8 serine protease family, as described in MEROPS - The Peptidase Data base (Rawlings et al., 2006). The term subtilisin includes the various Bacillus subtilisin proteases that have been identified and sequenced, such as subtilisin 168, subtilisin BPN′, subtilisin Carlsberg, etc., and includes mutant (variant) proteases derived therefrom, etc.
[0112] In certain one or more embodiments, exemplary subtilisin reporters include, but are not limited to, native Bacillus clausii subtilisin and its functional variants, native Bacillus gibsonii subtilisin and its functional variants, native Bacillus lentus subtilisin and its functional variants, native Bacillus licheniformis subtilisin (AprL) and its functional variants, native Bacillus subtilis subtilisin (AprE) and its functional variants, native Bacillus amyloliquefaciens subtilisin (BPN′) and its functional variants, etc. In certain aspects, exemplary Bacillus clausii, Bacillus gibsonii, and / or Bacillus lentus subtilisin reporters may be referred to as alkaline proteases. For example, alkaline subtilisins typically have an isoelectric point (pI) of about 9.5, while Bacillus licheniformis, Bacillus subtilis, and Bacillus amyloliquefaciens subtilisins have a pI of about 6.5.
[0113] In certain embodiments, the present disclosure relates to one or more variant subtilisins that are derived from a parental (native) subtilisin sequence, such as native Bacillus subtilis subtilisin (e.g., 168), native Bacillus amyloliquefaciens (e.g., BPN′), native Bacillus licheniformis subtilisin (e.g., Carlsberg), native Bacillus lentus subtilisin (e.g., 309), Bacillus alcalophilus subtilisin (e.g., PB92), etc. Those skilled in the art can readily design, construct, screen, and identify functional subtilisin variants using conventional methods known in the art. In particular, PCT Publication Nos. WO 2010 / 056634, WO 2011 / 130222, WO 2015 / 089447, WO 2016 / 202839, WO 2017 / 207762, and WO 2023 / 114936 (each incorporated herein by reference in its entirety) describe suitable methods and compositions for constructing functional subtilisin variants derived from native Bacillus clausii subtilisin, functional subtilisin variants derived from native Bacillus amyloliquefaciens subtilisin, functional subtilisin variants derived from native Bacillus gibsonii subtilisin, etc.
[0114] As used herein, "suitable regulatory sequences" refers to nucleotide sequences that are located upstream (5' non-coding sequence), internal, or downstream (3' non-coding sequence) of a coding sequence and that affect the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences can include promoters, transcriptional leader sequences, RNA processing sites, effector binding sites, and stem-loop structures.
[0115] As used herein, "host cell" refers to a cell that has the capacity to serve as a host or expression vehicle for newly introduced DNA sequences. Thus, in certain embodiments of the present disclosure, the host cell is a Gram-positive cell (e.g., Bacillus species) and / or a Gram-negative cell (e.g., Escherichia coli).
[0116] As used herein, "modified cell" refers to a recombinant cell that contains at least one genetic modification that is not present in the parental cell, reference cell, or control cell from which the modified cell was derived.
[0117] As used herein, when comparing the expression and / or production of a protein of interest (POI) in a recombinant (modified) cell to the expression and / or production of the same POI in an unmodified (control) cell, it is understood that the modified and unmodified cells are grown / cultured / fermented under the same conditions (e.g., the same conditions such as medium, temperature, pH, etc.).
[0118] As used herein, when used in the phrase "a recombinant cell 'expresses / produces an increased amount' of a protein of interest relative to an unmodified (control) cell", the "increased amount" specifically refers to the "increased amount" of the protein of interest (POI) expressed / produced in the recombinant cell, which "increased amount" is always relative to an unmodified (control) cell that expresses / produces the same POI, where the modified and unmodified cells are grown / cultured / fermented under the same conditions.
[0119] As used herein, "increasing" protein production or "increased" protein production means an increase in the amount of protein produced (e.g., the protein of interest). The protein can be produced within the host cell or secreted (or transported) into the culture medium. In certain embodiments, the protein of interest is produced (secreted) into the culture medium. Increased protein production can be detected as, for example, a higher maximum level of protein or enzyme activity (such as amylase activity) or total extracellular protein produced as compared to the parental host cell.
[0120] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) introducing, substituting, or removing one or more nucleotides in a gene (or its ORF), or introducing, substituting, or removing one or more nucleotides in a regulatory element required for transcription or translation of a gene or its ORF, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) specific mutagenesis of any one or more of the genes disclosed herein, and / or (g) random mutagenesis.
[0121] As used herein, as used in phrases such as "introducing a 'gene', 'polynucleotide', 'open reading frame (ORF)', 'gene coding sequence', 'vector', 'expression cassette' into a Gram - positive bacterial cell", the term "introducing" includes methods known in the art for introducing polynucleotides (DNA) into cells, including but not limited to protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation, etc.
[0122] As used herein, "transformed" or "transformation" means transforming a cell by using recombinant DNA technology. Transformation typically occurs by inserting one or more nucleotide sequences (e.g., polynucleotide, ORF, or gene) into the cell. The inserted nucleotide sequence can be a heterologous nucleotide sequence (i.e., a sequence that is not naturally present in the cell to be transformed). Thus, transformation generally refers to the introduction of foreign DNA into a host cell such that the DNA remains as a chromosomal integrant or a self - replicating extrachromosomal vector.
[0123] As used herein, "transforming DNA", "transformation sequence", and "DNA construct" refer to DNA used to introduce a sequence into a host cell or organism. The transforming DNA is DNA used to introduce a sequence into a host cell or organism. The DNA can be generated in vitro by PCR or any other suitable technique. In some embodiments, the transforming DNA comprises an input sequence, while in other embodiments, it further comprises an input sequence flanked by homology boxes. In still other embodiments, the transforming DNA comprises additional non-homologous sequences (i.e., filler sequences or flanks) added to the ends. The ends can be closed such that the transforming DNA forms a closed loop, such as when inserted into a vector.
[0124] As used herein, "disruption of a gene" or "gene disruption" are used interchangeably and broadly refer to any genetic modification that substantially prevents a host cell from producing a functional gene product (e.g., a protein). Thus, as used herein, gene disruption includes, but is not limited to, frameshift mutations, premature stop codons (i.e., such that a functional protein is not produced), substitutions that eliminate or reduce activity of an internal deletion of a protein (such that a functional protein is not produced), insertions that disrupt the coding sequence, mutations that remove the operable linkage between the native promoter required for transcription and the open reading frame, etc.
[0125] As used herein, "input sequence" refers to a DNA sequence introduced into the chromosome of a bacterial cell. In some embodiments, the input sequence is part of a DNA construct. In other embodiments, the input sequence encodes one or more proteins of interest. In some embodiments, the input sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it can be a homologous or heterologous sequence). In some embodiments, the input sequence encodes one or more proteins of interest, genes, and / or mutated or modified genes. In alternative embodiments, the input sequence encodes a functional wild-type gene or operon, a functional mutated gene or operon, or a non-functional gene or operon. In some embodiments, a non-functional sequence can be inserted into a gene to disrupt its function. In another embodiment, the input sequence includes a selectable marker. In additional embodiments, the input sequence includes two homology boxes.
[0126] As used herein, a "homeobox" refers to a nucleic acid sequence that is homologous to a sequence in the chromosome of a bacterial cell. More particularly, according to the present invention, a homeobox is an upstream or downstream region that has a sequence identity of between about 80% and 100%, between about 90% and 100%, or between about 95% and 100% with the directly flanking coding regions of a gene or a portion of a gene to be deleted, disrupted, inactivated, downregulated, etc. These sequences direct the integration position of a DNA construct in the bacterial cell chromosome and direct which part of the chromosome is replaced by the input sequence. Although not intended to limit the disclosure, a homeobox can include between about 1 base pair (bp) and 200 kilobases (kb). Preferably, the homeobox includes between about 1 bp and 10.0 kb; between 1 bp and 5.0 kb; between 1 bp and 2.5 kb; between 1 bp and 1.0 kb; and between 0.25 kb and 2.5 kb. A homeobox can also include about 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb, and 0.1 kb. In some embodiments, the 5' and 3' ends of a selectable marker are flanked by homeoboxes, where the homeoboxes contain nucleic acid sequences that are closely flanked to the coding regions of a gene.
[0127] As used herein, a host cell "genome", a bacterial (host) cell "genome", or a Bacillus species (host) cell "genome" includes chromosomal genes and extrachromosomal genes.
[0128] As used herein, the terms "plasmid", "vector", and "cassette" refer to extrachromosomal elements that typically carry genes that are not part of the central metabolism of a cell and are typically in the form of circular double-stranded DNA molecules. Such elements can be linear or circular self-replicating sequences, genomic integration sequences, phages, or nucleotide sequences of single-stranded or double-stranded DNA or RNA derived from any source, where multiple nucleotide sequences have been ligated or recombined into a single construct that is capable of introducing a promoter fragment and a DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences, into a cell.
[0129] As used herein, the term "plasmid" refers to a circular double-stranded (ds) DNA construct that is used as a cloning vector and forms an extrachromosomal self-replicating genetic element in many bacteria and some eukaryotes. In some embodiments, the plasmid is incorporated into the genome of a host cell. In some embodiments, the plasmid is present in a parental cell and is lost in daughter cells.
[0130] As used herein, a "transformation cassette" refers to a specific vector that contains a gene (or its ORF) and, in addition to the foreign gene, has elements that facilitate the transformation of a specific host cell.
[0131] As used herein, the term "vector" refers to any nucleic acid that can replicate (propagate) in a cell and can carry a new gene or DNA segment into a cell. Thus, the term refers to a nucleic acid construct designed for transfer between different host cells. Vectors include viruses, bacteriophages, proviruses, plasmids, phagemids, transposons, and artificial chromosomes such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), PLACs (plant artificial chromosomes), etc. that are "episomes" (i.e., they replicate autonomously or can integrate into the chromosome of the host organism).
[0132] An "expression vector" is a vector that has the ability to incorporate and express heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and are known to those of skill in the art. The selection of an appropriate expression vector is within the knowledge of those of skill in the art.
[0133] As used herein, the terms "expression cassette" and "expression vector" refer to a nucleic acid construct that is recombinantly or synthetically produced and has a series of specified nucleic acid elements (i.e., these are vectors or vector elements as described above) that permit transcription of a specific nucleic acid in a target cell. A recombinant expression cassette can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes (among other sequences) the nucleic acid sequence to be transcribed and a promoter. In some embodiments, the DNA construct also includes a series of specified nucleic acid elements that permit transcription of a specific nucleic acid in a target cell. In certain embodiments, the DNA constructs of the present disclosure contain selectable markers and inactivated chromosomes, or genes, or DNA segments as defined herein.
[0134] As used herein, a "targeting vector" is a vector that includes a polynucleotide sequence that is homologous to a region in the chromosome of the host cell into which the targeting vector is transformed and that can drive homologous recombination at that region. For example, a targeting vector can be used to introduce a mutation into the chromosome of a host cell by homologous recombination. In some embodiments, the targeting vector contains, for example, additional non-homologous sequences (i.e., filler sequences or flanking sequences) added to the ends. The ends can be closed such that the targeting vector forms a closed loop, such as in an insertion vector. For example, in certain embodiments, a (host) cell is modified (e.g., transformed) by introducing one or more "targeting vectors" into a parental Bacillus licheniformis cell.
[0135] As used herein, the term "protein of interest" or "POI" refers to a polypeptide of interest that is desired to be expressed in a modified (recombinant) Gram-positive host cell, wherein the POI is preferably expressed at an increased level (i.e., relative to an "unmodified" (parental or control) cell). Thus, as used herein, the POI can be an enzyme, a substrate-binding protein, a surfactant protein, a structural protein, a receptor protein, etc. In certain embodiments, the modified cells of the present disclosure produce an increased amount of a heterologous protein of interest relative to control cells. In particular embodiments, the increased amount of the protein of interest produced by the modified cells of the present disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or more than a 5.0% increase relative to control cells.
[0136] Similarly, as defined herein, the "gene of interest" or "GOI" refers to a nucleic acid sequence (e.g., polynucleotide, gene, or ORF) encoding a POI. The "gene of interest" encoding the "protein of interest" can be a naturally occurring gene, a mutated gene, or a synthetic gene.
[0137] As used herein, the terms "polypeptide" and "protein" are used interchangeably and refer to any length polymer comprising amino acid residues linked by peptide bonds. The conventional one (1)-letter or three (3)-letter codes for amino acid residues are used herein. The polypeptide can be linear or branched, it can contain modified amino acids, and it can be interrupted by non-amino acids. The term polypeptide also encompasses amino acid polymers that have been modified either naturally or by intervention; e.g., disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within this definition are, for example, polypeptides containing one or more amino acid analogs (including, e.g., non-natural amino acids, etc.) and other modifications known in the art.
[0138] In certain embodiments, the genes of the present disclosure encode proteins for commercial and industrial purposes, such as enzymes (e.g., acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof).
[0139] As used herein, a "variant" polypeptide refers to a polypeptide typically derived from a parent (or reference) polypeptide by substitution, addition, or deletion of one or more amino acids through recombinant DNA techniques. Variant polypeptides may differ from the parent polypeptide by a small number of amino acid residues and can be defined by their level of amino acid sequence homology / identity to the parent (reference) polypeptide. Preferably, the variant polypeptide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% amino acid sequence identity to the parent (reference) polypeptide sequence.
[0140] As used herein, a "variant" polynucleotide refers to a polynucleotide having a specified degree of sequence homology / identity to a parent polynucleotide, or a polynucleotide that hybridizes to the parent polynucleotide (or its complementary sequence) under stringent hybridization conditions. Preferably, the variant polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% nucleotide sequence identity to the parent (reference) polynucleotide sequence.
[0141] As used herein, a "mutation" refers to any change or alteration in a nucleic acid sequence. There are several types of mutations, including point mutations, deletion mutations, silent mutations, frameshift mutations, splicing mutations, etc. Mutations can be made specifically (e.g., via site-directed mutagenesis) or randomly (e.g., via chemical agents, by passage through a repair-minus bacterial strain).
[0142] As used herein, in the context of a polypeptide or its sequence, the term "substituted" means that one amino acid is replaced by another amino acid (i.e., substituted).
[0143] As used herein, the term "homology" relates to homologous polynucleotides or polypeptides. If two or more polynucleotides or two or more polypeptides are homologous, this means that the homologous polynucleotides or polypeptides have a "degree of identity" of at least 50%, at least 60%, more preferably at least 70%, even more preferably at least 85%, still more preferably at least 90%, more preferably at least 95%, and most preferably at least 98%.
[0144] The degree of homology between sequences can be determined using any suitable method known in the art (see, e.g., Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; programs in the Wisconsin Genetics Software Package (Genetics Computer Group, Madison, WI), such as GAP, BESTFIT, FASTA, and TFASTA; and Devereux et al., 1984). For the purposes of the present invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970) as implemented by the Needle program in the EMBOSS package (Rice et al., 2000) (preferably version 3.0.0 or later) is used to determine the degree of identity between two amino acid sequences. The optional parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EBLOSUM62 (the EMBOSS version of BLOSUM62) substitution matrix. The output of Needle labeled "longest identity" (obtained using the non-simplified option) is used as the percentage identity and calculated as follows:
[0145] (Number of identical residues x 100) / (Alignment length - Total number of gaps in the alignment)
[0146] For the purposes of the present invention, the degree of identity between two deoxyribonucleotide sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, ibid.) as implemented by the needle program as in the EMBOSS package (Rice et al., 2000, ibid.) (preferably version 3.0.0 or later). The optional parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EDNAFULL (EMBOSS version of NCBINUC4.4) substitution matrix. The output of needle labeled "longest identity" (obtained using the non-simplified option) is used as the percentage identity and is calculated as follows:
[0147] (Number of identical deoxyribonucleotides x 100) / (Alignment length - total number of gaps in the alignment)
[0148] As used herein, in the case of at least two nucleic acids or polypeptides, the phrases "substantially similar" and "substantially identical" typically mean that the polynucleotide or polypeptide contains a sequence having at least about 40% identity, at least about 50% identity, at least about 60% identity, at least about 70% identity, at least about 75% identity, at least about 80% identity, at least about 85% identity, at least about 90% identity, at least about 91% identity, at least about 92% identity, at least about 93% identity, at least about 94% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, or even at least about 99% identity, or higher identity compared to a reference (i.e., wild-type) sequence. Sequence identity can be determined using known programs such as BLAST, ALIGN, and CLUSTAL using standard parameters.
[0149] As used herein, the term "percent identity (%)" refers to the level of nucleic acid or amino acid sequence identity between the nucleic acid sequence encoding a polypeptide or the amino acid sequence of a polypeptide when aligned using a sequence alignment program.
[0150] As used herein, "specific productivity" is the total amount of protein produced per cell per time in a given period.
[0151] As used herein, the terms "purified", "isolated", or "enriched" mean that a biomolecule (e.g., a polypeptide or polynucleotide) has been altered from its native state by separating it from some or all of the naturally occurring components with which it is associated in nature. Such separation or purification can be accomplished by separation techniques well known in the art such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulfate precipitation or other protein salt precipitation, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis, or gradient separation to remove unwanted whole cells, cell debris, impurities, foreign proteins, or enzymes from the final composition. Components that provide additional benefits, such as activators, anti-inhibitors, desired ions, pH controlling compounds, or other enzymes or chemicals, can then be further added to the purified or isolated biomolecule composition.
[0152] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include: (a) introducing, substituting, or removing one or more nucleotides in a gene (or its ORF), or introducing, substituting, or removing one or more nucleotides in a regulatory element required for transcription or translation of a gene or its ORF, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-specific mutagenesis of any one or more of the genes disclosed herein, and / or (g) random mutagenesis.
[0153] II. Novel Promoter Region Mutant Sequences
[0154] Certain Gram-positive bacterial promoters suitable for expressing a target protein have been described (e.g., Kim et al., 2008, U.S. Patent No. 4,559,300, PCT Publication No. WO 2013 / 086219, and PCT Publication No. WO 2017 / 152169). As presented herein and illustrated in the examples below, the applicant designed and constructed a site evaluation library (SEL) to test / screen for genetic modifications (mutations) to improve recombinant protein productivity.
[0155] Specifically, the promoter and 5'-UTR regions of an expression construct encoding an exemplary reporter protein were genetically modified, where the SEL construct was introduced into and expressed in recombinant Bacillus species cells. More specifically, as generally described in the examples, the mature subtilisin reporter protein was expressed in Bacillus subtilis under the control of variant (mutant) promoter / 5'-UTR region sequences (see Tables 2 and SEQ ID NOs: 8-46) constructed by site-scanning mutagenesis libraries (see Table 1), or under the control of the reference (control) rrnI-P2 promoter / 5'-UTR region of SEQ ID NO: 1 (Example 1).
[0156] As described in Example 2, a recombinant strain expressing a subtilisin reporter protein under the control of a mutant promoter / 5′-UTR region sequence (Table 2) was compared to a reference strain expressing the same subtilisin reporter under the control of the (reference) rrnI-P2 promoter / 5′-UTR region sequence (SEQ ID NO:1). Similarly, as described in Example 3, the performance index (PI) values (Table 3) of recombinant strains that exhibited increased protein productivity after seventy-two (72) hours of growth were compared to a control strain.
[0157] Specifically, as described in Example 3, compared to the reference rrnI-P2 promoter / 5′-UTR region of SEQ ID NO:1, approximately 50% of the constructed SEL variants contained mutations upstream (5′) of the UP element, while approximately 22% of the SEL variants contained mutations within the UP element (see Figure 2 A and Figure 3 A). Compared to the reference rrnI-P2 / 5′-UTR region of SEQ ID NO:1, approximately 22% of the SEL variants contained mutations around the Shine Dalgarno (SD) region (see Figure 2 B and Figure 3 B). In contrast, approximately 72% of the SEL variants that contained mutations in the -35 and / or -10 promoter elements had performance index (PI) values between 0.1 and 0.5 after 24, 48, and 72 hours of growth, and an additional 21% of the SEL variants that contained mutations in the SD region had PI values between 0.1 and 0.5 after 24, 48, and 72 hours of growth (data not shown).
[0158] Thus, as generally described above, certain aspects of the present disclosure relate to the novel mutant promoter / 5′-UTR region nucleic acids described herein. Accordingly, certain embodiments relate to such novel mutant promoter / 5′-UTR region nucleic acid (DNA) sequences for gene coding sequences (CDS) suitable for expressing a protein of interest.
[0159] III. Recombinant Polynucleotides and Molecular Biology
[0160] As generally described above, certain embodiments of the present disclosure relate to novel mutant promoter / 5′-UTR region sequences suitable for expressing the gene CDS encoding a protein of interest. In a related aspect, the present disclosure provides recombinant polynucleotides comprising one or more mutant promoter / 5′-UTR region nucleic acid (DNA) sequences. Accordingly, certain embodiments relate to recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant Gram-positive bacterial cells / strains expressing a protein of interest, and the like. In certain aspects, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells (strains) for enhancing the production of a protein of interest. In certain aspects, the polynucleotide constructs of the present disclosure are referred to as expression cassettes, wherein the cassette comprises, in a 5′ to 3′ direction and in operable combination, at least an upstream (5′) promoter / 5′-UTR region DNA sequence, which is linked to a downstream (3′) gene CDS encoding a mature protein of interest (POI).
[0161] In certain aspects, the expression cassette comprises a variant promoter / 5′-UTR region of the present disclosure operably linked to a downstream gene CDS encoding a mature POI. In certain other aspects, the expression cassettes of the present disclosure comprise one or more DNA sequence elements, including but not limited to DNA sequence elements encoding protein / peptide signal (secretion) sequences (SS), DNA sequence elements encoding propeptide (pro-region) amino acid residues (PRO), DNA sequence elements comprising transcription terminator sequences (term), DNA sequence elements comprising 5′-UTR, 3′-UTR, and the like.
[0162] As generally described above, certain embodiments of the present disclosure relate to novel mutant pro-region sequences. In a related aspect, the present disclosure provides recombinant polynucleotides comprising one or more mutant pro-region nucleic acid (DNA) sequences. Accordingly, certain embodiments relate to recombinant polynucleotides (e.g., vectors, plasmids, expression cassettes, etc.), recombinant Gram-positive bacterial cells / strains expressing a protein of interest, and the like. In certain aspects, the present disclosure provides polynucleotide constructs suitable for introduction into recombinant Gram-positive bacterial cells (strains) for enhancing the production of a protein of interest. In certain aspects, the polynucleotide constructs of the present disclosure are referred to as expression cassettes, wherein the cassette comprises, in a 5′ to 3′ direction and in operable combination, at least an upstream (5′) pro-region DNA sequence, which is linked to a downstream (3′) gene CDS encoding a mature protein of interest (POI).
[0163] For example, one or more nucleic acid sequences described herein can be generated by using any suitable synthetic, manipulative, and / or isolation techniques or combinations thereof. For example, one or more polynucleotides described herein can be generated using standard nucleic acid synthesis techniques well known to those of skill in the art, such as solid-phase synthesis techniques. In such techniques, typically fragments of up to fifty (50) or more nucleotide bases are synthesized and then ligated (e.g., by enzymatic or chemical ligation methods) to substantially form any desired continuous nucleic acid sequence. The synthesis of one or more polynucleotides described herein can also be facilitated by any suitable methods known in the art, including but not limited to chemical synthesis using the following methods: the classical phosphoramidite method (e.g., Beaucage and Caruthers, 1981), or the method described by Matthes et al. (1984) as typically practiced in automated synthesis methods. One or more polynucleotides described herein can also be generated using an automated DNA synthesizer. Custom nucleic acids can be ordered from a variety of commercial sources (e.g., ATUM (DNA 2.0), Newark, CA, USA; Life Tech (GeneArt), Carlsbad, CA, USA; GenScript, Ontario, Canada; Base Clear B.V., Leiden, Netherlands; Integrated DNA Technologies, Skokie, IL, USA; Ginkgo Bioworks (Gen9), Boston, MA, USA; and Twist Bioscience, San Francisco, CA, USA). Other techniques and related principles for synthesizing nucleic acids are described and known in the art.
[0164] Recombinant DNA techniques for modifying nucleic acids are well known in the art, such as restriction endonuclease digestion, ligation, reverse transcription and cDNA production, and polymerase chain reaction (e.g., PCR). One or more of the polynucleotides described herein can also be obtained by screening a cDNA library using one or more oligonucleotide probes that can hybridize to or PCR amplify polynucleotides encoding one or more of the variants described herein. Procedures for screening and isolating cDNA clones and PCR amplification procedures are well known to those skilled in the art and are described in standard references known to those skilled in the art. One or more of the polynucleotides described herein can be obtained by altering a naturally occurring polynucleotide backbone (e.g., encoding one or more of the precursor region sequences of the variants described herein) by, for example, known mutagenesis procedures (e.g., site-directed mutagenesis, site saturation mutagenesis, and in vitro recombination). A variety of methods suitable for generating modified polynucleotides described herein encoding one or more of the variants described herein are known in the art, including but not limited to, for example, site saturation mutagenesis, scanning mutagenesis, insertional mutagenesis, deletion mutagenesis, random mutagenesis, site-directed mutagenesis and directed evolution, and various other recombination methods.
[0165] As generally described above and further described in the examples below, certain embodiments of the present disclosure relate to recombinant (modified) Gram-positive cells capable of producing increased amounts of heterologous target proteins. Accordingly, certain embodiments relate to methods for constructing such recombinant Gram-positive cells with increased protein production capabilities. In certain embodiments, one or more expression cassettes encoding a target protein are introduced into the Gram-positive cells of the present disclosure. In an exemplary embodiment, these cassettes are integrated into the genome of the cell. Accordingly, certain embodiments relate to nucleic acid molecules, polynucleotides (e.g., vectors, plasmids, expression cassettes), regulatory elements, etc. suitable for constructing recombinant (modified) Gram-positive host cells.
[0166] Thus, as presented in the examples and generally described herein, the recombinant cells of the present disclosure can be constructed by those skilled in the art using standard and conventional recombinant DNA and molecular cloning techniques well known in the art. Methods for genetic modification include but are not limited to (a) introducing, substituting, or removing one or more nucleotides in a gene, or introducing, substituting, or removing one or more nucleotides in a regulatory element required for transcription or translation of a gene, (b) gene disruption, (c) gene conversion, (d) gene deletion, (e) gene downregulation, (f) site-specific mutagenesis, and / or (g) random mutagenesis.
[0167] In certain embodiments, modified cells of the present disclosure can be constructed by reducing or eliminating the expression of a gene by using methods well known in the art (e.g., insertion, disruption, substitution, or deletion). The portion of the gene to be modified or inactivated can be, for example, the coding region or regulatory elements required for the expression of the coding region.
[0168] Examples of such regulatory or control sequences can be a promoter sequence or a functional part thereof (i.e., a part sufficient to affect the expression of a nucleic acid sequence). Other control sequences for modification include, but are not limited to, a leader sequence, a propeptide sequence, a signal sequence, a transcription terminator, a transcriptional activator, and the like.
[0169] In certain other embodiments, modified cells are constructed by gene deletion to eliminate or reduce the expression of a gene. Gene deletion techniques enable the partial or complete removal of one or more genes, thereby eliminating their expression or expressing non-functional (or reduced-activity) protein products. In such a method, the deletion of a gene can be accomplished by homologous recombination using a plasmid that has been constructed to continuously contain the 5' and 3' regions flanking the gene. The contiguous 5' and 3' regions can be introduced into the cell, for example, on a temperature-sensitive plasmid, in combination with a second selectable marker at the permissive temperature to allow the plasmid to establish in the cell. The cell is then transferred to the non-permissive temperature to select cells that have integrated the plasmid into one of the chromosomal homologous flanking regions. The selection of plasmid integration is effected by selecting the second selectable marker. After integration, the recombination event at the second homologous flanking region is stimulated by transferring the cells to the permissive temperature for several generations without selection. The cells are plated to obtain single colonies, and the colonies are examined for the loss of both selectable markers. Thus, those skilled in the art can readily identify nucleotide regions (suitable for complete or partial deletion) in the coding sequence and / or non-coding sequence of the gene.
[0170] In other embodiments, modified cells are constructed by introducing, substituting, or removing one or more nucleotides in the gene or regulatory elements required for its transcription or translation. For example, nucleotides can be inserted or removed so as to result in the introduction of a stop codon, the removal of a start codon, or a frameshift of the reading frame. Such modifications can be accomplished by site-directed mutagenesis or mutagenesis generated by PCR according to methods known in the art. Thus, in certain embodiments, the genes of the present disclosure are inactivated by complete or partial deletion.
[0171] In another embodiment, a modified cell is constructed by a gene conversion process. For example, in a gene conversion method, a nucleic acid sequence corresponding to one or more genes is mutagenized in vitro to produce a defective nucleic acid sequence, which is then transformed into a parental cell to produce a defective gene. By homologous recombination, the defective nucleic acid sequence replaces the endogenous gene. It may be desirable that the defective gene or gene fragment also encodes a marker that can be used to select for transformants containing the defective gene. For example, the defective gene can be associated with a selectable marker and introduced on a non-replicating or temperature-sensitive plasmid. Selection for plasmid integration is effected by selection for the marker under conditions that do not permit plasmid replication. Selection for the second recombination event leading to gene replacement is effected by examining the colonies for loss of the selectable marker and for acquisition of the mutated gene. Alternatively, the defective nucleic acid sequence can contain an insertion, substitution, or deletion of one or more nucleotides of the gene, as described below.
[0172] In other embodiments, a modified cell is constructed by established antisense techniques using a nucleotide sequence complementary to the nucleic acid sequence of a gene. More particularly, the expression of a gene in a Gram-positive cell can be reduced (downregulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which can be transcribed in the cell and is capable of hybridizing to the mRNA produced in the cell. Under conditions that permit the complementary antisense nucleotide sequence to hybridize to the mRNA, the amount of the translated protein is thus reduced or eliminated. Such antisense methods include, but are not limited to, RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), antisense oligonucleotides, etc., all of which are well known to those of skill in the art.
[0173] In other embodiments, modified cells are generated / constructed via CRISPR-Cas9 editing. For example, a gene encoding a protein of interest can be edited or disrupted (or deleted or downregulated) by means of a nucleic acid-guided endonuclease that finds its target DNA by binding a guide RNA (e.g., Cas9) and Cpf1 or a guide DNA (e.g., NgAgo), which recruits the endonuclease to the target sequence on the DNA, where the endonuclease can create single-stranded or double-stranded breaks in the DNA. This targeted DNA break becomes a substrate for DNA repair and can recombine with a provided editing template to cause gene disruption or deletion. For example, a gene encoding a nucleic acid-guided endonuclease (for this purpose, Cas9 from Streptococcus pyogenes) or a codon-optimized gene encoding the Cas9 nuclease can be operably linked to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells, thereby generating a Gram-positive cell Cas9 expression cassette. Similarly, one or more target sites specific to the gene of interest can be readily identified by those skilled in the art. For example, to construct a DNA construct encoding a gRNA - directed to a target site within the gene of interest, the variable targeting domain (VT) will contain the nucleotides of the target site that are 5' of the protospacer adjacent motif (PAM) (TGG) and are fused to the DNA encoding the Cas9 endonuclease recognition domain (CER) of Streptococcus pyogenes Cas9. The DNA encoding the VT domain and the DNA encoding the CER domain are combined, thereby generating the DNA encoding the gRNA. Thus, a Gram-positive expression cassette for the gRNA is generated by operably linking the DNA encoding the gRNA to a promoter active in Gram-positive cells and a terminator active in Gram-positive cells.
[0174] In certain embodiments, the DNA breaks induced by the endonuclease are repaired / replaced with an input sequence. For example, to precisely repair the DNA breaks generated by the above-described Cas9 expression cassette and gRNA expression cassette, a nucleotide editing template is provided such that the cell's DNA repair machinery can utilize the editing template. For example, approximately 500 bp of the 5' of the target gene can be fused to approximately 500 bp of the 3' of the target gene to generate an editing template that is used by the Gram-positive host machinery to repair the DNA breaks generated by the RGEN.
[0175] Many different methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence) can be used to co-deliver the Cas9 expression cassette, the gRNA expression cassette, and the editing template to filamentous fungal cells. Transformed cells are screened by amplifying the locus of the gene by PCR using forward and reverse primers. These primers can amplify the wild-type locus or the modified locus that has been edited by the RGEN. These fragments are then sequenced using sequencing primers to identify the edited colonies.
[0176] In still other embodiments, modified cells are constructed by random or specific mutagenesis using methods well known in the art, including but not limited to chemical mutagenesis and transposition. Modification of a gene can be carried out by subjecting the parental cells to mutagenesis and screening for mutant cells in which the gene expression has been reduced or eliminated. Mutagenesis, which can be specific or random, can be carried out, for example, by using a suitable physical or chemical mutagen, by using a suitable oligonucleotide, or by subjecting a DNA sequence to mutagenesis generated by PCR. In addition, mutagenesis can be carried out by using any combination of these mutagenesis methods.
[0177] Examples of physical or chemical mutagens suitable for the purposes of the present invention include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl-N'-nitrosoguanidine (NTG), O-methylhydroxylamine, nitrous acid, ethyl methane sulfonate (EMS), sodium bisulfite, formic acid, and nucleotide analogs. When using such reagents, mutagenesis is typically carried out by incubating the parental cells to be mutagenized in the presence of the selected mutagen under suitable conditions and selecting mutant cells that exhibit reduced or no expression of the gene.
[0178] PCT Publication No. WO 2003 / 083125 discloses methods for modifying Gram-positive (Bacillus) cells, such as using PCR fusion to generate Bacillus deletion strains and DNA constructs to bypass Escherichia coli. PCT Publication No. WO 2002 / 14490 discloses methods for modifying Bacillus cells, which include (1) construction and transformation of an integrating plasmid (pComK), (2) random mutagenesis of the coding sequence, signal sequence, and propeptide sequence, (3) homologous recombination, (4) increasing transformation efficiency by adding non-homologous flanks to the transforming DNA, (5) optimizing double crossover integration, (6) site-directed mutagenesis, and (7) markerless deletion.
[0179] Those skilled in the art are familiar with suitable methods for introducing polynucleotide sequences into bacterial cells (e.g., Gram-negative cells, Gram-positive cells). In fact, methods such as transformation, including protoplast transformation and mid-plate assembly, transduction, and protoplast fusion, are known and suitable for the present disclosure. The transformation method is particularly preferably used to introduce the DNA constructs of the present disclosure into host cells.
[0180] In addition to the common methods, in some embodiments, the host cells are directly transformed (i.e., without using an intermediate cell to amplify the DNA construct or otherwise process the DNA construct before introducing it into the host cell). Introducing the DNA construct into the host cell includes those physical and chemical methods known in the art for introducing DNA into the host cell without inserting a plasmid or vector. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, liposomes, etc. In additional embodiments, the DNA construct is co-transformed with a plasmid without inserting into the plasmid. In additional embodiments, a selectable marker is deleted or substantially excised from the modified Bacillus strain by methods known in the art. In some embodiments, the vector is disassembled from the host chromosome, leaving the flanking regions on the chromosome while removing the native chromosomal regions.
[0181] Promoters and promoter sequence regions for expressing genes, their coding sequences (CDS), open reading frames (ORF), and / or variant sequences in Gram-positive cells are generally known to those skilled in the art. The promoter sequences of the present disclosure are typically selected such that they function in Gram-positive cells. For example, promoters that can be used to drive gene expression in Bacillus cells include, but are not limited to, the Bacillus subtilis alkaline protease (aprE) promoter, the α-amylase promoter (amyE) of Bacillus subtilis, the α-amylase promoter (amyL) of Bacillus licheniformis, the α-amylase promoter of Bacillus amyloliquefaciens, the neutral protease (nprE) promoter from Bacillus subtilis, the mutant aprE promoter, or any other promoter from Bacillus licheniformis or other related Bacillus species. Methods for screening and generating a library of promoters with a range of activities (promoter strength) in Bacillus cells are described in published patent application WO 2002 / 14490.
[0182] IV. Fermenting Gram-Positive Cells for Protein Production
[0183] As generally described above, certain embodiments relate to compositions and methods for constructing and obtaining Gram-positive cells with an increased protein production phenotype. Thus, certain embodiments relate to methods for producing a protein of interest in Gram-positive cells by fermenting the cells in a suitable medium. Fermentation methods well known in the art can be used to ferment the Gram-positive cells of the present disclosure.
[0184] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system where the composition of the medium is set at the start of fermentation and does not change during fermentation. At the start of fermentation, the medium is inoculated with one or more desired organisms. In this method, fermentation occurs without adding any components to the system. Typically, batch fermentation qualifies as "batch" with respect to the addition of carbon source, and attempts are often made to control factors such as pH and oxygen concentration. The metabolite and biomass composition of the batch system changes continuously until fermentation stops. In a typical batch culture, cells can progress through a static lag phase to a high-growth logarithmic phase, and finally enter a stationary phase where the growth rate decreases or stops. If untreated, cells in the stationary phase will eventually die. Generally, cells in the logarithmic phase are responsible for the bulk production of products.
[0185] A suitable variation of the standard batch system is the "fed-batch" fermentation system. In this variant of the typical batch system, substrate is added incrementally as fermentation progresses. Fed-batch systems are useful when catabolite repression may inhibit the metabolism of the cells and when a limited amount of substrate is desired in the medium. Measurement of the actual substrate concentration in a fed-batch system is difficult and thus it is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of off-gases (such as CO 2 )). Batch and fed-batch fermentations are commonly used and known in the art.
[0186] Continuous fermentation is an open system where a defined fermentation medium is continuously added to a bioreactor while an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation typically maintains the culture at a constant high density where the cells are mainly in the logarithmic phase of growth. Continuous fermentation allows for the regulation of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient (such as a carbon source or a nitrogen source) is maintained at a fixed rate and all other parameters are allowed to be adjusted. In other systems, many factors that affect growth can be continuously changed while the cell concentration measured by the turbidity of the medium remains constant. Continuous systems strive to maintain steady-state growth conditions. Thus, the loss of cells due to the withdrawal of the medium should be balanced with the cell growth rate in the fermentation. Methods for regulating nutrients and growth factors for continuous fermentation processes and techniques for maximizing the rate of product formation are well known in the field of industrial microbiology.
[0187] In certain embodiments, the desired protein expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures that include separating the host cells from the culture medium by centrifugation or filtration, or, if desired, disrupting the cells and removing the supernatant from the cell fractions and debris. Typically, after clarification, the proteinaceous components of the supernatant or filtrate are precipitated by means of salts (e.g., ammonium sulfate). The precipitated protein is then solubilized and can be purified by a variety of chromatographic procedures (e.g., ion exchange chromatography, gel filtration).
[0188] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. Classical batch fermentation is a closed system in which the composition of the culture medium is set at the beginning of the fermentation and the composition does not change during the fermentation. At the start of the fermentation, the culture medium is inoculated with one or more desired organisms. In this method, the fermentation is allowed to occur without adding any components to the system. Typically, batch fermentation qualifies as a "batch" with respect to the addition of the carbon source and often attempts are made to control factors such as pH and oxygen concentration. The metabolite and biomass composition of the batch system changes continuously until the fermentation stops. In a typical batch culture, the cells can progress through a static lag phase to a high-growth logarithmic phase and finally into a stationary phase where the growth rate decreases or stops. If untreated, the cells in the stationary phase eventually die. Usually, the cells in the logarithmic phase are responsible for the bulk production of the product.
[0189] A suitable variation of the standard batch system is the "fed-batch" fermentation system. In this variant of the typical batch system, the substrate is added incrementally as the fermentation progresses. The fed-batch system is useful when catabolite repression might inhibit the metabolism of the cells and when it is desirable to have a limited amount of substrate in the culture medium. Measurement of the actual substrate concentration in the fed-batch system is difficult and thus it is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of off-gases (e.g., CO 2 )). Batch and fed-batch fermentations are commonly used and known in the art.
[0190] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation typically maintains the culture at a constant high density, where the cells are mainly in the logarithmic growth phase. Continuous fermentation allows for the regulation of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient (such as a carbon source or a nitrogen source) is maintained at a fixed rate, and all other parameters are allowed to be adjusted. In other systems, many factors that affect growth can be continuously changed while the cell concentration measured by the turbidity of the medium remains constant. The continuous system endeavors to maintain steady-state growth conditions. Therefore, the cell loss caused by the withdrawal of the medium should be balanced with the cell growth rate in the fermentation. Methods for regulating nutrients and growth factors for continuous fermentation processes and techniques for maximizing the product formation rate are well known in the field of industrial microbiology.
[0191] In certain embodiments, the protein of interest expressed / produced by the Gram-positive cells of the present disclosure can be recovered from the culture medium by conventional procedures, which include separating the host cells from the culture medium by centrifugation or filtration, or, if desired, disrupting the cells and removing the supernatant from the cell fractions and debris. Typically, after clarification, the protein component of the supernatant or filtrate is precipitated by means of a salt (such as ammonium sulfate). The precipitated protein is then dissolved and can be purified by a variety of chromatographic procedures (such as ion exchange chromatography, gel filtration).
[0192] V. Protein of Interest
[0193] The protein of interest (POI) of the present disclosure can be any endogenous or heterologous protein, and it can be a variant of such a POI. The protein can contain one or more disulfide bridges, or it can be a protein whose functional form is monomeric or polymeric, i.e., the protein has a quaternary structure and is composed of multiple identical (homologous) or different (heterologous) subunits, where the POI or its variant POI is preferably a POI with the desired characteristics.
[0194] For example, in certain embodiments, the modified Gram-positive cells of the present disclosure produce at least about 0.1%, at least about 0.5%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more of the POI more than their unmodified (parent or control) cells.
[0195] In certain embodiments, the modified Gram-positive cells of the present disclosure exhibit an increase in the specific productivity (Qp) of the POI relative to control cells. For example, the detection of specific productivity (Qp) is a suitable method for evaluating protein production. The specific productivity (Qp) can be determined using the following equation:
[0196] “Qp = gP / gDCW·hr”
[0197] Wherein, “gP” is the number of grams of protein produced in the tank; “gDCW” is the number of grams of dry cell weight (DCW) in the tank; and “hr” is the fermentation time in hours starting from the inoculation time, which includes the production time and the growth time.
[0198] Thus, in certain other embodiments, the modified Gram-positive cells of the present disclosure have an increase in specific productivity (Qp) of at least about 0.1%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more relative to the unmodified (parent) cells.
[0199] In certain embodiments, the POI or its variant POI is selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, ligase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0200] Thus, in certain embodiments, the POI or its variant POI is an enzyme selected from Enzyme Commission (EC) numbers EC 1, EC 2, EC 3, EC 4, EC 5, or EC 6.
[0201] There are various assays known to those of ordinary skill in the art for detecting and measuring the activity of proteins expressed intracellularly and extracellularly.
[0202] VI. Exemplary Embodiments
[0203] Non-limiting examples of the compositions and methods disclosed herein are as follows:
[0204] 1. A variant nucleic acid (DNA) sequence comprising at least one mutation shown in any one of SEQ ID NOs: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant nucleic acid sequence are numbered according to SEQ ID NO: 1.
[0205] 2. A variant nucleic acid (DNA) sequence comprising any one of SEQ ID NOs: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant sequence are numbered according to SEQ ID NO: 1.
[0206] 3. The variant nucleic acid according to Example 1 or Example 2, which has at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO: 1.
[0207] 4. The variant nucleic acid according to any one of Examples 1 - 3, which is further defined as a variant rrnI - P2 promoter and 5′ - untranslated region (5′ - UTR) sequence (variant rrnI - P2 / 5′ - UTR sequence).
[0208] 5. A polynucleotide comprising the variant nucleic acid sequence according to any one of Examples 1 - 4.
[0209] 6. A polynucleotide comprising the variant nucleic acid sequence according to any one of Examples 1 - 4, wherein the variant nucleic acid sequence is operably linked to a downstream (3′) nucleic acid sequence encoding a protein of interest (POI).
[0210] 7. A polynucleotide comprising the variant nucleic acid sequence according to any one of Examples 1 - 4, wherein the variant nucleic acid sequence is operably linked to a downstream (3′) nucleic acid sequence encoding a proregion (PRO), and the downstream (3′) nucleic acid sequence is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).
[0211] 8. A polynucleotide comprising the variant nucleic acid sequence according to any one of Examples 1 - 4, wherein the variant nucleic acid sequence is operably linked to a downstream (3′) nucleic acid sequence encoding a protein signal (secretion) sequence (SS), and the downstream (3′) nucleic acid sequence is operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).
[0212] 9. A polynucleotide comprising a variant nucleic acid sequence as described in any one of Examples 1-4, the variant nucleic acid sequence being operably linked to a downstream (3′) nucleic acid sequence encoding a protein signal (secretion) sequence (SS), the downstream (3′) nucleic acid sequence being operably linked to a downstream nucleic acid sequence encoding a pro-region (PRO) sequence, and the downstream nucleic acid sequence being operably linked to a downstream nucleic acid sequence encoding a protein of interest (POI).
[0213] 10. The polynucleotide as described in any one of Examples 5-9, further comprising a downstream (3′) terminator (term) sequence operably linked to the nucleic acid encoding the POI protein.
[0214] 11. The polynucleotide as described in any one of Examples 5-9, wherein the POI is selected from the group consisting of: enzymes, antibodies, receptor proteins, lectins, and regulatory proteins.
[0215] 12. The polynucleotide as described in any one of Examples 5-9, wherein the POI is an enzyme selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0216] 13. The polynucleotide as described in Example 12, wherein the enzyme is a protease.
[0217] 14. The polynucleotide as described in Example 13, wherein the enzyme is subtilisin.
[0218] 15. The polynucleotide as described in Example 14, wherein the subtilisin has at least about 80% to 100% identity with SEQ ID NO:2 or SEQ ID NO:3.
[0219] 16. The polynucleotide as described in Example 7 or Example 9, wherein the proregion (PRO) sequence has at least about 80% to 100% identity with SEQ ID NO: 5.
[0220] 17. The polynucleotide as described in Example 8 or Example 9, wherein the DNA encoding the protein signal (secretion) sequence (SS) has at least about 80% to 100% identity with SEQ ID NO: 4.
[0221] 18. The polynucleotide as described in Example 10, wherein the DNA encoding the terminator (term) sequence has at least about 80% to 100% identity with SEQ ID NO: 6.
[0222] 19. An expression cassette comprising the polynucleotide as described in any one of Examples 5 - 18.
[0223] 20. A Gram - positive bacterial cell comprising the introduced cassette as described in Example 19.
[0224] 21. The Gram - positive cell as described in Example 20, wherein the cassette is integrated into the genome of the cell.
[0225] 22. The Gram - positive cell as described in Example 20, wherein the cell is a Bacillus sp. cell.
[0226] 23. The Gram - positive cell as described in Example 22, wherein the cell is a Bacillus sp. cell selected from the group consisting of: Bacillus subtilis, Bacillus licheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus clausii, Bacillus halodurans, Bacillus megaterium, Bacillus coagulans, Bacillus circulans, Bacillus lautus, and Bacillus thuringiensis.
[0227] 25. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising (a) introducing into the Gram-positive bacterial cell a polynucleotide comprising an upstream (5′) variant promoter and a 5′-untranslated region (5-UTR) nucleic acid sequence, the nucleic acid sequence (i) comprising at least one mutation shown in any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, or (ii) comprising any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, the nucleic acid sequence being operably linked to a downstream (3′) open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for the production of the POI.
[0228] 26. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising (a) introducing into the Gram-positive bacterial cell a polynucleotide comprising an upstream (5′) variant promoter and a 5′-untranslated region (5-UTR) nucleic acid sequence, the nucleic acid sequence (i) comprising at least one mutation shown in any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, or (ii) comprising any one of SEQ ID NO: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, the nucleic acid sequence being operably linked to a downstream (3′) nucleic acid sequence encoding a pro region (PRO), the downstream (3′) nucleic acid sequence being operably linked to a downstream open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for the production of the POI.
[0229] 27. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising (a) introducing into the Gram-positive bacterial cell a polynucleotide comprising an upstream (5′) variant promoter and a 5′-untranslated region (5-UTR) nucleic acid sequence, the nucleic acid sequence (i) comprising at least one mutation shown in any one of SEQ ID NOs: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, or (ii) comprising any one of SEQ ID NOs: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, the nucleic acid sequence being operably linked to a downstream (3′) nucleic acid sequence encoding a protein signal (secretion) sequence (SS), the downstream (3′) nucleic acid sequence being operably linked to a downstream open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
[0230] 28. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising (a) introducing into the Gram-positive bacterial cell a polynucleotide comprising an upstream (5′) variant promoter and a 5′-untranslated region (5-UTR) nucleic acid sequence, the nucleic acid sequence (i) comprising at least one mutation shown in any one of SEQ ID NOs: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, or (ii) comprising any one of SEQ ID NOs: 8 to SEQ ID NO: 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, the nucleic acid sequence being operably linked to a downstream (3′) nucleic acid encoding a protein signal (secretion) sequence (SS), the downstream (3′) nucleic acid being operably linked to a downstream nucleic acid sequence encoding a pro-region (PRO) sequence, the downstream nucleic acid sequence being operably linked to a downstream open reading frame (ORF) encoding the protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
[0231] 29. The method according to any one of embodiments 25 - 28, wherein the variant promoter / 5′-UTR sequence has at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:1.
[0232] 30. The method according to any one of embodiments 25 - 29, further comprising an operably linked downstream (3′) terminator (term) sequence to the ORF encoding the POI.
[0233] 31. The method according to any one of embodiments 25 - 29, wherein the modified cell produces an increased amount of the POI relative to control cells cultured under the same conditions.
[0234] 32. The method according to any one of embodiments 25 - 29, wherein after culturing for about seventy - two (72) hours, the modified cell produces an increased amount of the POI relative to the control cells.
[0235] 33. The method according to any one of embodiments 25 - 32, wherein the POI is selected from the group consisting of: enzymes, antibodies, receptor proteins, lectins, and regulatory proteins.
[0236] 34. The method according to any one of embodiments 25 - 32, wherein the POI is an enzyme selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α - galactosidase, β - galactosidase, α - glucanase, glucan lyase, endo - β - glucanase, glucoamylase, glucose oxidase, α - glucosidase, β - glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetylesterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter protein, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
[0237] 35. The method according to embodiment 34, wherein the enzyme is a protease.
[0238] 36. The method according to embodiment 35, wherein the protease is subtilisin.
[0239] 37. The method according to embodiment 36, wherein the subtilisin has at least about 80% to 100% identity with SEQ ID NO:2 or SEQ ID NO:3.
[0240] 38. The method according to embodiment 26 or embodiment 28, wherein the pro region (PRO) sequence has at least about 80% to 100% identity with SEQ ID NO:5.
[0241] 39. The method according to embodiment 27 or embodiment 28, wherein the DNA encoding the protein signal (secretion) sequence (SS) has at least about 80% to 100% identity with SEQ ID NO:4.
[0242] 40. The method according to embodiment 30, wherein the DNA encoding the terminator (term) sequence has at least about 80% to 100% identity with SEQ ID NO:6.
[0243] 41. The method according to any one of embodiments 25 - 29, wherein the polynucleotide is integrated into the genome of the cell.
[0244] 42. The method according to any one of embodiments 25 - 29, wherein the Gram - positive cell is a Bacillus species cell.
[0245] 43. The method according to embodiment 42, wherein the Bacillus species cell is selected from the group consisting of: Bacillus subtilis, Bacillus licheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus clausii, Bacillus halodurans, Bacillus megaterium, Bacillus coagulans, Bacillus circulans, Bacillus lautus, and Bacillus thuringiensis.
[0246] Examples
[0247] Certain aspects of the present disclosure can be further understood from the following examples, which should not be construed as limiting. Modifications to the materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989).
[0248] Example 1
[0249] Insertion-Deletion Site Assessment of Promoter / 5′-Untranslated Region Sequences
[0250] In this example, the variant Bacillus gibsonii (BG46 variant) subtilisin (SEQ ID NO:3) was used as a reporter to monitor protein expression as described herein. More particularly, a DNA fragment containing the upstream (5′) aprE gene flanking region (5′ aprE gene FR; SEQ ID NO:49) was operably linked to a polynucleotide construct (expression cassette) containing the upstream (5′) Bacillus subtilis variant (control) rrnI-P2 promoter / 5′-UTR region DNA sequence (SEQ ID NO:1), which upstream DNA sequence (SEQ ID NO:1) was operably linked to a DNA sequence encoding the AprE signal sequence (SEQ ID NO:4), which DNA sequence (SEQ ID NO:4) was operably linked to a DNA sequence encoding the Bacillus lentus propeptide sequence (SEQ ID NO:5), which DNA sequence (SEQ ID NO:5) was operably linked to a DNA sequence encoding the mature BG46 variant reporter protein (SEQ ID NO:3), which DNA sequence was operably linked to the Bacillus amyloliquefaciens BPN′ terminator DNA sequence (SEQ ID NO:6), and the polynucleotide construct was operably linked to the downstream (3′) aprE gene flanking region (3′ aprE gene FR; SEQ ID NO:50) sequence, which flanking region sequence included the downstream kanamycin (kan) gene expression cassette. More particularly, the DNA fragment was assembled using standard molecular biology techniques and used as a template to develop linear DNA expression cassettes containing one or more promoter region SEL mutations as described herein.
[0251] Insertion-Deletion Site Assessment Library
[0252] As generally described herein, a seventy-five (75) insertion-deletion (In-Del) site assessment library (SEL) was designed / constructed for the reference (control) rrnI-P2 promoter / 5′-UTR region sequence (SEQ ID NO:1) as generally described in Table 1, and developed as a 4.4 kb fragment by Twist Bioscience HQ (South San Francisco). More particularly, the linear DNA of the expression cassette (Table 2) was used to transform competent Bacillus subtilis cells, where the transformation mixture was plated onto LA plates containing 1.8 ppm kanamycin and incubated overnight at 37°C. Single colonies were picked and grown in Luria broth at 37°C under antibiotic selection.
[0253] DNA sequence analysis was performed to identify unique (variant) promoter / 5′-UTR region sequences, and these sequences were carefully selected into 96-well microtiter plates (MTPs). For example, as Figure 1 shown, the reference rrnI-P2 promoter / 5′-UTR region (SEQ ID NO:1) contains 149 nucleotides, where nucleotide positions are numbered 1-149 in the 5′ to 3′ direction. In particular, as Figure 1 shown, the reference promoter / 5′-UTR region contains nucleotide positions 1-149 of SEQ ID NO:1, and this SEL generates a library of seventy-five (75) unique promoter / 5′-UTR regions by altering (modifying) two (2) adjacent nucleotide positions of SEQ ID NO:1.
[0254] For example, Table 1 listed below shows thirty-one (31) possible variants at the first (1st) site of the reference promoter / 5′-UTR region (i.e., adjacent nucleotide positions 1 (guanine, "G") and 2 (cytosine, "C") of SEQ ID NO:1), where the first column (Table 1) shows the two (2) nucleotide positions at the same position relative to the reference rrnI-P2 p promoter / 5′-UTR region where two (2) nucleotide changes ("mutations") occur (Table 1, 2nd column; "results with GC as reference")
[0255] Table 1 shows the insertion-deletion SELs of 31 variants
[0256]
[0257]
[0258] Similarly, Table 2 below presents the reference rrnI-P2 promoter / 5′-UTR region and forty (40) mutant promoter / 5′-UTR region sequences identified in the SEL (Table 2; UTR-00664, SEQ ID NO:8 to UTR-00752, SEQ ID NO:46). As further discussed in Example 2, in the reporter protein expression experiment, the transformed cells were grown in a medium (enriched semi-defined medium based on MOP buffer, with urea as the main nitrogen source, maltodextrin as the main carbon source, supplemented with 3% soy peptone for robust cell growth, containing antibiotic selection) in 96-well MTPs in an orbital shaker at 32 °C, 300 rpm, and 80% humidity for three (3) days, and then centrifuged and filtered. The clarified culture supernatant was used to measure (assay) the reporter protease activity to determine the productivity level, where samples were collected after the 72-hour time point (Example 3, Table 3). The reporter protease activity assay is further described in Example 2 below.
[0259] Table 2 Variants in the rrnI-P2 Promoter Region
[0260]
[0261] Table 2 (continued) Variants in the rrnI-P2 Promoter Region
[0262]
[0263] Table 2 (continued) Variants in the rrnI-P2 Promoter Region
[0264]
[0265] Table 2 (continued) Variants in the rrnI-P2 Promoter Region
[0266]
[0267] Table 2 (continued) Variants in the rrnI-P2 Promoter Region
[0268]
[0269] Table 2 (continued) Variants in the rrnI-P2 Promoter Region
[0270]
[0271] Table 2 (continued) Variants in the rrnI-P2 Promoter Region
[0272]
[0273] Example 2
[0274] Determination of Protease Activity
[0275] The protease activity of the reporter protein (BG46 variant) was determined by measuring the hydrolysis of the synthetic suc-AAPF-pNA peptide substrate. For the AAPF assay, the reagent solution used was: 100 mM Tris pH 8.6, 10 mM CaCl 2 , 0.005% -80 (Tris / Ca buffer) and 160 mM suc-AAPF-pNA in DMSO (suc-AAPF-pNA stock solution; Sigma: S-7388). To prepare the working solution, one (1) mL of the suc-AAPF-pNA stock solution was added to 100 mL of Tris / Ca buffer and mixed. Enzyme samples were added to a microtiter plate (MTP) containing one (1) mg / mL suc-AAPF-pNA working solution, and activity was measured kinetically at 405 nm for three to five (3-5) minutes at room temperature using a SpectraMax plate reader. Protease activity was expressed as mOD / minute. Specifically, the protease activity of each variant constructed was measured and compared to a reference construct (rrnI-P2 promoter / 5′-UTR region; SEQ ID NO:1) grown in the same plate. The performance index (PI) was measured and presented (Table 3) by dividing the value of the reference sample by the value of the variant sample, as described in Example 3.
[0276] Example 3
[0277] Promoter and 5′-untranslated region modifications for enhanced protein productivity
[0278] As described above, the mature Bacillus subtilis protease reporter protein (BG46 variant) was expressed in Bacillus subtilis under the control of variant (mutant) promoter / 5′-UTR region sequences constructed by site-saturation mutagenesis library (SEL, Example 1), where the reporter protein activity of the constructs was determined (Example 2). As Figure 2 and Figure 3 shown in the DNA sequence alignment of, variant promoter / 5′-UTR region sequences that showed increased productivity after 72-hour growth and had a performance index (PI) greater than (>) 1.2 compared to the reference (control) rrnI-P2 promoter / 5′-UTR region sequence (SEQ ID NO:1) are listed in Table 3 below.
[0279] For example, compared to the reference rrnI-P2 promoter / 5′-UTR region (SEQ ID NO:1), approximately 50% of the SEL variants constructed contained mutations upstream of the UP element, while approximately 22% of the SEL variants contained mutations in the UP element ( Figure 2 A / Figure 3 A). Similarly, compared to the reference rrnI-P2 promoter / 5′-UTR region (SEQ ID NO:1), approximately 22% of the SEL variants contained mutations in the 5′-UTR around the Shine Dalgarno (SD) region ( Figure 2 B / Figure 3 B).
[0280] In contrast, approximately 72% of the SEL variants containing a mutation in the -35 / -10 promoter region had PI values between 0.1 and 0.5 after 72 hours of growth, and an additional 21% of the SEL variants containing a mutation in the 5′-UTR around the SD region had PI values between 0.1 and 0.5 after 72 hours of growth (data not shown).
[0281] Table 3 Performance Index (PI) Reporting Protein Productivity
[0282]
[0283] Table 3 (continued) Performance Index (PI) Reporting Protein Productivity
[0284]
[0285] Table 3 (continued) Performance Index (PI) Reporting Protein Productivity
[0286]
[0287] References
[0288] U.S. Patent No. 4,559,300
[0289] PCT Publication No. WO 2002 / 14490
[0290] PCT Publication No. WO 2003 / 083125
[0291] PCT Publication No. WO 2013 / 086219
[0292] PCT Publication No. WO 2017 / 152169
[0293] Ausubel et al., “Current Protocols in Molecular Biology”, published by Greene Publishing Assoc. and Wiley-Interscience (1987).
[0294] Kim et al., “Comparison of PaprE, PamyE, and PP43 promoter strength for β-galactosidase and staphylokinase expression in Bacillus subtilis”, Biotechnology and Bioprocess Engineering, 13:313, 2008.
[0295] Sambrook et al., “Molecular Cloning: A Laboratory Manual” Cold Spring Harbor Laboratory: Cold Spring Harbor, N.Y. (1989), (2001) and (2012).
Claims
1. A variant nucleic acid comprising at least one mutation shown in any one of SEQ ID NO:8 to SEQ ID NO:46, wherein the nucleotide positions of the variant nucleic acid sequence are numbered according to SEQ ID NO:
1.
2. The variant nucleic acid according to claim 2, having at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:
1.
3. A polynucleotide comprising the variant nucleic acid sequence according to claim 1.
4. The polynucleotide according to claim 3, operably linked to a downstream nucleic acid encoding a protein of interest (POI).
5. The polynucleotide according to claim 3, operably linked to a downstream nucleic acid encoding a pro-region, the downstream nucleic acid encoding the pro-region being operably linked to a downstream nucleic acid encoding a protein of interest (POI).
6. The polynucleotide according to claim 3, operably linked to a downstream nucleic acid encoding a protein signal sequence, the downstream nucleic acid encoding the protein signal sequence being operably linked to a downstream nucleic acid encoding a protein of interest (POI).
7. The polynucleotide according to claim 3, operably linked to a downstream nucleic acid encoding a protein signal sequence, the downstream nucleic acid encoding the protein signal sequence being operably linked to a downstream nucleic acid encoding a pro-region sequence, the downstream nucleic acid encoding the pro-region sequence being operably linked to a downstream nucleic acid encoding a protein of interest (POI).
8. The polynucleotide according to claim 3, further comprising a downstream terminator sequence operably linked to the nucleic acid encoding the POI protein.
9. The polynucleotide according to claim 3, wherein the POI is selected from the group consisting of: enzymes, antibodies, receptor proteins, lectins, and regulatory proteins.
10. The polynucleotide according to claim 11, wherein the POI is an enzyme selected from the group consisting of: acetyl esterase, aminopeptidase, amylase, arabinase, arabinofuranosidase, carbonic anhydrase, carboxypeptidase, catalase, cellulase, chitinase, chymosin, cutinase, deoxyribonuclease, epimerase, esterase, α-galactosidase, β-galactosidase, α-glucanase, glucan lyase, endo-β-glucanase, glucoamylase, glucose oxidase, α-glucosidase, β-glucosidase, glucuronidase, glycosyl hydrolase, hemicellulase, hexose oxidase, hydrolase, invertase, isomerase, laccase, lipase, lyase, mannosidase, oxidase, oxidoreductase, pectate lyase, pectin acetyl esterase, pectin depolymerase, pectin methyl esterase, pectinolytic enzyme, perhydrolase, polyol oxidase, peroxidase, phenol oxidase, phytase, polygalacturonase, protease, peptidase, rhamnogalacturonase, ribonuclease, transferase, transporter, transglutaminase, xylanase, hexose oxidase, and combinations thereof.
11. The polynucleotide according to claim 12, wherein the enzyme is a protease.
12. The polynucleotide according to claim 13, wherein the protease is subtilisin.
13. An expression cassette comprising the polynucleotide according to claim 3.
14. A Gram-positive bacterial cell comprising the introduced cassette according to claim 15.
15. The Gram-positive cell according to claim 16, wherein the cassette is integrated into the genome of the cell.
16. A method for producing a protein of interest (POI) in a Gram-positive bacterial cell, the method comprising: (a) introducing into a Gram-positive bacterial cell a polynucleotide comprising an upstream variant promoter and a 5′-untranslated region (5-UTR) nucleic acid sequence, the nucleic acid sequence: (i) comprising at least one mutation shown in any one of SEQ ID NOs: 8 to 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, or (ii) comprising any one of SEQ ID NOs: 8 to 46, wherein the nucleotide positions of the variant promoter / 5′-UTR sequence are numbered according to SEQ ID NO: 1, the nucleic acid sequence being operably linked to a downstream open reading frame (ORF) encoding a protein of interest (POI); and (b) culturing the modified cell under conditions suitable for producing the POI.
17. The method according to claim 18, wherein the variant promoter / 5′-UTR sequence has at least about 97.5%, 97.6%, 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to SEQ ID NO:
1.
18. The method according to claim 18, wherein the modified cell produces an increased amount of the POI relative to control cells cultured under the same conditions, wherein the control cells comprise an introduced polynucleotide comprising an upstream promoter / 5′-UTR sequence of SEQ ID NO:1, the upstream promoter / 5′-UTR sequence being operably linked to a downstream ORF encoding the same POI as the modified cell.
19. The method according to claim 18, wherein after culturing for at least seventy-two (72) hours, the modified cell produces an increased amount of the POI relative to the control cells.
20. The method according to claim 21, wherein after culturing for at least seventy-two (72) hours, the increased amount of the POI is at least 5% relative to the control cells.
21. The method according to claim 18, wherein the POI is selected from the group consisting of: an enzyme, an antibody, a receptor protein, a lectin, and a regulatory protein.
22. The method according to claim 18, wherein the Gram-positive bacterial cell is a Bacillus sp. cell.
Citation Information
Patent Citations
Trouser remover
US2840285A
Information storage units
US2860287A
Method for using an homologous bacillus promoter and associated natural or modified ribosome binding site-containing DNA sequence in streptomyces
US4559300A
Bacillus transformation, transformants and mutant libraries
WO2002014490A2
Ehanced protein expression in bacillus
WO2003083125A1